Search
Testing & QA · For data scientists
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release. | R6410418/ | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier. | langchain-ai/ | 1.3k | — | ~4k | Automated safety check: Pass | MIT | 2 days ago |
| 3 | Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor. | smixs/ | 179 | — | ~6.6k | Automated safety check: Pass | MIT | 2 mo ago |
| 4 | Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim. | JetBrains/ | 10k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 5 | browser-based page capture and text extraction for public-opinion research. | 123321kk/ | 107 | — | ~631 | Automated safety check: Pass | No licence | 6 mo ago |
| 6 | Browser automation, debugging, and performance analysis using Puppeteer CLI scripts. | einverne/ | 121 | 1 repo | ~1.6k | Automated safety check: Notes | Apache-2.0 | 1 mo ago |
| 7 | A skill your agent uses when adding support for a new model to VeOmni. | ByteDance-Seed/ | 2.2k | — | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app. | HaolemeApp/ | 157 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | 1 mo ago |
| 9 | Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately. | ai-evals-course/ | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 10 | Image-to-code replication pipeline. An agent skill from Yu-369/VibeCurb. | Yu-369/ | 979 | — | ~8.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 11 | Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 12 | Turn a user-owned coverage CSV and reviewed press-clip captures into an honest, source-linked earned-media dashboard and a motion-designed highlight reel (MP4) that scrolls each real article to the… | elvisun/ | 1.5k | — | ~3.9k | Automated safety check: Pass | MIT | today |
| 13 | Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics. | Orchestra-Research/ | 13k | 4 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 14 | Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data. | ai-evals-course/ | 1.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 15 | Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks. | amd/ | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 16 | Benchmark Claude Code skill performance with variance analysis, tracking pass rate, execution time, and token usage across iterations. | AgriciDaniel/ | 179 | — | ~1.4k | Automated safety check: Pass | MIT | 6 mo ago |
| 17 | Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI. | diegosouzapw/ | 75k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 18 | Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines. | wshobson/ | 40k | 11 repos | ~1.1k | Automated safety check: Pass | MIT | 6 days ago |
| 19 | Guide a tester through hipfire bring-up, serve smoke, claim-scoped harnesses, and benchmark reporting on AMD RDNA/CDNA GPUs. | warpfront/ | 658 | — | ~1.5k | Automated safety check: Pass | Unknown | yesterday |
| 20 | A skill your agent uses when a Token Meter issue or approved feature needs diagnosis or implementation in the repository. | splunk/ | 116 | — | ~1.1k | Automated safety check: Pass | MIT | 2 days ago |
| 21 | Add or modify a Lizard language reader. An agent skill from terryyin/lizard. | terryyin/ | 2.5k | — | ~1.1k | Automated safety check: Pass | Unknown | today |
| 22 | Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships. | timothywarner-org/ | 224 | — | ~696 | Automated safety check: Notes | MIT | 2 mo ago |
| 23 | 23.Test Macafm Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis. | scouzi1966/ | 346 | — | ~7k | Automated safety check: Pass | MIT | yesterday |
| 24 | 24.Vibe To Spec Convert messy voice input, imperfect speech recognition, half-formed product ideas, and iterative corrections into an implementation-ready software specification for Codex, Claude Code, Copilot… | FAIRY123456789/ | 103 | — | ~997 | Automated safety check: Pass | MIT | yesterday |
| 25 | Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification. | paperclipai/ | 100k | — | ~839 | Automated safety check: Pass | MIT | today |
| 26 | Gate quantitative Mira conclusions by requiring reproducible data, formulas, calculation ledgers, or explicit downgrades when numbers drive judgment. | byteseek/ | 275 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 27 | Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats. | glebis/ | 391 | — | ~2.9k | Automated safety check: Pass | MIT | 3 days ago |
| 28 | 28.Verifier Web Evidence-capture protocol for verifying web/dashboard/backoffice/checkout changes in the Polar local stack by driving the real UI with Playwright. | polarsource/ | 10k | — | ~3k | Automated safety check: Notes | MIT | 2 days ago |
| 29 | Data quality validation and analysis accuracy verification. An agent skill from liangdabiao/claude-data-analysis-ultra-main. | liangdabiao/ | 290 | — | ~728 | Automated safety check: Pass | No licence | 5 mo ago |
| 30 | A skill your agent uses when extending phx.gen.auth — adding registration fields, custom user attributes, extra migrations alongside generated auth, fixture updates. | j-morgan6/ | 167 | — | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 31 | Project Guidelines Skill (Example) | davila7/ | 33k | 6 repos | ~2.2k | Automated safety check: Notes | MIT | today |
| 32 | 32.Functional Functional programming patterns with immutable data. An agent skill from citypaul/.dotfiles. | citypaul/ | 740 | — | ~3.6k | Automated safety check: Pass | Unknown | 2 days ago |
| 33 | Review an AlbumentationsX transform for correctness, public API coherence, performance, documentation, and test coverage. | albumentations-team/ | 567 | — | ~973 | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 34 | A skill your agent uses when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows. | alirezarezvani/ | 28k | — | ~3.4k | Automated safety check: Notes | MIT | 1 mo ago |
| 35 | 35.Challenge Stress-test a finding against the choices you did not make. An agent skill from pedrohcgs/claude-code-my-workflow. | pedrohcgs/ | 1.7k | — | ~1.9k | Automated safety check: Notes | MIT | 13 days ago |
| 36 | Example template for project-specific skill files covering architecture, patterns, testing, and deployment. | vibeeval/ | 532 | 3 repos | ~2.2k | Automated safety check: Notes | MIT | 2 mo ago |
| 37 | 用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。 | affaan-m/ | 276k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 38 | 5-stage kernel correctness verification protocol for Triton and CUDA kernels. | ZJLi2013/ | 102 | — | ~702 | Automated safety check: Pass | No licence | 6 mo ago |
| 39 | Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 40 | Generate PR descriptions for SDK pod packages following template and format rules. | tetherto/ | 685 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 41 | Use before running any confirmatory analysis or looking at outcome data, when testing a hypothesis, computing a p-value, or about to claim an effect - locks predictions and decision rules before… | K-Dense-AI/ | 350 | — | ~3.3k | Automated safety check: Pass | Unknown | 28 days ago |
| 42 | Implement an approved repository change with explicit ownership, minimal scope, test-first evidence, and scoped verification. | rapidaai/ | 745 | — | ~661 | Automated safety check: Pass | Unknown | 4 days ago |
| 43 | 43.LLM Patterns AI-first application patterns, LLM testing, prompt management | alinaqi/ | 707 | — | ~2.1k | Automated safety check: Pass | MIT | 17 days ago |
| 44 | Stage 1 of Clinical ASR Flywheel. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 45 | Create Earth2Studio diagnostic model wrappers for single-step data transformations, including simple derived diagnostics, packaged AutoModel diagnostics, and generative or diffusion diagnostics. | NVIDIA/ | 3.6k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 46 | A skill your agent uses when starting empirical analysis, creating a data pipeline, generating results, or when data or model specifications change. | brycewang-stanford/ | 4.6k | — | ~1.7k | Automated safety check: Pass | Unknown | 6 days ago |
| 47 | 47.Firecrawl QA QA test a live website with Firecrawl browser and scrape evidence. | firecrawl/ | 117 | — | ~587 | Automated safety check: Pass | ISC | 2 days ago |
| 48 | Extract battery features for degradation analysis and health monitoring in MATLAB. | matlab/ | 1.1k | — | ~5.9k | Automated safety check: Pass | Unknown | 2 days ago |