Search

Testing & QA · For data scientists

68 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

R6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT3 mo ago
2

Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

langchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT2 days ago
3

Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor.

smixs/skill-conductor179—~6.6kAutomated safety check: PassMIT2 mo ago
4

Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

JetBrains/ideavim10k—~1.3kAutomated safety check: PassMITtoday
5

browser-based page capture and text extraction for public-opinion research.

123321kk/opinion-agent-ultimate107—~631Automated safety check: PassNo licence6 mo ago
6

Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

einverne/dotfiles1211 repo~1.6kAutomated safety check: NotesApache-2.01 mo ago
7

A skill your agent uses when adding support for a new model to VeOmni.

ByteDance-Seed/VeOmni2.2k—~2kAutomated safety check: PassApache-2.0yesterday
8

Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.

HaolemeApp/Haoleme157—~1.3kAutomated safety check: PassAGPL-3.01 mo ago
9

Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

ai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.016 days ago
10

Image-to-code replication pipeline. An agent skill from Yu-369/VibeCurb.

Yu-369/VibeCurb979—~8.7kAutomated safety check: PassMIT2 mo ago
11

Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

ai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.016 days ago
12

Turn a user-owned coverage CSV and reviewed press-clip captures into an honest, source-linked earned-media dashboard and a motion-designed highlight reel (MP4) that scrolls each real article to the…

elvisun/newsjack1.5k—~3.9kAutomated safety check: PassMITtoday
13

Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.9kAutomated safety check: PassMIT3 mo ago
14

Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

ai-evals-course/evals-skills1.5k—~2.2kAutomated safety check: PassApache-2.016 days ago
15

Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

amd/gaia1.6k—~1.8kAutomated safety check: PassMITtoday
16

Benchmark Claude Code skill performance with variance analysis, tracking pass rate, execution time, and token usage across iterations.

AgriciDaniel/skill-forge179—~1.4kAutomated safety check: PassMIT6 mo ago
17

Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

diegosouzapw/OmniRoute75k—~1.3kAutomated safety check: PassMITtoday
18

Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.

wshobson/agents40k11 repos~1.1kAutomated safety check: PassMIT6 days ago
19

Guide a tester through hipfire bring-up, serve smoke, claim-scoped harnesses, and benchmark reporting on AMD RDNA/CDNA GPUs.

warpfront/hipfire658—~1.5kAutomated safety check: PassUnknownyesterday
20

A skill your agent uses when a Token Meter issue or approved feature needs diagnosis or implementation in the repository.

splunk/token-meter116—~1.1kAutomated safety check: PassMIT2 days ago
21

Add or modify a Lizard language reader. An agent skill from terryyin/lizard.

terryyin/lizard2.5k—~1.1kAutomated safety check: PassUnknowntoday
22

Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.

timothywarner-org/claude-code224—~696Automated safety check: NotesMIT2 mo ago
23

Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis.

scouzi1966/maclocal-api346—~7kAutomated safety check: PassMITyesterday
24

Convert messy voice input, imperfect speech recognition, half-formed product ideas, and iterative corrections into an implementation-ready software specification for Codex, Claude Code, Copilot…

FAIRY123456789/human-edge-agent-skills103—~997Automated safety check: PassMITyesterday
25

Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

paperclipai/paperclip100k—~839Automated safety check: PassMITtoday
26

Gate quantitative Mira conclusions by requiring reproducible data, formulas, calculation ledgers, or explicit downgrades when numbers drive judgment.

byteseek/Mira275—~1.1kAutomated safety check: PassApache-2.01 mo ago
27

Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.

glebis/claude-skills391—~2.9kAutomated safety check: PassMIT3 days ago
28

Evidence-capture protocol for verifying web/dashboard/backoffice/checkout changes in the Polar local stack by driving the real UI with Playwright.

polarsource/polar10k—~3kAutomated safety check: NotesMIT2 days ago
29

Data quality validation and analysis accuracy verification. An agent skill from liangdabiao/claude-data-analysis-ultra-main.

liangdabiao/claude-data-analysis-ultra-main290—~728Automated safety check: PassNo licence5 mo ago
30

A skill your agent uses when extending phx.gen.auth — adding registration fields, custom user attributes, extra migrations alongside generated auth, fixture updates.

j-morgan6/elixir-phoenix-guide167—~2kAutomated safety check: PassMIT3 mo ago
31davila7/claude-code-templates33k6 repos~2.2kAutomated safety check: NotesMITtoday
32

Functional programming patterns with immutable data. An agent skill from citypaul/.dotfiles.

citypaul/.dotfiles740—~3.6kAutomated safety check: PassUnknown2 days ago
33

Review an AlbumentationsX transform for correctness, public API coherence, performance, documentation, and test coverage.

albumentations-team/AlbumentationsX567—~973Automated safety check: PassAGPL-3.0yesterday
34

A skill your agent uses when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows.

alirezarezvani/claude-skills28k—~3.4kAutomated safety check: NotesMIT1 mo ago
35

Stress-test a finding against the choices you did not make. An agent skill from pedrohcgs/claude-code-my-workflow.

pedrohcgs/claude-code-my-workflow1.7k—~1.9kAutomated safety check: NotesMIT13 days ago
36

Example template for project-specific skill files covering architecture, patterns, testing, and deployment.

vibeeval/vibecosystem5323 repos~2.2kAutomated safety check: NotesMIT2 mo ago
37

用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。

affaan-m/ECC276k—~1.4kAutomated safety check: PassMITyesterday
38

5-stage kernel correctness verification protocol for Triton and CUDA kernels.

ZJLi2013/awesome-kernel-skills102—~702Automated safety check: PassNo licence6 mo ago
39

Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
40

Generate PR descriptions for SDK pod packages following template and format rules.

tetherto/qvac685—~4.4kAutomated safety check: PassApache-2.0yesterday
41

Use before running any confirmatory analysis or looking at outcome data, when testing a hypothesis, computing a p-value, or about to claim an effect - locks predictions and decision rules before…

K-Dense-AI/science-superpowers350—~3.3kAutomated safety check: PassUnknown28 days ago
42

Implement an approved repository change with explicit ownership, minimal scope, test-first evidence, and scoped verification.

rapidaai/voice-ai745—~661Automated safety check: PassUnknown4 days ago
43

AI-first application patterns, LLM testing, prompt management

alinaqi/maggy707—~2.1kAutomated safety check: PassMIT17 days ago
44

Stage 1 of Clinical ASR Flywheel. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.02 days ago
45

Create Earth2Studio diagnostic model wrappers for single-step data transformations, including simple derived diagnostics, packaged AutoModel diagnostics, and generative or diffusion diagnostics.

NVIDIA/skills3.6k—~3.7kAutomated safety check: PassApache-2.02 days ago
46

A skill your agent uses when starting empirical analysis, creating a data pipeline, generating results, or when data or model specifications change.

brycewang-stanford/Auto-Empirical-Research-Skills4.6k—~1.7kAutomated safety check: PassUnknown6 days ago
47

QA test a live website with Firecrawl browser and scrape evidence.

firecrawl/skills117—~587Automated safety check: PassISC2 days ago
48

Extract battery features for degradation analysis and health monitoring in MATLAB.

matlab/matlab-agentic-toolkit1.1k—~5.9kAutomated safety check: PassUnknown2 days ago