Repository
gaasher/Agent-Loop-Skills agent skills
- skills
- 21
- GitHub stars
- 174
GitHub description: “Loop until it's better — drop-in agentic loops (autoresearch, scientific writing, data analysis, code/SQL/prompt optimization, red-teaming) as open-standard Agent Skills. Verification-gated; native on Claude Code, portable across Codex, Cursor & other Skills hosts.”
- Stars
- 174 (19 forks)
- Licence
- MIT
- Last push
- Jun 2026
- Created
- Jun 2026
- agent-skills
- agentic-loops
- ai-agents
- autoresearch
- claude-code
- data-analysis
- llm-agents
- open-source
- red-teaming
- skills
- subagents
- ml-autoresearch
- agentic-workflows
- anthropic
- claude
- literature-review
Install all skills
npx skills add gaasher/Agent-Loop-SkillsAdd --skill <name> for a single skill and -a <agent> to choose the agent (see the agent guides).
Skills in gaasher/Agent-Loop-Skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when the user wants to evolve an ML model/program through population-based search rather than a single sequential refine loop — a generational evolution where parallel… | gaasher/ | 174 | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | 2.Karpathy A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g. | gaasher/ | 174 | 1 repo | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 3 | A skill your agent uses when the user wants an autonomous ML research loop that pressure-tests competing ideas before spending compute — several research subagents each propose one architecture… | gaasher/ | 174 | 1 repo | ~3k | Automated safety check: Pass | MIT | 3 mo ago |
| 4 | A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g. | gaasher/ | 174 | 1 repo | ~2.6k | Automated safety check: Warn | MIT | 3 mo ago |
| 5 | A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed. | gaasher/ | 174 | — | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing… | gaasher/ | 174 | — | ~3.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | A skill your agent uses when the user wants an iterative, self-checking exploratory analysis of a dataset — surfacing findings that are each verified by re-running the computation, not asserted. | gaasher/ | 174 | — | ~1.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 8 | A skill your agent uses when the user wants to generate and literature-vet a pool of novel, testable research hypotheses for a question or domain. | gaasher/ | 174 | — | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 9 | A skill your agent uses when the user wants a structured, saturating literature survey on a question — not a one-shot summary, but an evidence/contradiction matrix (sources × claims) built by… | gaasher/ | 174 | — | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 10 | 10.Plan Loop A skill your agent uses when the user has a coding or engineering prompt and wants it refined into a detailed, executable plan before any code is written — the planning stage of a prompt → plan →… | gaasher/ | 174 | — | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 11 | A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it… | gaasher/ | 174 | — | ~2.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 12 | A skill your agent uses when the user has a prompt that feeds a system they can already score, and wants that prompt automatically improved to raise the score against their own evaluation command. | gaasher/ | 174 | — | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 13 | 13.Purple Team A skill your agent uses when the user wants to automatically harden a guardrail, classifier, content filter, prompt, or API they own by running attack and defense together as a closed loop, not just… | gaasher/ | 174 | — | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 14 | A skill your agent uses when the user has a research proposal (problem + proposed methodology + planned experiments) and wants it iteratively strengthened until it clears a passing grade. | gaasher/ | 174 | — | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 15 | A skill your agent uses when the user has a vague topic or area of interest and wants it sharpened into a few strong, novel, feasible research questions. | gaasher/ | 174 | — | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 16 | A skill your agent uses when the user has scientific data (or a prompt alluding to scientific data) and wants a publication-quality figure made from it. | gaasher/ | 174 | — | ~3.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | A skill your agent uses when the user has a scientific draft (with its dataset, figures, and optional analysis code) and wants it iteratively revised until it clears a quality bar. | gaasher/ | 174 | — | ~3.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 18 | A skill your agent uses when the user has a messy tabular data dump (CSV/TSV/parquet/Excel/JSON) and wants it iteratively cleaned to an inferred data contract — a checklist of deterministic… | gaasher/ | 174 | — | ~4k | Automated safety check: Pass | MIT | 3 mo ago |
| 19 | A skill your agent uses when a loop needs scholarly literature — paper discovery, novelty checks, full-text snippet search, citation-graph traversal, single-paper reads, or experimental-result… | gaasher/ | 174 | — | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 20 | A skill your agent uses when the user wants an autonomous ML research loop that does more than blindly try changes. | gaasher/ | 174 | 1 repo | ~5.2k | Automated safety check: Warn | MIT | 3 mo ago |
| 21 | A skill your agent uses when the user wants to iteratively improve an artifact under a hard correctness bound while minimizing a measured cost — refactoring a code module to cut complexity while its… | gaasher/ | 174 | — | ~2.4k | Automated safety check: Warn | MIT | 3 mo ago |
Questions, answered from the data.
What is the best skill in gaasher/Agent-Loop-Skills?
Alpha Evolve from gaasher/Agent-Loop-Skills ranks first of the 21 skills in gaasher/Agent-Loop-Skills listed here, with the highest score: its repository has 174 GitHub stars, 1 other GitHub owner carry a copy, its SKILL.md loads about 3.4k tokens and it passes the automated safety check with no findings. Next come Karpathy and Tournament Autoresearch.
Are the skills in gaasher/Agent-Loop-Skills official?
None yet. All 21 skills in gaasher/Agent-Loop-Skills listed here come from community repositories; a skill counts as official when the product's own GitHub organization publishes it.
How do I install all skills from gaasher/Agent-Loop-Skills?
Run npx skills add gaasher/Agent-Loop-Skills in your project: the open-source skills CLI installs the repository's skills into your coding agent's skills folder. To install a single skill, open its page here for the exact command.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.