Deep Agents to Pydantic AI Migration
pydantic/pydantic-ai
Migrates Python LangChain Deep Agents applications to Pydantic AI and Pydantic AI Harness while preserving the application's observed behavior.
Decision protocol for wiring a verify-then-fix loop around a code-editing LLM agent.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-test-fix-loop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-test-fix-loop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-test-fix-loop .claude/skills/agentsop-test-fix-loop && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-test-fix-loop" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-test-fix-loop into .claude/skills/agentsop-test-fix-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-test-fix-loop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-test-fix-loopType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-test-fix-loop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-test-fix-loop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-test-fix-loop .agents/skills/agentsop-test-fix-loop && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-test-fix-loop" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-test-fix-loop into .agents/skills/agentsop-test-fix-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-test-fix-loop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-test-fix-loop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-test-fix-loop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-test-fix-loop .cursor/skills/agentsop-test-fix-loop && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-test-fix-loop" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-test-fix-loop into .cursor/skills/agentsop-test-fix-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-test-fix-loop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-test-fix-loop--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-test-fix-loop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-test-fix-loop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-test-fix-loop .gemini/skills/agentsop-test-fix-loop && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-test-fix-loop" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-test-fix-loop into .gemini/skills/agentsop-test-fix-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-test-fix-loop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-test-fix-loopInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-test-fix-loop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-test-fix-loop .github/skills/agentsop-test-fix-loop && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-test-fix-loop" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-test-fix-loop into .github/skills/agentsop-test-fix-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-test-fix-loop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-test-fix-loop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-test-fix-loop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-test-fix-loop .opencode/skills/agentsop-test-fix-loop && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-test-fix-loop" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-test-fix-loop into .opencode/skills/agentsop-test-fix-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-test-fix-loop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-test-fix-loopDecision protocol for wiring a verify-then-fix loop around a code-editing LLM agent.
Agentsop Test Fix Loop is an agent skill from agentsope/SkillAlchemy. Decision protocol for wiring a verify-then-fix loop around a code-editing LLM agent. The agent edits → runs lint/test → reads the output → fixes → re-runs, bounded by an iteration cap and an escalation rule. Activates whenever a coder agent has a verifiable success criterion (exit code, type-checker output, failing assertion) and the user wants the agent to converge to "green" on its own. Framework-agnostic — wraps Aider's --auto-lint/--auto-test, an OpenHands SWE-Bench loop, a manual LangGraph cycle, or Claude…
Its SKILL.md is about 7.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-material.md`).
It sits in AI & LLM Engineering, covering Linting and formatting and Building AI agents. It works with Bash, LangGraph and pytest. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6ea799f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitpytestruffnpmprettiercargoeslintmypytscvitestgopnpmFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
aider.chatgithub.comdocs.langchain.comdocs.cline.botcline.botdocs.pytest.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Test Fix Loop loads about 7.2k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 144 tokens; SKILL.md has 3,343 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit 6ea799f, republished under its MIT licence (© agentsope). 3,343 words, ~7,200 tokens.
.claude/skills/agentsop-test-fix-loop/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.One-liner: The test result IS the next prompt. Wiring the verifier is 20% of the work; framing its output as a useful feedback message is 80%.
Activate this skill when any of the following triggers fire:
aider --auto-test,
cline --yes, or an OpenHands-style headless agent.Do not activate when:
pytest -x -k changed) in
the loop and gate the slow suite at PR review.The agent's next turn is conditioned almost entirely on the message you
inject between edit-N and edit-N+1. That message — formatted from
stdout, stderr, exit_code — is the prompt. The framework labels it
"tool result" or "verifier output" but mechanically it is a user-role message
the LM consumes verbatim.
⇒ Framing the feedback dominates the model choice. A 4000-line raw pytest dump prompts a worse fix than a 30-line "first failing test, traceback, the diff you just applied" digest, regardless of the model behind it.
+-----------------+ +-----------------+ +-----------------+ +-----------------+
| 1. Verifier | | 2. Capture | | 3. Format | | 4. Iteration |
| command | | (stdout + | | feedback | | bound |
| | | stderr + | | message | | |
| - pytest -x | | exit_code) | | - first error | | - max N tries |
| - ruff check | | - timeout cap | | - last K lines | | - escalate / |
| - mypy --strict | | - byte cap | | - drop noise | | commit / skip |
| - eslint . | | - kill on hang | | - keep colors=0 | | |
+-----------------+ +-----------------+ +-----------------+ +-----------------+Drop any one of these and the loop fails:
[oh/6357] is the canonical failure case.Naively: "let the agent run pytest and read the output". This breaks because:
The loop is a contract: *verifier wiring + output capture + feedback framing
| Verifier returns | Interpretation | Next action |
|---|---|---|
exit 0, no diagnostics | True success | Commit + exit loop |
exit 0, warnings | Soft success | Commit + log; optionally surface to user |
exit != 0, parseable error | Actionable failure | Format → feed back → next iter |
exit != 0, unparseable (e.g. segfault, OOM) | Environment / infra failure | Escalate; do not re-prompt the LM |
| Timeout / hang | Likely infinite loop in code | Kill, format as timeout error, escalate after 1 retry |
Pick the cheapest verifier that catches the class of bug you care about. Cascade from fastest to slowest:
| Stage | Command (concrete) | Catches | Typical latency |
|---|---|---|---|
| 1. Format | ruff format --check . / prettier --check . | Style | <1 s |
| 2. Lint | ruff check . / eslint . | Style + obvious bugs | 1–5 s |
| 3. Type | mypy --strict src/ / tsc --noEmit | Type errors | 5–30 s |
| 4. Test | pytest -x --ff / vitest run --bail 1 | Behavioural | 10 s–min |
| 5. Build | cargo build / go build ./... / npm run build | Link / compile | 10 s–min |
Rule: bind --lint-cmd and --test-cmd to stages 1–4 combined into one
shell command (ruff check . && pytest -x). This way one feedback message
covers all signals; you don't loop separately on lint then on tests.
For Aider:
aider --auto-lint --lint-cmd "ruff check ." \
--auto-test --test-cmd "pytest -x --tb=short"For Claude Code / generic agent:
result = subprocess.run(
["bash", "-c", "ruff check . && pytest -x --tb=short"],
capture_output=True, text=True, timeout=120
)result = subprocess.run(
cmd, capture_output=True, text=True, timeout=120, env={**os.environ, "NO_COLOR": "1"}
)
captured = {
"exit_code": result.returncode,
"stdout": result.stdout,
"stderr": result.stderr,
"timed_out": False,
}Common mistakes:
tsc / cargo go to stderr. Always capture both.NO_COLOR=1 — ANSI escapes burn tokens and confuse the model.cargo build log kills your context window.The single biggest lever in this skill. Don't paste raw output. Distill to:
The verifier failed (exit 1, pytest -x --tb=short).
FIRST FAILING TEST:
tests/test_auth.py::test_jwt_expiry — AssertionError: expected 401, got 200
TRACEBACK (last frame):
File "src/auth.py", line 47, in verify_token
if exp < now: return None
TypeError: '<' not supported between instances of 'NoneType' and 'datetime'
YOUR LAST EDIT touched src/auth.py:40-50.
Hypothesis: `exp` is None when the JWT lacks an `exp` claim. Either default
it or guard the comparison.Formatting recipe:
src/auth.py:40-50" makes
the model attribute the failure correctly.Two limits, both required:
MAX_ITERS = 5 is a sane default; Aider uses ~3, OpenHands
uses 50–100 for SWE-Bench).seen_errors = []
for i in range(MAX_ITERS):
edit = agent.propose_edit(feedback if i else initial_task)
apply_edit(edit)
git_commit(f"agent: iter {i+1}") # always commit each iter
verifier = run_verifier()
if verifier["exit_code"] == 0:
return Success(iters=i+1)
feedback = format_feedback(verifier, last_edit=edit)
if feedback in seen_errors[-1:]: # exact repeat
return Stall(reason="same error twice", last=feedback)
seen_errors.append(feedback)
return Escalate(reason=f"exhausted {MAX_ITERS} iters", last=feedback)After every edit, before the verifier runs, commit with a structured message:
git commit -am "agent[iter 3/5]: tighten exp guard in verify_token"Why mandatory:
git log --oneline | head -5.git diff HEAD~1 gives the formatter a precise "what you just changed" anchor.Aider's --auto-commits (on by default) does this. For non-Aider agents,
wrap the loop in commit logic yourself.
| Signal | Good or false-positive? |
|---|---|
exit 0 from full verifier command | Good |
exit 0 but stderr contains "warning" | Soft success; surface to user, don't loop |
exit 0 because no tests collected (pytest returns 5) | False positive — check pytest --collect-only count |
exit 0 from a || true-swallowed command | False positive — strip suppression from --test-cmd |
exit 0 but agent disabled / skipped tests to pass | Critical — diff for pytest.skip, @pytest.mark.skip, xfail added in last iter |
The agent disabling tests to "pass" is the most common pathological success.
Add a post-success diff check: git log -p -1 | grep -E '(skip|xfail|@disable)'.
When the loop exits without success:
exhausted, stalled, env_failure,
timeout. The user's fix differs per cause.Format: Trigger → Action → Output → Evidence.
<lint> && <type> && <test> once, capture three-tuple
(stdout, stderr, exit_code).[aider/lint-test] "Aider will try and fix any errors if the
command returns a non-zero exit code."git diff HEAD~1 --name-only -U0. Strip ANSI, coverage,
deprecation warnings. Hard byte cap.[aider/edit-errors] "Above about 25k tokens of context,
most models start to become distracted." Each iteration adds context; keep
the per-iter delta tiny.MAX_ITERS (3–5 interactive, 50–100 SWE-Bench), detect
stall (same error twice = break), enforce total wall-clock cap.while True.[oh/6357] OpenHands infinite-loop bug + [langgraph/recursion]
"Hitting recursion_limit indicates an underlying design flaw" — same lesson.git add -A && git commit -m "agent[iter N]: <one-line>".
Never --amend.[aider/git] per-edit auto-commit; [cline/auto-approve]
Cline mirrors the same "edit→commit→test" rhythm.pytest exit 5 ≠
success), (b) no test was newly skipped/xfailed in the last commit, (c) no
|| true suppression in the verifier command itself.[aider/lint-test] formatter wrapper
caveat (auto-formatters that rewrite + return non-zero need double-run).ImportError,
command not found, OOM, network 503, ConnectionRefused to test DB.env_failure. The agent cannot fix pytest: command not found by editing source.ruff format or prettier --write modify files and return
non-zero on first pass (means "I changed something").[aider/lint-test] explicit guidance on formatter wrappers.pytest -x (--exitfirst) so the test runner itself stops at
the first failure. The output is naturally bounded.-x for the agent loop, once with full output
captured into a side-file for the human report. Don't conflate the
two streams.git diff HEAD~1 --name-only:
"your last edit touched X; the first failure is in a test of Y." The
anchor breaks the "fix the last thing I read" bias.-x for the loop, full run for the human.ImportError: No module named psycopg2. The
agent obediently rewrites from psycopg2 import ... to import psycopg,
next iter: No module named psycopg. Iter 3: it removes the DB layer
entirely. The loop has hit its cap; the codebase is now broken.apt-get install libpq-dev.ImportError).ENV_PATTERNS = [
r"No module named",
r"command not found",
r"OSError: \[Errno 28\]", # disk full
r"ConnectionRefusedError", # service down
r"libpq.so", # missing system lib
]env_failure.env_failure: don't call agent.propose_edit(...). Exit the
loop immediately with a message to the user: "Verifier failed with
what looks like an environment issue (No module named psycopg2). The
agent has not edited files; please fix the environment and re-run."@pytest.mark.skip"exit 0. You celebrate. Then the user runs the
tests themselves and discovers the failing test now has @pytest.mark.skip
added by the agent. Technically green; pathologically wrong.skip outright — there are legitimate skips.git log -p $(git merge-base HEAD origin/main)..HEAD -- '*.py' \
| grep -E '^\+.*(skip|xfail|@disabled|pass # TODO)' && echo "POSSIBLE CHEAT"pytest.skip annotations. Review
the diff." Loop exit tag: suspicious_pass.tests_pre.count() == tests_post.count(). Any
reduction = cheat-suspect.AssertionError. The
agent edited different lines each time but the error didn't change. You
have 2 iters left in your budget. Push through, or break early?git diff touched a different file, that's exploration;
give it one more iter.[aider/edit-errors]. Format first.[oh/6357] is the textbook case — even mature frameworks get this wrong.exit 0 as ground truth. Check for (a) tests actually ran,
(b) no skips added this iter, (c) no || true swallowed.--amend between iterations. You lose the bisect trail. Each
iter is its own commit.pytest live in stdout; for
mypy, tsc, cargo they live in stderr. You need both.ModuleNotFoundError
for a missing system lib will never be fixed by editing source. Classify
and escalate.tests/, surface for review — agents fix code by weakening
tests more often than humans like to admit.pytest -x -k <changed> or
--testmon for the loop; gate the full suite at PR time.| Scenario | Use instead |
|---|---|
| Success is subjective (writing, UX, design) | Human-in-the-loop / pairwise eval |
| Verifier takes >5 min and you need interactive UX | Async/CI runner with a notification, not an in-loop wait |
| Multi-step verifier with branching (deploy → smoke → rollback) | A state graph (LangGraph) — the loop is not enough |
| You don't have git | Wrap in any other VCS or filesystem snapshot — the per-iter rollback is non-negotiable |
| The agent has no ability to read structured tool results | Use a framework that does (Aider, LangGraph, Claude Code tool use) — naked text-completion loops won't carry the feedback |
--no-auto-commits disables the per-iter commit. Don't turn it
off "to keep history clean" — git rebase -i after the loop is the right
cleanup. [aider/git]--ignore-missing-imports can mask real import errors;
prefer --strict in the loop, relax for general use.ruff --fix rewrites files. Either commit before re-running, or use
ruff check (no --fix) in the loop and let the agent do the fixing.[aider/sonnet-not-lazy].Aider --auto-lint/--auto-test | OpenHands SWE-Bench harness | Cline auto-approve | Claude Code (bash + read) | Manual LangGraph cycle | |
|---|---|---|---|---|---|
| Verifier wiring | --lint-cmd, --test-cmd flags | eval_config.json per instance | allowlist + run command | Bash tool the agent calls | Tool node returns stdout/stderr/exit |
| Iteration bound | ~3 internal retries on lint/test fail | max_iterations (50–100) | none built-in; user-set timeout | model-controlled (no hard cap) | recursion_limit + retry counter in state |
| Output formatting | Strips ANSI, sends to chat verbatim if non-zero | Raw observation injected into history | Raw terminal output to chat | Raw bash output (no compaction) | User-implemented in tool node |
| Per-iter commit | Yes (--auto-commits on) | Optional (eval mode) | Manual / via terminal tool | Manual (agent calls git) | Manual node |
| Escalation hook | "gives up after sensible tries" (silent) | Returns failure obs to harness | Stops on cap; user resumes | Returns to user | Conditional edge to END |
| Env-failure detection | Limited (treats all non-zero same) | Limited; SWE-Gym extends with infra setup phase | None | None | User-implemented |
| Sweet spot | Interactive pair-programming with one verifier | Batch evaluation; high iter budget | VS Code interactive | Generic agent harness | Custom workflows with non-trivial topology |
--auto-lint --auto-test --auto-commits is the minimum-effort win.
[aider/lint-test]max_iterations per instance.
Beware the context-overflow infinite-loop pattern. [oh/6357]npm test, npm run lint, pnpm build) is the
ergonomic shape. [cline/auto-approve]interrupt() at the deploy step. The loop is
the inner node, the graph is the orchestration. [langgraph/persistence]pytest output.
Formatting is engineering work, not cosmetics.apt-get"
incident. Classify before re-prompting.[aider/lint-test] = https://aider.chat/docs/usage/lint-test.html[aider/edit-errors] = https://aider.chat/docs/troubleshooting/edit-errors.html[aider/git] = https://aider.chat/docs/git.html[aider/sonnet-not-lazy] = https://aider.chat/2024/07/01/sonnet-not-lazy.html[oh/6357] = https://github.com/All-Hands-AI/OpenHands/issues/6357 (SWE-Bench infinite loop on context overflow)[oh/swe-bench] = https://github.com/All-Hands-AI/OpenHands/blob/main/evaluation/benchmarks/swe_bench/README.md[swe-gym] = https://github.com/SWE-Gym/SWE-Gym/blob/main/docs/OpenHands.md[cline/auto-approve] = https://docs.cline.bot/features/auto-approve[cline/cli] = https://cline.bot/blog/introducing-cline-cli-2-0[langgraph/recursion] = https://docs.langchain.com/oss/python/langgraph/errors (GRAPH_RECURSION_LIMIT)[langgraph/persistence] = https://docs.langchain.com/oss/python/langgraph/persistence[pytest/exit-codes] = https://docs.pytest.org/en/stable/reference/exit-codes.html© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/agentsop-test-fix-loop of agentsope/SkillAlchemy.
Open the folder on GitHubat commit 6ea799f
Agentsop Test Fix Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Test Fix Loop this skillagentsope/SkillAlchemy | 457 | — | ~7.2k | Automated safety check: Pass | MIT | |
| Deep Agents to Pydantic AI Migrationpydantic/pydantic-ai | 20k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Code Reviewlangchain-ai/langchain-azure | 147 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Edgeone Makers ToolsTencentEdgeOne/edgeone-makers-tools | 1.9k | 1 repos | ~516 | Automated safety check: Pass | MIT | |
| Langgraph Testing Evaluationsoba-labs/langchain-agent-skills | 107 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Chat Gun Backend ContractHsienW/chat-gun | 143 | — | ~2.1k | Automated safety check: Pass | Custom licence |
pydantic/pydantic-ai
Migrates Python LangChain Deep Agents applications to Pydantic AI and Pydantic AI Harness while preserving the application's observed behavior.
langchain-ai/langchain-azure
Reviews changes in the langchain-azure monorepo using package-specific knowledge of langchain-azure-ai, langchain-azure-compute, langchain-azure-cosmosdb, langchain-azure-postgresql…
TencentEdgeOne/edgeone-makers-tools
EdgeOne Makers platform development router — the single entry point for building, storing data, and deploying on Tencent EdgeOne Makers.
soba-labs/langchain-agent-skills
A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…
HsienW/chat-gun
Apply when creating, modifying, refactoring, debugging, testing, or reviewing TypeScript, LangGraph JS, LangChain, provider adapter, tool, MCP, prompt, state, checkpoint, runtime event, or backend…
aws/agent-toolkit-for-aws
A skill your agent uses when a developer wants to create a new agent project or get started with AgentCore.
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Decision protocol for wiring a verify-then-fix loop around a code-editing LLM agent. Agentsop Test Fix Loop is an agent skill from agentsope/SkillAlchemy. Decision protocol for wiring a verify-then-fix loop around a code-editing LLM agent.
Agentsop Test Fix Loop fits situations like: tasks that involve Linting and formatting; tasks that involve Building AI agents.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-test-fix-loop -a claude-code`. Or copy the skill folder (skills/agentsop-test-fix-loop in agentsope/SkillAlchemy) into .claude/skills/agentsop-test-fix-loop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-test-fix-loop -a codex`. Or copy the skill folder (skills/agentsop-test-fix-loop in agentsope/SkillAlchemy) into .agents/skills/agentsop-test-fix-loop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-test-fix-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-test-fix-loop, .gemini/skills/agentsop-test-fix-loop, .github/skills/agentsop-test-fix-loop and .opencode/skills/agentsop-test-fix-loop in your project.
Going by SKILL.md and its folder, Agentsop Test Fix Loop needs the command-line tools its instructions call (git, pytest, ruff, npm, prettier and cargo). Our summary lists: Python 3.
SKILL.md names 6 domains. As links in the text: aider.chat, github.com, docs.langchain.com, docs.cline.bot, cline.bot and docs.pytest.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Test Fix Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Test Fix Loop: Deep Agents to Pydantic AI Migration (pydantic/pydantic-ai, 20k stars), Code Review (langchain-ai/langchain-azure, 147 stars), Edgeone Makers Tools (TencentEdgeOne/edgeone-makers-tools, 1.9k stars) and Langgraph Testing Evaluation (soba-labs/langchain-agent-skills, 107 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 457 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 2, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.