Agent skill

Verified Coding Workflow

by amd in amd/gaia

Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.

MITAuto-check passedDevelopment

Install Verified Coding Workflow

skills CLI
$ npx skills add amd/gaia --skill coding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/gaia coding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/hub/skills/coding .claude/skills/coding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coding
GitHub stars
1.6k
Token cost
~2.1k tokens
SKILL.md length
1,252 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.

  • Fixing a bug in an unfamiliar codebase without reproducing it blindly
  • SKILL.md covers Find the code before you…, Read the file before you edit it, Reproduce it first, then fix it and Prove it, then say it, plus 6 more sections
  • Calls python, git and pytest
  • Making a focused edit that must not silently change unrelated code

What it does

The skill treats changing code as riskier than answering a question, since a wrong edit can look plausible for hours. It splits search into two tools: search_file_content for an exact token like a function name or error message, and search_code_index, a semantic search run after index_codebase once per repository, for when the behavior is known but not its name.

Before any edit it insists on reading the part of the file being changed, plus enough surrounding code to see what depends on it, warning that editing from memory of a similar file is how a special case quietly disappears. If edit_file's exact string replacement fails, the rule is to re-read the file rather than guess again, and never to fall back to rewriting the whole file, which silently reverts unrelated parts and rewrites every line ending.

Bugs are reproduced and the real output read before any change, so there is a clear before state to compare against. Fixes are only considered proven once pytest or python -m pytest actually runs against the suite; tracing logic mentally is treated as the same reasoning that produced the bug in the first place, not as verification.

When your agent uses it

  • Fixing a bug in an unfamiliar codebase without reproducing it blindly
  • Making a focused edit that must not silently change unrelated code
  • Proving a test passes before reporting that a fix works

Example prompts

  • “Fix the discount calculation bug in test_cart.py, reproducing it first.”
  • “Find every place that calls loadModel using an exact search.”
  • “Where do we decide which model to load? I don't know the function name.”

Requirements

  • pytest and python on PATH, for running tests

What it can do on your machine

Read from SKILL.md and the folder at commit 6c3bb5c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git
    • pytest
    • python3
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verified Coding Workflow loads about 2.1k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 1,252 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/gaia at commit 6c3bb5c, republished under its MIT licence (© amd). 1,252 words, ~2,147 tokens.

Download SKILL.mdSave it as .claude/skills/coding/SKILL.md (or your agent's skills folder).
name
coding
description
Work on a codebase — read, search, edit and verify source files. Use when the user asks to fix a bug, add a feature, refactor, explain code, make a test pass, or change anything in a repository. Covers finding the right file, editing safely, and proving the change works before reporting it.
license
MIT
version
0.2.1

Coding

Changing code is different from answering a question: you can be confidently wrong for hours because the answer looked plausible. Everything below exists to close that gap.

Find the code before you change it

You have two searches and they answer different questions:

  • search_file_content — grep. Use it when you know the string: a function name, an error message, a config key. Fastest way to find every call site.
  • search_code_index — semantic. Use it when you know the behaviour but not the name: "where do we decide which model to load". Run index_codebase once for the repo first; get_index_status says whether it is ready.

Grep first when you have an exact token — it is instant and exhaustive. Reach for the semantic index when grep returns nothing useful because you are guessing at names.

Read the file before you edit it

Not the top of it — use read_file on the part you are changing, plus enough around it to see what else depends on it. An edit written from memory of a similar file is how you delete someone's special case.

edit_file replaces an exact string. If it fails, the file is not what you thought: re-read it, do not retry with a guess. Never fall back to rewriting the whole file to get past a failed edit — that silently reverts anything else in it and rewrites every line's endings, turning a two-line change into a whole-file diff.

Reproduce it first, then fix it

Run the failing thing and read the actual output before changing anything. A bug you have not reproduced is a bug you are guessing at, and the fix for a guess usually lands somewhere real code was fine.

This also gives you the before state. Without it you cannot tell whether you fixed the problem or merely changed the symptom.

Prove it, then say it

A test you did not run is not a test that passed. Tracing the logic in your head is not verification — it is the same reasoning that produced the bug.

This skill grants pytest and python when they are installed on PATH, so run the suite directly with run_shell_command. Prefer the python -m spelling — it puts the project's own directory on sys.path, so it works on a checkout that was never installed, where bare pytest fails to import the project:

python -m pytest -q tests/
python -m pytest -x -k discount tests/test_cart.py
pytest -q tests/                     # same grant, same rules; for an installed project
python util/lint.py --all --fix      # or whatever the project's own runner is

python -m adds the current directory to the import path, not src/. For a project whose package lives under src/, scope the path to that one command:

PYTHONPATH=src python -m pytest -q tests/

Where only python3 exists, python3 -m pytest works the same way.

pytest often lives only in the project's virtualenv, so it is not on PATH. The skill still loads, but without the pytest grant — and python -m pytest is judged as pytest, so that spelling is refused too. Run the suite with execute_python_file instead, which works wherever pytest is importable:

python
import sys, pytest
sys.exit(pytest.main(["-q", "tests/"]))

Loading this skill grants pytest and python <script.py> execution without another prompt. Tests and scripts are trusted project code: they can write files, access the network, and launch other programs, including commands the direct CLI policy refuses. These grants do not sandbox their effects. The separate execute_python_file tool still requires per-call approval.

python -c "..." is refused because the grant requires a reviewable file in the checkout. Write new code to a file first so the diff shows it, and review what it does before executing it. For a one-off calculation that should not land in the repository at all, run_python takes the snippet through the tool path instead — it asks for approval on every call rather than riding this grant.

The rest of the grant is narrow on purpose. --pdb would hang waiting for a debugger nobody can answer, -p <plugin> imports arbitrary code, --junitxml writes outside the run. python -m pytest --pdb is refused for exactly the same reason pytest --pdb is — -m is not a way around a rule.

For a suite pytest cannot drive — npm, go, cargo, make — you have no grant by default. GAIA ships policies for npm, go, uv, pip, black, isort and ruff, so a project skill can declare the one it needs (shell:execute:npm) and get the same three-tier treatment. make has no policy and will not get one: its argument is a target in a file, so "run make test" means "run whatever the Makefile says", which no approval prompt can honestly describe.

If you could not run something, say so plainly — "I could not execute the suite, so this is unverified" — rather than implying it passed.

Show full SKILL.md (504 more words)Show less

Do not break what was already working

Run the WHOLE suite, not just the test you were asked about. A fix that repairs one test and breaks two is a worse state than you started in, and you will not notice if you only look at the one.

If something else fails, that is now your problem, whether or not you caused it. Say which of the two it is.

Committing: yours to propose, theirs to approve

You have git. Reads — status, diff, log, show, branch — run straight through, and you should use them constantly: git diff before you report a change is the cheapest possible check that the diff is only what you meant.

add, commit, checkout, switch, restore and stash each stop and show the user the exact command before running. That is not a formality to click past — write the commit message as if it is the only thing the reviewer reads, because for a squashed PR it is.

The direct Git policy refuses push, reset --hard, clean, rebase, commit --amend, git config. Publishing and history-rewriting are the user's to run, and destroying uncommitted work has no undo anywhere. When the work is committed and wants pushing, say so and stop — "committed on fix/discount; run git push -u origin fix/discount and I will open the PR" — rather than looking for a spelling that gets through.

Keep the diff to the change

  • Fix the bug you were asked about. Do not reformat, rename, reorder imports, or "clean up while you're here" — every unrelated line is a line the reviewer has to check.
  • Match the file's existing style, even where you would write it differently.
  • Do not add comments narrating what the code does. A comment earns its place by explaining a why that is not obvious from the code.

When you fix one instance, look for the others

The same mistake is rarely alone. A bad subprocess call, a missing encoding, a wrong default — grep for the pattern once you understand it, and say what you found even if you only fixed one. "Fixed here; three more in X, Y, Z" is far more useful than a silent single fix.

Scratch files are yours, not theirs

Runner scripts and scratch output go in the system temp directory — write them there and run them with execute_python_file, which does not care where the file lives. python does: its grant stops at the checkout, so a scratch script outside it is execute_python_file's job, not the shell's.

Writing create_doc.py and temp/ into the root of someone's repository leaves them in the next git status, and they did not ask for them.

Reporting a code change

Lead with whether it works, then what changed:

Both failing tests pass now, and the other six still do.

  • apply_discount was subtracting amount * percent instead of amount * percent / 100.
  • format_money used str(round(...)), which drops the trailing zero in $5.50.

Name the file and line only when the reader needs it to act. Never claim a test run you did not do.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in hub/skills/coding of amd/gaia.

Open the folder on GitHubat commit 6c3bb5c

Compare with similar skills

Verified Coding Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verified Coding Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verified Coding Workflow this skillamd/gaia1.6k—~2.1kAutomated safety check: PassMIT
Adk Setupgoogle/adk-python22k—~993Automated safety check: NotesApache-2.0
Flowfile Debugging PlaybookEdwardvaneechoud/Flowfile370—~6.3kAutomated safety check: PassMIT
Error Explanation GeneratorArabelaTso/Skills-4-SE253—~3.8kAutomated safety check: PassApache-2.0
Kedro Babysitkedro-org/kedro11k—~4kAutomated safety check: PassCustom licence
Python Project Creatorhaddock-development/claude-reflect-system379—~1.4kAutomated safety check: NotesNone

Similar skills

  • Adk Setup

    google/adk-python

    Official

    Sets up a local ADK Python development environment in a git clone of the open-source adk-python repository: a uv virtual environment, all dependency extras, pre-commit hooks, and a first unit-test…

    22k GitHub stars~993 tokensUpdated today
    DevelopmentAuto-check: notes
  • Flowfile Debugging Playbook

    Edwardvaneechoud/Flowfile

    Symptom-to-cause triage playbook for Flowfile (core/worker/kernel/frontend/AI) — covers "no such table" DB cascades (two distinct causes), import-time Alembic migration corruption, silent…

    370 GitHub stars~6.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Error Explanation Generator

    ArabelaTso/Skills-4-SE

    Explains test failures and provides actionable debugging guidance.

    253 GitHub stars~3.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Kedro Babysit

    kedro-org/kedro

    Run Kedro's local lint / format / type-check / tests on changed files (uses the project's pre-commit hooks, ruff, mypy, pytest, lint-imports, detect-secrets, Make targets — in the right venv), or…

    11k GitHub stars~4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Python Project Creator

    haddock-development/claude-reflect-system

    Creates Python projects with proper structure, virtual environments, and dependency management.

    379 GitHub stars~1.4k tokensUpdated 8 mo ago
    DevelopmentAuto-check: notes
  • Mastering Python Skill

    SpillwaveSolutions/agent-brain

    Modern Python coaching covering language foundations through advanced production patterns.

    120 GitHub stars~1.4k tokensUpdated 17 days ago
    DevelopmentAuto-check: notes

More from amd/gaia

All 44 skills in this repo
  • Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

    1.6k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.

    1.6k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.

    1.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Turns a source document such as a README or spec into an executive slide deck as one self-contained HTML file that prints to PDF, one slide per page.

    1.6k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Verified Coding Workflow

What does Verified Coding Workflow do?

Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run. The skill treats changing code as riskier than answering a question, since a wrong edit can look plausible for hours. It splits search into two tools: search_file_content for an exact token like a function name or error message, and search_code_index, a semantic search run after index_codebase once per repository, for when the behavior is known but not its name.

When should I use Verified Coding Workflow?

Verified Coding Workflow fits situations like: fixing a bug in an unfamiliar codebase without reproducing it blindly; making a focused edit that must not silently change unrelated code; proving a test passes before reporting that a fix works.

How do I install Verified Coding Workflow in Claude Code?

Run `npx skills add amd/gaia --skill coding -a claude-code`. Or copy the skill folder (hub/skills/coding in amd/gaia) into .claude/skills/coding in your project. Claude Code loads it when a task matches its description.

How do I install Verified Coding Workflow in Codex?

Run `npx skills add amd/gaia --skill coding -a codex`. Or copy the skill folder (hub/skills/coding in amd/gaia) into .agents/skills/coding in your project. Codex loads it when a task matches its description.

Can I use Verified Coding Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill coding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coding, .gemini/skills/coding, .github/skills/coding and .opencode/skills/coding in your project.

What does Verified Coding Workflow need to run?

Going by SKILL.md and its folder, Verified Coding Workflow needs the command-line tools its instructions call (python, git, pytest, python3 and make). Our summary lists: pytest and python on PATH, for running tests.

Does Verified Coding Workflow access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Verified Coding Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verified Coding Workflow use?

Verified Coding Workflow is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verified Coding Workflow use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verified Coding Workflow?

Skills that share tags, products or a category with Verified Coding Workflow: Adk Setup (google/adk-python, 22k stars), Flowfile Debugging Playbook (Edwardvaneechoud/Flowfile, 370 stars), Error Explanation Generator (ArabelaTso/Skills-4-SE, 253 stars) and Kedro Babysit (kedro-org/kedro, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verified Coding Workflow?

amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.