Agent skill

Autocontext

by greyhaven-ai in greyhaven-ai/autocontext

Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Autocontext

skills CLI
$ npx skills add greyhaven-ai/autocontext --skill autocontext -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install greyhaven-ai/autocontext autocontext --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/greyhaven-ai/autocontext.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pi/skills/autocontext .claude/skills/autocontext && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autocontext
GitHub stars
1.3k
Token cost
~892 tokens
SKILL.md length
260 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.

  • Works in 4 steps: Evaluate output quality → Improve output iteratively → Queue background tasks → …
  • Scoring a draft against a quality rubric before accepting it
  • SKILL.md covers Available Tools, Quick Start, Configuration and CLI Companion
  • Calls npm; needs AUTOCONTEXT_AGENT_API_KEY and AUTOCONTEXT_API_KEY

What it does

This skill scores agent output against a rubric, returning a 0-1 score with reasoning and a per-dimension breakdown through its judge tool. A separate improve tool automates the loop: it judges a draft, revises it based on the feedback, and rejudges until the output clears the quality bar or the round limit runs out.

A queue tool hands a named scenario to a background task runner for asynchronous evaluation, and a status tool checks on runs and queued tasks later. A scenarios tool lists what evaluation scenarios and families exist, and a runtime-snapshot tool inspects a run's artifacts, package provenance, branchable session lineage, and recent event-stream entries. Configuration is read from a project's .autoctx.json or from provider and API-key environment variables.

When your agent uses it

  • Scoring a draft against a quality rubric before accepting it
  • Running an automated revise-and-rejudge loop on agent output
  • Queuing an evaluation scenario for background processing
  • Checking the status or lineage of a past evaluation run

Example prompts

  • “Judge this function against the rubric for correctness and style.”
  • “Run the improve loop on this draft until it clears the quality bar.”
  • “Queue the agent_task scenario for background evaluation.”
  • “Show me the runtime snapshot for run_123.”

Requirements

  • Provider API key set via AUTOCONTEXT_API_KEY or AUTOCONTEXT_PROVIDER
  • Pre-approved tools (allowed-tools): autocontext_judge, autocontext_improve, autocontext_status, autocontext_scenarios, autocontext_queue, autocontext_runtime_snapshot

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Evaluate output quality
  2. Improve output iteratively
  3. Queue background tasks
  4. Discover scenarios

What it can do on your machine

Read from SKILL.md and the folder at commit f72c154. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • autocontext_judge
    • autocontext_improve
    • autocontext_status
    • autocontext_scenarios
    • autocontext_queue
    • autocontext_runtime_snapshot

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AUTOCONTEXT_AGENT_API_KEY
    • AUTOCONTEXT_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autocontext loads about 892 tokens when it runs. Until then it costs about 86 tokens; SKILL.md has 260 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~892

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from greyhaven-ai/autocontext at commit f72c154, republished under its Apache-2.0 licence (© greyhaven-ai). 260 words, ~892 tokens.

Download SKILL.mdSave it as .claude/skills/autocontext/SKILL.md (or your agent's skills folder).
name
autocontext
description
Iterative strategy generation and evaluation system. Use when the user wants to evaluate agent output quality, run improvement loops, queue tasks for background evaluation, check run status, inspect runtime artifacts and session branch lineage, or discover available scenarios. Provides LLM-based judging with rubric-driven scoring.
allowed-tools
autocontext_judge, autocontext_improve, autocontext_status, autocontext_scenarios, autocontext_queue, autocontext_runtime_snapshot

autocontext

autocontext is an iterative strategy generation and evaluation system that uses LLM-based judging to score and improve agent outputs.

Available Tools

  • autocontext_judge — Evaluate agent output against a rubric. Returns a 0–1 score with reasoning and per-dimension breakdowns.
  • autocontext_improve — Run a multi-round improvement loop. The agent output is judged, revised based on feedback, and re-evaluated until the quality threshold is met or max rounds are exhausted.
  • autocontext_queue — Enqueue a task for background evaluation by the task runner daemon.
  • autocontext_status — Check the status of runs and queued tasks.
  • autocontext_scenarios — List available evaluation scenarios and their families.
  • autocontext_runtime_snapshot — Inspect run artifacts, package provenance, branchable session lineage, and recent event-stream entries.

Quick Start

1. Evaluate output quality

Use autocontext_judge with a task prompt, the agent's output, and a rubric:

autocontext_judge(
  task_prompt="Write a Python function to parse CSV files",
  agent_output="def parse_csv(path): ...",
  rubric="Correctness, error handling, edge cases, documentation"
)
2. Improve output iteratively

Use autocontext_improve to automatically revise output through judge-guided feedback loops:

autocontext_improve(
  task_prompt="Write a Python function to parse CSV files",
  initial_output="def parse_csv(path): ...",
  rubric="Correctness, error handling, edge cases, documentation",
  max_rounds=5,
  quality_threshold=0.85
)
3. Queue background tasks

Use autocontext_queue with a scenario name to enqueue evaluation tasks for asynchronous processing:

autocontext_queue(spec_name="my_scenario")

Check results later with autocontext_status.

For deeper context, use autocontext_runtime_snapshot with the run ID. Add session_id when you need the active branch path before continuing work:

autocontext_runtime_snapshot(run_id="run_123", session_id="sess_123")
4. Discover scenarios

Use autocontext_scenarios to see what evaluation scenarios are available:

autocontext_scenarios()
autocontext_scenarios(family="agent_task")

Configuration

The extension auto-detects configuration from these sources:

  1. Project config — .autoctx.json in the working directory (created via autoctx init)
  2. Environment variables:
    • AUTOCONTEXT_AGENT_PROVIDER or AUTOCONTEXT_PROVIDER — Provider type
    • AUTOCONTEXT_AGENT_API_KEY or AUTOCONTEXT_API_KEY — Provider API key
    • AUTOCONTEXT_AGENT_DEFAULT_MODEL or AUTOCONTEXT_MODEL — Model override
    • AUTOCONTEXT_DB_PATH — SQLite database path override
  3. Pi provider — Falls back to Pi's configured LLM provider

CLI Companion

For standalone usage outside Pi, install the autoctx CLI:

bash
npm install -g autoctx
autoctx init
autoctx solve "your problem" --iterations 5
autoctx simulate --description "your simulation" --runs 3

© greyhaven-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in pi/skills/autocontext of greyhaven-ai/autocontext.

Open the folder on GitHubat commit f72c154

Compare with similar skills

Autocontext next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autocontext compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autocontext this skillgreyhaven-ai/autocontext1.3k—~892Automated safety check: PassApache-2.0
Benchflowbenchflow-ai/benchflow355—~1.9kAutomated safety check: NotesApache-2.0
Windmill AI Evalswindmill-labs/windmill18k—~969Automated safety check: NotesCustom licence
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Octocode Benchmark Runnerbgauryy/octocode949—~2.1kAutomated safety check: PassMIT
SWE Benchmark Task Adderory/lumen307—~497Automated safety check: PassCustom licence

Similar skills

  • Benchflow

    benchflow-ai/benchflow

    Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

    355 GitHub stars~1.9k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check: notes
  • Windmill AI Evals

    windmill-labs/windmill

    Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

    18k GitHub stars~969 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

    307 GitHub stars~497 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from greyhaven-ai/autocontext

  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated 2 days ago
    Auto-check passed
  • Autocontext Knowledge Reader

    greyhaven-ai/autocontext

    Reads and moves the playbooks and lessons that Autocontext has already learned, using the autoctx CLI and plain files on disk.

    1.3k GitHub stars~934 tokensUpdated 2 days ago
    Auto-check passed
  • Autocontext Knowledge Creator

    greyhaven-ai/autocontext

    Runs the `autoctx` CLI to improve an approach to a task over several generations, score or refine a single output and inspect what a run produced.

    1.3k GitHub stars~964 tokensUpdated 2 days ago
    Auto-check passed
  • Grid CTF Strategy Playbook

    greyhaven-ai/autocontext

    Operational notes for generating, evaluating and debugging strategies in the autocontext grid_ctf scenario, with tier rules and parameter ranges that worked or failed.

    1.3k GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed

Questions about Autocontext

What does Autocontext do?

Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met. This skill scores agent output against a rubric, returning a 0-1 score with reasoning and a per-dimension breakdown through its judge tool. A separate improve tool automates the loop: it judges a draft, revises it based on the feedback, and rejudges until the output clears the quality bar or the round limit runs out.

When should I use Autocontext?

Autocontext fits situations like: scoring a draft against a quality rubric before accepting it; running an automated revise-and-rejudge loop on agent output; queuing an evaluation scenario for background processing; checking the status or lineage of a past evaluation run.

How do I install Autocontext in Claude Code?

Run `npx skills add greyhaven-ai/autocontext --skill autocontext -a claude-code`. Or copy the skill folder (pi/skills/autocontext in greyhaven-ai/autocontext) into .claude/skills/autocontext in your project. Claude Code loads it when a task matches its description.

How do I install Autocontext in Codex?

Run `npx skills add greyhaven-ai/autocontext --skill autocontext -a codex`. Or copy the skill folder (pi/skills/autocontext in greyhaven-ai/autocontext) into .agents/skills/autocontext in your project. Codex loads it when a task matches its description.

Can I use Autocontext in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add greyhaven-ai/autocontext --skill autocontext -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autocontext, .gemini/skills/autocontext, .github/skills/autocontext and .opencode/skills/autocontext in your project.

What does Autocontext need to run?

Going by SKILL.md and its folder, Autocontext needs the command-line tools its instructions call (npm) and credentials named AUTOCONTEXT_AGENT_API_KEY and AUTOCONTEXT_API_KEY. Our summary lists: Provider API key set via AUTOCONTEXT_API_KEY or AUTOCONTEXT_PROVIDER. Its frontmatter pre-approves these tools: autocontext_judge, autocontext_improve, autocontext_status, autocontext_scenarios, autocontext_queue, autocontext_runtime_snapshot.

Does Autocontext access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Autocontext safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Autocontext use?

Autocontext is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Autocontext use?

About 892 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Autocontext?

Skills that share tags, products or a category with Autocontext: Benchflow (benchflow-ai/benchflow, 355 stars), Windmill AI Evals (windmill-labs/windmill, 18k stars), Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars) and Octocode Benchmark Runner (bgauryy/octocode, 949 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autocontext?

greyhaven-ai (a GitHub organization) maintains it in greyhaven-ai/autocontext, which has 1,304 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.

Source: greyhaven-ai/autocontext on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.