Official agent skill

Docs Grounding Verifier

by microsoft in microsoft/apm

A skill your agent uses to verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code.

OfficialMITAuto-check passedResearch & Science

Install Docs Grounding Verifier

skills CLI
$ npx skills add microsoft/apm --skill docs-grounding-verifier -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/apm docs-grounding-verifier --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/apm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.apm/skills/docs-grounding-verifier .claude/skills/docs-grounding-verifier && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
docs-grounding-verifier
GitHub stars
4k
Token cost
~1.9k tokens
SKILL.md length
662 words
Files
113 (incl. scripts, assets)
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses to verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code.

  • Works in 6 steps: SCOPE → EXTRACT (parallel) → RETRIEVE (deterministic, batched) → …
  • Verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code
  • SKILL.md covers Sibling contract, When to activate, When NOT to activate and Architecture…, plus 9 more sections
  • Runs Shell scripts from its folder

What it does

Docs Grounding Verifier is an agent skill from microsoft/apm, published by the product's own GitHub organization. Use this skill to verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code. Activate when you have specific pages to check for factual accuracy -- not when sweeping a whole corpus (use docs-corpus-audit for that) and not when triaging a PR diff (use docs-sync for that). Trigger nouns: "is this doc accurate", "verify the page against the code", "fact-check this section", "any claims that drifted from source", "fact-checking", "grounding audit", "drift hunt", "claim…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 121 other files, including scripts and assets (for example `assets/judge-prompt.md`, `evals/content-evals.json` and `evals/run-evals.sh`).

It sits in Research & Science, covering Fact-checking and source verification. The repository describes itself as: Agent Package Manager. The licence is MIT.

When your agent uses it

  • Verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code
  • Nouns: is this doc accurate
  • Verify the page against the code
  • Fact-check this section

Example prompts

  • “is this doc accurate”
  • “verify the page against the code”
  • “fact-check this section”
  • “/docs-grounding-verifier”

Requirements

  • A Bash shell

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. SCOPE
  2. EXTRACT (parallel)
  3. RETRIEVE (deterministic, batched)
  4. JUDGE (parallel)
  5. SYNTHESIZE
  6. ALIGNMENT LOOP (A8)

What it can do on your machine

Read from SKILL.md and the folder at commit 280b8a7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Docs Grounding Verifier loads about 1.9k tokens when it runs. Until then it costs about 227 tokens; SKILL.md has 662 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~227
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from microsoft/apm at commit 280b8a7, republished under its MIT licence (© microsoft). 662 words, ~1,881 tokens.

Download SKILL.mdSave it as .claude/skills/docs-grounding-verifier/SKILL.md (or your agent's skills folder). This skill also uses 112 other files; get the full folder from GitHub.
name
docs-grounding-verifier
description
Use this skill to verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code. Activate when you have specific pages to check for factual accuracy -- not when sweeping a whole corpus (use docs-corpus-audit for that) and not when triaging a PR diff (use docs-sync for that). Trigger nouns: "is this doc accurate", "verify the page against the code", "fact-check this section", "any claims that drifted from source", "fact-checking", "grounding audit", "drift hunt", "claim verification". Returns per-claim verdicts (GROUNDED | PARTIAL | CONTRADICTED | UNSUPPORTED) with file:line evidence citations. Catches paragraph-level inaccuracies that page-level audit averages over -- e.g. a paragraph with 5 claims where 4 are grounded and 1 is fabricated. Does NOT modify files (returns advisory only); does NOT re-architect the docs; does NOT triage PRs.

docs-grounding-verifier

CLAIM-LEVEL grounding verification. Adapts the RAGAS faithfulness-eval pattern (proven in RAG literature) to docs/code instead of generated- answers/retrieved-context. Source code is the ground truth; docs paragraphs are the candidate text under audit.

python-architect persona doc-writer persona

Sibling contract

This skill is a SIBLING of docs-corpus-audit and docs-sync. The boundary is load-bearing:

SkillTriggerScopeGranularity
docs-syncPR opened/synchronizedPR diff onlyPage-level
docs-corpus-auditMaintainer asks for whole-corpus passEntire corpusPage-level
docs-grounding-verifierVerify specific pages factually1..N pagesCLAIM-level

docs-corpus-audit invokes this skill in its VERIFY phase on the highest-risk pages of each wave. docs-sync can invoke it on the specific pages in a PR diff. The skill is also runnable standalone.

When to activate

  • Maintainer says "verify <page> against the code".
  • An audit wave wants per-claim grounding scores for its highest-risk pages.
  • A PR review wants to confirm that prose changes are not just plausible but actually consistent with the implementation.
  • A "fact-check" or "grounding" or "drift hunt" request.

When NOT to activate

  • Whole-corpus sweep with no specific page list -> use docs-corpus-audit.
  • PR review with mixed code+docs diff -> use docs-sync.
  • Editorial / tone review -> use editorial-owner persona directly.

Architecture (PIPELINE-of-PANELS)

PARENT
  -> [Stage 1: EXTRACT claims, fan-out PANEL]
       per page -> LLM extracts atomic factual claims as JSON
       script: scripts/extract-claims.py
  -> [Stage 2: RETRIEVE evidence, deterministic S7]
       per claim -> grep over src/ via keywords + hints
       script: scripts/retrieve-evidence.sh   (NO LLM)
  -> [Stage 3: JUDGE grounding, adversarial A7]
       per (claim, evidence) -> LLM rules GROUNDED|PARTIAL|CONTRADICTED|UNSUPPORTED
       asset: assets/judge-prompt.md
  -> [Stage 4: SYNTHESIZE]
       aggregate ungrounded -> doc-writer for fix
       re-verify after fix (A8 ALIGNMENT LOOP)

Stage 2 is the load-bearing design choice: evidence retrieval is DETERMINISTIC (grep + AST hints), not LLM. The judge in Stage 3 can only rule on evidence it actually receives -- it cannot hallucinate support that the retriever did not find. This is the structural guard against the failure mode "the LLM convinces itself the docs match the code."

Phase 1: SCOPE

Input: list of page paths to verify (1..N). If a risk_class is attached (e.g. "high-stakes"), prefer it; otherwise treat all as equal.

Out-of-scope:

  • Pages outside docs/src/content/docs/ or packages/apm-guide/.apm/skills/apm-usage/.
  • Pages with no factual claims (pure editorial / landing). Skip rather than force-extract.

Phase 2: EXTRACT (parallel)

For each page, dispatch ONE claim-extractor agent:

  • Prompt template: scripts/extract-claims.py <page> produces the prompt and embeds the page content.
  • Returns: JSON {"page", "claims":[{"id","text","section","keywords", "expected_source_areas"}]} capped at 15 claims per page.

Parallel safe; no shared state between extractors.

Phase 3: RETRIEVE (deterministic, batched)

For each claim, pipe to scripts/retrieve-evidence.sh:

  • Uses keywords + expected_source_areas to grep src/.
  • Returns one-line JSON: {"claim_id","claim_text","evidence":[...], "evidence_count"}.

Sequential is fine (grep is fast). No LLM. Diagnostics on stderr, data on stdout.

Phase 4: JUDGE (parallel)

For each (claim, evidence) tuple, dispatch ONE grounding-judge agent:

  • Load assets/judge-prompt.md.
  • Send the prompt + the tuple.
  • Returns: JSON verdict per the schema in judge-prompt.md.

Batching across claims-of-one-page into a single judge call is fine (prompt with all tuples at once). Across pages, fan out.

Show full SKILL.md (246 more words)Show less

Phase 5: SYNTHESIZE

Aggregate verdicts. Materialize the report:

{
  "summary": {
    "pages_verified": N,
    "claims_total": N,
    "grounded": N, "partial": N, "contradicted": N, "unsupported": N,
    "grounding_rate": N/total
  },
  "actionable": [
    {"page", "claim", "verdict", "evidence_cited", "fix_suggestion"}
  ]
}

CONTRADICTED and PARTIAL are doc-writer work items. UNSUPPORTED is split: if retrieval_fix_suggestion is plausible, retry retrieval with the suggested keywords; if still empty, treat as CONTRADICTED.

Phase 6: ALIGNMENT LOOP (A8)

Hand actionable items to doc-writer (one subagent per page). After edits, RE-RUN the pipeline on the same pages. The grounding_rate must MONOTONICALLY INCREASE between iterations or the loop has diverged -- stop and escalate to the operator.

Ship gate

  • grounding_rate >= 0.9 on each verified page after the alignment loop.
  • Every CONTRADICTED claim cited a specific code file:line that disproves it -- not vague "the code doesn't say that".
  • The eval-runner (see evals/) passes on the trigger evals and the content evals before the skill is treated as production-ready.

Bundled assets

  • scripts/extract-claims.py -- Stage 1 prompt builder. --help, --schema.
  • scripts/retrieve-evidence.sh -- Stage 2 retriever. Deterministic. --help.
  • scripts/verify-page.sh -- end-to-end orchestrator. --help.
  • assets/judge-prompt.md -- Stage 3 adversarial judge prompt.
  • evals/trigger-evals.json -- 20 dispatch queries (10 should, 10 shouldn't).
  • evals/content-evals.json -- seeded-drift recall scenarios.
  • evals/run-evals.sh -- the eval-runner that turns JSON into metrics.

Failure modes guarded against

  • Hallucinated grounding: Stage 2 is deterministic; judge sees only real evidence.
  • Adversarial weakness: Stage 3 prompt defaults to SKEPTICAL.
  • Page-level averaging: claim-level granularity surfaces partials.
  • Bundle leakage: design notes / one-time scripts stay in session state, never in references/.
  • Phantom dependency: SKILL.md links its persona deps via relative paths; A9 PROBE before invoking docs-corpus-audit's substrate.
  • Dispatch collision with sibling skills: trigger-eval validation split is the ship gate (must distinguish from docs-sync / docs-corpus-audit triggers).

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 112 other files (scripts, assets) in .apm/skills/docs-grounding-verifier of microsoft/apm.

  • SKILL.md
  • assets/judge-prompt.md
  • evals/content-evals.json
  • evals/run-evals.sh
  • evals/runs/20260527-194228/seeded-corpus/drift-install-flag/page.md
  • evals/runs/20260527-194228/seeded-corpus/drift-install-flag/scenario.json
  • evals/runs/20260527-194228/seeded-corpus/drift-policy-reject/page.md
  • evals/runs/20260527-194228/seeded-corpus/drift-policy-reject/scenario.json
  • evals/runs/20260527-194228/seeded-corpus/drift-registry-resolver/page.md
  • evals/runs/20260527-194228/seeded-corpus/drift-registry-resolver/scenario.json
  • evals/runs/20260527-194228/trigger-prompts.txt
  • evals/runs/proof/claims
  • … and 101 more

Open the folder on GitHubat commit 280b8a7

Compare with similar skills

Docs Grounding Verifier next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Docs Grounding Verifier compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Docs Grounding Verifier this skillmicrosoft/apm4k—~1.9kAutomated safety check: PassMIT
Perplexity Web Searchdavila7/claude-code-templates32k12 repos~3.5kAutomated safety check: NotesMIT
Citation Verification GuideGalaxy-Dawn/claude-scholar5.7k3 repos~1.9kAutomated safety check: PassMIT
Article Fact Checkerdigoal/blog8.6k—~939Automated safety check: PassGPL-2.0
Deep Research Agent TeamImbad0202/academic-research-skills51k—~13kAutomated safety check: PassCustom licence
Fact Checkingbradygaster/squad3.3k—~503Automated safety check: PassMIT

Similar skills

  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Research & ScienceAuto-check: notes
  • Citation Verification Guide

    Galaxy-Dawn/claude-scholar

    Reference guidance for checking every citation in academic writing against canonical sources such as DOI, arXiv, CrossRef and Semantic Scholar, to catch fake or wrong references.

    5.7k GitHub starsUsed in 3 repos~1.9k tokens
    Research & ScienceAuto-check passed
  • 三层审查模型,逐段逐句验证文章真伪、证据链与逻辑结构。Use when the user asks to fact-check, verify, audit, or evaluate the credibility of an article, essay, report, opinion piece, social-media post, or any written claim —…

    8.6k GitHub stars~939 tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Deep Research Agent Team

    Imbad0202/academic-research-skills

    Runs a 13-agent pipeline for rigorous academic research, from forming the question through systematic search, synthesis, bias checks and an APA 7.0 report.

    51k GitHub stars~13k tokensUpdated 4 days ago
    Research & ScienceAuto-check passed
  • Fact Checking

    bradygaster/squad

    Review and validate claims using counter-hypothesis testing.

    3.3k GitHub stars~503 tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Citation Verification

    Light0305/Light-skills

    Verifies that every reference in a manuscript is real, correctly identified and actually supports its claim, and produces a citation registry for typesetting.

    640 GitHub stars~3.4k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed

More from microsoft/apm

All 27 skills in this repo
  • Cut Release

    microsoft/apm

    Official

    A skill your agent uses to cut an APM release from the current worktree: assess whether the cycle since the last tag warrants a patch or minor bump (semver discipline against the…

    4k GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Docs Corpus Audit

    microsoft/apm

    Official

    A skill your agent uses to run a holistic regrounding pass on the entire microsoft/apm documentation corpus against current source code, page-by-page, and emit surgical fixes for stale claims.

    4k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Official

    A skill your agent uses to write the PR description (PR body) for any pull request opened against microsoft/apm.

    4k GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Official

    A skill your agent uses to implement ONE microsoft/apm issue already selected by autopilot-issue-delivery-scheduler.

    4k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Drive ONE already selected open pull request in microsoft/apm to mergeable.

    4k GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Apm Spec Guardian

    microsoft/apm

    Official

    A skill your agent uses to run a four-panel adversarial advisory review on any pull request that touches the OpenAPM specification artifact (docs/src/content/docs/specs/openapm-.md), its inline /…

    4k GitHub stars~4.9k tokensUpdated yesterday
    Auto-check passed

Questions about Docs Grounding Verifier

What does Docs Grounding Verifier do?

A skill your agent uses to verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code. Docs Grounding Verifier is an agent skill from microsoft/apm, published by the product's own GitHub organization. Use this skill to verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code.

When should I use Docs Grounding Verifier?

Docs Grounding Verifier fits situations like: verify CLAIM-LEVEL grounding of a documentation page (or set of pages) against the source code; nouns: is this doc accurate; verify the page against the code; fact-check this section.

How do I install Docs Grounding Verifier in Claude Code?

Run `npx skills add microsoft/apm --skill docs-grounding-verifier -a claude-code`. Or copy the skill folder (.apm/skills/docs-grounding-verifier in microsoft/apm) into .claude/skills/docs-grounding-verifier in your project. Claude Code loads it when a task matches its description.

How do I install Docs Grounding Verifier in Codex?

Run `npx skills add microsoft/apm --skill docs-grounding-verifier -a codex`. Or copy the skill folder (.apm/skills/docs-grounding-verifier in microsoft/apm) into .agents/skills/docs-grounding-verifier in your project. Codex loads it when a task matches its description.

Can I use Docs Grounding Verifier in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/apm --skill docs-grounding-verifier -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/docs-grounding-verifier, .gemini/skills/docs-grounding-verifier, .github/skills/docs-grounding-verifier and .opencode/skills/docs-grounding-verifier in your project.

What does Docs Grounding Verifier need to run?

Going by SKILL.md and its folder, Docs Grounding Verifier needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Docs Grounding Verifier access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Docs Grounding Verifier safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Docs Grounding Verifier use?

Docs Grounding Verifier is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Docs Grounding Verifier use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Docs Grounding Verifier?

Skills that share tags, products or a category with Docs Grounding Verifier: Perplexity Web Search (davila7/claude-code-templates, 32k stars), Citation Verification Guide (Galaxy-Dawn/claude-scholar, 5.7k stars), Article Fact Checker (digoal/blog, 8.6k stars) and Deep Research Agent Team (Imbad0202/academic-research-skills, 51k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Docs Grounding Verifier?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/apm, which has 3,968 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 6, 2026.

Source: microsoft/apm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.