Agent skill

Self Verify Beacon In Sandbox

by Asymptote-Labs in Asymptote-Labs/agent-beacon

Verify a Beacon change end to end by running a real Claude Code session inside a disposable Linux cloud sandbox and checking that Beacon captured what the agent actually did.

MITAuto-check passedDevelopment

Install Self Verify Beacon In Sandbox

skills CLI
$ npx skills add Asymptote-Labs/agent-beacon --skill self-verify-beacon-in-sandbox -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Asymptote-Labs/agent-beacon self-verify-beacon-in-sandbox --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Asymptote-Labs/agent-beacon.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/self-verify-beacon-in-sandbox .claude/skills/self-verify-beacon-in-sandbox && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
self-verify-beacon-in-sandbox
GitHub stars
1.8k
Token cost
~1.2k tokens
SKILL.md length
564 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Verify a Beacon change end to end by running a real Claude Code session inside a disposable Linux cloud sandbox and checking that Beacon captured what the agent actually did.

  • Works in 3 steps: check prerequisites → resolve what doctor reports → run it
  • Asked to verify
  • SKILL.md covers Step 1: check prerequisites, Step 2: resolve what doctor…, Step 3: run it and Four things not to get wrong
  • Calls go, modal and make; needs MODAL_TOKEN_SECRET and ANTHROPIC_API_KEY

What it does

Self Verify Beacon In Sandbox is an agent skill from Asymptote-Labs/agent-beacon. Verify a Beacon change end to end by running a real Claude Code session inside a disposable Linux cloud sandbox and checking that Beacon captured what the agent actually did. Use when asked to verify, validate, test, or prove that a Beacon change works for real rather than just compiling; when asked whether telemetry, event capture, commands, file paths, prompts, tokens, or approvals are still recorded correctly; when investigating a suspected capture gap; or when preparing a Beacon pull request that touches the…

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Pull requests. It works with Linux. The repository describes itself as: The cross-harness, self-improving memory layer for AI agents. Join our community: https://discord.com/invite/zdNChS2fBu. The licence is MIT.

When your agent uses it

  • Asked to verify
  • Prove that a Beacon change works for real rather than just compiling
  • Asked whether telemetry
  • Approvals are still recorded correctly

Example prompts

  • “/self-verify-beacon-in-sandbox”

Requirements

  • Python 3
  • A credential in MODAL_TOKEN_SECRET
  • A credential in ANTHROPIC_API_KEY

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. check prerequisites
  2. resolve what doctor reports
  3. run it

What it can do on your machine

Read from SKILL.md and the folder at commit 2462839. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • go
    • modal
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MODAL_TOKEN_SECRET
    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Self Verify Beacon In Sandbox loads about 1.2k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 564 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Asymptote-Labs/agent-beacon at commit 2462839, republished under its MIT licence (© Asymptote-Labs). 564 words, ~1,243 tokens.

Download SKILL.mdSave it as .claude/skills/self-verify-beacon-in-sandbox/SKILL.md (or your agent's skills folder).
name
self-verify-beacon-in-sandbox
description
Verify a Beacon change end to end by running a real Claude Code session inside a disposable Linux cloud sandbox and checking that Beacon captured what the agent actually did. Use when asked to verify, validate, test, or prove that a Beacon change works for real rather than just compiling; when asked whether telemetry, event capture, commands, file paths, prompts, tokens, or approvals are still recorded correctly; when investigating a suspected capture gap; or when preparing a Beacon pull request that touches the CLI, hooks, or the collector exporter.
metadata.internal
true

Verify Beacon in a sandbox

This repository ships beacon-sandbox, which rents a disposable Linux sandbox from Modal, installs the Beacon under test, runs a real Claude Code session inside it, and checks whether Beacon recorded what the agent did.

Read beacon-sandbox/AGENTS.md for the full operating manual — scenario selection, how to interpret a verdict, and the failure modes that look like Beacon bugs but are not. The commands below are enough to start; that file is what stops you misreading the result.

Step 1: check prerequisites

Always start here. It is free, needs no sandbox, and prints the exact fix for anything missing:

bash
cd beacon-sandbox
go run ./cmd/beacon-sandbox doctor

Add --json for machine-readable output with a top-level ready boolean.

Step 2: resolve what doctor reports

Apply the fix: line it prints, then rerun doctor. The two build artifacts you can resolve yourself:

bash
cd cli/beacon && make build-linux-amd64      # if beacon_binary FAILs
cd beacon-sandbox && go run ./cmd/beacon-sandbox doctor --fix   # downloads the collector

Two prerequisites you must NOT try to resolve yourself:

  1. Modal authentication. modal token new completes through an authenticated web session, so it opens a browser and blocks — running it yourself will hang until timeout. If modal_auth FAILs, ask the user to run it, suggesting they type ! pip install modal && modal token new so the output lands in the conversation. In CI, MODAL_TOKEN_ID and MODAL_TOKEN_SECRET work non-interactively instead.
  2. The Anthropic credential. Never invent, echo, or write a key. If anthropic_credential FAILs, ask the user which of the three paths they want: ANTHROPIC_API_KEY, --api-key-command CMD, or --modal-secret NAME.

Do not proceed to step 3 while doctor reports any FAIL — the run will fail later and more confusingly.

Step 3: run it

bash
go run ./cmd/beacon-sandbox run --scenario s02-bash-command   # one scenario -- do this while iterating
go run ./cmd/beacon-sandbox run                               # the whole suite, ~30 min

Pick the scenario matching what changed:

You changedScenario
Command capture / exporter tool handlings02-bash-command
File read or write signalss03-file-write or s04-file-read
Prompt, session, token, or cost captures01-hello
Approval or permission handlings07-denied-tool
endpoint install, config paths, service startupi01-install-supervised
The systemd backend, unit files, Linux system modei02-install-systemd
Something broad, or preparing a PRthe whole suite (bare run)

The s0* scenarios collect through beacon ci exec, a temporary collector. The i0* scenarios install Beacon first, so they are the only ones that cover installation and service management. i02 runs inside a nested privileged container, because systemd will not start unless it is PID 1 and the sandbox provider's own init holds that slot — expect it to take several minutes longer.

Other flags: --repeat N to tell flaky from broken, --keep-sandbox to leave the instance up for debugging.

Show full SKILL.md (167 more words)Show less

Four things not to get wrong

  1. A run costs real money — about $0.06 of sandbox plus a few cents of API per scenario, on the user's account. Say what you are about to run and roughly what it costs before spending it, and prefer a single --scenario while iterating. If you only changed what counts as correct, use verify <run-dir> instead — it re-judges collected artifacts offline and free.
  2. A change under collector-builder/ needs the collector rebuilt. The telemetry normalization compiles into beacon-otelcol, not the beacon CLI, so otherwise you verify the wrong binary and get a meaningless pass. doctor warns about this as collector_freshness.
  3. INCONCLUSIVE means the model never did the work — retry the scenario; do not investigate Beacon.
  4. A FAIL may be a stale assertion, not a bug. Read the failing expectation's why field against the verdict's action histogram before concluding anything.

beacon-sandbox/AGENTS.md explains each of these properly, plus diff, scenario authoring, and how to self-test that the checks still have teeth.

© Asymptote-Labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/self-verify-beacon-in-sandbox of Asymptote-Labs/agent-beacon.

Open the folder on GitHubat commit 2462839

Compare with similar skills

Self Verify Beacon In Sandbox next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Self Verify Beacon In Sandbox compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Self Verify Beacon In Sandbox this skillAsymptote-Labs/agent-beacon1.8k—~1.2kAutomated safety check: PassMIT
Prosestatic-web-server/static-web-server2.4k—~1.2kAutomated safety check: PassApache-2.0
Debug Os Failure On GitHubstrands-agents/box315—~1.3kAutomated safety check: NotesApache-2.0
Neru Create PRy3owk1n/neru805—~1.9kAutomated safety check: PassMIT
Nixpkgs Reviewryan4yin/nix-config2.1k—~2.6kAutomated safety check: PassMIT
Choruz PRinclusionAI/Choruz1k—~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Prose

    static-web-server/static-web-server

    Write or edit human-readable text for the Static Web Server (SWS) project — commit messages, CHANGELOG entries, PR descriptions, issue bodies, rustdoc comments, CLI help text, READMEs, and user docs…

    2.4k GitHub stars~1.2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Debug Os Failure On GitHub

    strands-agents/box

    Debug a CI failure on an OS you are not on (you are on Linux, it fails on macos-latest, or the reverse) without opening a pull request per attempt.

    315 GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check: notes
  • Neru Create PR

    y3owk1n/neru

    Commit working changes and open a Neru pull request the maintainer's way: conventional commit subjects, a PR title written for the changelog, the just ci gate, and the repo PR template filled…

    805 GitHub stars~1.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Nixpkgs Review

    ryan4yin/nix-config

    A skill your agent uses when reviewing an upstream NixOS/nixpkgs pull request before it is merged -- a PR number or NixOS/nixpkgs123 link -- including its package changes, passthru tests…

    2.1k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Choruz PR

    inclusionAI/Choruz

    Use before opening or completing a pull request in this repository — classify the change, add the tests its type requires, run exactly what CI will run, label it, merge only on a green "CI (linux)…

    1k GitHub stars~2.6k tokensUpdated 4 days ago
    DevelopmentAuto-check passed
  • Neru File Issue

    y3owk1n/neru

    File a Neru bug report or feature request that matches the repo's issue forms: duplicate check first, every required field filled with real diagnostics, correct labels.

    805 GitHub stars~890 tokensUpdated today
    Testing & QAAuto-check passed

More from Asymptote-Labs/agent-beacon

  • Beacon Memory Distill

    Asymptote-Labs/agent-beacon

    Turn recorded agent sessions (Beacon traces from Claude Code, Cursor, Codex, OpenCode, and other harnesses) into reviewed, reusable project memory.

    1.8k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Beacon Lens Create

    Asymptote-Labs/agent-beacon

    Create, revise, debug or validate a Beacon lens, a single HTML file that renders one agent trace as a purpose-built view (a per-file review, a cost breakdown, a timeline of risky commands, a map of…

    1.8k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Beacon Memory Promote

    Asymptote-Labs/agent-beacon

    Install approved Beacon project memory as an Agent Skill in the repository (.agents/skills/<slug/SKILL.md), so every skill-capable harness loads the lesson automatically without a memory lookup.

    1.8k GitHub stars~927 tokensUpdated today
    Auto-check passed
  • Beacon Memory Recall

    Asymptote-Labs/agent-beacon

    Retrieve reviewed project memory that Beacon distilled from earlier agent sessions in any harness (Claude Code, Cursor, Codex, OpenCode, and others) before starting work.

    1.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Self Verify Beacon In Sandbox

What does Self Verify Beacon In Sandbox do?

Verify a Beacon change end to end by running a real Claude Code session inside a disposable Linux cloud sandbox and checking that Beacon captured what the agent actually did. Self Verify Beacon In Sandbox is an agent skill from Asymptote-Labs/agent-beacon. Verify a Beacon change end to end by running a real Claude Code session inside a disposable Linux cloud sandbox and checking that Beacon captured what the agent actually did.

When should I use Self Verify Beacon In Sandbox?

Self Verify Beacon In Sandbox fits situations like: asked to verify; prove that a Beacon change works for real rather than just compiling; asked whether telemetry; approvals are still recorded correctly.

How do I install Self Verify Beacon In Sandbox in Claude Code?

Run `npx skills add Asymptote-Labs/agent-beacon --skill self-verify-beacon-in-sandbox -a claude-code`. Or copy the skill folder (.claude/skills/self-verify-beacon-in-sandbox in Asymptote-Labs/agent-beacon) into .claude/skills/self-verify-beacon-in-sandbox in your project. Claude Code loads it when a task matches its description.

How do I install Self Verify Beacon In Sandbox in Codex?

Run `npx skills add Asymptote-Labs/agent-beacon --skill self-verify-beacon-in-sandbox -a codex`. Or copy the skill folder (.claude/skills/self-verify-beacon-in-sandbox in Asymptote-Labs/agent-beacon) into .agents/skills/self-verify-beacon-in-sandbox in your project. Codex loads it when a task matches its description.

Can I use Self Verify Beacon In Sandbox in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Asymptote-Labs/agent-beacon --skill self-verify-beacon-in-sandbox -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/self-verify-beacon-in-sandbox, .gemini/skills/self-verify-beacon-in-sandbox, .github/skills/self-verify-beacon-in-sandbox and .opencode/skills/self-verify-beacon-in-sandbox in your project.

What does Self Verify Beacon In Sandbox need to run?

Going by SKILL.md and its folder, Self Verify Beacon In Sandbox needs the command-line tools its instructions call (go, modal and make) and credentials named MODAL_TOKEN_SECRET and ANTHROPIC_API_KEY. Our summary lists: Python 3; A credential in MODAL_TOKEN_SECRET; A credential in ANTHROPIC_API_KEY.

Does Self Verify Beacon In Sandbox access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Self Verify Beacon In Sandbox safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Self Verify Beacon In Sandbox use?

Self Verify Beacon In Sandbox is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Self Verify Beacon In Sandbox use?

About 1.2k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Self Verify Beacon In Sandbox?

Skills that share tags, products or a category with Self Verify Beacon In Sandbox: Prose (static-web-server/static-web-server, 2.4k stars), Debug Os Failure On GitHub (strands-agents/box, 315 stars), Neru Create PR (y3owk1n/neru, 805 stars) and Nixpkgs Review (ryan4yin/nix-config, 2.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Self Verify Beacon In Sandbox?

Asymptote-Labs (a GitHub organization) maintains it in Asymptote-Labs/agent-beacon, which has 1,810 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 11, 2026.

Source: Asymptote-Labs/agent-beacon on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.