Geo Fundamentals
wasp-lang/wasp
Generative Engine Optimization for AI search engines (ChatGPT, Claude, Perplexity).
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-rl-behavior --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyze-rl-behavior .claude/skills/analyze-rl-behavior && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "analyze-rl-behavior" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-rl-behavior into .claude/skills/analyze-rl-behavior/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-rl-behavior", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-rl-behaviorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-rl-behavior --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/analyze-rl-behavior .agents/skills/analyze-rl-behavior && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "analyze-rl-behavior" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-rl-behavior into .agents/skills/analyze-rl-behavior/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-rl-behavior", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-rl-behavior --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/analyze-rl-behavior .cursor/skills/analyze-rl-behavior && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "analyze-rl-behavior" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-rl-behavior into .cursor/skills/analyze-rl-behavior/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-rl-behavior", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/analyze-rl-behavior--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-rl-behavior --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/analyze-rl-behavior .gemini/skills/analyze-rl-behavior && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "analyze-rl-behavior" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-rl-behavior into .gemini/skills/analyze-rl-behavior/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-rl-behavior", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-rl-behaviorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/analyze-rl-behavior .github/skills/analyze-rl-behavior && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "analyze-rl-behavior" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-rl-behavior into .github/skills/analyze-rl-behavior/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-rl-behavior", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-rl-behavior --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/analyze-rl-behavior .opencode/skills/analyze-rl-behavior && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "analyze-rl-behavior" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-rl-behavior into .opencode/skills/analyze-rl-behavior/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-rl-behavior", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
analyze-rl-behaviorRun the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
Analyze Rl Behavior is an agent skill from open-thoughts/OpenThoughts-Agent. Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL impact. Use when asked to "analyze RL behavior", "compare pre/post RL", "what did RL change", or to produce the Q1–Q4 behavioral report + GPT-5 judge for an laion/... (or any) RL checkpoint. Runs LOCALLY on the Mac (no GPU).
Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It works with OpenAI. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythoncurlpython3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENSUPABASE_SERVICE_ROLE_KEYOPENAI_API_KEYLITELLM_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Analyze Rl Behavior loads about 4.2k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 1,720 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 1,720 words, ~4,163 tokens.
.claude/skills/analyze-rl-behavior/SKILL.md (or your agent's skills folder).Orchestrates scripts/analysis/analyze_rl_behavior.py — a local pipeline that pulls a
trained RL model's eval traces + training logs from HF/Supabase and answers four research
questions, each writing into --output-dir/<step>/:
behavioral_delta (macro metrics + behavioral features) + llm_judge_diff (GPT-5 same-task pairwise) + optional annotate_failure_modes.temporal_trace_analysis + parse_skyrl_metrics (RL reward/KL/grad-norm over time).eval_temporal_overlay + trace_pair_render (side-by-side same-task pairs).solve_rate_by_context.Before running, verify the model's artifacts are all present — a missing one silently downgrades the run (skipped Q2/Q3) or wastes a full pass. For laion/<MODEL>:
source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"
# (a) model repo exists + has weights + training_logs + README
curl -s -H "Authorization: Bearer $HF_TOKEN" "https://huggingface.co/api/models/laion/<MODEL>" \
| python3 -c "import sys,json;d=json.load(sys.stdin);s=[x['rfilename'] for x in d.get('siblings',[])] if 'error' not in d else None;print('MISSING/404') if s is None else print('files',len(s),'| safetensors',sum(f.endswith('.safetensors') for f in s),'| training_logs',sum(f.startswith('training_logs/') for f in s),'| README','README.md' in s)"Checklist (decide BEFORE launching):
.safetensors — else the model itself never landed (an RL-cleanup Step-6 miss); fix that first (re-upload weights from the Jupiter export), don't analyze a 404.training_logs/ present — required for Q2 parse_skyrl_metrics. If absent, either complete RL-cleanup Step 9 first (upload training_logs) or accept Q2-metrics will skip.<job_name> from the model repo's rl_config.json, check penfever/<job_name> exists on HF (/api/datasets/penfever/<job_name>). If yes → pass --rl-traces penfever/<job_name> (enables Q2-temporal + Q3-overlay). If 404 → those two steps just won't plan (fine, note it).--list-evals resolves a baseline/post-RL pair (run it — confirms Supabase has the eval jobs; pick/pin the benchmark if needed).--annotate-failure-modes — if the eval repos are under an org you can't write (e.g. DCAgent2/3 as penfever), OMIT that flag (it 403s, wasted; see Cost section).Only proceed to the run once 1 is satisfied; 2–3 determine which --rl-traces/Q2 steps you'll get; 5 determines whether to include --annotate-failure-modes.
Run from the repo root /Users/benjaminfeuer/Documents/OpenThoughts-Agent, otagent env, secrets sourced:
source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"
# 0. Preview what auto-resolve will pick (exits without running, no API spend):
/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python -m scripts.analysis.analyze_rl_behavior \
--model-repo laion/<MODEL> \
--list-evals \
--output-dir /Users/benjaminfeuer/Documents/notes/RL/<run>/<MODEL>
# 1. Dry-run (confirm the planned step list resolves cleanly — still no spend):
# same as the full command below + --dry-run
# 2. FULL run (cost-incurring steps ON by default here — see "Cost" to disable):
/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python -m scripts.analysis.analyze_rl_behavior \
--model-repo laion/<MODEL> \
--rl-traces penfever/<RL_TRACE_DATASET> \
--annotate-failure-modes --llm-judge \
--llm-judge-max-pairs 30 --llm-judge-concurrent 4 \
--output-dir /Users/benjaminfeuer/Documents/notes/RL/<run>/<MODEL> \
> /Users/benjaminfeuer/Documents/notes/RL/<run>/<MODEL>/_run.log 2>&1/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python -m scripts.analysis.analyze_rl_behavior (the symlinked python doesn't work in the sandbox).--output-dir. Running multiple models into the same dir collides their <step>/ outputs — give each its own subdir.source "$DC_AGENT_SECRET_ENV" first (NOT ~/secrets.env). The pipeline needs:
SUPABASE_URL + SUPABASE_SERVICE_ROLE_KEY — --model-repo auto-resolution (models + sandbox_jobs tables).HF_TOKEN — eval-trace datasets + training_logs/ snapshots.OPENAI_API_KEY — BOTH GPT-5 steps (no LITELLM_API_KEY needed; llm_judge falls back to the OpenAI SDK).--model-repo auto-resolution (the clean entry point)Given just --model-repo, the orchestrator calls scripts.analysis.auto_resolve.resolve() and autofills:
--post-rl-eval, --baseline-eval, their *-ts, and --training-log-dir (snapshotted from the model repo's training_logs/ on HF). Explicit CLI values win on conflict. It does NOT resolve --rl-traces.
--eval-selection=largest-delta picks, among matched benchmark pairs, the one with the biggest positive post−baseline score gain. Other modes: largest-abs-delta (catches regressions), latest, benchmark (pin via --eval-benchmark).--list-evals prints the matched / post-only / baseline-only pairs and exits — run it first to see (and, if needed, pin) the benchmark.training_end in Supabase.ended_at (the authoritative "when-evaluated" field; falls back to started_at then created_at, which is registration/backfill order and can disagree — so it's only a last-resort tiebreak). auto_resolve._eval_jobs_for_model sorts most-recent-first in Python (a null timestamp sorts as oldest, not newest — fixing PostgREST's NULLS-FIRST desc default), and every selection mode dedups by taking the first (= newest) per benchmark. This is a per-eval-job picker (newest wins); distinct from the ablation-table aggregation rule in crud-otagent-supabase (which averages identical-setting complete reruns).⚠️ CROSS-MODEL COMPARISON: PIN ONE BENCHMARK (
--eval-benchmark) — do NOT use the default
largest-deltais correct for single-model analysis ("what's this model's best eval pair?"), but it is WRONG for comparing models to each other: it picks a different benchmark per model — each model's best-looking one — so you end up comparing models on different yardsticks (and silently flattering each: a model that regressed on the shared benchmark can be surfaced via a different benchmark where it happened to gain). This is a real glitch that corrupted a 9-model study (2026-06-13):arm0-tis-15showed +0.0156 on its cherry-picked dev_set_v2 but was −0.0300 on swebench_verified_random_100, while the hero was compared on swebench — not the same axis at all.Rule: whenever you run
analyze_rl_behavioracross ≥2 models for comparison, pin every model to the same benchmark:--eval-selection benchmark --eval-benchmark <uuid>Pick a benchmark all the models share (run--list-evalsper model to confirm coverage) and that is binary pass/fail (clean binomial SE + paired McNemar) — e.g.swebench_verified_random_100(cc1aca76-98f5-4964-8d0b-efcb716b39c5) orterminal_bench_2(34ab93c4-…); avoid partial-credit sets likedev_set_v2(no clean SE). If no benchmark is universal, pin the max-coverage one and explicitly list the excluded models. Compare each model's delta only to its OWN benchmark's noise floor; never mix benchmarks in one ranking.
--rl-traces is NOT auto-resolved — pass it for Q2/Q3Without --rl-traces, the Q2 temporal_trace_analysis and Q3 eval_temporal_overlay steps are silently not planned (only parse_skyrl_metrics covers Q2, and only if training_logs/ exists in the model repo). Find the right RL-trace dataset from the model repo's rl_config.json (job_name field) → penfever/<job_name>; verify it exists on HF (some models have none → 404, then Q2-temporal/Q3-overlay just won't run). Don't guess from HF search — same-recipe older runs have similar names.
Two GPT-5 steps. To disable, simply omit the flag (both are opt-in):
--llm-judge — GPT-5 pairwise same-task classification. Robust (per-pair JSON-repair fallback). Default model openai/gpt-5-2025-08-07, --llm-judge-max-pairs 30, --llm-judge-concurrent 4. Caches per-pair verdicts to <out>/Q1_llm_judge_diff/llm_judge_cache.json → re-runs are free. ~30 calls, ~3 min. Keep this on — it's the headline Q1 signal and cheap.
--annotate-failure-modes — GPT-5 (update_hf_failure_modes, hardcoded default gpt-5.1) annotates failure modes on the baseline + post-RL eval rows, then pushes the annotations back to the eval HF repo. Populates behavioral_delta's "Failure-mode distribution" section.
⚠️ Two real failure modes (observed on all three 2026-06-12 ablation runs):
DCAgent2/DCAgent3 eval repos — the auth'd HF user is penfever, which can't write those orgs (and they're over public-storage quota). The GPT-5 work completes, the --push 403s, and because the step is optional=True it's logged "failed (rc=1) — non-fatal" and skipped — so the annotations are never persisted and behavioral_delta shows 0% failure-mode coverage. The full GPT-5 annotation budget (~tens-to-100+ batch calls over ~640 rows) is spent for zero usable output. It only pays off for a model whose eval repos YOU own/can write.update_hf_failure_modes.py:~191 does json.loads(content) with no per-batch try/except and no client timeout; one bad response (e.g. Invalid \escape) aborts the whole step. (A guard + retry + timeout there would fix both #1's wasted-spend visibility and this.)Recommendation: include --annotate-failure-modes only when the eval repos are writable by the authed HF user; otherwise omit it (saves the bulk of the runtime + cost; the failure-mode diff won't populate anyway).
| Step | Needs | Output |
|---|---|---|
| Q0.annotate_failure_modes.{baseline,post-rl} | --annotate-failure-modes + writable eval repo | annotations pushed to eval repo; local Q0_failure_mode_*/done.txt only on rc=0 |
| Q1.behavioral_delta | always | Q1_behavioral_delta/report.{md,json} |
| Q1.llm_judge_diff | --llm-judge | Q1_llm_judge_diff/report.{md,json}, llm_judge_cache.json |
| Q2.parse_skyrl_metrics | training_logs/ in model repo (auto-snapshotted) | Q2_skyrl_metrics/ CSVs + report + reward_vs_steps.png |
| Q2.temporal_trace_analysis | --rl-traces | temporal plots |
| Q3.eval_temporal_overlay | --rl-traces | overlay.png |
| Q3.trace_pair_render | always | Q3_trace_pairs/pairs.html (multi-MB) |
| Q4.solve_rate_by_context | always | Q4_solve_rate_by_context/solve_rate.png |
Top-level always: INDEX.md (cross-links every step — written LAST; it is the reliable completion marker), pipeline_plan.json, auto_resolve.json, _orchestrator_run.log.
Common case: the first run skipped Q2 (parse_skyrl_metrics, temporal_trace_analysis) and/or Q3 (eval_temporal_overlay) because training_logs/ wasn't in the model repo yet or --rl-traces wasn't passed. Once those inputs land (e.g. the RL-cleanup Step-9 upload finishes, or you locate the trace dataset), re-run the same command with the missing inputs supplied to fill the gaps:
--rl-traces was absent; parse_skyrl_metrics when no training_logs/) wrote no marker, so they run on the re-run automatically — no --force needed. Just pass --rl-traces <hf-id> and make sure training_logs/ now exists in the model repo.behavioral_delta, llm_judge_diff, trace_pair_render, solve_rate_by_context) are skipped (marker exists) — fine, they don't depend on the late inputs. llm_judge_diff re-hits its cache (free) if it does re-run.--force ONLY to refresh a step whose marker exists but whose inputs changed — chiefly behavioral_delta after a successful annotate_failure_modes (the stale-cache trap below).--annotate-failure-modes — on eval repos you can't write (e.g. DCAgent2/DCAgent3 as penfever) it only re-burns GPT-5 budget and 403s without populating anything. Keep --llm-judge (cached → free on re-run).Output-dir durability: write --output-dir to a dedicated per-model subdir, NOT the ~/Documents/notes/... root of a shared folder. A root-level run on this Mac (iCloud-synced ~/Documents) was observed to not persist its Q-dirs/INDEX.md even after the orchestrator reported success — use .../ablation_exploration_in_rl/<model>/ per model.
--output-dir INDEX.md file appearing. Do NOT rely on process-liveness — on this Mac (~/Documents iCloud + sandbox /tmp namespace) pgrep/kill -0/ps from background/monitor shells return phantom "process gone", and tqdm/logging stdout is block-buffered and lags minutes. Poll for INDEX.md (or the terminal Q4_solve_rate_by_context/solve_rate.png) via a foreground loop, not the Monitor tool.grep/tail — that triggers auto-backgrounding and the output is lost. Use a plain > log 2>&1 redirect.--output-dir/eval repo — they share the OpenAI budget + HF --resume state and clobber the same log.Q1_behavioral_delta/report.md already exists, behavioral_delta is skipped — so a later successful annotation won't refresh the failure-mode diff without --force (or deleting the report).--llm-judge only; ~10–45 min with --annotate-failure-modes (it annotates ALL eval rows sequentially via GPT-5 — the slow part, even when it ultimately 403s).Three pymethods2test ablation checkpoints, each --model-repo laion/<m> --annotate-failure-modes --llm-judge into its own subdir under notes/RL/ablation_exploration_in_rl/:
ablation-pymethods2test-seqmean-arm0-tis-15-8B → judge 66.7% post-RL win (Q2 ran — had training_logs).ablation-pymethods2test-shaped-45-8B → judge 53.3% post-RL win.ablation-pymethods2test-seqmean-arm0-30-8B → judge 50/50; behavioral_delta showed reward dipped 0.457→0.379.
All three: --annotate-failure-modes 403'd on the DCAgent2/3 eval repos (failure-mode section empty); --llm-judge succeeded 30/30. --rl-traces had to be passed/located via rl_config.json; absent for two → Q2-temporal/Q3-overlay skipped.teacher_hint PRM (prm/teacher_hint.py) injects hints into the student's observation text — NOT a separate steps[].source in the ATIF trajectory. Grep for the literal [HINT FROM TEACHER]: (wrapped by \n\n[HINT FROM TEACHER]: … \n\n). Appears in agent/trajectory.json (substring of an agent-source step), agent/episode-N/prompt.txt (the prompt AFTER the hint fired), and maybe episode-N/debug.json. Fires every check_interval turns (default 5; prod used 8), skipped if turn < min_turns (default 3; prod 4), and silently returns None if the teacher engine fails to init/generate — so absence of the marker doesn't distinguish "not eligible" from "engine failed"; cross-reference turn count + trial.log/exception.txt.© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/analyze-rl-behavior of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Analyze Rl Behavior next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Analyze Rl Behavior this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Geo Fundamentalswasp-lang/wasp | 19k | 9 repos | ~861 | Automated safety check: Pass | MIT | |
| AI SDKvercel-labs/ai-facts | 168 | 21 repos | ~1.2k | Automated safety check: Pass | None | |
| AI Image Generation and Editingzhayujie/CowAgent | 47k | — | ~1.3k | Automated safety check: Pass | MIT | |
| ModLens Image Vision Bridgeliustack/modlens | 4.2k | 1 repos | ~1.3k | Automated safety check: Notes | MIT | |
| PR Design DocOpenHands/OpenHands | 90k | — | ~2.4k | Automated safety check: Pass | MIT |
wasp-lang/wasp
Generative Engine Optimization for AI search engines (ChatGPT, Claude, Perplexity).
vercel-labs/ai-facts
Answer questions about the AI SDK and help build AI-powered features.
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
liustack/modlens
Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.
OpenHands/OpenHands
For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…
JimmyLv/BibiGPT-v1
BibiGPT CLI for summarizing videos, audio, and podcasts directly in the terminal.
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
open-thoughts/OpenThoughts-Agent
Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format.
Works with
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…. Analyze Rl Behavior is an agent skill from open-thoughts/OpenThoughts-Agent.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL impact.
Analyze Rl Behavior fits situations like: asked to analyze RL behavior; compare pre/post RL; what did RL change; produce the Q1–Q4 behavioral report + GPT-5 judge for an laion/..
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a claude-code`. Or copy the skill folder (.agents/skills/analyze-rl-behavior in open-thoughts/OpenThoughts-Agent) into .claude/skills/analyze-rl-behavior in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a codex`. Or copy the skill folder (.agents/skills/analyze-rl-behavior in open-thoughts/OpenThoughts-Agent) into .agents/skills/analyze-rl-behavior in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-rl-behavior, .gemini/skills/analyze-rl-behavior, .github/skills/analyze-rl-behavior and .opencode/skills/analyze-rl-behavior in your project.
Going by SKILL.md and its folder, Analyze Rl Behavior needs the command-line tools its instructions call (python, curl and python3) and credentials named HF_TOKEN, SUPABASE_SERVICE_ROLE_KEY, OPENAI_API_KEY and LITELLM_API_KEY. Our summary lists: Python 3; A credential in SUPABASE_SERVICE_ROLE_KEY; A credential in OPENAI_API_KEY.
SKILL.md names 1 domain. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Analyze Rl Behavior is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Analyze Rl Behavior: Geo Fundamentals (wasp-lang/wasp, 19k stars), AI SDK (vercel-labs/ai-facts, 168 stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars) and ModLens Image Vision Bridge (liustack/modlens, 4.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.