Agent skill

Analyze Rl Behavior

by open-thoughts in open-thoughts/OpenThoughts-Agent

Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

Apache-2.0Auto-check passed

Install Analyze Rl Behavior

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-rl-behavior --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyze-rl-behavior .claude/skills/analyze-rl-behavior && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyze-rl-behavior
GitHub stars
301
Token cost
~4.2k tokens
SKILL.md length
1,720 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

  • Works in 5 steps: Repo exists + ≥1 .safetensors — else the… → training_logs/ present — required for Q2… → RL-trace dataset exists — find from the… → …
  • Asked to analyze RL behavior
  • SKILL.md covers Step 0 — preflight artifact…, TL;DR invocation, Environment / secrets… and --model-repo auto-resolution…, plus 6 more sections
  • Calls python, curl and python3; reaches huggingface.co; needs HF_TOKEN and SUPABASE_SERVICE_ROLE_KEY

What it does

Analyze Rl Behavior is an agent skill from open-thoughts/OpenThoughts-Agent. Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL impact. Use when asked to "analyze RL behavior", "compare pre/post RL", "what did RL change", or to produce the Q1–Q4 behavioral report + GPT-5 judge for an laion/... (or any) RL checkpoint. Runs LOCALLY on the Mac (no GPU).

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with OpenAI. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Asked to analyze RL behavior
  • Compare pre/post RL
  • What did RL change
  • Produce the Q1–Q4 behavioral report + GPT-5 judge for an laion/..

Example prompts

  • “analyze RL behavior”
  • “compare pre/post RL”
  • “what did RL change”
  • “/analyze-rl-behavior”

Requirements

  • Python 3
  • A credential in SUPABASE_SERVICE_ROLE_KEY
  • A credential in OPENAI_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Repo exists + ≥1 .safetensors — else the model itself never landed (an RL-cleanup Step-6 miss); fix that first (re-upload weights from the…
  2. training_logs/ present — required for Q2 parse_skyrl_metrics. If absent, either complete RL-cleanup Step 9 first (upload training_logs) or…
  3. RL-trace dataset exists — find from the model repo's rl_config.json, check penfever/ exists on HF (/api/datasets/penfever/). If yes → pass…
  4. list-evals resolves a baseline/post-RL pair (run it — confirms Supabase has the eval jobs; pick/pin the benchmark if needed).
  5. Eval-repo write access for --annotate-failure-modes — if the eval repos are under an org you can't write (e.g. DCAgent2/3 as penfever)…

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • curl
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • SUPABASE_SERVICE_ROLE_KEY
    • OPENAI_API_KEY
    • LITELLM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyze Rl Behavior loads about 4.2k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 1,720 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 1,720 words, ~4,163 tokens.

Download SKILL.mdSave it as .claude/skills/analyze-rl-behavior/SKILL.md (or your agent's skills folder).
name
analyze-rl-behavior
description
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyze_rl_behavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL impact. Use when asked to "analyze RL behavior", "compare pre/post RL", "what did RL change", or to produce the Q1–Q4 behavioral report + GPT-5 judge for an `laion/...` (or any) RL checkpoint. Runs LOCALLY on the Mac (no GPU).

analyze-rl-behavior

Orchestrates scripts/analysis/analyze_rl_behavior.py — a local pipeline that pulls a trained RL model's eval traces + training logs from HF/Supabase and answers four research questions, each writing into --output-dir/<step>/:

  • Q1 (what changed): behavioral_delta (macro metrics + behavioral features) + llm_judge_diff (GPT-5 same-task pairwise) + optional annotate_failure_modes.
  • Q2 (attribution): temporal_trace_analysis + parse_skyrl_metrics (RL reward/KL/grad-norm over time).
  • Q3 (persistence): eval_temporal_overlay + trace_pair_render (side-by-side same-task pairs).
  • Q4 (eval impact): solve_rate_by_context.

Step 0 — preflight artifact check (ALWAYS do this FIRST)

Before running, verify the model's artifacts are all present — a missing one silently downgrades the run (skipped Q2/Q3) or wastes a full pass. For laion/<MODEL>:

bash
source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"
# (a) model repo exists + has weights + training_logs + README
curl -s -H "Authorization: Bearer $HF_TOKEN" "https://huggingface.co/api/models/laion/<MODEL>" \
 | python3 -c "import sys,json;d=json.load(sys.stdin);s=[x['rfilename'] for x in d.get('siblings',[])] if 'error' not in d else None;print('MISSING/404') if s is None else print('files',len(s),'| safetensors',sum(f.endswith('.safetensors') for f in s),'| training_logs',sum(f.startswith('training_logs/') for f in s),'| README','README.md' in s)"

Checklist (decide BEFORE launching):

  1. Repo exists + ≥1 .safetensors — else the model itself never landed (an RL-cleanup Step-6 miss); fix that first (re-upload weights from the Jupiter export), don't analyze a 404.
  2. training_logs/ present — required for Q2 parse_skyrl_metrics. If absent, either complete RL-cleanup Step 9 first (upload training_logs) or accept Q2-metrics will skip.
  3. RL-trace dataset exists — find <job_name> from the model repo's rl_config.json, check penfever/<job_name> exists on HF (/api/datasets/penfever/<job_name>). If yes → pass --rl-traces penfever/<job_name> (enables Q2-temporal + Q3-overlay). If 404 → those two steps just won't plan (fine, note it).
  4. --list-evals resolves a baseline/post-RL pair (run it — confirms Supabase has the eval jobs; pick/pin the benchmark if needed).
  5. Eval-repo write access for --annotate-failure-modes — if the eval repos are under an org you can't write (e.g. DCAgent2/3 as penfever), OMIT that flag (it 403s, wasted; see Cost section).

Only proceed to the run once 1 is satisfied; 2–3 determine which --rl-traces/Q2 steps you'll get; 5 determines whether to include --annotate-failure-modes.

TL;DR invocation

Run from the repo root /Users/benjaminfeuer/Documents/OpenThoughts-Agent, otagent env, secrets sourced:

bash
source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"

# 0. Preview what auto-resolve will pick (exits without running, no API spend):
/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python -m scripts.analysis.analyze_rl_behavior \
  --model-repo laion/<MODEL> \
  --list-evals \
  --output-dir /Users/benjaminfeuer/Documents/notes/RL/<run>/<MODEL>

# 1. Dry-run (confirm the planned step list resolves cleanly — still no spend):
#    same as the full command below + --dry-run

# 2. FULL run (cost-incurring steps ON by default here — see "Cost" to disable):
/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python -m scripts.analysis.analyze_rl_behavior \
  --model-repo laion/<MODEL> \
  --rl-traces penfever/<RL_TRACE_DATASET> \
  --annotate-failure-modes --llm-judge \
  --llm-judge-max-pairs 30 --llm-judge-concurrent 4 \
  --output-dir /Users/benjaminfeuer/Documents/notes/RL/<run>/<MODEL> \
  > /Users/benjaminfeuer/Documents/notes/RL/<run>/<MODEL>/_run.log 2>&1
  • Always run from repo root, with /Users/benjaminfeuer/miniconda3/envs/otagent/bin/python -m scripts.analysis.analyze_rl_behavior (the symlinked python doesn't work in the sandbox).
  • One model per --output-dir. Running multiple models into the same dir collides their <step>/ outputs — give each its own subdir.

Environment / secrets (mandatory)

source "$DC_AGENT_SECRET_ENV" first (NOT ~/secrets.env). The pipeline needs:

  • SUPABASE_URL + SUPABASE_SERVICE_ROLE_KEY — --model-repo auto-resolution (models + sandbox_jobs tables).
  • HF_TOKEN — eval-trace datasets + training_logs/ snapshots.
  • OPENAI_API_KEY — BOTH GPT-5 steps (no LITELLM_API_KEY needed; llm_judge falls back to the OpenAI SDK).

--model-repo auto-resolution (the clean entry point)

Given just --model-repo, the orchestrator calls scripts.analysis.auto_resolve.resolve() and autofills: --post-rl-eval, --baseline-eval, their *-ts, and --training-log-dir (snapshotted from the model repo's training_logs/ on HF). Explicit CLI values win on conflict. It does NOT resolve --rl-traces.

  • Default --eval-selection=largest-delta picks, among matched benchmark pairs, the one with the biggest positive post−baseline score gain. Other modes: largest-abs-delta (catches regressions), latest, benchmark (pin via --eval-benchmark).
  • --list-evals prints the matched / post-only / baseline-only pairs and exits — run it first to see (and, if needed, pin) the benchmark.
  • Baseline-eval timestamp defaults to the base model's training_end in Supabase.
  • Duplicate evals on the same benchmark → MOST-RECENT wins. A model often has >1 traces-bearing eval on the same benchmark (reruns). The resolver picks the one with the latest ended_at (the authoritative "when-evaluated" field; falls back to started_at then created_at, which is registration/backfill order and can disagree — so it's only a last-resort tiebreak). auto_resolve._eval_jobs_for_model sorts most-recent-first in Python (a null timestamp sorts as oldest, not newest — fixing PostgREST's NULLS-FIRST desc default), and every selection mode dedups by taking the first (= newest) per benchmark. This is a per-eval-job picker (newest wins); distinct from the ablation-table aggregation rule in crud-otagent-supabase (which averages identical-setting complete reruns).
⚠️ CROSS-MODEL COMPARISON: PIN ONE BENCHMARK (--eval-benchmark) — do NOT use the default

largest-delta is correct for single-model analysis ("what's this model's best eval pair?"), but it is WRONG for comparing models to each other: it picks a different benchmark per model — each model's best-looking one — so you end up comparing models on different yardsticks (and silently flattering each: a model that regressed on the shared benchmark can be surfaced via a different benchmark where it happened to gain). This is a real glitch that corrupted a 9-model study (2026-06-13): arm0-tis-15 showed +0.0156 on its cherry-picked dev_set_v2 but was −0.0300 on swebench_verified_random_100, while the hero was compared on swebench — not the same axis at all.

Rule: whenever you run analyze_rl_behavior across ≥2 models for comparison, pin every model to the same benchmark: --eval-selection benchmark --eval-benchmark <uuid> Pick a benchmark all the models share (run --list-evals per model to confirm coverage) and that is binary pass/fail (clean binomial SE + paired McNemar) — e.g. swebench_verified_random_100 (cc1aca76-98f5-4964-8d0b-efcb716b39c5) or terminal_bench_2 (34ab93c4-…); avoid partial-credit sets like dev_set_v2 (no clean SE). If no benchmark is universal, pin the max-coverage one and explicitly list the excluded models. Compare each model's delta only to its OWN benchmark's noise floor; never mix benchmarks in one ranking.

--rl-traces is NOT auto-resolved — pass it for Q2/Q3

Without --rl-traces, the Q2 temporal_trace_analysis and Q3 eval_temporal_overlay steps are silently not planned (only parse_skyrl_metrics covers Q2, and only if training_logs/ exists in the model repo). Find the right RL-trace dataset from the model repo's rl_config.json (job_name field) → penfever/<job_name>; verify it exists on HF (some models have none → 404, then Q2-temporal/Q3-overlay just won't run). Don't guess from HF search — same-recipe older runs have similar names.

Cost-incurring steps (ON in the TL;DR; how to disable)

Two GPT-5 steps. To disable, simply omit the flag (both are opt-in):

  • --llm-judge — GPT-5 pairwise same-task classification. Robust (per-pair JSON-repair fallback). Default model openai/gpt-5-2025-08-07, --llm-judge-max-pairs 30, --llm-judge-concurrent 4. Caches per-pair verdicts to <out>/Q1_llm_judge_diff/llm_judge_cache.json → re-runs are free. ~30 calls, ~3 min. Keep this on — it's the headline Q1 signal and cheap.

  • --annotate-failure-modes — GPT-5 (update_hf_failure_modes, hardcoded default gpt-5.1) annotates failure modes on the baseline + post-RL eval rows, then pushes the annotations back to the eval HF repo. Populates behavioral_delta's "Failure-mode distribution" section.

    ⚠️ Two real failure modes (observed on all three 2026-06-12 ablation runs):

    1. Write-back 403 on DCAgent2/DCAgent3 eval repos — the auth'd HF user is penfever, which can't write those orgs (and they're over public-storage quota). The GPT-5 work completes, the --push 403s, and because the step is optional=True it's logged "failed (rc=1) — non-fatal" and skipped — so the annotations are never persisted and behavioral_delta shows 0% failure-mode coverage. The full GPT-5 annotation budget (~tens-to-100+ batch calls over ~640 rows) is spent for zero usable output. It only pays off for a model whose eval repos YOU own/can write.
    2. Crashes on malformed GPT-5 JSON — update_hf_failure_modes.py:~191 does json.loads(content) with no per-batch try/except and no client timeout; one bad response (e.g. Invalid \escape) aborts the whole step. (A guard + retry + timeout there would fix both #1's wasted-spend visibility and this.)

    Recommendation: include --annotate-failure-modes only when the eval repos are writable by the authed HF user; otherwise omit it (saves the bulk of the runtime + cost; the failure-mode diff won't populate anyway).

Show full SKILL.md (648 more words)Show less

Step / Q mapping (what to expect)

StepNeedsOutput
Q0.annotate_failure_modes.{baseline,post-rl}--annotate-failure-modes + writable eval repoannotations pushed to eval repo; local Q0_failure_mode_*/done.txt only on rc=0
Q1.behavioral_deltaalwaysQ1_behavioral_delta/report.{md,json}
Q1.llm_judge_diff--llm-judgeQ1_llm_judge_diff/report.{md,json}, llm_judge_cache.json
Q2.parse_skyrl_metricstraining_logs/ in model repo (auto-snapshotted)Q2_skyrl_metrics/ CSVs + report + reward_vs_steps.png
Q2.temporal_trace_analysis--rl-tracestemporal plots
Q3.eval_temporal_overlay--rl-tracesoverlay.png
Q3.trace_pair_renderalwaysQ3_trace_pairs/pairs.html (multi-MB)
Q4.solve_rate_by_contextalwaysQ4_solve_rate_by_context/solve_rate.png

Top-level always: INDEX.md (cross-links every step — written LAST; it is the reliable completion marker), pipeline_plan.json, auto_resolve.json, _orchestrator_run.log.

Re-running to fill skipped steps (partial-run / late-arriving inputs)

Common case: the first run skipped Q2 (parse_skyrl_metrics, temporal_trace_analysis) and/or Q3 (eval_temporal_overlay) because training_logs/ wasn't in the model repo yet or --rl-traces wasn't passed. Once those inputs land (e.g. the RL-cleanup Step-9 upload finishes, or you locate the trace dataset), re-run the same command with the missing inputs supplied to fill the gaps:

  • Steps that were never planned (Q2-temporal / Q3-overlay when --rl-traces was absent; parse_skyrl_metrics when no training_logs/) wrote no marker, so they run on the re-run automatically — no --force needed. Just pass --rl-traces <hf-id> and make sure training_logs/ now exists in the model repo.
  • Steps that already produced output (behavioral_delta, llm_judge_diff, trace_pair_render, solve_rate_by_context) are skipped (marker exists) — fine, they don't depend on the late inputs. llm_judge_diff re-hits its cache (free) if it does re-run.
  • Use --force ONLY to refresh a step whose marker exists but whose inputs changed — chiefly behavioral_delta after a successful annotate_failure_modes (the stale-cache trap below).
  • For pure fill re-runs, omit --annotate-failure-modes — on eval repos you can't write (e.g. DCAgent2/DCAgent3 as penfever) it only re-burns GPT-5 budget and 403s without populating anything. Keep --llm-judge (cached → free on re-run).

Output-dir durability: write --output-dir to a dedicated per-model subdir, NOT the ~/Documents/notes/... root of a shared folder. A root-level run on this Mac (iCloud-synced ~/Documents) was observed to not persist its Q-dirs/INDEX.md even after the orchestrator reported success — use .../ablation_exploration_in_rl/<model>/ per model.

Operational gotchas

  • Completion signal = the per---output-dir INDEX.md file appearing. Do NOT rely on process-liveness — on this Mac (~/Documents iCloud + sandbox /tmp namespace) pgrep/kill -0/ps from background/monitor shells return phantom "process gone", and tqdm/logging stdout is block-buffered and lags minutes. Poll for INDEX.md (or the terminal Q4_solve_rate_by_context/solve_rate.png) via a foreground loop, not the Monitor tool.
  • Do NOT pipe the orchestrator stdout through grep/tail — that triggers auto-backgrounding and the output is lost. Use a plain > log 2>&1 redirect.
  • Do NOT launch duplicate concurrent runs against the same --output-dir/eval repo — they share the OpenAI budget + HF --resume state and clobber the same log.
  • Resumable: skip-if-output-marker-exists per step + the llm_judge cache. Killing and relaunching is safe; a re-run skips completed steps and the judge is 30/30 cache hits (free). Stale-cache trap: if Q1_behavioral_delta/report.md already exists, behavioral_delta is skipped — so a later successful annotation won't refresh the failure-mode diff without --force (or deleting the report).
  • Runtime: ~3–10 min with --llm-judge only; ~10–45 min with --annotate-failure-modes (it annotates ALL eval rows sequentially via GPT-5 — the slow part, even when it ultimately 403s).

Worked example (2026-06-12)

Three pymethods2test ablation checkpoints, each --model-repo laion/<m> --annotate-failure-modes --llm-judge into its own subdir under notes/RL/ablation_exploration_in_rl/:

  • ablation-pymethods2test-seqmean-arm0-tis-15-8B → judge 66.7% post-RL win (Q2 ran — had training_logs).
  • ablation-pymethods2test-shaped-45-8B → judge 53.3% post-RL win.
  • ablation-pymethods2test-seqmean-arm0-30-8B → judge 50/50; behavioral_delta showed reward dipped 0.457→0.379. All three: --annotate-failure-modes 403'd on the DCAgent2/3 eval repos (failure-mode section empty); --llm-judge succeeded 30/30. --rl-traces had to be passed/located via rl_config.json; absent for two → Q2-temporal/Q3-overlay skipped.

Operating notes (folded from memory 2026-06-14)

  • teacher_hint marker: the teacher_hint PRM (prm/teacher_hint.py) injects hints into the student's observation text — NOT a separate steps[].source in the ATIF trajectory. Grep for the literal [HINT FROM TEACHER]: (wrapped by \n\n[HINT FROM TEACHER]: … \n\n). Appears in agent/trajectory.json (substring of an agent-source step), agent/episode-N/prompt.txt (the prompt AFTER the hint fired), and maybe episode-N/debug.json. Fires every check_interval turns (default 5; prod used 8), skipped if turn < min_turns (default 3; prod 4), and silently returns None if the teacher engine fails to init/generate — so absence of the marker doesn't distinguish "not eligible" from "engine failed"; cross-reference turn count + trial.log/exception.txt.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/analyze-rl-behavior of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Analyze Rl Behavior next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyze Rl Behavior compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyze Rl Behavior this skillopen-thoughts/OpenThoughts-Agent301—~4.2kAutomated safety check: PassApache-2.0
Geo Fundamentalswasp-lang/wasp19k9 repos~861Automated safety check: PassMIT
AI SDKvercel-labs/ai-facts16821 repos~1.2kAutomated safety check: PassNone
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
ModLens Image Vision Bridgeliustack/modlens4.2k1 repos~1.3kAutomated safety check: NotesMIT
PR Design DocOpenHands/OpenHands90k—~2.4kAutomated safety check: PassMIT

Similar skills

  • Geo Fundamentals

    wasp-lang/wasp

    Generative Engine Optimization for AI search engines (ChatGPT, Claude, Perplexity).

    19k GitHub starsUsed in 9 repos~861 tokens
    Marketing & SEOAuto-check passed
  • AI SDK

    vercel-labs/ai-facts

    Official

    Answer questions about the AI SDK and help build AI-powered features.

    168 GitHub starsUsed in 21 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.2k GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check: notes
  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    90k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Bibi

    JimmyLv/BibiGPT-v1

    BibiGPT CLI for summarizing videos, audio, and podcasts directly in the terminal.

    6.2k GitHub starsUsed in 1 repo~885 tokens
    Media & CreativeAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 9 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed
  • Commit

    open-thoughts/OpenThoughts-Agent

    Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format.

    301 GitHub stars~2.2k tokensUpdated 9 days ago
    Auto-check: notes

Works with

Questions about Analyze Rl Behavior

What does Analyze Rl Behavior do?

Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…. Analyze Rl Behavior is an agent skill from open-thoughts/OpenThoughts-Agent.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL impact.

When should I use Analyze Rl Behavior?

Analyze Rl Behavior fits situations like: asked to analyze RL behavior; compare pre/post RL; what did RL change; produce the Q1–Q4 behavioral report + GPT-5 judge for an laion/..

How do I install Analyze Rl Behavior in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a claude-code`. Or copy the skill folder (.agents/skills/analyze-rl-behavior in open-thoughts/OpenThoughts-Agent) into .claude/skills/analyze-rl-behavior in your project. Claude Code loads it when a task matches its description.

How do I install Analyze Rl Behavior in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a codex`. Or copy the skill folder (.agents/skills/analyze-rl-behavior in open-thoughts/OpenThoughts-Agent) into .agents/skills/analyze-rl-behavior in your project. Codex loads it when a task matches its description.

Can I use Analyze Rl Behavior in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-rl-behavior -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-rl-behavior, .gemini/skills/analyze-rl-behavior, .github/skills/analyze-rl-behavior and .opencode/skills/analyze-rl-behavior in your project.

What does Analyze Rl Behavior need to run?

Going by SKILL.md and its folder, Analyze Rl Behavior needs the command-line tools its instructions call (python, curl and python3) and credentials named HF_TOKEN, SUPABASE_SERVICE_ROLE_KEY, OPENAI_API_KEY and LITELLM_API_KEY. Our summary lists: Python 3; A credential in SUPABASE_SERVICE_ROLE_KEY; A credential in OPENAI_API_KEY.

Does Analyze Rl Behavior access the network?

SKILL.md names 1 domain. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Analyze Rl Behavior safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyze Rl Behavior use?

Analyze Rl Behavior is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyze Rl Behavior use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyze Rl Behavior?

Skills that share tags, products or a category with Analyze Rl Behavior: Geo Fundamentals (wasp-lang/wasp, 19k stars), AI SDK (vercel-labs/ai-facts, 168 stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars) and ModLens Image Vision Bridge (liustack/modlens, 4.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyze Rl Behavior?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.