Paper Orchestra
Ar9av/PaperOrchestra
Orchestrate the full PaperOrchestra (Song et al., 2026, arXiv:2604.05018) five-agent pipeline to turn unstructured research materials (idea, experimental log, LaTeX template, conference guidelines…
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-id-eval-ranking --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyze-id-eval-ranking .claude/skills/analyze-id-eval-ranking && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "analyze-id-eval-ranking" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-id-eval-ranking into .claude/skills/analyze-id-eval-ranking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-id-eval-ranking", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-id-eval-rankingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-id-eval-ranking --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/analyze-id-eval-ranking .agents/skills/analyze-id-eval-ranking && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "analyze-id-eval-ranking" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-id-eval-ranking into .agents/skills/analyze-id-eval-ranking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-id-eval-ranking", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-id-eval-ranking --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/analyze-id-eval-ranking .cursor/skills/analyze-id-eval-ranking && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "analyze-id-eval-ranking" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-id-eval-ranking into .cursor/skills/analyze-id-eval-ranking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-id-eval-ranking", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/analyze-id-eval-ranking--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-id-eval-ranking --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/analyze-id-eval-ranking .gemini/skills/analyze-id-eval-ranking && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "analyze-id-eval-ranking" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-id-eval-ranking into .gemini/skills/analyze-id-eval-ranking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-id-eval-ranking", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-id-eval-rankingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/analyze-id-eval-ranking .github/skills/analyze-id-eval-ranking && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "analyze-id-eval-ranking" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-id-eval-ranking into .github/skills/analyze-id-eval-ranking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-id-eval-ranking", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-id-eval-ranking --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/analyze-id-eval-ranking .opencode/skills/analyze-id-eval-ranking && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "analyze-id-eval-ranking" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-id-eval-ranking into .opencode/skills/analyze-id-eval-ranking/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-id-eval-ranking", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
analyze-id-eval-rankingGiven a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
Analyze Id Eval Ranking is an agent skill from open-thoughts/OpenThoughts-Agent. Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100, OT-TBLite=devsetv2, Terminal-Bench-2.0=tb2), HF links to each eval's trace dataset, and a NORMALIZED column = average per-benchmark z-score, ranked. Normalization matches the OpenThoughts-Agent paper (otagent-paper/02arXiv/otagent.tex §Pipeline): per-benchmark z over the candidate set, averaged. Read-only. Use when asked to rank models…
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering LaTeX, LLM evaluation and Database schema design. It works with Supabase and arXiv. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Analyze Id Eval Ranking loads about 3.1k tokens when it runs. Until then it costs about 151 tokens; SKILL.md has 1,113 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 1,113 words, ~3,062 tokens.
.claude/skills/analyze-id-eval-ranking/SKILL.md (or your agent's skills folder).Produce the paper's ID ranking table for an arbitrary set of models: raw scores on the three
in-distribution agentic benchmarks + HF trace links + a normalized average z-score column,
sorted by the normalized score. This reproduces the ranking method in
otagent-paper/02_arXiv/otagent.tex (§Pipeline / App. task-gen tables). Read-only — it never
writes Supabase.
| paper name | Supabase benchmarks.name | task count N (for SE) |
|---|---|---|
| SWE-Bench Verified (100) | swebench-verified-random-100-folders | 100 |
| OT-TBLite | dev_set_v2 (partial-credit) | — |
| Terminal-Bench 2.0 | terminal_bench_2 | 89 |
⚠ Mapping traps: OT-TBLite IS
dev_set_v2(not a separate benchmark). SWE-Bench-100 is the-random-100-folderssubset, NOT fullswebench-verified(500, which is OOD).terminal_bench_2runs attimeout_multiplier 2.0→ resolve its family +dev_set_v2's family viaduplicate_of(seecrud-otagent-supabase§GOTCHA 2/3).dev_set_v2is partial-credit → its raw % still enters the mean and its z-score, but it has no clean binomialN.
"We compute the z-score of every candidate strategy's accuracy across the stage's full candidate set (subtracting the per-benchmark mean and dividing by its standard deviation), then average the three resulting per-benchmark z-scores."
So, with the candidate set = the input model list (this is the population for mean/std — NOT a global population):
b, over all candidate models with a score on b: mean_b, std_b.z[m,b] = (acc[m,b] − mean_b) / std_b.normalized[m] = mean(z[m,b] over the 3 benchmarks the model has).normalized descending.std uses population std (ddof=0, numpy.std default) — the candidate set IS the full
population being compared. (Document this if you switch to sample std; it changes the magnitudes,
not the ordering, when all models have all 3 benchmarks.) Equal per-benchmark weight is the whole
point — don't weight by N.
PREREQUISITE — read
.agents/skills/crud-otagent-supabase/SKILL.mdFIRST. It is the source of truth for HOW to poll this Supabase and, critically, how to handle duplicate / multiple candidate evals for a (model, benchmark). This skill depends on it for four things:
- Connect + query —
crud-otagent-supabase§0 (local Mac,otagentenv,DC_AGENT_SECRET_ENV, service-role key for reads) and §Schema (sandbox_jobs= one row per model×benchmark eval;model_id/benchmark_id/metrics/stats/job_status/hf_traces_link). PAGINATE (>1000 rows).get_metricshape-robust helper (§GOTCHA 1) —metricsis list-OR-dict; NEVER index it directly. Also pullsaccuracy_stderrfor the SE subscript.- Duplicate/sibling pulls (§GOTCHA 2) — the SAME model can have (a) multiple
sandbox_jobsrows per benchmark [aPending/Startedrow AND aFinishedrow, or reruns], and (b) multiplemodelsrows [trainer auto-push + a manual-<step>-<size>row, or a duplicate]. So query models byilikeon a name stub, not exact match, and UNIONsandbox_jobsacross all siblingmodel_ids. And benchmark FAMILIES (§GOTCHA 3) resolve viaduplicate_of.- Which candidate eval to use when there are several (§GOTCHA 2 rule 1 — the selection rule this skill lives or dies by): keep only
Finishedrows with a non-null accuracy (get_metric); among ≥2 COMPLETE entries with IDENTICAL evaluation settings, AVERAGE them — do NOT pick max, do NOT pick first. Entries with DIFFERENT settings (a differentn_rep_evalor harness) are NOT "identical settings" → do not average across them; keep the canonical one (the terminus-2, n=3 ID-eval setting the paper uses).crud-otagent-supabase'sget_model_scores()recipe implements exactly this union+average — mirror it.
Input = a list of model name stubs. Either passed directly, or derived from an experiment dir:
read its tracker (~/Documents/experiments/*/<name>/*tracker*.md / DESIGN.md / the HF-upload log)
for the model HF names/stubs that ablation produced (laion/…, DCAgent*/…, bare run-names).
import numpy as np
ID = {"swebench-verified-random-100-folders":"SWE-Bench-100",
"dev_set_v2":"OT-TBLite", "terminal_bench_2":"Terminal-Bench-2.0"}
bm = {b["id"]: b for b in c.table("benchmarks").select("id,name,duplicate_of").execute().data}
name2canon = {} # benchmark name -> canonical ID-set name (via duplicate_of)
for b in bm.values():
canon = b; seen=set()
while canon.get("duplicate_of") and canon["duplicate_of"] in bm and canon["id"] not in seen:
seen.add(canon["id"]); canon = bm[canon["duplicate_of"]]
if canon["name"] in ID: name2canon[b["name"]] = canon["name"]
if b["name"] in ID: name2canon[b["name"]] = b["name"]
def id_scores(stub):
"""-> {canon_bench: {'acc':float,'se':float|None,'trace':url|None}} averaging Finished repeats."""
mods = c.table("models").select("id,name").ilike("name", f"%{stub}%").execute().data # sibling rows
perb = {} # canon bench -> list of (acc, se, trace)
for m in mods:
for j in c.table("sandbox_jobs").select("benchmark_id,metrics,job_status,hf_traces_link") \
.eq("model_id", m["id"]).execute().data:
canon = name2canon.get(bm.get(j["benchmark_id"],{}).get("name"))
if canon is None: continue # not one of the 3 ID benchmarks
acc = get_metric(j["metrics"])
if j["job_status"] != "Finished" or acc is None: continue # real score only
se = get_metric(j["metrics"], "accuracy_stderr")
perb.setdefault(canon, []).append((acc, se, j.get("hf_traces_link")))
out = {}
for canon, entries in perb.items(): # AVERAGE identical-setting complete repeats
accs=[e[0] for e in entries]
out[canon] = {"acc": sum(accs)/len(accs),
"se": next((e[1] for e in entries if e[1] is not None), None),
"trace": next((e[2] for e in entries if e[2]), None)} # first non-null trace link
return out, mods
scores = {stub: id_scores(stub) for stub in MODEL_STUBS}In practice a (model, benchmark) often has several Finished rows that are not identical-setting,
so the "average identical repeats" branch does NOT apply — you must pick the canonical clean
measurement (per crud-otagent-supabase §GOTCHA 2 rule 1's "different settings → keep the canonical
one"). Detect and EXCLUDE the non-canonical ones (validated grid-exact on the RL ablation, 2026-07-09):
stats.evals.*.exception_stats.SummarizationTimeoutError scored lower because of the summarization
bug, not the model. Drop it in favor of the post-fix clean run.dev_set_v2 1.0–1.7% when the clean grid value is ~12%). Drop.dev_set_v2 20.5%@2026-06-29 vs 9.8%@2026-07-08). Do NOT average across generations; keep the
study's canonical measurement (the earliest clean post-fix run, matching the experiment's
id_eval_grid.md / ABLATION_DEFINITIONS.md). Averaging here would mix generations and desync from
the grid.timeout_multiplier=2.0, n=3).Cross-check the result against the experiment's own grid (id_eval_grid.md /
COMPARISON_*.md) — every ranked cell should reproduce it exactly; a mismatch means you picked a
non-canonical run. If the clean/canonical value the grid cites is not present in sandbox_jobs
(only superseded pre-fix rows exist), treat that benchmark as MISSING for §2 (flag it) rather than
substituting a deflated row.
A model is ID-valid only if it has a Finished score on all three ID benchmarks. Report (do
NOT silently drop) any input model missing ≥1 — the normalization population must be the models that
actually have the benchmark (partial models distort mean_b/std_b). Decide explicitly: rank only
the fully-ID-complete models (default), and list the incomplete ones separately with their gaps.
complete = {s:(sc,_m) for s,(sc,_m) in scores.items() if all(b in sc for b in ID)}
acc = {b: {s: complete[s][0][b]["acc"] for s in complete} for b in ID} # per-benchmark accs
z = {}
for b in ID:
vals = np.array(list(acc[b].values()), float)
mu, sd = vals.mean(), vals.std(ddof=0) # population std
z[b] = {s: (acc[b][s]-mu)/sd if sd>0 else 0.0 for s in acc[b]}
norm = {s: float(np.mean([z[b][s] for b in ID])) for s in complete}
raw = {s: float(np.mean([acc[b][s] for b in ID])) for s in complete}
ranking = sorted(complete, key=lambda s: norm[s], reverse=True)Columns (match the paper's layout): Rank · Model · SWE-Bench-100 (%) · OT-TBLite (%) ·
Terminal-Bench-2.0 (%) · Raw avg (%) · Normalized (z) · Trace links. Per-benchmark cell = raw
accuracy % (append ±SE from accuracy_stderr when present). The Trace links column carries the
per-benchmark hf_traces_link URLs (swe / v2 / tb2) — the same field the leaderboard uses; a missing
link → note "—". Sort by Normalized desc; number the ranks.
<experiment>/id_eval_ranking.md. Also print a one-line summary (N models ranked, N flagged
incomplete).normalized to 2 decimals with sign (e.g. +0.49, −0.57) like the paper; raw % to 2 dp.crud-otagent-supabase
§hf_traces_link — a different, write task.)N or by raw range.models-aware (ilike)duplicate_of) per crud-otagent-supabase §GOTCHA 2/3. Don't pick max/first.dev_set_v2; SWE-Bench-100=-random-100-folders (NOT full 500);
tb2=terminal_bench_2. Getting SWE wrong silently swaps an OOD benchmark into the ID ranking.crud-otagent-supabase — the schema, get_metric, sibling/family resolution, hf_traces_link,
the ID/OOD master list. This skill is a read-only consumer of it.otagent-paper/02_arXiv/otagent.tex — the normalization source of truth (§Pipeline, App.
task-gen full tables). Re-read if the method changes.© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/analyze-id-eval-ranking of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Analyze Id Eval Ranking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Analyze Id Eval Ranking this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Paper OrchestraAr9av/PaperOrchestra | 677 | 1 repos | ~3.5k | Automated safety check: Pass | Custom licence | |
| Section Writing AgentAr9av/PaperOrchestra | 677 | 1 repos | ~3.3k | Automated safety check: Pass | Custom licence | |
| SearchMuuuun/luxas | 1.2k | — | ~2.3k | Automated safety check: Warn | MIT | |
| Oleafly Pre SubmissionOleafly/Oleafly | 209 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Anmath Submissionfranklee16/academic-research-skills | 223 | 1 repos | ~1.1k | Automated safety check: Pass | None |
Ar9av/PaperOrchestra
Orchestrate the full PaperOrchestra (Song et al., 2026, arXiv:2604.05018) five-agent pipeline to turn unstructured research materials (idea, experimental log, LaTeX template, conference guidelines…
Ar9av/PaperOrchestra
Step 4 of the PaperOrchestra pipeline (arXiv:2604.05018). An agent skill from Ar9av/PaperOrchestra.
Muuuun/luxas
Unified academic paper search, citation chains, paper download (arXiv LaTeX/PDF, Sci-Hub), figure extraction from papers, LaTeX source reading, BibTeX fetching, web search, and browser automation…
Oleafly/Oleafly
Run a pre-flight pass over the project before uploading to arXiv or a venue, and write a pass or fail checklist.
franklee16/academic-research-skills
A skill your agent uses when running the final pre-submission preflight for an Annals of Mathematics manuscript — AMS-LaTeX compile, theorem environments, abstract and MSC, references, arXiv…
wentorai/research-plugins
Download and parse LaTeX source files from arXiv preprints. An agent skill from wentorai/research-plugins.
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
open-thoughts/OpenThoughts-Agent
Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format.
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…. Analyze Id Eval Ranking is an agent skill from open-thoughts/OpenThoughts-Agent.0=tb2), HF links to each eval's trace dataset, and a NORMALIZED column = average per-benchmark z-score, ranked.
Analyze Id Eval Ranking fits situations like: asked to rank models / ablation arms by their ID evals the way the paper does; tasks that involve LaTeX; tasks that involve LLM evaluation.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a claude-code`. Or copy the skill folder (.agents/skills/analyze-id-eval-ranking in open-thoughts/OpenThoughts-Agent) into .claude/skills/analyze-id-eval-ranking in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a codex`. Or copy the skill folder (.agents/skills/analyze-id-eval-ranking in open-thoughts/OpenThoughts-Agent) into .agents/skills/analyze-id-eval-ranking in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-id-eval-ranking, .gemini/skills/analyze-id-eval-ranking, .github/skills/analyze-id-eval-ranking and .opencode/skills/analyze-id-eval-ranking in your project.
SKILL.md names no scripts, command-line tools or credentials: Analyze Id Eval Ranking is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Analyze Id Eval Ranking is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Analyze Id Eval Ranking: Paper Orchestra (Ar9av/PaperOrchestra, 677 stars), Section Writing Agent (Ar9av/PaperOrchestra, 677 stars), Search (Muuuun/luxas, 1.2k stars) and Oleafly Pre Submission (Oleafly/Oleafly, 209 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.