Agent skill

Analyze Id Eval Ranking

by open-thoughts in open-thoughts/OpenThoughts-Agent

Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

Apache-2.0Auto-check passedResearch & Science

Install Analyze Id Eval Ranking

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-id-eval-ranking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyze-id-eval-ranking .claude/skills/analyze-id-eval-ranking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyze-id-eval-ranking
GitHub stars
301
Token cost
~3.1k tokens
SKILL.md length
1,113 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

  • Works in 5 steps: Connect (read-only) + the model list → Pull each model's 3 ID scores (sibling-… → Validity gate — flag models missing any… → …
  • Asked to rank models / ablation arms by their ID evals the way the paper does
  • SKILL.md covers The three ID benchmarks (and…, The normalization (must match…, 0. Connect (read-only) + the… and 1. Pull each model's 3 ID…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Analyze Id Eval Ranking is an agent skill from open-thoughts/OpenThoughts-Agent. Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100, OT-TBLite=devsetv2, Terminal-Bench-2.0=tb2), HF links to each eval's trace dataset, and a NORMALIZED column = average per-benchmark z-score, ranked. Normalization matches the OpenThoughts-Agent paper (otagent-paper/02arXiv/otagent.tex §Pipeline): per-benchmark z over the candidate set, averaged. Read-only. Use when asked to rank models…

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering LaTeX, LLM evaluation and Database schema design. It works with Supabase and arXiv. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Asked to rank models / ablation arms by their ID evals the way the paper does
  • Tasks that involve LaTeX
  • Tasks that involve LLM evaluation

Example prompts

  • “/analyze-id-eval-ranking”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Connect (read-only) + the model list
  2. Pull each model's 3 ID scores (sibling- + family-aware, averaged)
  3. Validity gate — flag models missing any ID benchmark
  4. Normalize + rank
  5. Emit the table

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyze Id Eval Ranking loads about 3.1k tokens when it runs. Until then it costs about 151 tokens; SKILL.md has 1,113 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~151
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 1,113 words, ~3,062 tokens.

Download SKILL.mdSave it as .claude/skills/analyze-id-eval-ranking/SKILL.md (or your agent's skills folder).
name
analyze-id-eval-ranking
description
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100, OT-TBLite=dev_set_v2, Terminal-Bench-2.0=tb2), HF links to each eval's trace dataset, and a NORMALIZED column = average per-benchmark z-score, ranked. Normalization matches the OpenThoughts-Agent paper (otagent-paper/02_arXiv/otagent.tex §Pipeline): per-benchmark z over the candidate set, averaged. Read-only. Use when asked to rank models / ablation arms by their ID evals the way the paper does.

analyze-id-eval-ranking

Produce the paper's ID ranking table for an arbitrary set of models: raw scores on the three in-distribution agentic benchmarks + HF trace links + a normalized average z-score column, sorted by the normalized score. This reproduces the ranking method in otagent-paper/02_arXiv/otagent.tex (§Pipeline / App. task-gen tables). Read-only — it never writes Supabase.

The three ID benchmarks (and their paper display names)

paper nameSupabase benchmarks.nametask count N (for SE)
SWE-Bench Verified (100)swebench-verified-random-100-folders100
OT-TBLitedev_set_v2 (partial-credit)—
Terminal-Bench 2.0terminal_bench_289

⚠ Mapping traps: OT-TBLite IS dev_set_v2 (not a separate benchmark). SWE-Bench-100 is the -random-100-folders subset, NOT full swebench-verified (500, which is OOD). terminal_bench_2 runs at timeout_multiplier 2.0 → resolve its family + dev_set_v2's family via duplicate_of (see crud-otagent-supabase §GOTCHA 2/3). dev_set_v2 is partial-credit → its raw % still enters the mean and its z-score, but it has no clean binomial N.

The normalization (must match the paper — otagent.tex §231)

"We compute the z-score of every candidate strategy's accuracy across the stage's full candidate set (subtracting the per-benchmark mean and dividing by its standard deviation), then average the three resulting per-benchmark z-scores."

So, with the candidate set = the input model list (this is the population for mean/std — NOT a global population):

  1. For each benchmark b, over all candidate models with a score on b: mean_b, std_b.
  2. z[m,b] = (acc[m,b] − mean_b) / std_b.
  3. normalized[m] = mean(z[m,b] over the 3 benchmarks the model has).
  4. Rank by normalized descending.

std uses population std (ddof=0, numpy.std default) — the candidate set IS the full population being compared. (Document this if you switch to sample std; it changes the magnitudes, not the ordering, when all models have all 3 benchmarks.) Equal per-benchmark weight is the whole point — don't weight by N.

0. Connect (read-only) + the model list

PREREQUISITE — read .agents/skills/crud-otagent-supabase/SKILL.md FIRST. It is the source of truth for HOW to poll this Supabase and, critically, how to handle duplicate / multiple candidate evals for a (model, benchmark). This skill depends on it for four things:

  • Connect + query — crud-otagent-supabase §0 (local Mac, otagent env, DC_AGENT_SECRET_ENV, service-role key for reads) and §Schema (sandbox_jobs = one row per model×benchmark eval; model_id/benchmark_id/metrics/stats/job_status/hf_traces_link). PAGINATE (>1000 rows).
  • get_metric shape-robust helper (§GOTCHA 1) — metrics is list-OR-dict; NEVER index it directly. Also pulls accuracy_stderr for the SE subscript.
  • Duplicate/sibling pulls (§GOTCHA 2) — the SAME model can have (a) multiple sandbox_jobs rows per benchmark [a Pending/Started row AND a Finished row, or reruns], and (b) multiple models rows [trainer auto-push + a manual -<step>-<size> row, or a duplicate]. So query models by ilike on a name stub, not exact match, and UNION sandbox_jobs across all sibling model_ids. And benchmark FAMILIES (§GOTCHA 3) resolve via duplicate_of.
  • Which candidate eval to use when there are several (§GOTCHA 2 rule 1 — the selection rule this skill lives or dies by): keep only Finished rows with a non-null accuracy (get_metric); among ≥2 COMPLETE entries with IDENTICAL evaluation settings, AVERAGE them — do NOT pick max, do NOT pick first. Entries with DIFFERENT settings (a different n_rep_eval or harness) are NOT "identical settings" → do not average across them; keep the canonical one (the terminus-2, n=3 ID-eval setting the paper uses). crud-otagent-supabase's get_model_scores() recipe implements exactly this union+average — mirror it.

Input = a list of model name stubs. Either passed directly, or derived from an experiment dir: read its tracker (~/Documents/experiments/*/<name>/*tracker*.md / DESIGN.md / the HF-upload log) for the model HF names/stubs that ablation produced (laion/…, DCAgent*/…, bare run-names).

1. Pull each model's 3 ID scores (sibling- + family-aware, averaged)

python
import numpy as np
ID = {"swebench-verified-random-100-folders":"SWE-Bench-100",
      "dev_set_v2":"OT-TBLite", "terminal_bench_2":"Terminal-Bench-2.0"}

bm  = {b["id"]: b for b in c.table("benchmarks").select("id,name,duplicate_of").execute().data}
name2canon = {}                                  # benchmark name -> canonical ID-set name (via duplicate_of)
for b in bm.values():
    canon = b; seen=set()
    while canon.get("duplicate_of") and canon["duplicate_of"] in bm and canon["id"] not in seen:
        seen.add(canon["id"]); canon = bm[canon["duplicate_of"]]
    if canon["name"] in ID: name2canon[b["name"]] = canon["name"]
    if b["name"] in ID:     name2canon[b["name"]] = b["name"]

def id_scores(stub):
    """-> {canon_bench: {'acc':float,'se':float|None,'trace':url|None}} averaging Finished repeats."""
    mods = c.table("models").select("id,name").ilike("name", f"%{stub}%").execute().data   # sibling rows
    perb = {}                                          # canon bench -> list of (acc, se, trace)
    for m in mods:
        for j in c.table("sandbox_jobs").select("benchmark_id,metrics,job_status,hf_traces_link") \
                   .eq("model_id", m["id"]).execute().data:
            canon = name2canon.get(bm.get(j["benchmark_id"],{}).get("name"))
            if canon is None: continue                  # not one of the 3 ID benchmarks
            acc = get_metric(j["metrics"])
            if j["job_status"] != "Finished" or acc is None: continue   # real score only
            se  = get_metric(j["metrics"], "accuracy_stderr")
            perb.setdefault(canon, []).append((acc, se, j.get("hf_traces_link")))
    out = {}
    for canon, entries in perb.items():                # AVERAGE identical-setting complete repeats
        accs=[e[0] for e in entries]
        out[canon] = {"acc": sum(accs)/len(accs),
                      "se":  next((e[1] for e in entries if e[1] is not None), None),
                      "trace": next((e[2] for e in entries if e[2]), None)}   # first non-null trace link
    return out, mods

scores = {stub: id_scores(stub) for stub in MODEL_STUBS}
Show full SKILL.md (547 more words)Show less
1a. Selecting the canonical eval when repeats are NOT identical-setting (load-bearing)

In practice a (model, benchmark) often has several Finished rows that are not identical-setting, so the "average identical repeats" branch does NOT apply — you must pick the canonical clean measurement (per crud-otagent-supabase §GOTCHA 2 rule 1's "different settings → keep the canonical one"). Detect and EXCLUDE the non-canonical ones (validated grid-exact on the RL ablation, 2026-07-09):

  • Summarization-buggy (deflated) runs — a run with non-trivial stats.evals.*.exception_stats.SummarizationTimeoutError scored lower because of the summarization bug, not the model. Drop it in favor of the post-fix clean run.
  • Degenerate broken-serving-batch runs — an implausibly low value from all-zero-reward batches (e.g. dev_set_v2 1.0–1.7% when the clean grid value is ~12%). Drop.
  • Drifted eval generations — the same clean setting re-run weeks apart can differ materially (e.g. dev_set_v2 20.5%@2026-06-29 vs 9.8%@2026-07-08). Do NOT average across generations; keep the study's canonical measurement (the earliest clean post-fix run, matching the experiment's id_eval_grid.md / ABLATION_DEFINITIONS.md). Averaging here would mix generations and desync from the grid.
  • Always prefer the canonical harness setting (terminus-2, timeout_multiplier=2.0, n=3).

Cross-check the result against the experiment's own grid (id_eval_grid.md / COMPARISON_*.md) — every ranked cell should reproduce it exactly; a mismatch means you picked a non-canonical run. If the clean/canonical value the grid cites is not present in sandbox_jobs (only superseded pre-fix rows exist), treat that benchmark as MISSING for §2 (flag it) rather than substituting a deflated row.

2. Validity gate — flag models missing any ID benchmark

A model is ID-valid only if it has a Finished score on all three ID benchmarks. Report (do NOT silently drop) any input model missing ≥1 — the normalization population must be the models that actually have the benchmark (partial models distort mean_b/std_b). Decide explicitly: rank only the fully-ID-complete models (default), and list the incomplete ones separately with their gaps.

3. Normalize + rank

python
complete = {s:(sc,_m) for s,(sc,_m) in scores.items() if all(b in sc for b in ID)}
acc = {b: {s: complete[s][0][b]["acc"] for s in complete} for b in ID}      # per-benchmark accs
z   = {}
for b in ID:
    vals = np.array(list(acc[b].values()), float)
    mu, sd = vals.mean(), vals.std(ddof=0)                                  # population std
    z[b] = {s: (acc[b][s]-mu)/sd if sd>0 else 0.0 for s in acc[b]}
norm = {s: float(np.mean([z[b][s] for b in ID])) for s in complete}
raw  = {s: float(np.mean([acc[b][s] for b in ID])) for s in complete}
ranking = sorted(complete, key=lambda s: norm[s], reverse=True)

4. Emit the table

Columns (match the paper's layout): Rank · Model · SWE-Bench-100 (%) · OT-TBLite (%) · Terminal-Bench-2.0 (%) · Raw avg (%) · Normalized (z) · Trace links. Per-benchmark cell = raw accuracy % (append ±SE from accuracy_stderr when present). The Trace links column carries the per-benchmark hf_traces_link URLs (swe / v2 / tb2) — the same field the leaderboard uses; a missing link → note "—". Sort by Normalized desc; number the ranks.

  • Emit markdown (and optionally a CSV alongside) to the experiment dir when run on one, e.g. <experiment>/id_eval_ranking.md. Also print a one-line summary (N models ranked, N flagged incomplete).
  • Report normalized to 2 decimals with sign (e.g. +0.49, −0.57) like the paper; raw % to 2 dp.

Guardrails

  • Read-only. Never write Supabase. (For trace-link repair, that's crud-otagent-supabase §hf_traces_link — a different, write task.)
  • Population = the candidate set (the input models), per-benchmark. Not a global mean. If the input list changes, the z-scores change — that is by design (relative ranking).
  • All three benchmarks equal weight — average the z-scores, never weight by N or by raw range.
  • Averaging repeats: average identical-setting Finished repeats; sibling-models-aware (ilike)
    • family-aware (duplicate_of) per crud-otagent-supabase §GOTCHA 2/3. Don't pick max/first.
  • Benchmark mapping: OT-TBLite=dev_set_v2; SWE-Bench-100=-random-100-folders (NOT full 500); tb2=terminal_bench_2. Getting SWE wrong silently swaps an OOD benchmark into the ID ranking.
  • Flag, don't drop, incomplete models — surface any input model lacking all 3 ID scores.
  • crud-otagent-supabase — the schema, get_metric, sibling/family resolution, hf_traces_link, the ID/OOD master list. This skill is a read-only consumer of it.
  • otagent-paper/02_arXiv/otagent.tex — the normalization source of truth (§Pipeline, App. task-gen full tables). Re-read if the method changes.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/analyze-id-eval-ranking of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Analyze Id Eval Ranking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyze Id Eval Ranking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyze Id Eval Ranking this skillopen-thoughts/OpenThoughts-Agent301—~3.1kAutomated safety check: PassApache-2.0
Paper OrchestraAr9av/PaperOrchestra6771 repos~3.5kAutomated safety check: PassCustom licence
Section Writing AgentAr9av/PaperOrchestra6771 repos~3.3kAutomated safety check: PassCustom licence
SearchMuuuun/luxas1.2k—~2.3kAutomated safety check: WarnMIT
Oleafly Pre SubmissionOleafly/Oleafly209—~2.9kAutomated safety check: PassMIT
Anmath Submissionfranklee16/academic-research-skills2231 repos~1.1kAutomated safety check: PassNone

Similar skills

  • Paper Orchestra

    Ar9av/PaperOrchestra

    Orchestrate the full PaperOrchestra (Song et al., 2026, arXiv:2604.05018) five-agent pipeline to turn unstructured research materials (idea, experimental log, LaTeX template, conference guidelines…

    677 GitHub starsUsed in 1 repo~3.5k tokens
    Research & ScienceAuto-check passed
  • Section Writing Agent

    Ar9av/PaperOrchestra

    Step 4 of the PaperOrchestra pipeline (arXiv:2604.05018). An agent skill from Ar9av/PaperOrchestra.

    677 GitHub starsUsed in 1 repo~3.3k tokens
    Research & ScienceAuto-check passed
  • Search

    Muuuun/luxas

    Unified academic paper search, citation chains, paper download (arXiv LaTeX/PDF, Sci-Hub), figure extraction from papers, LaTeX source reading, BibTeX fetching, web search, and browser automation…

    1.2k GitHub stars~2.3k tokensUpdated 1 mo ago
    Research & ScienceAuto-check: warnings
  • Oleafly Pre Submission

    Oleafly/Oleafly

    Run a pre-flight pass over the project before uploading to arXiv or a venue, and write a pass or fail checklist.

    209 GitHub stars~2.9k tokensUpdated today
    Research & ScienceAuto-check passed
  • Anmath Submission

    franklee16/academic-research-skills

    A skill your agent uses when running the final pre-submission preflight for an Annals of Mathematics manuscript — AMS-LaTeX compile, theorem environments, abstract and MSC, references, arXiv…

    223 GitHub starsUsed in 1 repo~1.1k tokens
    Research & ScienceAuto-check passed
  • Arxiv Latex Source

    wentorai/research-plugins

    Download and parse LaTeX source files from arXiv preprints. An agent skill from wentorai/research-plugins.

    298 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 11 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed
  • Commit

    open-thoughts/OpenThoughts-Agent

    Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format.

    301 GitHub stars~2.2k tokensUpdated 11 days ago
    Auto-check: notes

Works with

Questions about Analyze Id Eval Ranking

What does Analyze Id Eval Ranking do?

Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…. Analyze Id Eval Ranking is an agent skill from open-thoughts/OpenThoughts-Agent.0=tb2), HF links to each eval's trace dataset, and a NORMALIZED column = average per-benchmark z-score, ranked.

When should I use Analyze Id Eval Ranking?

Analyze Id Eval Ranking fits situations like: asked to rank models / ablation arms by their ID evals the way the paper does; tasks that involve LaTeX; tasks that involve LLM evaluation.

How do I install Analyze Id Eval Ranking in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a claude-code`. Or copy the skill folder (.agents/skills/analyze-id-eval-ranking in open-thoughts/OpenThoughts-Agent) into .claude/skills/analyze-id-eval-ranking in your project. Claude Code loads it when a task matches its description.

How do I install Analyze Id Eval Ranking in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a codex`. Or copy the skill folder (.agents/skills/analyze-id-eval-ranking in open-thoughts/OpenThoughts-Agent) into .agents/skills/analyze-id-eval-ranking in your project. Codex loads it when a task matches its description.

Can I use Analyze Id Eval Ranking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-id-eval-ranking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-id-eval-ranking, .gemini/skills/analyze-id-eval-ranking, .github/skills/analyze-id-eval-ranking and .opencode/skills/analyze-id-eval-ranking in your project.

What does Analyze Id Eval Ranking need to run?

SKILL.md names no scripts, command-line tools or credentials: Analyze Id Eval Ranking is instructions for the agent only. Our summary lists: Python 3.

Does Analyze Id Eval Ranking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Analyze Id Eval Ranking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyze Id Eval Ranking use?

Analyze Id Eval Ranking is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyze Id Eval Ranking use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyze Id Eval Ranking?

Skills that share tags, products or a category with Analyze Id Eval Ranking: Paper Orchestra (Ar9av/PaperOrchestra, 677 stars), Section Writing Agent (Ar9av/PaperOrchestra, 677 stars), Search (Muuuun/luxas, 1.2k stars) and Oleafly Pre Submission (Oleafly/Oleafly, 209 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyze Id Eval Ranking?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.