Agent skill

Recent Conjecture Evaluations

by morluto in morluto/jacobian

Evaluate Jacobian reliability using recently resolved conjectures as held-out probes.

MITAuto-check passedResearch & Science

Install Recent Conjecture Evaluations

skills CLI
$ npx skills add morluto/jacobian --skill recent-conjecture-evaluations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install morluto/jacobian recent-conjecture-evaluations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/morluto/jacobian.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/recent-conjecture-evaluations .claude/skills/recent-conjecture-evaluations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
recent-conjecture-evaluations
GitHub stars
211
Token cost
~816 tokens
SKILL.md length
375 words
Files
8 (incl. scripts, references)
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Evaluate Jacobian reliability using recently resolved conjectures as held-out probes.

  • Tasks that involve Math and symbolic computation
  • SKILL.md covers Choose the requested mode, Preserve the evidence boundary and Completion and stalls
  • Runs Python scripts from its folder; calls python

What it does

Recent Conjecture Evaluations is an agent skill from morluto/jacobian. Evaluate Jacobian reliability using recently resolved conjectures as held-out probes.

Its SKILL.md is about 820 tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/action-policy.md` and `references/evaluation-and-scoring.md`).

It sits in Research & Science, covering Math and symbolic computation. It works with Model Context Protocol. The repository describes itself as: Composable mathematics tools for agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Math and symbolic computation

Example prompts

  • “/recent-conjecture-evaluations”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9dc2aaf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Recent Conjecture Evaluations loads about 816 tokens when it runs, and up to ~4.1k if it reads all its reference files. Until then it costs about 29 tokens; SKILL.md has 375 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~816
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from morluto/jacobian at commit 9dc2aaf, republished under its MIT licence (© morluto). 375 words, ~816 tokens.

Download SKILL.mdSave it as .claude/skills/recent-conjecture-evaluations/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
recent-conjecture-evaluations
description
Evaluate Jacobian reliability using recently resolved conjectures as held-out probes.

Recent Conjecture Evaluations

Use recently resolved conjectures to probe Jacobian reliability. The outcome is an evidence-backed diagnosis, not a collection of solved examples. This skill owns source selection, deterministic replay, optional model comparisons, and attribution; Harbor packaging is a separate workflow.

Choose the requested mode

  • Probe: read probe workflow for one source cycle, including source gating, independent oracle, current-main replay, and reporting.
  • Review: inspect the completed cycle against source gating, the oracle and frozen inputs, and the applicable scoring rules. Use report fields to identify missing evidence. Reproduce disputed deterministic claims where useful; do not repeat model arms.
  • Coordinate: search saved reports, reservations, and issue/PR ownership for the source and root mechanism. Use the inventory helper shown below and the ownership rules in action policy. Do not begin a new probe merely to answer a coordination question.
sh
python .agents/skills/recent-conjecture-evaluations/scripts/search_inventory.py \
  "source or root-cause phrase" outputs benchmarks/results

Preserve the evidence boundary

Bind prompts, payloads, and oracles to the exact intended input. Reconstruct gold independently of evaluated model arms and distinguish source-supplied replay from an independent oracle. Record precise source status and dates. Computation, verification, imported theorems, and unproved claims establish different things. Timeouts, unavailable operations, and missing witnesses are non-conclusions. Never attribute a model's fallback or transcription error to Jacobian.

Audit new probes deterministically on current main before paid model calls. Run comparisons only when they resolve uncertainty that direct evidence cannot, with user-authorized cost boundaries and frozen control/treatment conditions. Do not weaken verification or create a benchmark-specific operation to make a probe pass. Before external action, apply the action policy within existing user authorization. This evaluation workflow produces at most a localized draft PR; merging is a separate user-authorized task.

Show full SKILL.md (104 more words)Show less

Completion and stalls

Complete one requested cycle by default, including rejected and deterministic-only outcomes. An oracle disagreement stops dependent model evaluation while allowing bounded diagnosis. A known unavailable operation skips dependent work, not the attribution and report. Preserve checkpoints and raw evidence at phase boundaries; inspect a phase without a checkpoint for 30 minutes.

Additional cycles or model calls must fit the operator's requested scope and cost limits. Retry a cancelled call only with new diagnostic evidence. Stop collecting examples for an owned root cause once the agreed independence threshold is met. Report unresolved evidence and a distinct next direction without silently starting it.

© morluto, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in .agents/skills/recent-conjecture-evaluations of morluto/jacobian.

  • SKILL.md
  • agents/openai.yaml
  • references/action-policy.md
  • references/evaluation-and-scoring.md
  • references/probe.md
  • references/report-schema.md
  • references/source-gating.md
  • scripts/search_inventory.py

Open the folder on GitHubat commit 9dc2aaf

Compare with similar skills

Recent Conjecture Evaluations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Recent Conjecture Evaluations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Recent Conjecture Evaluations this skillmorluto/jacobian211—~816Automated safety check: PassMIT
Research RefinezjYao36/Auto-Research-Refine1286 repos~6.9kAutomated safety check: NotesNone
Read GitHubAgentTeam-TaichuAI/ScienceClaw6712 repos~638Automated safety check: PassNone
Proof Run Orchestratorwanshuiyin/Auto-claude-code-research-in-sleep17k1 repos~4.7kAutomated safety check: PassMIT
FirstdataMLT-OSS/FirstData183—~3.1kAutomated safety check: PassMIT
Fin Generate Ideacsmar432/finai-research109—~2.5kAutomated safety check: PassMIT

Similar skills

  • Research Refine

    zjYao36/Auto-Research-Refine

    Turns a vague research direction into a focused, problem-anchored method plan through up to five review rounds with a second model.

    128 GitHub starsUsed in 6 repos~6.9k tokens
    Research & ScienceAuto-check: notes
  • Read GitHub

    AgentTeam-TaichuAI/ScienceClaw

    Read and search GitHub repository documentation via gitmcp.io MCP service.

    671 GitHub starsUsed in 2 repos~638 tokens
    Research & ScienceAuto-check passed
  • Proof Run Orchestrator

    wanshuiyin/Auto-claude-code-research-in-sleep

    Runs a mathematical proof project as a stateful pipeline of run directories: a local attempt first, then a manual GPT Pro handoff package, with an optional DeepSeek audit.

    17k GitHub starsUsed in 1 repo~4.7k tokens
    Research & ScienceAuto-check passed
  • Firstdata

    MLT-OSS/FirstData

    Find official portals, APIs, and download paths for authoritative primary data sources (governments, international organizations, research institutions, etc.).

    183 GitHub stars~3.1k tokensUpdated 6 days ago
    Research & ScienceAuto-check passed
  • Fin Generate Idea

    csmar432/finai-research

    针对经济金融研究方向的创意生成与评估。生成8-12个可发表的研究idea,过滤后在数据可行的情况下进行小规模实证验证,输出排序后的研究想法报告。

    109 GitHub stars~2.5k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Math Proof Solo

    anthropics/claude-plugins-official

    Official

    Solves one hard mathematics problem in a single session without subagents, keeping settled steps in a notes file and ending with a self-contained proof.md.

    38k GitHub stars~1.5k tokensUpdated today
    Research & ScienceAuto-check passed

More from morluto/jacobian

All 11 skills in this repo
  • Harbor Benchmarks

    morluto/jacobian

    Author, package, validate, or run mathematical evaluations as Jacobian Harbor datasets.

    211 GitHub stars~690 tokensUpdated 4 days ago
    Auto-check passed
  • Design or audit a Jacobian operation’s mathematical contract, boundedness, exact results, and composition.

    211 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check passed
  • Extract reusable Jacobian capabilities from a mathematical solution corpus, rather than one agent trajectory.

    211 GitHub stars~1.9k tokensUpdated 4 days ago
    Auto-check passed
  • Investigate MCP tool availability, discovery, and selection, including controlled adoption evaluations.

    211 GitHub stars~1k tokensUpdated 4 days ago
    Auto-check passed
  • Review mathematical agent trajectories for evidence-backed Jacobian improvements; do not resume solving.

    211 GitHub stars~848 tokensUpdated 4 days ago
    Auto-check passed
  • Verifier Evaluations

    morluto/jacobian

    Design, audit, or repair mathematical benchmark verifiers, submission contracts, and scoring.

    211 GitHub stars~661 tokensUpdated 4 days ago
    Auto-check passed

Questions about Recent Conjecture Evaluations

What does Recent Conjecture Evaluations do?

Evaluate Jacobian reliability using recently resolved conjectures as held-out probes. Recent Conjecture Evaluations is an agent skill from morluto/jacobian. Evaluate Jacobian reliability using recently resolved conjectures as held-out probes.

When should I use Recent Conjecture Evaluations?

Recent Conjecture Evaluations fits situations like: tasks that involve Math and symbolic computation.

How do I install Recent Conjecture Evaluations in Claude Code?

Run `npx skills add morluto/jacobian --skill recent-conjecture-evaluations -a claude-code`. Or copy the skill folder (.agents/skills/recent-conjecture-evaluations in morluto/jacobian) into .claude/skills/recent-conjecture-evaluations in your project. Claude Code loads it when a task matches its description.

How do I install Recent Conjecture Evaluations in Codex?

Run `npx skills add morluto/jacobian --skill recent-conjecture-evaluations -a codex`. Or copy the skill folder (.agents/skills/recent-conjecture-evaluations in morluto/jacobian) into .agents/skills/recent-conjecture-evaluations in your project. Codex loads it when a task matches its description.

Can I use Recent Conjecture Evaluations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add morluto/jacobian --skill recent-conjecture-evaluations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/recent-conjecture-evaluations, .gemini/skills/recent-conjecture-evaluations, .github/skills/recent-conjecture-evaluations and .opencode/skills/recent-conjecture-evaluations in your project.

What does Recent Conjecture Evaluations need to run?

Going by SKILL.md and its folder, Recent Conjecture Evaluations needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Recent Conjecture Evaluations access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Recent Conjecture Evaluations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Recent Conjecture Evaluations use?

Recent Conjecture Evaluations is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Recent Conjecture Evaluations use?

About 816 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.

What are the alternatives to Recent Conjecture Evaluations?

Skills that share tags, products or a category with Recent Conjecture Evaluations: Research Refine (zjYao36/Auto-Research-Refine, 128 stars), Read GitHub (AgentTeam-TaichuAI/ScienceClaw, 671 stars), Proof Run Orchestrator (wanshuiyin/Auto-claude-code-research-in-sleep, 17k stars) and Firstdata (MLT-OSS/FirstData, 183 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Recent Conjecture Evaluations?

morluto (a GitHub user) maintains it in morluto/jacobian, which has 211 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 5, 2026.

Source: morluto/jacobian on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.