Agent skill

Scienceworld Threshold Evaluator

by zjunlp in zjunlp/SkillNet

A skill your agent uses when the agent has just obtained a numerical measurement (temperature, weight, pH) and must compare it against a predefined threshold to determine a binary outcome.

MITAuto-check passed

Install Scienceworld Threshold Evaluator

skills CLI
$ npx skills add zjunlp/SkillNet --skill scienceworld-threshold-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zjunlp/SkillNet scienceworld-threshold-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zjunlp/SkillNet.git skills-src && mkdir -p .claude/skills && cp -r skills-src/experiments/src/skills/scienceworld/scienceworld-threshold-evaluator .claude/skills/scienceworld-threshold-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scienceworld-threshold-evaluator
GitHub stars
1.4k
Token cost
~796 tokens
SKILL.md length
317 words
Files
3 (incl. references)
Skills in repo
122
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the agent has just obtained a numerical measurement (temperature, weight, pH) and must compare it against a predefined threshold to determine a binary outcome.

  • Works in 4 steps: Extract the measurement -- Parse the… → Identify the threshold and condition --… → Evaluate the comparison -- Compare:… → …
  • The agent has just obtained a numerical measurement (temperature
  • SKILL.md covers Purpose, When to Use, Workflow and Examples, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Scienceworld Threshold Evaluator is an agent skill from zjunlp/SkillNet. Use when the agent has just obtained a numerical measurement (temperature, weight, pH) and must compare it against a predefined threshold to determine a binary outcome. This skill extracts the measured value, evaluates it against the threshold condition (above/below), and executes the corresponding branch action such as classification or placement.

Its SKILL.md is about 800 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/action_guide.md` and `references/usage_examples.md`).

The repository describes itself as: Create, Evaluate, and Connect AI Skills. The licence is MIT.

When your agent uses it

  • The agent has just obtained a numerical measurement (temperature
  • PH) and must compare it against a predefined threshold to determine a binary outcome

Example prompts

  • “/scienceworld-threshold-evaluator”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Extract the measurement -- Parse the numerical value from the observation (e.g., "the thermometer measures a temperature of 56 degrees…
  2. Identify the threshold and condition -- From the task instruction, determine the threshold value and comparison operator (e.g., "above…
  3. Evaluate the comparison -- Compare: measured_value > threshold or measured_value < threshold.
  4. Execute the correct branch -- Perform the action specified for the satisfied condition.

What it can do on your machine

Read from SKILL.md and the folder at commit 3fcebf8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scienceworld Threshold Evaluator loads about 796 tokens when it runs, and up to ~2.1k if it reads all its reference files. Until then it costs about 96 tokens; SKILL.md has 317 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~796
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from zjunlp/SkillNet at commit 3fcebf8, republished under its MIT licence (© zjunlp). 317 words, ~796 tokens.

Download SKILL.mdSave it as .claude/skills/scienceworld-threshold-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
scienceworld-threshold-evaluator
description
Use when the agent has just obtained a numerical measurement (temperature, weight, pH) and must compare it against a predefined threshold to determine a binary outcome. This skill extracts the measured value, evaluates it against the threshold condition (above/below), and executes the corresponding branch action such as classification or placement.

Skill: scienceworld-threshold-evaluator

Purpose

Compare a measured numerical value against a predefined threshold to determine which of two conditional actions to execute. This is the decision-making step that immediately follows a measurement in ScienceWorld experimental workflows.

When to Use

  • Immediately after obtaining a numerical measurement (e.g., temperature reading from a thermometer)
  • When the task includes a conditional instruction like "if above X, do A; if below X, do B"
  • When classifying or sorting objects based on measured properties

Workflow

  1. Extract the measurement -- Parse the numerical value from the observation (e.g., "the thermometer measures a temperature of 56 degrees celsius" yields 56).
  2. Identify the threshold and condition -- From the task instruction, determine the threshold value and comparison operator (e.g., "above 50.0 degrees" means threshold=50.0, operator=">").
  3. Evaluate the comparison -- Compare: measured_value > threshold or measured_value < threshold.
  4. Execute the correct branch -- Perform the action specified for the satisfied condition.

Examples

Example 1: Temperature-based classification

Task: "Measure the temperature. If above 50.0 degrees, move to the orange box. If below 50.0 degrees, move to the blue box."

> use thermometer on unknown substance B
The thermometer measures a temperature of 56 degrees celsius.

Evaluation: 56 > 50.0 is TRUE, so execute the "above" branch.

> move unknown substance B to orange box
You move the unknown substance B to the orange box.
Example 2: Weight-based sorting

Task: "If the object weighs more than 200 grams, place in the red bin. Otherwise, place in the green bin."

> use scale on rock sample
The scale measures a weight of 145 grams.

Evaluation: 145 > 200 is FALSE, so execute the "otherwise" branch.

> move rock sample to green bin
You move the rock sample to the green bin.

Key Principles

  • Immediate execution -- Do not perform other actions between obtaining the measurement and evaluating the threshold.
  • Precision -- Use the exact numerical value from the observation; do not estimate or round.
  • Binary decision -- The outcome is strictly one of two paths. If the measurement equals the threshold, re-examine the instruction for boundary guidance ("above" typically means >, not >=).

Common Pitfalls

  • Incorrect branching -- Executing the action for the opposite condition (e.g., blue box when value is above threshold).
  • Premature evaluation -- Attempting to evaluate before the measurement is complete and valid.
  • Action confusion -- Targeting the wrong object in the post-evaluation action.

© zjunlp, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in experiments/src/skills/scienceworld/scienceworld-threshold-evaluator of zjunlp/SkillNet.

  • SKILL.md
  • references/action_guide.md
  • references/usage_examples.md

Open the folder on GitHubat commit 3fcebf8

Compare with similar skills

Scienceworld Threshold Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scienceworld Threshold Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scienceworld Threshold Evaluator this skillzjunlp/SkillNet1.4k—~796Automated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k13 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence
Agent Evaluation Reportingsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 13 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check passed
  • Agent Evaluation Reporting

    sickn33/agentic-awesome-skills

    A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Official

    Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.

    40k GitHub stars~6.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from zjunlp/SkillNet

All 122 skills in this repo
  • Skillnet

    zjunlp/SkillNet

    Search, download, create, evaluate, analyze and route reusable agent skills with SkillNet.

    1.4k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Performs an initial scan of the ALFWorld environment to identify all visible objects and receptacles.

    1.4k GitHub starsUsed in 1 repo~786 tokens
    Auto-check passed
  • This skill searches for a specific receptacle (e.g., garbage can, cabinet) by systematically exploring the environment, checking multiple locations until found.

    1.4k GitHub stars~709 tokensUpdated yesterday
    Auto-check passed
  • Prepares a household appliance (microwave, oven, toaster, fridge) for use by ensuring it is in the correct open/closed state.

    1.4k GitHub stars~611 tokensUpdated yesterday
    Auto-check passed
  • Operates a device or appliance (like a desklamp, microwave, or fridge) to interact with another object.

    1.4k GitHub stars~916 tokensUpdated yesterday
    Auto-check passed
  • Uses a heating appliance (microwave, stoveburner, oven) to apply heat to a specified object.

    1.4k GitHub stars~744 tokensUpdated yesterday
    Auto-check passed

Questions about Scienceworld Threshold Evaluator

What does Scienceworld Threshold Evaluator do?

A skill your agent uses when the agent has just obtained a numerical measurement (temperature, weight, pH) and must compare it against a predefined threshold to determine a binary outcome. Scienceworld Threshold Evaluator is an agent skill from zjunlp/SkillNet. Use when the agent has just obtained a numerical measurement (temperature, weight, pH) and must compare it against a predefined threshold to determine a binary outcome.

When should I use Scienceworld Threshold Evaluator?

Scienceworld Threshold Evaluator fits situations like: the agent has just obtained a numerical measurement (temperature; PH) and must compare it against a predefined threshold to determine a binary outcome.

How do I install Scienceworld Threshold Evaluator in Claude Code?

Run `npx skills add zjunlp/SkillNet --skill scienceworld-threshold-evaluator -a claude-code`. Or copy the skill folder (experiments/src/skills/scienceworld/scienceworld-threshold-evaluator in zjunlp/SkillNet) into .claude/skills/scienceworld-threshold-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Scienceworld Threshold Evaluator in Codex?

Run `npx skills add zjunlp/SkillNet --skill scienceworld-threshold-evaluator -a codex`. Or copy the skill folder (experiments/src/skills/scienceworld/scienceworld-threshold-evaluator in zjunlp/SkillNet) into .agents/skills/scienceworld-threshold-evaluator in your project. Codex loads it when a task matches its description.

Can I use Scienceworld Threshold Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zjunlp/SkillNet --skill scienceworld-threshold-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scienceworld-threshold-evaluator, .gemini/skills/scienceworld-threshold-evaluator, .github/skills/scienceworld-threshold-evaluator and .opencode/skills/scienceworld-threshold-evaluator in your project.

What does Scienceworld Threshold Evaluator need to run?

SKILL.md names no scripts, command-line tools or credentials: Scienceworld Threshold Evaluator is instructions for the agent only.

Does Scienceworld Threshold Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scienceworld Threshold Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scienceworld Threshold Evaluator use?

Scienceworld Threshold Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scienceworld Threshold Evaluator use?

About 796 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Scienceworld Threshold Evaluator?

Skills that share tags, products or a category with Scienceworld Threshold Evaluator: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Evaluators (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scienceworld Threshold Evaluator?

zjunlp (a GitHub organization) maintains it in zjunlp/SkillNet, which has 1,393 GitHub stars. The repository holds 122 skills in this directory. The repository was last updated on October 7, 2026.

Source: zjunlp/SkillNet on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.