DeepTutor CLI
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
Decision rubric for when an LM agent should write-and-run code (Program-of-Thought / code interpreter) versus reason in natural language: classify each step as deterministic- computable (emit +…
$ npx skills add agentsope/SkillAlchemy --skill agentsop-code-execution-decision -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-code-execution-decision --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-code-execution-decision .claude/skills/agentsop-code-execution-decision && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-code-execution-decision" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-code-execution-decision into .claude/skills/agentsop-code-execution-decision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-code-execution-decision", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-code-execution-decisionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-code-execution-decision -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-code-execution-decision --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-code-execution-decision .agents/skills/agentsop-code-execution-decision && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-code-execution-decision" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-code-execution-decision into .agents/skills/agentsop-code-execution-decision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-code-execution-decision", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-code-execution-decision -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-code-execution-decision --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-code-execution-decision .cursor/skills/agentsop-code-execution-decision && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-code-execution-decision" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-code-execution-decision into .cursor/skills/agentsop-code-execution-decision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-code-execution-decision", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-code-execution-decision--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-code-execution-decision -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-code-execution-decision --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-code-execution-decision .gemini/skills/agentsop-code-execution-decision && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-code-execution-decision" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-code-execution-decision into .gemini/skills/agentsop-code-execution-decision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-code-execution-decision", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-code-execution-decisionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-code-execution-decision -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-code-execution-decision .github/skills/agentsop-code-execution-decision && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-code-execution-decision" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-code-execution-decision into .github/skills/agentsop-code-execution-decision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-code-execution-decision", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-code-execution-decision -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-code-execution-decision --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-code-execution-decision .opencode/skills/agentsop-code-execution-decision && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-code-execution-decision" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-code-execution-decision into .opencode/skills/agentsop-code-execution-decision/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-code-execution-decision", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-code-execution-decisionDecision rubric for when an LM agent should write-and-run code (Program-of-Thought / code interpreter) versus reason in natural language: classify each step as deterministic- computable (emit +…
Agentsop Code Execution Decision is an agent skill from agentsope/SkillAlchemy. Decision rubric for when an LM agent should write-and-run code (Program-of-Thought / code interpreter) versus reason in natural language: classify each step as deterministic- computable (emit + execute code, feed the result back) vs judgment (stay in prose). Use when designing or debugging an agent step that does arithmetic/parsing/data transforms, when prose reasoning hallucinates a computation (under-coding), or when a sandbox round- trip is wasted on a judgment task (over-coding). Search keywords: code…
Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).
It sits in Education, covering Quizzes and assessments. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6ea799f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Code Execution Decision loads about 6.1k tokens when it runs, and up to ~7.3k if it reads all its reference files. Until then it costs about 169 tokens; SKILL.md has 2,628 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit 6ea799f, republished under its MIT licence (© agentsope). 2,628 words, ~6,100 tokens.
.claude/skills/agentsop-code-execution-decision/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.One-liner: LMs are unreliable calculators but reliable coders. When the answer needs determinism and precision — arithmetic, exact data manipulation, deterministic transforms — emit code and run it. When the answer needs judgment, taste, or open-ended synthesis, reason in natural language. The cost of getting this gate wrong is silent: prose arithmetic hallucinates a plausible-looking wrong number, and over-coding a judgment task burns a sandbox round-trip for nothing.
This is an enhancement overlay. DSPy already gives you dspy.ProgramOfThought (PoT) —
the mechanism for write-then-execute. What it does not give you is the decision rubric for
when to reach for it. That rubric is this skill. Cross-link the sibling
[[agentsop-output-format-by-model]] (which decides how code-shaped content should be serialized) and
[[agentsop-test-fix-loop]] (which closes the execute → error → retry loop).
Activate this skill before committing a step to a reasoning strategy whenever the task has a verifiable, deterministic core — or whenever you catch an agent doing arithmetic in prose.
| Trigger | Signal |
|---|---|
| Arithmetic / math | "compute the compound interest", "what's 17.5% of $4,392.18", multi-step word problems, unit conversions, date deltas |
| Precise data manipulation | "sort these 240 rows by the third column", "dedupe and count", "join these two lists on id", "parse this CSV and sum column B" |
| Deterministic transforms | regex extraction, string reformatting, base conversion, hashing, sorting, set operations |
| Symbolic / combinatorial | "how many distinct permutations", "solve this system of equations", calendar/scheduling math |
| You see a model doing math in prose | "Let me add: 1,204 + 8,991 + ... = 10,195" — almost always worth a code check |
| Choosing a DSPy module | deciding between ChainOfThought and ProgramOfThought for a signature [dspy.ai/learn/programming/modules/] |
Anti-triggers (do NOT reach for code execution):
LMs are unreliable calculators but reliable coders.
A language model predicts the next token, not the correct value. When you ask it to add
48,217 + 9,884 in prose, it emits the most plausible-looking digit sequence — which is
frequently wrong, and wrong in a way that looks right. The same model can write
48217 + 9884 as a Python expression flawlessly, because emitting the program is a
pattern-matching task it is genuinely good at, and the Python interpreter is a deterministic
oracle. This decoupling — model writes the recipe, interpreter computes the result — is the
entire thesis of Program-of-Thought (PoT) [arXiv 2211.12588] and PAL [arXiv 2211.10435].
┌────────────────────────────────────────────────────────┐
│ THE GATE: does this answer need determinism/precision? │
└────────────────────────────────────────────────────────┘
│ │
YES (computable) NO (judgment)
│ │
▼ ▼
┌─────────────────────┐ ┌─────────────────────┐
│ EMIT CODE │ │ REASON IN PROSE │
│ model writes recipe │ │ model is the engine │
│ interpreter = oracle│ │ no oracle exists │
└─────────────────────┘ └─────────────────────┘
│
▼
sandbox → run → feed result back into LM → LM narrates/uses itTwo failure modes the gate prevents:
| Failure | Mechanism | Symptom |
|---|---|---|
| Under-coding (reason when you should compute) | LM hallucinates a calculation it cannot reliably perform | A confident, wrong, plausible-looking number; off-by-one counts; arithmetic that "looks" right |
| Over-coding (compute when you should reason) | LM wraps a judgment task in a sandbox round-trip that adds no determinism | Wasted latency + cost; brittle code that encodes a subjective rubric as if it were a formula; print("the tone is friendly") |
Why the asymmetry matters. Under-coding fails silently — the wrong number propagates downstream and no exception fires. Over-coding fails loudly and cheaply — you notice the useless sandbox call. So the default lean, when genuinely uncertain and a verifiable core exists, is toward code. But "genuinely uncertain" is the operative phrase: most judgment tasks are not close calls.
The format corollary (from [[agentsop-output-format-by-model]]): once you decide to emit code, emit it as code in a fenced block / single string field — never nested inside JSON sub-structure. Code-in-JSON measurably degrades the code itself (Aider: 61%→20% on GPT-4 Turbo). The execution decision and the serialization decision are two separate gates; pass both.
A three-step gate. Run it per step, not per task — one task can have computable steps and judgment steps interleaved.
Ask: "Is there a single correct answer that a program could verify?"
A useful tiebreaker for the "trivial computable" gray zone: if the model would be embarrassed to get it wrong and you'd reach for a calculator yourself, emit code. If you'd do it in your head without a second thought, prose is fine.
print/return the result. No network, no filesystem unless the task is I/O.ProgramOfThought uses a Python interpreter (Deno/PythonInterpreter sandbox in recent versions); OpenAI Code Interpreter runs in a managed container; Anthropic code-execution tool runs in a sandboxed VM; LangChain PythonREPLTool runs in-process and is unsandboxed — treat as untrusted-input-hostile.ProgramOfThought defaults to max_iters ≈ 3). After the bound, fall back to prose reasoning or surface the failure — do not loop forever.Exit criterion: the step produces either (a) a code-derived value the LM has consumed, or (b) a prose judgment, with the gate decision recorded so a reviewer can audit why code was or wasn't used.
Seven operations. Each row is a reusable move.
| # | Op | Trigger | Action | Output | Evidence |
|---|---|---|---|---|---|
| 1 | Computable-vs-judgment gate | Any step about to be reasoned | Apply §3 Step 1: single verifiable answer? | Route to code or prose | PoT premise: decouple compute from reasoning [arXiv 2211.12588] |
| 2 | Decompose mixed steps | Step has both a number and a narrative | Split into computable sub-step (code) + judgment sub-step (prose) | Two routed sub-steps | Mirrors mixed-content split in [[agentsop-output-format-by-model]] §5 Case B |
| 3 | Sandbox choice | Decided to emit code | Pick interpreter by trust + capability: DSPy PoT (Python sandbox), OpenAI Code Interpreter (managed container), Anthropic code-exec (VM), LangChain PythonREPLTool (unsandboxed, in-process) | Chosen runtime | DSPy modules [dspy.ai/learn/programming/modules/]; LangChain PythonREPLTool docs |
| 4 | Result-back-into-LM | Code produced a value | Inject stdout/return value into the next LM turn so the model narrates/uses it | LM-consumed result | PoT design: code computes, LM contextualizes [arXiv 2211.12588] |
| 5 | Error retry (bounded) | Code raised a traceback | Feed error to LM, regenerate, re-run; cap at max_iters (~3) then fall back | Fixed code or graceful fallback | DSPy ProgramOfThought max_iters; see [[agentsop-test-fix-loop]] |
| 6 | Precision escalation | Prose answer involves multi-step arithmetic | Re-route the arithmetic to code even if prose started it | Code-verified number | "LMs are unreliable calculators" — PAL [arXiv 2211.10435] |
| 7 | Over-coding veto | About to sandbox a judgment task | Stop: no deterministic core → reasoning, not code | Prose reasoning, no sandbox call | §5 Case B; avoids wasted round-trip |
In DSPy terms, op #1 is exactly the choice between dspy.ChainOfThought (prose reasoning) and
dspy.ProgramOfThought (emit+run) for a signature — the dspy skill lists the modules but this
overlay supplies the when.
Trigger: A finance-summary agent step: "Given these 14 line items, compute the total,
the 8.25% tax, and the grand total." The agent is a dspy.ChainOfThought module emitting prose.
Constraints:
Decision steps:
ChainOfThought to ProgramOfThought. The model now emits
subtotal = sum([...]); tax = round(subtotal * 0.0825, 2); total = subtotal + tax and the
interpreter computes it exactly.total = 4,217.93 and narrates the invoice
line. Code computed; LM contextualized.program field — code-in-JSON would degrade it.Outcome: The arithmetic is now deterministic and auditable. PoT-style code execution is the documented fix for exactly this class of error [arXiv 2211.12588, arXiv 2211.10435].
Extractable operation: Multi-step arithmetic in prose is a smell. Re-route it to code (op #6).
Trigger: A support-triage agent step: "Read this customer message and decide whether the
tone is hostile, neutral, or warm." An over-eager engineer wires it through ProgramOfThought
because "code is more reliable."
Constraints:
if "!!!" in msg: tone = "hostile" — a brittle
rule that encodes a subjective rubric as if it were a formula, and is worse than the model's
native judgment.Decision steps:
ChainOfThought (or Predict). The LM is the right engine for judgment.Outcome: No sandbox call. The judgment stays where judgment belongs. Over-coding is a real and common anti-pattern: not everything benefits from an interpreter, only things with a verifiable deterministic core.
Extractable operation: No verifiable answer → no code. Veto the sandbox round-trip (op #7).
Trigger: "Summarize this quarter's sales narrative and give me the exact total revenue."
Constraints: One sentence contains both a judgment (summary) and a computation (total).
Decision steps:
ProgramOfThought: code sums the figures, interpreter verifies.ChainOfThought: prose synthesis, no oracle exists.Outcome: Each sub-step uses its correct engine. This is the execution-decision analogue of the mixed-content two-pass pattern in [[agentsop-output-format-by-model]] §5 Case B.
Extractable operation: One task ≠ one strategy. Gate per step, decompose mixed steps.
max_iters ≈ 3) and fall back to prose or surface the failure. Infinite code-repair loops
burn cost. See [[agentsop-test-fix-loop]].PythonREPLTool as a safe sandbox. It runs in-process, unsandboxed.
Fine for trusted self-authored code; hostile to untrusted input. Use a real sandbox
(OpenAI Code Interpreter container, Anthropic code-exec VM, DSPy's interpreter) when inputs are untrusted.How major frameworks expose the emit-code-vs-reason mechanism. This skill operates at the decision layer; each framework supplies the mechanism.
ProgramOfThought [dspy.ai/learn/programming/modules/]The canonical declarative version. A signature compiled with dspy.ProgramOfThought(Sig)
makes the LM emit Python, runs it in an interpreter sandbox, and feeds the result back —
with bounded retry on error (max_iters). The sibling dspy skill lists ChainOfThought
vs ProgramOfThought as module choices but does not give the decision rubric; this overlay
is that rubric. Use ProgramOfThought exactly when §3 Step 1 returns "computable."
A managed container that the model can write Python into and execute, with files and state persisted across turns. Heavier and stateful — good for data-analysis sessions (load CSV, compute, plot). Same gate applies: route computable steps in, keep judgment in chat.
A sandboxed VM exposed as a tool; the model emits code, it runs, results return to the conversation. First-party, sandboxed — safe for untrusted inputs. The execution-decision gate maps directly: offer the tool, but the model/agent should only invoke it for computable steps.
PythonREPLToolA tool wrapping a Python REPL. Runs in-process and is unsandboxed — powerful and dangerous. Use only for trusted, self-authored computation; never expose it to untrusted input without an external sandbox. The decision rubric is identical; the safety profile is worst-in-class.
Framework | Mechanism | Sandbox | Result-back
---------------------|--------------------------|-------------|------------
DSPy PoT | ProgramOfThought module | Python intp | automatic (max_iters)
OpenAI Code Interp. | Assistants code tool | container | persisted state
Anthropic code-exec | code-execution tool | VM (safe) | into conversation
LangChain | PythonREPLTool | NONE (proc) | manual wiringEvery framework can be mis-invoked — pointed at a judgment task (over-code) or skipped for a computation (under-code). This overlay is the gate that decides invocation, regardless of which mechanism is underneath.
┌──────────────────────────────────────────────────────────────────────┐
│ EMIT-CODE-VS-REASON DECISION CARD │
├──────────────────────────────────────────────────────────────────────┤
│ Single verifiable answer a program could check? │
│ YES → EMIT CODE → sandbox → run → feed result back into LM │
│ NO → REASON IN PROSE (no oracle exists for judgment) │
│ MIXED → decompose; route each sub-step independently │
├──────────────────────────────────────────────────────────────────────┤
│ Arithmetic / parse / sort / count / regex / symbolic → CODE │
│ Tone / quality / summary / design / synthesis → PROSE │
│ "summary AND total" → SPLIT │
├──────────────────────────────────────────────────────────────────────┤
│ LMs are unreliable calculators but reliable coders. │
│ Under-coding fails SILENTLY (wrong plausible number). │
│ Over-coding fails LOUDLY+CHEAPLY (wasted sandbox round-trip). │
│ When genuinely uncertain AND a verifiable core exists → lean CODE. │
├──────────────────────────────────────────────────────────────────────┤
│ NEVER: │
│ • code a judgment task (over-coding veto) │
│ • do multi-step arithmetic in prose (under-coding) │
│ • forget to feed the code result back into the LM │
│ • nest generated code in JSON (see [[agentsop-output-format-by-model]]) │
│ • loop code-repair unbounded (see [[agentsop-test-fix-loop]]) │
│ • trust LangChain PythonREPLTool on untrusted input (unsandboxed) │
└──────────────────────────────────────────────────────────────────────┘Primary anchors:
ProgramOfThought module — [dspy.ai/learn/programming/modules/] — emit+run+retry mechanism.Framework / API docs:
PythonREPLTool — [python.langchain.com/docs/integrations/tools/python] (unsandboxed; in-process).Companion / overlaid skills:
dspy-sop-skill/SKILL.md — ships ProgramOfThought but not this decision rubric (the gap this overlay fills).d-output-format-by-model-skill/SKILL.md — sibling: once you emit code, how to serialize it (PoT for math/parse; code never nested in JSON). Cross-linked as [[agentsop-output-format-by-model]].test-fix-loop — the execute → error → retry loop this skill defers to for bounded code repair. Cross-linked as [[agentsop-test-fix-loop]].© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/agentsop-code-execution-decision of agentsope/SkillAlchemy.
Open the folder on GitHubat commit 6ea799f
Agentsop Code Execution Decision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Code Execution Decision this skillagentsope/SkillAlchemy | 459 | — | ~6.1k | Automated safety check: Pass | MIT | |
| DeepTutor CLIHKUDS/DeepTutor | 41k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch | 66k | — | ~2k | Automated safety check: Pass | MIT | |
| Codebase to Coursezarazhangrui/codebase-to-course | 5.7k | — | ~4.4k | Automated safety check: Pass | None | |
| AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch | 66k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Scholar EvaluationK-Dense-AI/claude-scientific-writer | 2.4k | 2 repos | ~2.9k | Automated safety check: Notes | MIT |
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
rohitg00/ai-engineering-from-scratch
Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.
zarazhangrui/codebase-to-course
Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.
rohitg00/ai-engineering-from-scratch
Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.
K-Dense-AI/claude-scientific-writer
Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.
guanyang/open-agent-hub
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Categories
Decision rubric for when an LM agent should write-and-run code (Program-of-Thought / code interpreter) versus reason in natural language: classify each step as deterministic- computable (emit +…. Agentsop Code Execution Decision is an agent skill from agentsope/SkillAlchemy. Decision rubric for when an LM agent should write-and-run code (Program-of-Thought / code interpreter) versus reason in natural language: classify each step as deterministic- computable (emit + execute code, feed the result back) vs judgment (stay in prose).
Agentsop Code Execution Decision fits situations like: debugging an agent step that does arithmetic/parsing/data transforms; prose reasoning hallucinates a computation (under-coding); A sandbox round- trip is wasted on a judgment task (over-coding).
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-code-execution-decision -a claude-code`. Or copy the skill folder (skills/agentsop-code-execution-decision in agentsope/SkillAlchemy) into .claude/skills/agentsop-code-execution-decision in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-code-execution-decision -a codex`. Or copy the skill folder (skills/agentsop-code-execution-decision in agentsope/SkillAlchemy) into .agents/skills/agentsop-code-execution-decision in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-code-execution-decision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-code-execution-decision, .gemini/skills/agentsop-code-execution-decision, .github/skills/agentsop-code-execution-decision and .opencode/skills/agentsop-code-execution-decision in your project.
Going by SKILL.md and its folder, Agentsop Code Execution Decision needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Code Execution Decision is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Code Execution Decision: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 459 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 2, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.