Agent skill

Capability Experiments

by jellydn in jellydn/my-ai-tools

Build an interactive report or experiment when the user asks to explore model capabilities.

MITAuto-check passedEducation

Install Capability Experiments

skills CLI
$ npx skills add jellydn/my-ai-tools --skill capability-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jellydn/my-ai-tools capability-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jellydn/my-ai-tools.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/capability-experiments .claude/skills/capability-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
capability-experiments
GitHub stars
123
Token cost
~2k tokens
SKILL.md length
393 words
Files
1
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Build an interactive report or experiment when the user asks to explore model capabilities.

  • Works in 3 steps: Decompose the Problem → Research and Validate → Recommend with Evidence
  • Asks to explore model capabilities
  • SKILL.md covers When to Use, What It Does, HTML Report Generation and Proactive Research, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Capability Experiments is an agent skill from jellydn/my-ai-tools. Build an interactive report or experiment when the user asks to explore model capabilities.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: cline, claude, opencode, amp, codex, gemini, cursor, pi

It sits in Education, covering Quizzes and assessments. The repository describes itself as: Comprehensive configuration management for AI coding tools - Replicate my complete setup for Claude Code, OpenCode, Amp, Li, Codex and Claude Code Switch with custom… The licence is MIT.

When your agent uses it

  • Asks to explore model capabilities
  • Tasks that involve Quizzes and assessments

Example prompts

  • “/capability-experiments”

Requirements

  • Compatibility (from SKILL.md): cline, claude, opencode, amp, codex, gemini, cursor, pi

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Decompose the Problem
  2. Research and Validate
  3. Recommend with Evidence

What it can do on your machine

Read from SKILL.md and the folder at commit 163951e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are html and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    cline, claude, opencode, amp, codex, gemini, cursor, pi

    From compatibility in the SKILL.md frontmatter.

Context cost

Capability Experiments loads about 2k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 393 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jellydn/my-ai-tools at commit 163951e, republished under its MIT licence (© jellydn). 393 words, ~2,046 tokens.

Download SKILL.mdSave it as .claude/skills/capability-experiments/SKILL.md (or your agent's skills folder).
name
capability-experiments
description
Build an interactive report or experiment when the user asks to explore model capabilities.
compatibility
cline, claude, opencode, amp, codex, gemini, cursor, pi
license
MIT
hint
Use for rich outputs, interactive forms, or exploring what next-gen models can do
user-invocable
true

Capability Experiments

When to Use

Use this skill when:

  • You need to present complex analysis in a readable format
  • Interactive questionnaires would improve user interaction
  • Exploring what's possible with next-generation model capabilities
  • Multi-step reasoning is needed for architectural decisions
  • You want to generate rich outputs (HTML, tables, interactive elements)
  • Standard text output isn't sufficient for the task

What It Does

Teaches techniques for leveraging advanced model capabilities — HTML generation, embedded interactive elements, proactive research, and multi-step reasoning. These patterns showcase what's newly possible with next-generation models and should be used freely.

HTML Report Generation

Advanced models can generate rich, self-contained HTML. This is useful for:

Analysis Reports

Generate structured HTML reports for complex findings:

html
<!DOCTYPE html>
<html>
<head><style>
  body { font-family: system-ui; max-width: 800px; margin: 2rem auto; }
  .finding { border-left: 4px solid #e74c3c; padding: 1rem; margin: 1rem 0; }
  .finding.fixed { border-color: #2ecc71; }
  .severity { font-weight: 600; font-size: 0.85rem; }
</style></head>
<body>
  <h1>Code Review: PR #288</h1>
  <div class="finding">
    <span class="severity">🔴 Critical</span>
    <p>Hardcoded path in config...</p>
  </div>
  ...
</body></html>

Use HTML reports when:

  • Comparing multiple options or decisions
  • Presenting structured analysis with severity levels
  • Creating interactive documentation
  • Showing progress or status dashboards
Embedded Questionnaires

Generate HTML questionnaires for spec interviews and quizzes:

html
<form id="quiz">
  <div class="question">
    <p>1. Why did we choose GitHub App Installation flow?</p>
    <label><input type="radio" name="q1" value="a"> OAuth is deprecated</label>
    <label><input type="radio" name="q1" value="b"> Org-level access ✓</label>
  </div>
  <button type="button" onclick="checkAnswers()">Check</button>
</form>
<script>
function checkAnswers() {
  const correct = { q1: 'b', q2: 'c' };
  // ... scoring logic
}
</script>

These work well with Fable's ability to render and execute embedded HTML in responses.

Decision Trees and Flowcharts

Use HTML/CSS to visualize decision processes:

html
<div class="decision-tree">
  <div class="node root">Feature Change</div>
  <div class="branch">
    <div class="node">Familiar code?</div>
    <div class="yes">→ Standard pattern</div>
    <div class="no">→ /blindspots first</div>
  </div>
</div>
<style>
  .decision-tree { font-family: monospace; }
  .node { background: #f0f0f0; padding: 8px; margin: 4px; }
  .yes { color: #2ecc71; }
  .no { color: #e74c3c; }
</style>

Proactive Research

Let the agent self-direct exploration rather than waiting for instructions:

Pattern: "Investigate and Report"

Instead of asking the user what to look at, proactively scan the codebase:

1. Scan recent changes with `fff` and `git log`
2. Identify areas of concern or interest
3. Analyze patterns and potential issues
4. Report findings with actionable recommendations
Pattern: "Unknown Unknowns Scan"

Before starting work, proactively look for gotchas:

1. Search git history for past bugs in related areas
2. Check qmd for relevant learnings
3. Search ctx for past agent discussions
4. Review existing ADRs for architectural context
5. Report findings before proposing implementation
Pattern: "Capability Probe"

When faced with a complex task, probe what's possible:

1. Generate multiple approaches (not just the obvious one)
2. For each approach, assess feasibility
3. Consider approaches that were hard with older models
4. Recommend the best approach with rationale

Multi-Step Reasoning

Break complex decisions into structured analysis:

Step 1: Decompose the Problem
markdown
## Analysis: Authentication Strategy

### Dimensions to Consider
1. Security requirements (OAuth 2.0, org-level access)
2. User experience (login flow, token management)
3. Maintenance (token refresh, error handling)
4. Scalability (multiple orgs, rate limits)

### Trade-offs
| Approach | Security | UX | Maintenance | Scalability |
|----------|----------|----|-------------|-------------|
| OAuth App | Medium | High | Low | Low |
| GitHub App | High | Medium | Medium | High |
| PAT | Low | Low | High | Medium |
Show full SKILL.md (157 more words)Show less
Step 2: Research and Validate

For each promising approach, gather evidence:

markdown
## Research: GitHub App Installation

### Findings
- ✓ Org-level repo access (required)
- ✓ Webhook-based events
- ✓ Token caching supported
- ✗ More complex setup
- ✗ Requires webhook endpoint

### Past Context
- Previous PR #123 attempted similar approach
- ADR-005 discusses webhook infrastructure
- qmd has learnings about token caching
Step 3: Recommend with Evidence
markdown
## Recommendation

**Approach**: GitHub App Installation Flow

**Rationale**:
1. Required for org-level access (blocker for OAuth App)
2. Token caching addresses UX concerns
3. Existing webhook infrastructure from PR #123
4. ADR-005 confirms architectural fit

**Risks**:
- Webhook endpoint needs high availability
- Token refresh handling adds complexity

Embedded Questionnaires

Interactive questionnaires improve spec interviews and knowledge verification:

For Spec Interviews

Generate structured questions with HTML forms:

html
<h3>Architecture Questions</h3>
<div class="question">
  <p><strong>Q1:</strong> Should we use a single shared database or per-tenant databases?</p>
  <details>
    <summary>Context</summary>
    <p>Current system uses shared DB. Per-tenant improves isolation but adds operational complexity.</p>
  </details>
  <textarea rows="2" placeholder="Your thinking..."></textarea>
</div>
For Quizzes

Generate self-assessment quizzes:

html
<h3>Implementation Quiz</h3>
<form id="quiz">
  <div class="q">
    <p>1. What caching strategy did we use for installation tokens?</p>
    <label><input type="radio" name="q1" value="r"> Redis with TTL</label>
    <label><input type="radio" name="q1" value="w"> In-memory cache</label>
  </div>
  <button type="button" onclick="grade()">Check Understanding</button>
</form>

Proactive Patterns Summary

When to Use Each Pattern:

┌──────────────────────────────┬──────────────────────────┐
│ Situation                     │ Use Pattern              │
├──────────────────────────────┼──────────────────────────┤
│ Complex analysis to present  │ HTML report              │
│ Spec is vague                │ Embedded questionnaire   │
│ Large, unfamiliar codebase   │ Proactive research scan  │
│ Hard architectural decision  │ Multi-step reasoning     │
│ After implementation         │ Embedded quiz            │
│ Exploring options            │ Capability probe         │
└──────────────────────────────┴──────────────────────────┘

Integration with Other Skills

  • spec-interview: Use embedded questionnaires for richer interviews
  • quiz-me: Generate HTML quizzes instead of plaintext
  • blindspot-pass: Present findings as structured HTML reports
  • implementation-logger: Generate HTML reports from logs
  • doc-search: Search for past HTML report patterns

Tips

  • Use HTML for structure, not flash: Clean, readable HTML is better than complex designs
  • Self-contained is key: Include styles inline so reports render anywhere
  • Interactive when useful: Forms and buttons engage the user; plain text is fine for simple cases
  • Probe before committing: Try one capability at a time, verify it works
  • Fall back gracefully: If HTML doesn't render well, plaintext always works
  • Document what works: New capabilities discovered should be added to this skill

© jellydn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/capability-experiments of jellydn/my-ai-tools.

Open the folder on GitHubat commit 163951e

Compare with similar skills

Capability Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Capability Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Capability Experiments this skilljellydn/my-ai-tools123—~2kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.3kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch65k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch65k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.3k tokensUpdated 3 days ago
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    65k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    65k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    973 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from jellydn/my-ai-tools

All 35 skills in this repo
  • Babysit PR

    jellydn/my-ai-tools

    A skill your agent uses when monitoring an open GitHub PR for CI failures, review feedback, mergeability, and safe retries or fixes.

    123 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Visual PR

    jellydn/my-ai-tools

    Posts a concise visual outline as a GitHub pull request comment.

    123 GitHub stars~764 tokensUpdated today
    Auto-check passed
  • Qmd Knowledge

    jellydn/my-ai-tools

    Manage project knowledge with qmd — captures learnings, decisions, and conventions

    123 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Prd

    jellydn/my-ai-tools

    Generate Product Requirements Documents from feature ideas — plans specs and requirements

    123 GitHub starsUsed in 4 repos~1.8k tokens
    Auto-check passed
  • PR Review

    jellydn/my-ai-tools

    Fix PR review comments by implementing requested changes. An agent skill from jellydn/my-ai-tools.

    123 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Ralph

    jellydn/my-ai-tools

    Convert PRDs to prd.json format for the Ralph autonomous agent system

    123 GitHub starsUsed in 3 repos~1.8k tokens
    Auto-check passed

Categories

Questions about Capability Experiments

What does Capability Experiments do?

Build an interactive report or experiment when the user asks to explore model capabilities. Capability Experiments is an agent skill from jellydn/my-ai-tools. Build an interactive report or experiment when the user asks to explore model capabilities.

When should I use Capability Experiments?

Capability Experiments fits situations like: asks to explore model capabilities; tasks that involve Quizzes and assessments.

How do I install Capability Experiments in Claude Code?

Run `npx skills add jellydn/my-ai-tools --skill capability-experiments -a claude-code`. Or copy the skill folder (skills/capability-experiments in jellydn/my-ai-tools) into .claude/skills/capability-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Capability Experiments in Codex?

Run `npx skills add jellydn/my-ai-tools --skill capability-experiments -a codex`. Or copy the skill folder (skills/capability-experiments in jellydn/my-ai-tools) into .agents/skills/capability-experiments in your project. Codex loads it when a task matches its description.

Can I use Capability Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jellydn/my-ai-tools --skill capability-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/capability-experiments, .gemini/skills/capability-experiments, .github/skills/capability-experiments and .opencode/skills/capability-experiments in your project.

What does Capability Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Capability Experiments is instructions for the agent only. Compatibility (from SKILL.md): cline, claude, opencode, amp, codex, gemini, cursor, pi.

Does Capability Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Capability Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Capability Experiments use?

Capability Experiments is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Capability Experiments use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Capability Experiments?

Skills that share tags, products or a category with Capability Experiments: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 65k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 65k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Capability Experiments?

jellydn (a GitHub user) maintains it in jellydn/my-ai-tools, which has 123 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 7, 2026.

Source: jellydn/my-ai-tools on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.