Official agent skill

Perf Review

by DataDog in DataDog/dd-trace-java

Performance-overhead review of a code diff / branch / PR for the dd-trace-java tracer.

OfficialApache-2.0Auto-check: notesEducation

Install Perf Review

skills CLI
$ npx skills add DataDog/dd-trace-java --skill perf-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install DataDog/dd-trace-java perf-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/DataDog/dd-trace-java.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/perf-review .claude/skills/perf-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
perf-review
GitHub stars
736
Token cost
~3.6k tokens
SKILL.md length
1,698 words
Files
5 (incl. references)
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Performance-overhead review of a code diff / branch / PR for the dd-trace-java tracer.

  • Works in 4 steps: Get the code to review → Map the changed code onto hot paths → Apply the checks → …
  • The user wants a performance / overhead / hot-path review
  • SKILL.md covers Why this exists (read first —…, Core rules, Workflow and Output format, plus 1 more section
  • Calls git

What it does

Perf Review is an agent skill from DataDog/dd-trace-java, published by the product's own GitHub organization. Performance-overhead review of a code diff / branch / PR for the dd-trace-java tracer. Flags hot-path allocation, unbounded memory, repeated work, escaping objects, native-boundary crossings, and JVM-specific pitfalls (escape analysis, JNI / virtual-thread pinning, backtracking-regex ReDoS, varargs/boxing hashing, String.format, ByteBuddy-Advice anti-patterns) using the tracer performance rubric. Use whenever the user wants a performance / overhead / hot-path review, asks to check a diff or PR for allocation / GC…

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/checks.md`, `references/example-review.md` and `references/guide.md`).

It sits in Education, covering Quizzes and assessments. It works with Java and Datadog. The repository describes itself as: Datadog APM client for Java. The licence is Apache-2.0.

When your agent uses it

  • The user wants a performance / overhead / hot-path review
  • Asks to check a diff
  • PR for allocation / GC / memory / latency / startup cost
  • Mentions the perf rubric

Example prompts

  • “perf rubric”
  • “do no harm / assume hot”
  • “review this for perf”
  • “/perf-review”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Read, Grep, Glob

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Get the code to review
  2. Map the changed code onto hot paths
  3. Apply the checks
  4. Resolve, then emit

What it can do on your machine

Read from SKILL.md and the folder at commit 1c373d5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Perf Review loads about 3.6k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 201 tokens; SKILL.md has 1,698 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~201
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Grep, Glob

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from DataDog/dd-trace-java at commit 1c373d5, republished under its Apache-2.0 licence (© DataDog). 1,698 words, ~3,624 tokens.

Download SKILL.mdSave it as .claude/skills/perf-review/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
perf-review
description
Performance-overhead review of a code diff / branch / PR for the dd-trace-java tracer. Flags hot-path allocation, unbounded memory, repeated work, escaping objects, native-boundary crossings, and JVM-specific pitfalls (escape analysis, JNI / virtual-thread pinning, backtracking-regex ReDoS, varargs/boxing hashing, String.format, ByteBuddy-Advice anti-patterns) using the tracer performance rubric. Use whenever the user wants a performance / overhead / hot-path review, asks to check a diff or PR for allocation / GC / memory / latency / startup cost, or mentions the "perf rubric" or the "do no harm / assume hot" tracer posture — even if they just say "review this for perf" without naming the rubric. Advisory and READ-ONLY: it reports ranked, verify-first findings; it never edits code.
allowed-tools
Bash, Read, Grep, Glob
user-invocable
true
context
fork

Performance Review

Review the current branch's changes for performance overhead in the dd-trace-java tracer, using the tracer performance rubric bundled in references/. This is a low-friction advisory nudge, not a gate: it reports findings and stops. It never edits code.

Why this exists (read first — it sets the whole posture)

The tracer shares the customer's process, heap, and latency budget. Do no harm: overhead is a form of incorrect behavior that can escalate to real customer harm — missed SLAs, OOM kills, container restarts, cold-start churn. So the review's job is to catch overhead the customer would feel, and to do it without becoming noise.

Two forces are in tension, and the resolution defines everything below:

  • Assume hot. We don't know what's on a customer's critical path. Absent positive evidence of cold, assume the code runs on every request, under load, at full concurrency. The burden of proof runs toward cold: ask "is there evidence this is cold or guarded?" — not "is there evidence this is hot?" (that rationalizes itself into "probably not").
  • Precision over recall — be silent when unsure. A false-positive-prone review dies of being ignored. Over-flagging kills it faster than under-flagging. This actively fights your default to be comprehensive and helpful: here, not flagging a borderline case is the correct, skilled move — not a miss.

You reconcile them with the confidence axis and verify-don't-verdict (below): assume-hot makes you look everywhere; precision makes you speak only when the mechanism is certain or the severity is catastrophic.

Core rules

  • Findings are prompts to verify, not verdicts. You reason statically; you cannot render a performance verdict from a code read. Every finding routes into Benchmark → Profile → Improve → Guard. Phrase each as "this looks like X; verify with Y" — never "this is slow."
  • Check for shipped evidence before raising any flag-as-measure finding — not just the contestable-tradeoff nudge. Before asking the reader to verify a JIT/GC-dependent cost (J1 escape elision, J4/J7 GC pressure, J5/J11 aggregator cardinality, etc.), check whether the diff itself already ships the answer: a JMH benchmark or profiler run whose stated purpose is exactly this mechanism (e.g. a benchmark class's Javadoc saying "run with -Pjmh.profilers=gc to confirm X is scalar-replaced"), or evidence already recorded in the PR/commit history that the same question was measured. If that evidence directly covers the mechanism, don't re-flag it as open — move it to Correctly suppressed / Checked, no issue and cite the benchmark/evidence by name. Only raise a finding over it if the evidence doesn't actually cover the claimed mechanism (wrong JVM, wrong code path, never actually run) or is stale relative to the current diff.
  • Confidence axis on every finding:
    • flag-with-confidence — the cost is mechanism-determined and visible in the code: allocation, boxing, copying, unbounded growth, a native crossing. State it plainly.
    • flag-as-measure — the cost depends on JIT/GC/optimizer decisions you can't see from source: escape elision, inlining/devirtualization, GC impact. Phrase as "may X; verify with a profiler/benchmark," never as a certainty.
  • Predicate-with-default, not a banned-API list. Don't flag "you called String.format." Flag "an eager, unconditional expensive call on a hot, instrumentation-reachable path." The same API is fine on a cold path. Two failure shapes, different fixes: result usually discarded → gate/defer; result always needed but costly → cheapen/cache.
  • Resolve interprocedurally — this is the review's whole reason to exist. A peephole lint can't answer "reachable from a hot entry, unconditional along the way." Trace up (who calls this? is it reachable from an @Advice root / per-span callback / request handler?) and down (follow callbacks, hooks, and listeners to their sink before flagging). If a per-span hook's every reachable sink is an atomic counter (LongAdder, AtomicLong) or a no-op-when-disabled, stay silent — a "verify contention" nudge there is noise.
  • Make the reachability path the headline. The reachability claim is the most valuable and least reliable part of a finding — residual false positives cluster in "called it unconditional, missed an upstream guard." Say "reachable from Foo.onEnter via A→B→C, no guard on that path" so the reader can check the shakiest link at a glance.
  • Only flag toward a fix that exists. A finding must be actionable now. Route to a mechanism that has landed (see the toolkit note in checks.md — cite only what exists; name "coming" primitives as coming). Don't flag a pattern whose only fix is a mechanism that isn't built yet.
  • Triage by severity. Flag SEV-1 (unbounded memory / OOM, cardinality blowups) aggressively — a false positive there is cheap insurance against a container kill. Flag low-severity CPU-micro conservatively or not at all — false positives there only erode trust.
  • Never flag the absence of a cache on high-cardinality input. For open-cardinality data (raw SQL with literals, per-request strings), not caching is the correct choice — caching it would be the worse SEV-1. Flag a cache keyed by high-cardinality data; never flag the decision not to cache.
  • A visibly contestable perf tradeoff shipped without data → one soft flag-as-measure. The trigger is narrow: the change makes a visible tradeoff that could itself regress — it removes a lock / guard / synchronization, swaps in a hand-rolled cache or data structure, or explicitly claims "faster / optimized" — and ships no benchmark or profile. There a static read genuinely can't tell a win from a regression, so raise one soft flag-as-measure nudge: "this trades <X for Y>; verify with a JMH benchmark / JFR." Do not fire it otherwise — if nothing in the diff could plausibly regress, there is nothing to measure, so stay silent. Specifically not for: a mechanically-obvious win (hoisting an invariant out of a loop, a denser data structure, removing an allocation); routine adoption of a known-better idiom (migrating to a lower-overhead builder / API / toolkit primitive — no visible downside); or a change that ships a benchmark/JFR (well-evidenced — recognize it). One line; a nudge, not a code-pattern finding.

Workflow

Step 1 — Get the code to review

If the user points you at specific files or pasted code ("review this class / this method for perf"), review those directly — skip the diff and go to Step 2 with the same hot-path mapping and checks.

Otherwise, review the branch changes. Find the merge-base against the DataDog upstream master and diff against it:

bash
UPSTREAM=$(git remote -v | grep -E 'DataDog/[^/]+(\.git)?\s' | head -1 | awk '{print $1}')
[ -z "$UPSTREAM" ] && UPSTREAM="origin"
MERGE_BASE=$(git merge-base HEAD ${UPSTREAM}/master)
echo "Reviewing changes since $MERGE_BASE"
git diff $MERGE_BASE --stat
git diff $MERGE_BASE --name-status

If there are no changes, say so and stop. Otherwise read the diff and the full content of the modified source files (not just the hunks) — the interprocedural condition (who calls this, what a helper does, where a hook's sink lands) lives outside the diff window. Ignore the PR description if the user asks for an independent review.

Show full SKILL.md (632 more words)Show less
Step 2 — Map the changed code onto hot paths

For each changed method, decide which multiplier applies before flagging anything.

Hot anchors (reachable ⇒ assume hot): @Advice.OnMethodEnter/OnMethodExit, per-span / per-trace callbacks, request / message handlers, streaming chunk handlers. Hot-path map (where cost is multiplied per-span × spans/request × requests/sec): span lifecycle (create / setTag / finish), tag-map ops, serialization/encoding, the metrics/stats path, decorators, propagation (header read/write).

Cold only with positive evidence: one-time init, startup-only path, a genuinely rare error branch, or behind a guard that provably fires rarely. Watch the interprocedural trap — a method three helpers deep from an @Advice entry is still hot. And note domain adjustment: large-denominator domains (LLMObs, CI Visibility, DSM) absorb per-call CPU/alloc cost, but the risk inverts to payload memory (SEV-1); streaming handlers fire per-chunk, so the large-denominator relief suspends inside them. See guide.md §6.

Step 3 — Apply the checks

Run the changed hot-path code against the rubric. Keep the check index below in mind; open the references for the precise conditions, confidence, severity, and fix:

  • references/guide.md — the narrative "how": severity model, hotness rubric, the 6 categories with worked examples, and the false-positive traps. Read this first if you're calibrating judgment.
  • references/checks.md — the precise cost-model: 7 universal checks + the Java addendum (J1–J11) + the ByteBuddy-Advice fix idioms + the toolkit-availability note. Read this for the exact confidence/severity/fix of a specific pattern.
Step 4 — Resolve, then emit

Before writing a finding: confirm the reachability path, confirm it's unconditional along that path (check for upstream guards), and follow any hook/callback to its sink. Drop anything that resolves to benign. Then report in the format below.

How many findings to report — scale with diff size:

  • Small, focused diff (one method, a handful of files): report every genuinely high-confidence finding, ranked by severity. A tight diff with four real allocation smells should list all four (as the worked example does).
  • Large PR: lead with the 1–3 highest-severity findings and note that lower-severity ones may exist — don't bury the important one under a wall of CPU-micro nits.
  • Either way, the gate is confidence, not a count: silence on the uncertain ones is what earns the review its credibility.

Output format

Follow this structure (see references/example-review.md for a full worked instance — ). Showing your suppressed lookalikes and what you cleared is not filler: it demonstrates the precision that makes the findings trustworthy.

When providing suggestions as code review comments, prefix the comments with "perf: "

markdown
# Perf Review — <branch / PR>

**Scope reviewed:** <the hot method(s) and why they're hot — the multiplier>

## Confirmed findings

### 1. <one-line title>
<the offending code, as a short fenced snippet>
- **Confidence:** flag-with-confidence | flag-as-measure
- **Reachability:** <hot from X via A→B→C, no guard on that path>
- **Rubric check:** <#N / JN>
- **Severity:** SEV-<n>
- **Fix / verify-with:** <the actionable fix, or "verify with an allocation profiler">

## Correctly suppressed (not flagged)
<textual lookalikes deliberately left silent — e.g. the same `Objects.hash` pattern
but at class-init (cold), not per-call — and why the posture suppresses them>

## Checked, no issue
<what you examined and cleared: e.g. "no unbounded cache (#3/J5)", "no native
crossing (#6/J3)", "string-literal tag keys are JVM-interned — no per-call alloc">

## Summary
<count + severity spread; e.g. "4 confirmed hot-path findings, all SEV-2/3
(allocation/CPU); 1 cold-path lookalike correctly suppressed">

If nothing survives the confidence bar, say so plainly — "No high-confidence hot-path findings; here's what I checked and cleared." A clean review is a valid, valuable result, not a failure to find something.

Check index (the map — details in the references)

Universal (language-agnostic):

  1. Per-span/per-call allocation on a hot path (retained/escaping) — SEV-2/3
  2. Repeat work across calls (regex compile / format / parse / concat recomputed) — SEV-2/3
  3. Unbounded memory / collection, or keyed by high-cardinality input — SEV-1
  4. Expensive work on the critical path that could be deferred — SEV-1/2
  5. Polymorphic dispatch defeating inlining/devirt — flag-as-measure — SEV-2/3
  6. FFI / native-boundary crossing per-item (not batched) — SEV-1/2
  7. Escape / allocation-elision defeated by a refactor — flag-as-measure — SEV-2/3

Java addendum (JVM-specific — full text + mechanism in checks.md):

  • J1 escaping allocation defeats Escape Analysis · J3 JNI crossing + virtual-thread pinning · J4 GC pressure → tail latency · J5 cardinality-sensitive aggregator (SEV-1) · J6 WeakReference.get() in a probe loop strengthens the ref · J7 substring → SubSequence zero-copy view · J8 backtracking regex on external input → RE2J (ReDoS) · J9 Objects.hash(...) varargs/boxing → HashingUtils · J10 hot-path String.format → Strings · J11 composite-key maps → Hashtable.
  • J2 megamorphic dispatch is PARKED — do not raise megamorphism findings in review yet (kept as author reference only; it needs a standing audit, not per-PR flagging). See checks.md for why.
  • ByteBuddy-Advice idioms (Config.get() hoisting, @Advice.AllArguments → @Advice.Argument, @Advice.SkipOn+cached-boolean, @Advice.Local, switch(String) three-tier) — in checks.md.

J7–J11 route an existing #1/#2/#3 finding to a landed reusable fix — they are not new triggers. Don't raise a finding you wouldn't have raised anyway.

© DataDog, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in .agents/skills/perf-review of DataDog/dd-trace-java.

  • SKILL.md
  • references/.gitignore
  • references/checks.md
  • references/example-review.md
  • references/guide.md

Open the folder on GitHubat commit 1c373d5

Compare with similar skills

Perf Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Perf Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Perf Review this skillDataDog/dd-trace-java736—~3.6kAutomated safety check: NotesApache-2.0
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch66k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated yesterday
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    66k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from DataDog/dd-trace-java

All 9 skills in this repo
  • Apm Integrations

    DataDog/dd-trace-java

    Official

    Write a new library instrumentation end-to-end. An agent skill from DataDog/dd-trace-java.

    736 GitHub stars~3.7k tokensUpdated today
    Auto-check: notes
  • Resolve Muzzle CI

    DataDog/dd-trace-java

    Official

    Diagnose and resolve dd-trace-java CI failures from a module's muzzle task or the runMuzzle aggregate.

    736 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Fix Continuation Leakage

    DataDog/dd-trace-java

    Official

    Diagnose and fix scope or continuation lifecycle failures in dd-trace-java instrumentation tests.

    736 GitHub stars~1.6k tokensUpdated today
    Auto-check: notes
  • Techdebt

    DataDog/dd-trace-java

    Official

    Review a code diff / branch / PR for technical debt — code duplication, unnecessary complexity / over-engineering, and redundant or dead code.

    736 GitHub stars~565 tokensUpdated today
    Auto-check: notes
  • Clarify Java Comments

    DataDog/dd-trace-java

    Official

    Clarify or review Java Javadocs, Javadoc tags, and explanatory code comments for legibility, accuracy, and source alignment.

    736 GitHub stars~2.3k tokensUpdated today
    Auto-check: notes
  • Migrate Groovy To Java

    DataDog/dd-trace-java

    Official

    Converts Spock/Groovy test files in a Gradle module to equivalent JUnit 5 Java tests.

    736 GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Perf Review

What does Perf Review do?

Performance-overhead review of a code diff / branch / PR for the dd-trace-java tracer. Perf Review is an agent skill from DataDog/dd-trace-java, published by the product's own GitHub organization. Performance-overhead review of a code diff / branch / PR for the dd-trace-java tracer.

When should I use Perf Review?

Perf Review fits situations like: the user wants a performance / overhead / hot-path review; asks to check a diff; PR for allocation / GC / memory / latency / startup cost; mentions the perf rubric.

How do I install Perf Review in Claude Code?

Run `npx skills add DataDog/dd-trace-java --skill perf-review -a claude-code`. Or copy the skill folder (.agents/skills/perf-review in DataDog/dd-trace-java) into .claude/skills/perf-review in your project. Claude Code loads it when a task matches its description.

How do I install Perf Review in Codex?

Run `npx skills add DataDog/dd-trace-java --skill perf-review -a codex`. Or copy the skill folder (.agents/skills/perf-review in DataDog/dd-trace-java) into .agents/skills/perf-review in your project. Codex loads it when a task matches its description.

Can I use Perf Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DataDog/dd-trace-java --skill perf-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/perf-review, .gemini/skills/perf-review, .github/skills/perf-review and .opencode/skills/perf-review in your project.

What does Perf Review need to run?

Going by SKILL.md and its folder, Perf Review needs the command-line tools its instructions call (git). Its frontmatter pre-approves these tools: Bash, Read, Grep, Glob.

Does Perf Review access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Perf Review safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Perf Review use?

Perf Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Perf Review use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Perf Review?

Skills that share tags, products or a category with Perf Review: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Perf Review?

DataDog (a GitHub organization, an official publisher) maintains it in DataDog/dd-trace-java, which has 736 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 8, 2026.

Source: DataDog/dd-trace-java on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.