Agent skill

Fory Performance Optimization

by apache in apache/fory

Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala).

Apache-2.0Auto-check passedMobile

Install Fory Performance Optimization

skills CLI
$ npx skills add apache/fory --skill fory-performance-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install apache/fory fory-performance-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/apache/fory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fory-performance-optimization .claude/skills/fory-performance-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fory-performance-optimization
GitHub stars
4.6k
Token cost
~2.2k tokens
SKILL.md length
1,070 words
Files
6 (incl. references)
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala).

  • Improving serialize/deserialize throughput
  • SKILL.md covers Mission, Operating Principles, Enforce Hard Constraints and Execute Workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Recovering regressions against a reference commit

What it does

Fory Performance Optimization is an agent skill from apache/fory. Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala). Use when improving serialize/deserialize throughput or latency, recovering regressions against a reference commit, diagnosing flamegraphs, fixing perf-related CI failures, or porting proven optimizations across languages without protocol or API regressions.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `agents/openai.yaml`, `references/bottleneck-playbook.md` and `references/language-command-matrix.md`).

It sits in Mobile, covering Performance optimization, Failing and flaky tests and Android development. It works with C++, Java, Rust and JavaScript. The repository describes itself as: A blazingly fast multi-language serialization framework for idiomatic domain objects, schema IDL, and cross-language data exchange. The licence is Apache-2.0.

When your agent uses it

  • Improving serialize/deserialize throughput
  • Recovering regressions against a reference commit
  • Diagnosing flamegraphs
  • Fixing perf-related CI failures

Example prompts

  • “/fory-performance-optimization”

What it can do on your machine

Read from SKILL.md and the folder at commit dcf8d24. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fory Performance Optimization loads about 2.2k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 1,070 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from apache/fory at commit dcf8d24, republished under its Apache-2.0 licence (© apache). 1,070 words, ~2,186 tokens.

Download SKILL.mdSave it as .claude/skills/fory-performance-optimization/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
fory-performance-optimization
description
Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C#, JavaScript/TypeScript, Dart, Kotlin, Scala). Use when improving serialize/deserialize throughput or latency, recovering regressions against a reference commit, diagnosing flamegraphs, fixing perf-related CI failures, or porting proven optimizations across languages without protocol or API regressions.

Fory Performance Optimization

Mission

Deliver measurable performance improvements in Apache Fory without protocol drift, correctness regressions, benchmark-shape tricks, or accidental API rollback.

Operating Principles

  • Start from data, not intuition.
  • Profile before changing hot code.
  • Change one bottleneck at a time.
  • Benchmark sequentially on the same machine state (one benchmark process at a time).
  • Compare old/new benchmark results case-by-case in adjacent pairs: run one case on apache/main, then immediately run that same case on the current branch before moving to the next case.
  • Under high or variable host load, run multiple short adjacent baseline/current pairs. Keep each process short and alternate sides instead of lengthening one run or batching all baseline runs.
  • Keep only measured wins or explicitly requested architecture cleanups.
  • Revert speculative changes that do not pay off.
  • During each optimization round, benchmark only cases directly affected by the change. Defer unrelated controls, full-suite sanity benchmarks, and full comparison matrices to final verification of the whole optimization task unless the user explicitly requests them earlier.
  • Align with reference runtimes (usually C++ first, then Rust/Java) when behavior and ownership models differ.

Enforce Hard Constraints

  • Preserve wire protocol unless explicitly requested.
  • Preserve cross-language semantics and xlang compatibility.
  • Never run two benchmarks at the same time on one host; run exactly one benchmark command at a time.
  • Do not optimize by changing benchmark payload definitions, field encodings, or benchmark methodology.
  • Do not add payload-identity or repeated-input caches that depend on benchmark shape.
  • Do not restore removed APIs/legacy wrappers when the user forbids it.
  • Do not preserve legacy/dead code or stale docs in optimization rounds; remove them when touched.
  • Keep API surface minimal: do not add new API unless required by protocol/correctness or explicitly requested.
  • Never add public hacky API for performance shortcuts; keep optimization helpers internal/private and conceptually clean.
  • Do not hide regressions behind unsafe compiler flags or benchmark-only code paths.
  • Keep optimization surfaces nested-safe; avoid root-only shortcuts unless they are architecturally valid and requested.
  • Do not add reader-side validation solely to produce an earlier or more precise malformed-input error. A necessary crash, panic, undefined-behavior, out-of-bounds, resource-amplification, no-progress, state-pollution, type, or policy guard must keep its hot success path to a primitive branch and move exception allocation and message formatting into a cold no-inline helper when supported. If an existing bounds-safe downstream operation already raises a controlled root error, do not duplicate its validation on the hot path.

Execute Workflow

  1. Read context and constraints.
  • Read tasks/perf_optimization_rounds.md and tasks/lessons.md.
  • Read the relevant spec in docs/specification/ for any path that may affect wire behavior.
  • Record explicit user constraints (forbidden APIs, naming, architecture, protocol rules).
  1. Define target and baseline.
  • Identify one primary KPI (for example Struct Serialize ns/op or ops/sec).
  • Benchmark current HEAD.
  • If a reference commit is provided, persist its built benchmark artifact and commit identity. Treat stored numbers as historical context, not a substitute for an adjacent baseline run in each comparison pair.
  1. Profile the hotspot.
  • Capture a flamegraph or sampled stacks on the exact benchmark command.
  • Quantify top costs by bucket (runtime bookkeeping, dispatch, allocation/copy, map/cache operations, buffer growth, metadata parse/validation).
  • Tie each bucket to concrete file/line ownership before proposing changes.
  1. Form one round hypothesis.
  • State one bottleneck and one expected effect.
  • Prefer structural fixes over micro-tweaks.
  • If another runtime already solved the same bottleneck, port its design shape first.
  1. Implement minimal change.
  • Touch the smallest surface that can validate the hypothesis.
  • Keep invariants explicit: protocol bytes, ownership, cache lifetime, reference semantics, nullability, schema-compatible behavior.
  1. Verify correctness.
  • Run language-local build/test/lint for the touched implementation.
  • Run cross-language checks when runtime/type/protocol behavior can affect xlang.
  • Confirm serialized sizes and compatibility expectations where applicable.
  • For Java performance work, defer local GraalVM image builds and executions to the final verification of the whole optimization task. Do not repeat them in individual rounds unless the user explicitly requests an earlier run.
  1. Benchmark and compare.
Show full SKILL.md (427 more words)Show less
  • Run targeted benchmark at least twice sequentially.
  • Pair each baseline case with the matching current-branch case before starting another case, so both measurements see closer machine load conditions.
  • When host load is high or pair results conflict, use several short baseline/current pairs with the same warmup and measurement settings. Run baseline, current, baseline, current as separate processes; never run all baseline samples before all current samples.
  • Record every pair while it runs. Exclude a pair only with objective contamination evidence such as a competing process, load spike, interruption, or throughput collapse; record the exclusion and do not cherry-pick by direction.
  • Compare paired deltas using their median and dispersion. Do not optimize from a single pair, non-adjacent samples, or a contaminated result. If the retained pairs do not establish a stable signal, stop and wait for a cleaner window instead of changing code against the apparent result.
  • At final verification of the whole optimization task, run the deferred full-suite sanity benchmark and matched comparison matrix to check collateral regressions. Do not repeat these benchmarks in individual optimization rounds.
  1. Decide keep or revert.
  • Keep only if gain is repeatable or cleanup is explicitly requested and accepted with measured tradeoff.
  • Revert if performance regresses or gain is within noise and complexity increases.
  • If a required cleanup regresses, redesign inside the new architecture instead of restoring banned patterns.
  1. Log every round.
  • Append one round entry to tasks/perf_optimization_rounds.md before starting the next round.
  • Include hypothesis, code change, exact commands, before/after numbers, and keep/revert decision.
  • Commit retained non-trivial rounds immediately.
  1. Re-plan on instability.
  • Stop and re-plan when benchmark runs conflict, machine contention is suspected, or profile does not match hypothesis.
  • On a busy machine, re-plan the measurement schedule to multiple short adjacent pairs before forming an optimization hypothesis from benchmark deltas.
  • Re-ground on current HEAD after reset/rebase/checkout events before making further changes.

Apply Decision Rules

  • Treat <1-2% movement as noise unless repeated under controlled runs.
  • Require explicit proof for complexity-increasing optimizations.
  • Prefer deleting dead APIs and dead state quickly after refactors.
  • Keep naming/API cleanup only if performance remains in band.
  • Never run before/after comparisons in parallel.

Use References

Produce Output

When finishing an optimization task, report:

  • Baseline command and numbers.
  • Final command and numbers.
  • Net delta on primary KPI.
  • Correctness and compatibility verification run.
  • Kept vs reverted rounds and rationale.

© apache, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in .agents/skills/fory-performance-optimization of apache/fory.

  • SKILL.md
  • agents/openai.yaml
  • references/bottleneck-playbook.md
  • references/language-command-matrix.md
  • references/round-template.md
  • references/workflow-checklist.md

Open the folder on GitHubat commit dcf8d24

Compare with similar skills

Fory Performance Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fory Performance Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fory Performance Optimization this skillapache/fory4.6k—~2.2kAutomated safety check: PassApache-2.0
Build Teaql Appteaql/teaql-agent-kit2.8k—~4.6kAutomated safety check: PassMIT
Dbgtheodo-group/debug-that158—~1.9kAutomated safety check: PassMIT
Code Revieweralirezarezvani/claude-skills28k1 repos~1.6kAutomated safety check: PassMIT
MCP Debuggerdebugmcp/mcp-debugger171—~3.8kAutomated safety check: PassMIT
MCP SDK Tier Auditmodelcontextprotocol/conformance129—~4.4kAutomated safety check: PassCustom licence

Similar skills

  • Build Teaql App

    teaql/teaql-agent-kit

    Build or change a TeaQL application in Java, Rust, Go, Swift, Python, C/.NET, or TypeScript, including Kotlin/JVM applications that consume Java-generated libraries.

    2.8k GitHub stars~4.6k tokensUpdated 10 days ago
    MobileAuto-check passed
  • Dbg

    theodo-group/debug-that

    Debug applications using the dbg CLI debugger. An agent skill from theodo-group/debug-that.

    158 GitHub stars~1.9k tokensUpdated 4 mo ago
    DevelopmentAuto-check passed
  • Code Reviewer

    alirezarezvani/claude-skills

    Code review automation for TypeScript, JavaScript, Python, Go, Swift, Kotlin, C, .NET, Java, C, C++, Rust, Ruby, PHP, and Dart/Flutter.

    28k GitHub starsUsed in 1 repo~1.6k tokens
    DevelopmentAuto-check passed
  • MCP Debugger

    debugmcp/mcp-debugger

    A skill your agent uses when investigating a bug, failing test, or unexpected runtime behavior and the mcp-debugger MCP server is available — drives real step-through debuggers (breakpoints, stack…

    171 GitHub stars~3.8k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • MCP SDK Tier Audit

    modelcontextprotocol/conformance

    Official

    Comprehensive tier assessment for an MCP SDK repository against SEP-1730.

    129 GitHub stars~4.4k tokensUpdated 2 days ago
    MobileAuto-check passed
  • Official

    Audits existing tests in any language using formal, research-backed test smell names and the testsmells.org 19-smell academic taxonomy.

    5.6k GitHub starsUsed in 1 repo~2.5k tokens
    MobileAuto-check passed

More from apache/fory

  • Fory Release

    apache/fory

    Prepare an Apache Fory release candidate from a clean release branch, including the version bump, RC tag, JVM staging, ASF source artifacts, SVN upload, and vote email.

    4.6k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Bump Apache Fory release or post-release development versions across Java, Kotlin, Scala, Python, Rust, Go, C++, C, Dart, JavaScript, Swift, integration tests, examples, and source docs.

    4.6k GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Categories

Questions about Fory Performance Optimization

What does Fory Performance Optimization do?

Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala). Fory Performance Optimization is an agent skill from apache/fory. Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala).

When should I use Fory Performance Optimization?

Fory Performance Optimization fits situations like: improving serialize/deserialize throughput; recovering regressions against a reference commit; diagnosing flamegraphs; fixing perf-related CI failures.

How do I install Fory Performance Optimization in Claude Code?

Run `npx skills add apache/fory --skill fory-performance-optimization -a claude-code`. Or copy the skill folder (.agents/skills/fory-performance-optimization in apache/fory) into .claude/skills/fory-performance-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Fory Performance Optimization in Codex?

Run `npx skills add apache/fory --skill fory-performance-optimization -a codex`. Or copy the skill folder (.agents/skills/fory-performance-optimization in apache/fory) into .agents/skills/fory-performance-optimization in your project. Codex loads it when a task matches its description.

Can I use Fory Performance Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add apache/fory --skill fory-performance-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fory-performance-optimization, .gemini/skills/fory-performance-optimization, .github/skills/fory-performance-optimization and .opencode/skills/fory-performance-optimization in your project.

What does Fory Performance Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: Fory Performance Optimization is instructions for the agent only.

Does Fory Performance Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Fory Performance Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fory Performance Optimization use?

Fory Performance Optimization is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fory Performance Optimization use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Fory Performance Optimization?

Skills that share tags, products or a category with Fory Performance Optimization: Build Teaql App (teaql/teaql-agent-kit, 2.8k stars), Dbg (theodo-group/debug-that, 158 stars), Code Reviewer (alirezarezvani/claude-skills, 28k stars) and MCP Debugger (debugmcp/mcp-debugger, 171 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fory Performance Optimization?

apache (a GitHub organization) maintains it in apache/fory, which has 4,566 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.

Source: apache/fory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.