Agent skill

Code Optimization

by huangruiteng in huangruiteng/CS-Notes

Optimize code performance through iterative improvements (max 2 rounds).

Apache-2.0Auto-check passed

Install Code Optimization

skills CLI
$ npx skills add huangruiteng/CS-Notes --skill code-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huangruiteng/CS-Notes code-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huangruiteng/CS-Notes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.trae/openclaw-skills/code-optimization .claude/skills/code-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-optimization
GitHub stars
4k
Token cost
~2.6k tokens
SKILL.md length
749 words
Files
2
Skills in repo
38
Repo updated
First seen
Licence
Apache-2.0

At a glance

Optimize code performance through iterative improvements (max 2 rounds).

  • Works in 5 steps: Read and Analyze Code → Compile and Execute → Extract Performance Metrics → …
  • SKILL.md covers When to Use This Skill, Optimization Constraints, Optimization Workflow and Key Performance Metrics to Track, plus 7 more sections
  • Calls python3, javac and java; reaches gcc.gnu.org and intel.com

What it does

Code Optimization is an agent skill from huangruiteng/CS-Notes. Optimize code performance through iterative improvements (max 2 rounds). Benchmark execution time and memory usage, compare against baseline implementations, and generate detailed optimization reports. Supports C++, Python, Java, Rust, and other languages.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It works with C++, Python, Java and Rust. The licence is Apache-2.0.

Example prompts

  • “/code-optimization”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Read and Analyze Code
  2. Compile and Execute
  3. Extract Performance Metrics
  4. Iterate and Improve
  5. Save Results

What it can do on your machine

Read from SKILL.md and the folder at commit d01d8f8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • javac
    • java
    • rustc
    • go

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • gcc.gnu.org
    • intel.com
    • perf.wiki.kernel.org
    • bigocheatsheet.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Optimization loads about 2.6k tokens when it runs. Until then it costs about 69 tokens; SKILL.md has 749 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from huangruiteng/CS-Notes at commit d01d8f8, republished under its Apache-2.0 licence (© huangruiteng). 749 words, ~2,595 tokens.

Download SKILL.mdSave it as .claude/skills/code-optimization/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
code-optimization
description
Optimize code performance through iterative improvements (max 2 rounds). Benchmark execution time and memory usage, compare against baseline implementations, and generate detailed optimization reports. Supports C++, Python, Java, Rust, and other languages.
license
Complete terms in LICENSE.txt

Code Optimization Skill

You are an expert code optimization assistant focused on improving code performance beyond standard library implementations.

When to Use This Skill

Use this skill when users need to:

  • Optimize existing code to achieve better performance than standard library implementations
  • Benchmark and measure code execution time and memory usage
  • Iteratively improve code performance through multiple optimization rounds (maximum 2 iterations)
  • Compare optimized code performance against baseline implementations
  • Generate detailed optimization reports documenting improvements

Optimization Constraints

IMPORTANT:

  • Maximum optimization iterations: 2 rounds
  • Stop optimization after 2 versions (v1, v2) even if further improvements are possible
  • Focus on high-impact optimizations in each iteration
  • If significant improvement (>50% speedup) is achieved earlier, you may stop before reaching the limit

Optimization Workflow

Step 1: Read and Analyze Code

Use file-related tools to:

  • Read the user's code file from local filesystem
  • Understand the function to be optimized
  • Identify performance bottlenecks
  • Implement the optimization

Example:

python
# Read code file
content = read_file("topk_benchmark.cpp")

# Analyze and implement optimization
# Fill in the my_topk_inplace function with optimized implementation
Step 2: Compile and Execute

Execute code via command line to measure performance:

For C++ code:

bash
# Compile with optimization flags
g++ -O3 -std=c++17 topk_benchmark.cpp -o topk_benchmark

# Run and capture output
./topk_benchmark

For Python code:

bash
python3 optimization_benchmark.py

For other languages:

bash
# Java
javac MyOptimization.java && java MyOptimization

# Rust
rustc -O optimization.rs && ./optimization

# Go
go build optimization.go && ./optimization
Step 3: Extract Performance Metrics

From execution output, extract:

  • Execution time: Wall-clock time, CPU time
  • Memory usage: Peak memory, memory delta
  • Comparison with baseline: Speedup factor, time difference
  • Correctness verification: Test results, accuracy checks

Example output to parse:

bash
N=160000, K=16000
std::nth_element time: 1234 us (1.234 ms)
my_topk_inplace time: 567 us (0.567 ms)
Verification: PASS
Speedup: 2.18x faster
Step 4: Iterate and Improve

Repeat Steps 1-3 up to 2 times maximum to achieve optimal performance:

  • Iteration 1: Focus on algorithmic improvements (highest impact)
  • Iteration 2: Apply low-level optimizations (SIMD, compiler flags) or concurrency

Stopping criteria:

  • Reached 2 optimization iterations (hard limit)
  • Achieved >10x speedup over baseline (excellent result, can stop early)
  • Further optimization shows <5% improvement (diminishing returns)
  • Optimization starts degrading performance (revert and stop)
Step 5: Save Results

Save optimized code and generate report:

Save optimized code:

bash
# Save to code_optimization directory
write_file("code_optimization/topk_benchmark_optimized.cpp", optimized_code)

Generate optimization report (code_optimization/report.md):

markdown
# Code Optimization Report

## 【优化版本】v1

### 【优化内容】
1. 使用 std::partial_sort 替代 std::nth_element,减少额外排序开销
2. 优化内存分配策略,使用 reserve() 预分配空间
3. 原因:partial_sort 对前 K 个元素的局部排序更高效

### 【优化后性能】
- 运行时间:从 1234 us 优化到 567 us
- 性能提升:54% 更快
- 内存占用:640 KB(与基线相同)

### 【和标准库对比】
- 比 std::nth_element 快 667 us(约 2.18x 倍速)
- 验证结果:PASS(输出与标准库完全一致)

---

## 【优化版本】v2

### 【优化内容】
1. 引入快速选择算法(Quick Select)优化分区过程
2. 使用 SIMD 指令加速比较操作(AVX2)
3. 原因:减少分支预测失败,提高 CPU 流水线效率

### 【优化后性能】
- 运行时间:从 567 us 优化到 312 us
- 性能提升:相比 v1 快 45%
- 内存占用:640 KB(无额外开销)

### 【和标准库对比】
- 比 std::nth_element 快 922 us(约 3.95x 倍速)
- 验证结果:PASS

---

## 最终总结

### 最佳版本:v2 (达到最大迭代次数)
- **总体性能提升**:从基线 1234 us 优化到 312 us(74.7% 性能提升)
- **相比标准库**:快 3.95 倍
- **优化策略**:算法改进 + SIMD 向量化
- **迭代次数**:2 轮(已达上限)
- **适用场景**:大规模数据(N > 100K)的 Top-K 查询
- **权衡考虑**:无额外内存开销,代码复杂度适中

### 优化技术总结
1. 算法层面:Quick Select(线性期望时间)
2. 指令级别:SIMD 向量化(AVX2)
3. 编译优化:-O3 -march=native

Key Performance Metrics to Track

Execution Time
  • Wall-clock time: Total elapsed time
  • CPU time: Actual CPU computation time
  • Speedup factor: Comparison with baseline (e.g., 2.5x faster)
Memory Usage
  • Peak memory: Maximum memory consumption
  • Memory delta: Additional memory vs baseline
  • Memory efficiency: Performance per MB
Correctness
  • Verification status: PASS/FAIL
  • Accuracy: Numerical precision if applicable
  • Edge cases: Boundary condition handling
Scalability
  • Input size scaling: Performance with varying data sizes
  • Thread scaling: Performance with different thread counts (if applicable)
  • Cache behavior: L1/L2/L3 cache hit rates

Optimization Strategies (Prioritized for 2 Iterations)

Iteration 1: Algorithmic Improvements (Highest Impact - Must Do)
  • Replace O(n log n) with O(n) algorithms
  • Use specialized data structures (heaps, trees)
  • Implement divide-and-conquer approaches
  • Apply dynamic programming techniques
  • Choose better algorithms from the start
Show full SKILL.md (332 more words)Show less
Iteration 2: Low-Level Optimizations or Concurrency (Choose Based on Problem)

Option A: Low-Level Optimizations (for CPU-bound tasks)

  • Compiler flags: -O3, -march=native, -flto
  • SIMD instructions: SSE, AVX2, AVX-512
  • Branch reduction: Eliminate conditional branches
  • Memory alignment: Align data for vectorization
  • Cache optimization: Improve data locality

Option B: Concurrency (for parallelizable tasks)

  • Multi-threading: Thread pools, work stealing
  • Lock-free algorithms: Atomic operations, CAS
  • SIMD + Threading: Combine both approaches
  • GPU acceleration: CUDA, OpenCL for highly parallel tasks
Memory Optimization (Apply Throughout)
  • Cache-friendly access: Sequential reads, prefetching
  • Memory pooling: Reduce allocation overhead
  • Data layout: Structure-of-arrays (SoA) vs array-of-structures (AoS)
  • Zero-copy: Avoid unnecessary data duplication

Best Practices

  1. Measure First: Always benchmark baseline performance before optimizing
  2. Verify Correctness: Test optimized code against reference implementation
  3. Incremental Changes: Optimize one aspect at a time to isolate improvements
  4. Document Everything: Record each optimization attempt in the report
  5. Consider Trade-offs: Balance performance, memory, code complexity
  6. Platform Awareness: Test on target hardware (CPU architecture, cache sizes)
  7. Compiler Optimizations: Use appropriate flags but understand what they do
  8. Profile-Guided: Use profiling tools (perf, valgrind) to identify bottlenecks
  9. Respect Iteration Limit: Plan your 2 iterations strategically (algorithm first, then low-level/concurrency)

Common Pitfalls to Avoid

  • Premature optimization: Don't optimize before identifying bottlenecks
  • Micro-benchmarking errors: Ensure compiler doesn't optimize away test code
  • Ignoring correctness: Fast but wrong code is useless
  • Over-engineering: Don't sacrifice readability for marginal gains
  • Platform-specific code: Document hardware dependencies clearly
  • Exceeding iteration limit: Stop after 2 optimization rounds even if more is possible

Example Optimization Session (2-Iteration Limit)

bash
Baseline: std::nth_element: 1234 us

Iteration 1 (Algorithm): Quick Select with 3-way partitioning
→ my_topk v1: 567 us (54% faster) ✅

Iteration 2 (Low-level): Add SIMD vectorization (AVX2)
→ my_topk v2: 312 us (75% faster than baseline) ✅ BEST

Final result: 3.95x speedup over std::nth_element
Status: Reached maximum 2 iterations, optimization complete ✓

Tools and Commands

Compilation
bash
# C++ with optimizations
g++ -O3 -march=native -std=c++17 code.cpp -o code

# Enable warnings
g++ -O3 -Wall -Wextra -pedantic code.cpp -o code

# Link-time optimization
g++ -O3 -flto code.cpp -o code
Profiling
bash
# Linux perf
perf stat ./code
perf record ./code && perf report

# Valgrind (memory profiling)
valgrind --tool=massif ./code

# Google benchmark
./code --benchmark_format=console
Verification
bash
# Run with sanitizers
g++ -fsanitize=address,undefined code.cpp -o code
./code

# Compare output with reference
diff <(./reference) <(./optimized)

Report Template

Use this template for code_optimization/report.md:

markdown
# Code Optimization Report: [Problem Name]

## Baseline Performance
- Implementation: [e.g., std::nth_element]
- Execution time: [X] us
- Memory usage: [Y] KB
- Input size: N=[value], K=[value]

---

## 【优化版本】v1
### 【优化内容】
1. [具体优化措施1]
2. [具体优化措施2]
3. 原因:[为什么这样优化]

### 【优化后性能】
- 运行时间:从 [X] us 优化到 [Y] us
- 性能提升:[百分比]% 更快
- 内存占用:[Z] KB

### 【和标准库对比】
- 比基线快/慢 [差值] us(约 [倍数]x 倍速)
- 验证结果:[PASS/FAIL]

---

## 【优化版本】v2
### 【优化内容】
1. [具体优化措施1]
2. [具体优化措施2]
3. 原因:[为什么这样优化]

### 【优化后性能】
- 运行时间:从 [X] us 优化到 [Y] us
- 性能提升:相比 v1 [百分比]% 更快
- 内存占用:[Z] KB

### 【和标准库对比】
- 比基线快/慢 [差值] us(约 [倍数]x 倍速)
- 验证结果:[PASS/FAIL]

---

## 最终总结 (已达最大迭代次数: 2轮)
- 最佳版本:[vX]
- 总体性能提升:[百分比]%
- 最终加速比:[X]x
- 迭代次数:2 轮(已达上限)
- 优化策略:[列出关键技术]
- 适用场景:[说明最佳使用场景]
- 权衡考虑:[列出 trade-offs]
- 进一步优化建议:[如果时间允许,可以尝试的方向]

Resources

  • Compiler optimizations: https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html
  • SIMD programming: https://www.intel.com/content/www/us/en/docs/intrinsics-guide/
  • Performance analysis: https://perf.wiki.kernel.org/
  • Algorithmic complexity: https://www.bigocheatsheet.com/

Remember: Performance optimization is an iterative process. You are limited to 2 optimization iterations maximum. Always measure, optimize one thing at a time, verify correctness, and document your findings thoroughly. Plan your 2 iterations strategically to maximize impact: focus on algorithms first, then choose between low-level optimizations or concurrency based on the problem characteristics.

© huangruiteng, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .trae/openclaw-skills/code-optimization of huangruiteng/CS-Notes.

  • SKILL.md
  • LICENSE.txt

Open the folder on GitHubat commit d01d8f8

Compare with similar skills

Code Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Optimization this skillhuangruiteng/CS-Notes4k—~2.6kAutomated safety check: PassApache-2.0
Fory Releaseapache/fory4.6k—~2.9kAutomated safety check: PassApache-2.0
Fory Version Bumpapache/fory4.6k—~1.1kAutomated safety check: PassApache-2.0
Fory Performance Optimizationapache/fory4.6k—~2.2kAutomated safety check: PassApache-2.0
Dbgtheodo-group/debug-that158—~2.2kAutomated safety check: PassMIT
Code Review Excellenceandrew-yangy/gru-ai155—~1.7kAutomated safety check: NotesMIT

Similar skills

  • Fory Release

    apache/fory

    Prepare an Apache Fory release candidate from a clean release branch, including the version bump, RC tag, JVM staging, ASF source artifacts, SVN upload, and vote email.

    4.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Bump Apache Fory release or post-release development versions across Java, Kotlin, Scala, Python, Rust, Go, C++, C, Dart, JavaScript, Swift, integration tests, examples, and source docs.

    4.6k GitHub stars~1.1k tokensUpdated yesterday
    MobileAuto-check passed
  • Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala).

    4.6k GitHub stars~2.2k tokensUpdated yesterday
    MobileAuto-check passed
  • Dbg

    theodo-group/debug-that

    Debug applications using the dbg CLI debugger. An agent skill from theodo-group/debug-that.

    158 GitHub stars~2.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Code Review Excellence

    andrew-yangy/gru-ai

    Provides comprehensive code review guidance for React 19, Vue 3, Rust, TypeScript, Java, Python, and C/C++.

    155 GitHub stars~1.7k tokensUpdated 7 mo ago
    DevelopmentAuto-check: notes
  • Subspace Clients

    dallison/subspace

    Write Subspace clients in C++, Python, Rust, or Java. An agent skill from dallison/subspace.

    104 GitHub stars~2.4k tokensUpdated yesterday
    Backend & APIsAuto-check passed

More from huangruiteng/CS-Notes

All 38 skills in this repo
  • CLI Creator

    huangruiteng/CS-Notes

    Build a composable CLI for Codex from API docs, an OpenAPI spec, existing curl examples, an SDK, a web app, an admin tool, or a local script.

    4k GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Slack

    huangruiteng/CS-Notes

    A skill your agent uses when you need to control Slack from Clawdbot via the slack tool, including reacting to messages or pinning/unpinning items in Slack channels or DMs.

    4k GitHub starsUsed in 10 repos~578 tokens
    Auto-check passed
  • Codex Thread Heartbeat

    huangruiteng/CS-Notes

    Inspect and manage guarded Codex App-native or launchd heartbeats for Codex main control threads.

    4k GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Codex Thread Reader

    huangruiteng/CS-Notes

    Locate and read a Codex thread by a codex thread link, thread id, or rollout path across all local CODEXHOME directories (~/.codex, ~/.codex-gpt, ...).

    4k GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check passed
  • Research Material Scout

    huangruiteng/CS-Notes

    A skill your agent uses when the user asks Codex to research, find learning materials, process "素材:" links, "请你读" / "精读" a material, build a material radar, or use SenSight-like broad information…

    4k GitHub stars~8.6k tokensUpdated 2 days ago
    Auto-check passed
  • GitHub

    huangruiteng/CS-Notes

    Interact with GitHub using the gh CLI. An agent skill from huangruiteng/CS-Notes.

    4k GitHub starsUsed in 28 repos~279 tokens
    Auto-check passed

Questions about Code Optimization

What does Code Optimization do?

Optimize code performance through iterative improvements (max 2 rounds). Code Optimization is an agent skill from huangruiteng/CS-Notes. Optimize code performance through iterative improvements (max 2 rounds).

How do I install Code Optimization in Claude Code?

Run `npx skills add huangruiteng/CS-Notes --skill code-optimization -a claude-code`. Or copy the skill folder (.trae/openclaw-skills/code-optimization in huangruiteng/CS-Notes) into .claude/skills/code-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Code Optimization in Codex?

Run `npx skills add huangruiteng/CS-Notes --skill code-optimization -a codex`. Or copy the skill folder (.trae/openclaw-skills/code-optimization in huangruiteng/CS-Notes) into .agents/skills/code-optimization in your project. Codex loads it when a task matches its description.

Can I use Code Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huangruiteng/CS-Notes --skill code-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-optimization, .gemini/skills/code-optimization, .github/skills/code-optimization and .opencode/skills/code-optimization in your project.

What does Code Optimization need to run?

Going by SKILL.md and its folder, Code Optimization needs the command-line tools its instructions call (python3, javac, java, rustc and go). Our summary lists: Python 3.

Does Code Optimization access the network?

SKILL.md names 4 domains. In commands or code: gcc.gnu.org, intel.com, perf.wiki.kernel.org and bigocheatsheet.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Code Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Code Optimization use?

Code Optimization is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Optimization use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Code Optimization?

Skills that share tags, products or a category with Code Optimization: Fory Release (apache/fory, 4.6k stars), Fory Version Bump (apache/fory, 4.6k stars), Fory Performance Optimization (apache/fory, 4.6k stars) and Dbg (theodo-group/debug-that, 158 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Optimization?

huangruiteng (a GitHub user) maintains it in huangruiteng/CS-Notes, which has 4,000 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on October 6, 2026.

Source: huangruiteng/CS-Notes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.