Agent skill

Analyze Traj

by xlang-ai in xlang-ai/OSWorld-V2

Analyze OSWorld-V2 agent trajectory logs and task results to produce actionable insights.

Apache-2.0Auto-check passedProductivity & Automation

Install Analyze Traj

skills CLI
$ npx skills add xlang-ai/OSWorld-V2 --skill analyze-traj -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xlang-ai/OSWorld-V2 analyze-traj --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xlang-ai/OSWorld-V2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/analyze-traj .claude/skills/analyze-traj && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyze-traj
GitHub stars
359
Used in
1 other repo
Token cost
~298 tokens
SKILL.md length
113 words
Files
3
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze OSWorld-V2 agent trajectory logs and task results to produce actionable insights.

  • The user wants to understand agent performance on OSWorld tasks — including analyzing trajectories
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Reviewing task results
  • Finding error patterns

What it does

Analyze Traj is an agent skill from xlang-ai/OSWorld-V2. Analyze OSWorld-V2 agent trajectory logs and task results to produce actionable insights. Use this skill whenever the user wants to understand agent performance on OSWorld tasks — including analyzing trajectories, reviewing task results, finding error patterns, comparing code vs GUI strategies, identifying which tools/commands the agent used, or deciding which task types to scale up in the benchmark.

Its SKILL.md is about 300 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `prompts/analyze-full-run.md` and `prompts/analyze-single-traj.md`).

It sits in Productivity & Automation. The repository describes itself as: OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. The licence is Apache-2.0.

When your agent uses it

  • The user wants to understand agent performance on OSWorld tasks — including analyzing trajectories
  • Reviewing task results
  • Finding error patterns
  • Comparing code vs GUI strategies

Example prompts

  • “/analyze-traj”

What it can do on your machine

Read from SKILL.md and the folder at commit acdd349. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyze Traj loads about 298 tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 113 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~298

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from xlang-ai/OSWorld-V2 at commit acdd349, republished under its Apache-2.0 licence (© xlang-ai). 113 words, ~298 tokens.

Download SKILL.mdSave it as .claude/skills/analyze-traj/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
analyze-traj
description
Analyze OSWorld-V2 agent trajectory logs and task results to produce actionable insights. Use this skill whenever the user wants to understand agent performance on OSWorld tasks — including analyzing trajectories, reviewing task results, finding error patterns, comparing code vs GUI strategies, identifying which tools/commands the agent used, or deciding which task types to scale up in the benchmark.

If only one task is issued, analyze it directly with instruction: analyze-single-traj.md.

If multiple tasks or a whole results directory are issued, use subagents to analyze them in parallel (one agent for each task). Do not analyze them sequentially by yourself. DO NOT tell it what to do. Just ask the subagent to analyze the task in target directory and use this skill (analyze-traj) to do the analysis. Pass any user instructions to every subagent.

After the per-task reports are ready:

  • Do nothing but report to the user that the analysis is done and where to find the reports.
  • Ask user if they want to synthesize a run-level summary, if yes use: analyze-full-run.md.

© xlang-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .codex/skills/analyze-traj of xlang-ai/OSWorld-V2.

  • SKILL.md
  • prompts/analyze-full-run.md
  • prompts/analyze-single-traj.md

Open the folder on GitHubat commit acdd349

Used in 1 other repository

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in xlang-ai/OSWorld-V2, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Analyze Traj next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyze Traj compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyze Traj this skillxlang-ai/OSWorld-V23591 repos~298Automated safety check: PassApache-2.0
Agent Browserquran/quran.com-frontend-next1.9k40 repos~3.3kAutomated safety check: PassNone
Dependency Watchtelegramdesktop/tdesktop33k1 repos~2.2kAutomated safety check: PassGPL-3.0
Perform Tasktelegramdesktop/tdesktop33k2 repos~3kAutomated safety check: PassGPL-3.0
Brave Searchbadlogic/pi-skills2.6k5 repos~592Automated safety check: PassMIT
Garden Inboxpaperclipai/paperclip99k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Agent Browser

    quran/quran.com-frontend-next

    Automates browser interactions for web testing, form filling, screenshots, and data extraction.

    1.9k GitHub starsUsed in 40 repos~3.3k tokens
    Productivity & AutomationAuto-check passed
  • Dependency Watch

    telegramdesktop/tdesktop

    Audit Telegram Desktop dependencies on freshly fetched origin/dev for releases and security fixes, including upstream lag and backport candidates in patched forks.

    33k GitHub starsUsed in 1 repo~2.2k tokens
    Productivity & AutomationAuto-check passed
  • Perform Task

    telegramdesktop/tdesktop

    Resolve, start or resume, implement, review, test, and publish exactly one existing ai-tdesktop task by short slug or full dated id, including rare blocked retries and split-required results.

    33k GitHub starsUsed in 2 repos~3k tokens
    Productivity & AutomationAuto-check passed
  • Brave Search

    badlogic/pi-skills

    Web search and content extraction via Brave Search API. An agent skill from badlogic/pi-skills.

    2.6k GitHub starsUsed in 5 repos~592 tokens
    Productivity & AutomationAuto-check passed
  • Garden Inbox

    paperclipai/paperclip

    Scan a Paperclip user's Mine inbox, classify reversible archive candidates, request checkbox confirmation, and archive only accepted selections.

    99k GitHub stars~1.1k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Continue

    telegramdesktop/tdesktop

    Continue autonomous Telegram Desktop development from the shared ai-tdesktop repository.

    33k GitHub starsUsed in 2 repos~9.4k tokens
    Productivity & AutomationAuto-check passed

More from xlang-ai/OSWorld-V2

  • Setup Osworld

    xlang-ai/OSWorld-V2

    Provision and verify an OSWorld-V2 checkout after clone. An agent skill from xlang-ai/OSWorld-V2.

    359 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check: notes
  • Migrate Osworld Agent

    xlang-ai/OSWorld-V2

    Migrate an agent from upstream OSWorld into this OSWorld-V2 repository, add matching evaluation entrypoints, and verify the integration.

    359 GitHub starsUsed in 1 repo~546 tokens
    Auto-check passed
  • Analyze Task

    xlang-ai/OSWorld-V2

    Check OSWorld tasks. An agent skill from xlang-ai/OSWorld-V2.

    359 GitHub starsUsed in 1 repo~223 tokens
    Auto-check passed

Questions about Analyze Traj

What does Analyze Traj do?

Analyze OSWorld-V2 agent trajectory logs and task results to produce actionable insights. Analyze Traj is an agent skill from xlang-ai/OSWorld-V2. Analyze OSWorld-V2 agent trajectory logs and task results to produce actionable insights.

When should I use Analyze Traj?

Analyze Traj fits situations like: the user wants to understand agent performance on OSWorld tasks — including analyzing trajectories; reviewing task results; finding error patterns; comparing code vs GUI strategies.

How do I install Analyze Traj in Claude Code?

Run `npx skills add xlang-ai/OSWorld-V2 --skill analyze-traj -a claude-code`. Or copy the skill folder (.codex/skills/analyze-traj in xlang-ai/OSWorld-V2) into .claude/skills/analyze-traj in your project. Claude Code loads it when a task matches its description.

How do I install Analyze Traj in Codex?

Run `npx skills add xlang-ai/OSWorld-V2 --skill analyze-traj -a codex`. Or copy the skill folder (.codex/skills/analyze-traj in xlang-ai/OSWorld-V2) into .agents/skills/analyze-traj in your project. Codex loads it when a task matches its description.

Can I use Analyze Traj in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xlang-ai/OSWorld-V2 --skill analyze-traj -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-traj, .gemini/skills/analyze-traj, .github/skills/analyze-traj and .opencode/skills/analyze-traj in your project.

What does Analyze Traj need to run?

SKILL.md names no scripts, command-line tools or credentials: Analyze Traj is instructions for the agent only.

Does Analyze Traj access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Analyze Traj safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyze Traj use?

Analyze Traj is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyze Traj use?

About 298 tokens (SKILL.md is roughly 1.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyze Traj?

Skills that share tags, products or a category with Analyze Traj: Agent Browser (quran/quran.com-frontend-next, 1.9k stars), Dependency Watch (telegramdesktop/tdesktop, 33k stars), Perform Task (telegramdesktop/tdesktop, 33k stars) and Brave Search (badlogic/pi-skills, 2.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyze Traj?

xlang-ai (a GitHub organization) maintains it in xlang-ai/OSWorld-V2, which has 359 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 1, 2026.

Source: xlang-ai/OSWorld-V2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.