Agent skill

Workflow Authoring

by letta-ai in letta-ai/letta-code

Reference for writing a Workflow tool script (script API and gotchas, pipeline-vs-barrier rules, quality patterns, worked examples).

Apache-2.0Auto-check passedAgent Workflows

Install Workflow Authoring

skills CLI
$ npx skills add letta-ai/letta-code --skill workflow-authoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install letta-ai/letta-code workflow-authoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/letta-ai/letta-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/builtin/workflow-authoring .claude/skills/workflow-authoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
workflow-authoring
GitHub stars
3.6k
Token cost
~3.7k tokens
SKILL.md length
1,928 words
Files
1
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reference for writing a Workflow tool script (script API and gotchas, pipeline-vs-barrier rules, quality patterns, worked examples).

  • Agent Workflows work in your project
  • SKILL.md covers What a subagent is, Script body hooks, Pipeline vs barrier and Quality patterns — pick by…, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Workflow Authoring is an agent skill from letta-ai/letta-code. Reference for writing a Workflow tool script (script API and gotchas, pipeline-vs-barrier rules, quality patterns, worked examples). Load before authoring a script for a workflow the user already opted into; it does not itself authorize running one.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows. The repository describes itself as: Stateful agents that are like people, with memory, identity, and the ability to learn and adapt. The licence is Apache-2.0.

When your agent uses it

  • Agent Workflows work in your project

Example prompts

  • “/workflow-authoring”

What it can do on your machine

Read from SKILL.md and the folder at commit 253a3bc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Workflow Authoring loads about 3.7k tokens when it runs. Until then it costs about 67 tokens; SKILL.md has 1,928 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from letta-ai/letta-code at commit 253a3bc, republished under its Apache-2.0 licence (© letta-ai). 1,928 words, ~3,709 tokens.

Download SKILL.mdSave it as .claude/skills/workflow-authoring/SKILL.md (or your agent's skills folder).
name
workflow-authoring
description
Reference for writing a Workflow tool script (script API and gotchas, pipeline-vs-barrier rules, quality patterns, worked examples). Load before authoring a script for a workflow the user already opted into; it does not itself authorize running one.

Workflow authoring reference

A workflow structures work across many subagents — to be comprehensive (decompose and cover in parallel), to be confident (independent perspectives and adversarial checks before committing), or to take on scale one context can't hold (migrations, audits, broad sweeps). The script is where you encode that structure: what fans out, what verifies, what synthesizes.

When you do call the Workflow tool, the right move is often hybrid: scout inline first (list the files, scope the diff, find the modules) to discover the work-list, then call Workflow to pipeline over it, passing the list via args. You don't need to know the shape before the task — only before the orchestration step.

Common single-phase workflows you can chain across turns:

  • Understand — parallel readers over relevant subsystems → structured map
  • Design — judge panel of N independent approaches → scored synthesis
  • Review — dimensions → find → adversarially verify
  • Research — multi-modal sweep → deep-read → synthesize
  • Migrate — discover sites → transform each → verify

For larger work, run several in sequence — read each result before deciding the next phase. You stay in the loop; each workflow is one well-scoped fan-out.

Every script must begin with export const meta = {...}:

export const meta = {
  name: 'find-flaky-tests',                       // kebab-case, required
  description: 'Find flaky tests and propose fixes',  // required
  phases: [                                       // one entry per phase() call
    { title: 'Scan', detail: 'grep test logs for retries' },
    { title: 'Fix', detail: 'one agent per flaky test' },
  ],
}
// script body starts here — use agent()/parallel()/pipeline()/phase()/log()

The meta object must be a PURE LITERAL — no variables, function calls, spreads, or template interpolation.

What a subagent is

Every agent() call is one agent-free ephemeral conversation linked to the invoking parent agent (is_subagent: true, named after the call's label). It starts with nothing but its prompt: no memory, no conversation context, no skills. Put ALL context a stage needs in the prompt — file paths, the rule it should apply, what shape to return.

Subagents are told their final text IS the return value (not a human-facing message), so they return raw data. Workers inherit the invoking session's tools and permission mode; use allowedTools to restrict a workflow or stage. For stages whose input is entirely in the prompt (synthesis, judging, scoring) pass allowedTools: [] — a model that can still read files tends to wander, and a subagent that re-issues an identical tool call three times is stopped and rejects with the failure detail.

Their model defaults to the invoking conversation's model. opts.model (or the tool's model input) accepts any handle or alias listed by letta model list; an unknown value rejects that call. Use a cheaper model for mechanical stages only when you know a valid handle.

Use the invoking backend for workflow workers. Local execution requires an Agent SDK and App Server with local conversation support; decide() still uses the Cloud decisions service.

Script body hooks

  • agent(prompt, opts?) → Promise. Spawn one subagent. Resolves to its final text, or with schema (JSON Schema) to a validated object. Prefer schema for shaped results: invalid or missing output is retried, then rejects with validation detail in the error and journal. json: true still parses without validating; schema wins if both are set. Await each call; failures reject with the cause, callIndex, and conversationId when available. Catch errors explicitly to recover. Options: label (display name), phase (progress group — use this inside concurrent stages), schema, json, model, effort ('low' for mechanical stages, higher for the hardest verify/judge stages), allowedTools, systemPrompt (extra system prompt for this subagent), timeoutMs (default 10 minutes), maxToolCalls (positive safe integer; default 1000 unique tool calls for this subagent), conversationId (resume a worker — see "Diagnosing a run").

  • pipeline(items, stage1, stage2, ...) → run each item through all stages independently, NO barrier between stages. Item A can be in stage 3 while item B is still in stage 1. This is the DEFAULT for multi-stage work. Wall-clock = slowest single-item chain, not sum-of-slowest-per-stage. Every stage callback receives (prevResult, originalItem, index) — use originalItem/index in later stages to label work without threading context through stage 1's return value. A stage that throws skips that item's remaining stages; siblings finish, then the helper rejects.

  • parallel(thunks) → run zero-arg functions concurrently. This is a BARRIER: it awaits all thunks, then rejects if any threw. For best-effort processing, catch errors inside each thunk and return explicit success/error results. Use ONLY when you genuinely need all results together.

  • phase(title) — start a new phase; subsequent agent() calls are grouped under this title in progress output.

  • log(message) — emit a progress message to the user.

  • decide(state, questions, opts?) → Promise; not a subagent call. Ask a calibrated Jev model (chosen internally — no model option) questions about state. questions is a non-empty OBJECT keyed by question id — never an array. Each question needs instructions and a type: choice (criteria map of id → description, ≤255), score (criteria array of strings, unlike questions), noul (no criteria). Returns a response whose answers are keyed by the same ids and calibrated, or null after one retried invalid answer — guard if (!call); API errors throw.

    const call = await decide(evidence, {
      behavior: { type: 'choice', instructions: 'Is this behavior a bug?',
        criteria: { bad: 'Wrong or harmful', not_bad: 'Expected or harmless' } },
    })
    const verdict = call?.answers?.behavior?.choice  // 'bad' | 'not_bad'
    
  • tools.mcp__server__tool(args?) → Promise; not a subagent call. Calls one of the invoking agent's MCP tools directly (find names with letta mcp search / letta mcp tools before writing the script). Each call goes through the same permission rules as a normal MCP tool call; a tool that would ask for approval throws instead, so the user must allow it first. Resolves to the tool's structuredContent, else its text output (parsed when it is JSON), else the raw content array. Throws when the tool is unavailable, not allowed, or reports an error; parallel() and pipeline() propagate that rejection. Use it for deterministic fetches and writes whose arguments the script already knows; use agent() when a step needs judgment.

    const issues = await tools.mcp__linear__list_issues({ team: 'LET', limit: 20 })
    
  • args — the value passed as the tool's args input, verbatim. Pass arrays/objects as actual JSON values, NOT as a JSON-encoded string.

decide() sees only the state you pass; it cannot read files. When compact, bounded items are already prepared, pass them via args and call decide() on each directly — don't spawn agent() readers just to relay inputs. Raw large traces don't belong in args: prepare bounded state that preserves the user instruction, observed action, and outcome, and mark what was omitted.

Scripts are plain JavaScript, NOT TypeScript — type annotations, interfaces, and generics fail to parse. The script body runs in an async context — use await directly and return the final result. Standard JS built-ins (JSON, Math, Array, etc.) are available; the hooks are the only globals provided. The script runs inside the CLI process with the CLI's own privileges (the vm context is a scope, not a security boundary), and the user approves it by reading it. Keep the script to orchestration: decide what runs and combine results. Direct MCP calls belong in tools.mcp__*(); other reading, searching, and writing belongs in subagents, where the tool allowlist applies.

Show full SKILL.md (622 more words)Show less

Pipeline vs barrier

DEFAULT TO pipeline(). Only reach for a barrier (parallel between stages) when stage N needs cross-item context from all of stage N−1:

  • Dedup/merge across the full result set before expensive downstream work
  • Early-exit if the total count is zero ("0 bugs found → skip verification")
  • Stage N's prompt references "the other findings" for comparison

A barrier is NOT justified by:

  • "I need to flatten/map/filter first" — do it inside a pipeline stage: pipeline(items, stageA, r => transform([r]).flat(), stageB)
  • "The stages are conceptually separate" — that's what pipeline() models. Separate stages ≠ synchronized stages.
  • "It's cleaner code" — barrier latency is real. If 5 finders run and the slowest takes 3× the fastest, a barrier wastes 2/3 of the fast finders' idle time.

agent() and decide() share one pool of concurrent slots per run (maxConcurrent, default 16) — excess calls queue and run as slots free up, so passing 100 items is fine. Each has its own 1000-call lifetime backstop. A single parallel()/pipeline() call accepts at most 4096 items.

When a barrier IS correct — dedup across all findings before expensive verification:

const all = await parallel(DIMENSIONS.map(d => () => agent(d.prompt, {json: true})))
const deduped = dedupeByFileAndLine(all.filter(Boolean).flatMap(r => r.findings ?? []))  // needs ALL at once
const verified = await parallel(deduped.map(f => () => agent(verifyPrompt(f), {json: true})))

Loop-until-count pattern — accumulate to a target:

const bugs = []
let rounds = 0
while (bugs.length < 10 && rounds++ < 5) {
  const result = await agent('Find bugs in this codebase. Reply {"bugs": [{file, line, summary}]}', {json: true})
  bugs.push(...(result?.bugs ?? []))
  log(`${bugs.length}/10 found`)
}

Always bound a loop with a round counter as well as the target: a subagent that keeps returning nothing would otherwise run to the 1000-agent cap.

Composing patterns — exhaustive review (find → dedup vs seen → diverse-lens panel → loop-until-dry):

const seen = new Set(), confirmed = []
let dry = 0
while (dry < 2) {                                              // loop-until-dry
  const found = (await parallel(FINDERS.map(f => () =>          // barrier: collect all finders this round
    agent(f.prompt, {phase: 'Find', json: true})))).filter(Boolean).flatMap(r => r.bugs ?? [])
  const fresh = found.filter(b => !seen.has(key(b)))           // dedup vs ALL seen — plain code, not an agent
  if (!fresh.length) { dry++; continue }
  dry = 0; fresh.forEach(b => seen.add(key(b)))
  const judged = await parallel(fresh.map(b => () =>           // every fresh bug judged concurrently...
    parallel(['correctness','security','repro'].map(lens => () =>   // ...each by 3 distinct lenses
      agent(`Judge "${b.desc}" via the ${lens} lens — real? Reply {"real": boolean}`, {phase: 'Verify', json: true, allowedTools: []})))
      .then(vs => ({ b, real: vs.filter(Boolean).filter(v => v.real).length >= 2 }))))
  confirmed.push(...judged.filter(v => v.real).map(v => v.b))
}
return confirmed
// dedup vs `seen`, NOT `confirmed` — else judge-rejected findings reappear every round and it never converges.

Quality patterns — pick by task and compose freely

  • Adversarial verify: spawn N independent skeptics per finding, each prompted to REFUTE. Kill if ≥majority refute. Prevents plausible-but-wrong findings from surviving.

    const votes = await parallel(Array.from({length: 3}, () => () =>
      agent(`Try to refute: ${claim}. Default to refuted=true if uncertain. Reply {"refuted": boolean}`, {json: true})))
    const survives = votes.filter(Boolean).filter(v => !v.refuted).length >= 2
    
  • Perspective-diverse verify: when a finding can fail in more than one way, give each verifier a distinct lens (correctness, security, perf, does-it-reproduce) instead of N identical refuters — diversity catches failure modes redundancy can't.

  • Judge panel: generate N independent attempts from different angles (e.g. MVP-first, risk-first, user-first), score with parallel judges, synthesize from the winner while grafting the best ideas from runners-up. Beats one-attempt-iterated when the solution space is wide.

  • Loop-until-dry: for unknown-size discovery (bugs, issues, edge cases), keep spawning finders until K consecutive rounds return nothing new. Simple counters (while count < N) miss the tail.

  • Multi-modal sweep: parallel agents each searching a different way (by-container, by-content, by-entity, by-time). Each is blind to what the others surface; useful when one search angle won't find everything.

  • Completeness critic: a final agent that asks "what's missing — modality not run, claim unverified, source unread?" What it finds becomes the next round of work.

  • No silent caps: if a workflow bounds coverage (top-N, no-retry, sampling), log() what was dropped — silent truncation reads as "covered everything" when it didn't.

Scale to what the user asked for. "find any bugs" → a few finders, single-vote verify. "thoroughly audit this" or "be comprehensive" → larger finder pool, 3–5 vote adversarial pass, synthesis stage. When unsure, lean toward thoroughness for research/review/audit requests and toward brevity for quick checks.

Diagnosing a run

Every run persists its script, args, and a journal.jsonl with one line per completed subagent call (prompt, outcome, conversation id) under ~/.letta/workflows/executions/<id>/; the tool result names the paths. Before diagnosing why a workflow returned an empty or unexpected result, read that journal — it records each agent's actual return value and failure detail. A failed run is not resumable: fix the script and launch it again.

The script is never replayed, but one worker can continue: agent(prompt, {conversationId}) re-prompts it with history and model intact, using a journal ID and only once its last Run is terminal. Local workers can only continue inside the same workflow execution that observed their completed turn. Tools inherit the invoking session unless allowedTools is supplied; schema and effort are chosen per turn. Needs SDK 0.8.20+.

© letta-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/skills/builtin/workflow-authoring of letta-ai/letta-code.

Open the folder on GitHubat commit 253a3bc

Compare with similar skills

Workflow Authoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Workflow Authoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Workflow Authoring this skillletta-ai/letta-code3.6k—~3.7kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers297k3 repos~1.7kAutomated safety check: PassMIT
Grillingpietheinstrengholt/rssmonster56431 repos~510Automated safety check: PassMIT
CodeGraph Agent Evalcolbymchenry/codegraph74k—~950Automated safety check: PassMIT
Improvefossasia/eventyay-interpretation1.6k10 repos~3.7kAutomated safety check: WarnMIT
Kayba Stage 1 API Analysiskayba-ai/agentic-context-engine2.6k—~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    297k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Grilling

    pietheinstrengholt/rssmonster

    Grill the user relentlessly about a plan, decision, or idea.

    564 GitHub starsUsed in 31 repos~510 tokens
    Agent WorkflowsAuto-check passed
  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    74k GitHub stars~950 tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Improve

    fossasia/eventyay-interpretation

    Survey any codebase as a senior advisor and produce prioritized, self-contained implementation plans for OTHER models/agents to execute.

    1.6k GitHub starsUsed in 10 repos~3.7k tokens
    Agent WorkflowsAuto-check: warnings
  • Kayba Stage 1 API Analysis

    kayba-ai/agentic-context-engine

    Fetch pre-computed insights from the Kayba API and build a structured summary.

    2.6k GitHub stars~1.1k tokensUpdated 15 days ago
    Agent WorkflowsAuto-check passed
  • Agent QA Authoring

    vostride/agent-qa

    A skill your agent uses when creating, editing, validating, or running agent-qa tests, suites, or hooks.

    903 GitHub stars~569 tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed

More from letta-ai/letta-code

All 25 skills in this repo
  • Creating Skills

    letta-ai/letta-code

    Guide for creating effective skills. An agent skill from letta-ai/letta-code.

    3.6k GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Generating Mod Envs

    letta-ai/letta-code

    Generates and reviews mod learning env JSON files for Letta Code local mods.

    3.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Initializing Memory

    letta-ai/letta-code

    Comprehensive guide for initializing or reorganizing agent memory.

    3.6k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Self Configuration

    letta-ai/letta-code

    Inspect or modify Letta Code's own memory, model, context window, system prompt, compaction, permissions, toolsets, mods, skills, channels, schedules, agent secrets, and local runtime settings.

    3.6k GitHub stars~7k tokensUpdated today
    Auto-check passed
  • Browser Use

    letta-ai/letta-code

    Control a real browser to navigate pages, click, type, fill forms, inspect rendered UI, take screenshots, or record video.

    3.6k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Creating Mods

    letta-ai/letta-code

    Creates and edits trusted local Letta Code mods, including tools, slash commands, local-only model providers, lifecycle/turn events, scoped conversation helpers, panels, and capability-gated behavior.

    3.6k GitHub stars~2.5k tokensUpdated today
    Auto-check passed

Questions about Workflow Authoring

What does Workflow Authoring do?

Reference for writing a Workflow tool script (script API and gotchas, pipeline-vs-barrier rules, quality patterns, worked examples). Workflow Authoring is an agent skill from letta-ai/letta-code. Reference for writing a Workflow tool script (script API and gotchas, pipeline-vs-barrier rules, quality patterns, worked examples).

When should I use Workflow Authoring?

Workflow Authoring fits situations like: agent Workflows work in your project.

How do I install Workflow Authoring in Claude Code?

Run `npx skills add letta-ai/letta-code --skill workflow-authoring -a claude-code`. Or copy the skill folder (src/skills/builtin/workflow-authoring in letta-ai/letta-code) into .claude/skills/workflow-authoring in your project. Claude Code loads it when a task matches its description.

How do I install Workflow Authoring in Codex?

Run `npx skills add letta-ai/letta-code --skill workflow-authoring -a codex`. Or copy the skill folder (src/skills/builtin/workflow-authoring in letta-ai/letta-code) into .agents/skills/workflow-authoring in your project. Codex loads it when a task matches its description.

Can I use Workflow Authoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add letta-ai/letta-code --skill workflow-authoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/workflow-authoring, .gemini/skills/workflow-authoring, .github/skills/workflow-authoring and .opencode/skills/workflow-authoring in your project.

What does Workflow Authoring need to run?

SKILL.md names no scripts, command-line tools or credentials: Workflow Authoring is instructions for the agent only.

Does Workflow Authoring access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Workflow Authoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Workflow Authoring use?

Workflow Authoring is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Workflow Authoring use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Workflow Authoring?

Skills that share tags, products or a category with Workflow Authoring: Diagnosing Superpowers Sessions (obra/superpowers, 297k stars), Grilling (pietheinstrengholt/rssmonster, 564 stars), CodeGraph Agent Eval (colbymchenry/codegraph, 74k stars) and Improve (fossasia/eventyay-interpretation, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Workflow Authoring?

letta-ai (a GitHub organization) maintains it in letta-ai/letta-code, which has 3,552 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 9, 2026.

Source: letta-ai/letta-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.