Agent skill

Create Dataflow Approximation

by seqra in seqra/opentaint

Model a method's taint propagation as code-based dataflow approximation and refine it against a test project until the sample passes.

Apache-2.0Auto-check passedSecurity

Install Create Dataflow Approximation

skills CLI
$ npx skills add seqra/opentaint --skill create-dataflow-approximation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install seqra/opentaint create-dataflow-approximation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/seqra/opentaint.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/create-dataflow-approximation .claude/skills/create-dataflow-approximation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
create-dataflow-approximation
GitHub stars
162
Token cost
~1.9k tokens
SKILL.md length
950 words
Files
4 (incl. scripts, references)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Model a method's taint propagation as code-based dataflow approximation and refine it against a test project until the sample passes.

  • Works in 4 steps: Understand the propagation → Write the approximation → Test against the test project → …
  • A dropped method that requires code-based approximation
  • SKILL.md covers Inputs, Workflow, Output and Tracking, plus 1 more section
  • Runs Python scripts from its folder

What it does

Create Dataflow Approximation is an agent skill from seqra/opentaint. Model a method's taint propagation as code-based dataflow approximation and refine it against a test project until the sample passes. Use for a dropped method that requires code-based approximation

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/debugging.md`, `references/java.md` and `scripts/check-test-result.py`).

It sits in Security, covering Static analysis and SAST. The repository describes itself as: The open source taint analysis engine for the AI era. A formal dataflow analysis tool you can customize and self-host, built so AI agents drive your application security analysis… The licence is Apache-2.0.

When your agent uses it

  • A dropped method that requires code-based approximation
  • Tasks that involve Static analysis and SAST

Example prompts

  • “/create-dataflow-approximation”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Understand the propagation
  2. Write the approximation
  3. Test against the test project
  4. Escalate

What it can do on your machine

Read from SKILL.md and the folder at commit f945f92. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Create Dataflow Approximation loads about 1.9k tokens when it runs, and up to ~4.4k if it reads all its reference files. Until then it costs about 57 tokens; SKILL.md has 950 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from seqra/opentaint at commit f945f92, republished under its Apache-2.0 licence (© seqra). 950 words, ~1,868 tokens.

Download SKILL.mdSave it as .claude/skills/create-dataflow-approximation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
create-dataflow-approximation
description
Model a method's taint propagation as code-based dataflow approximation and refine it against a test project until the sample passes. Use for a dropped method that requires code-based approximation
license
Apache-2.0
metadata.author
opentaint
metadata.version
0.3.0

Skill: Create Dataflow Approximation

A dataflow approximation is code that expresses how data moves through a method the analyzer can't trace through — an opaque call where the engine loses taint because it can't see the body. You write a small stand-in that reproduces the method's real propagation from its inputs to its outputs, so the analyzer can follow taint through it. Run it against the prepared test project and refine until the sample passes.

Inputs

Provided by the caller, fall back to the default value when omitted. Ask back only when a required input is missing and has no sensible default

  • project-root (optional) — root of the target project. Opentaint keeps all analysis artifacts under the fixed <project-root>/.opentaint/ directory, so every .opentaint/... path below resolves there. Default: current directory
  • language (required) — target language for this project and language-specific instructions
  • batch (required) — the batch whose .opentaint/tracking/approximations/<batch>.yaml provides the dataflow methods to model and holds tracking state
  • methods (optional) — a specific subset of the batch's dataflow methods to (re)model; default all not yet in build.done

Workflow

1. Understand the propagation

Find and read each dataflow method's real source: take methods not yet in build.done, or the specific methods handed for repair even when already built. Leave built methods outside that explicit subset and their approximation source unchanged. An app-internal method sits in the project's own sources, a library method's source comes from its dependency (the language reference has how to get it). Read it to see how data moves from the method's inputs (receiver, arguments) to its outputs (return value, arguments it writes into, state it stores), gathering the full context needed to understand the function's behavior.

2. Write the approximation

Reproduce that propagation in the language-specific code form under .opentaint/dataflow/<batch>, following the language reference's artifact layout. Cover every assigned callable variant, repair an explicitly handed callable in the existing source, and add new ones there rather than rewriting the file. The engine is field-sensitive — taint is tracked per field — so route data field-to-field exactly as the source does rather than tainting the whole object. The test project's negative samples (if present) verify this by storing taint in one field and reading another, so an over-broad model makes them fire. The concrete constructs and patterns are in the language reference.

3. Test against the test project

Run the approximation test directly as a foreground, blocking command and wait for exit — never background it or use Monitor. Apply this batch's sources and iterate until the samples pass. Feedback loop: a failing sample might be caused by the model's target type/member or signature not matching what the analyzer sees, or by the body not routing taint from the real source to the modeled output — diagnose the mismatch, fix, and re-run, don't rationalize a non-result. When the cause isn't obvious, localize where taint dies with a fact-reachability trace before guessing further per references/debugging.md. On a pass, append the method only if it is not already present in build.done (per Tracking); a repaired method remains recorded there.

4. Escalate

When the sample won't converge after ~3 fixes — whether the trace shows a faithful model still can't propagate (taint dying at a plain instruction the engine should carry through, an engine limitation) or the cause stays unclear — don't add a new method to build.done or alter an existing repaired method's tracking entry. Report it with the brief cause you found (per Output), for the orchestrator to escalate. Don't retry further.

Show full SKILL.md (378 more words)Show less

Output

Artifacts
  • .opentaint/dataflow/<batch> — the language-specific code approximation artifacts that the scan consumes; report the path and the exact test command used
  • the passing methods present in the batch file's build.done (new methods appended; repaired methods already recorded, per Tracking)
Summary
  • the methods modeled and the test status (passing / non-converging)

Tracking

.opentaint/tracking/approximations/<batch>.yaml — one batch's callable classification, <batch> the plan's filename stem. Every callable sits in exactly one verdict bucket, keyed with the exact language-specific method and signature from the plan so distinct variants stay separate:

  • passthrough, dataflow — modeled carriers; each entry { method, signature }
  • skipped — terminal non-carriers; each { method, signature, reason }
  • engine_issues — a separate bucket for carriers the engine provably can't propagate (built but still dropped); each { method, signature, reason }. Terminal and treated just like skipped — the only difference is the reason. merge-skipped carries it into skipped.yaml as its own engine_issues group alongside the regular skipped methods.

dependencies lists the dependency identifiers a dataflow test project needs. The build block tracks the build — test_project records each dataflow method's test-project status (done if a sample was written into the batch's test project, failed if none could be written so the method was excluded from it), and done holds the finished { method, signature }. Keep it clear from comments

yaml
passthrough:
  - { method: "<qualified-member-a>", signature: "<language-signature-a>" }
dataflow:
  - { method: "<qualified-member-b>", signature: "<language-signature-b>" }
skipped:
  - { method: "<qualified-member-c>", signature: "<language-signature-c>", reason: "retains none of its input data" }
engine_issues: []
dependencies: []
build:
  test_project:
    - { method: "<qualified-member-b>", signature: "<language-signature-b>", status: done }
  done: []

This skill appends each method whose sample passes to build.done as { method, signature } when absent. A newly assigned method that still fails, or one the engine provably can't propagate, stays out and is reported (per Output); an explicitly repaired method leaves its existing entry unchanged. Don't touch the classification buckets (passthrough/dataflow/skipped/engine_issues) or edit an entry already in build.done.

Constraints

OpenTaint is a whole-program, interprocedural, field-sensitive alias analysis engine. It already propagates through visible application code, calls, aliases, and individual fields; custom rules and approximations model only the assigned source, sink, or opaque-method boundary. Compile-time constants and literals carry no taint, so a source or carrier whose output is only a constant introduces nothing.

  • Verify only with the approximation test on the test project
  • The test project's sample sources are a fixed input — never edit or recompile them to force a pass; if a faithful model can't pass, leave the method out and report it (per Output)
  • Model every dataflow method and overload the batch lists, not only the ones you have a sample for

© seqra, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/create-dataflow-approximation of seqra/opentaint.

  • SKILL.md
  • references/debugging.md
  • references/java.md
  • scripts/check-test-result.py

Open the folder on GitHubat commit f945f92

Compare with similar skills

Create Dataflow Approximation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Create Dataflow Approximation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Create Dataflow Approximation this skillseqra/opentaint162—~1.9kAutomated safety check: PassApache-2.0
Semgrepvigolium/piolium1401 repos~2.4kAutomated safety check: NotesMIT
C To AstNarwhal-Lab/MagicSkills316—~1.1kAutomated safety check: PassMIT
Semgrep Security Scantrailofbits/skills7.4k—~3.7kAutomated safety check: NotesCC-BY-SA-4.0
LLM Sast ScannerSunWeb3Sec/llm-sast-scanner286—~6.2kAutomated safety check: PassNone
Sast SemgrepAgentSecOps/SecOpsAgentKit2202 repos~2.4kAutomated safety check: PassCustom licence

Similar skills

  • Semgrep

    vigolium/piolium

    Run Semgrep static analysis scan on a codebase using parallel subagents.

    140 GitHub starsUsed in 1 repo~2.4k tokens
    SecurityAuto-check: notes
  • C To Ast

    Narwhal-Lab/MagicSkills

    Parse C source code into an Abstract Syntax Tree (AST). An agent skill from Narwhal-Lab/MagicSkills.

    316 GitHub stars~1.1k tokensUpdated 6 mo ago
    SecurityAuto-check passed
  • Semgrep Security Scan

    trailofbits/skills

    Official

    Detects languages, proposes rulesets for approval, then runs the approved Semgrep scan across a codebase and merges the output into one SARIF file.

    7.4k GitHub stars~3.7k tokensUpdated today
    SecurityAuto-check: notes
  • LLM Sast Scanner

    SunWeb3Sec/llm-sast-scanner

    General-purpose Static Application Security Testing (SAST) skill for code vulnerability analysis.

    286 GitHub stars~6.2k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Sast Semgrep

    AgentSecOps/SecOpsAgentKit

    Static application security testing (SAST) using Semgrep for vulnerability detection, security code review, and secure coding guidance with OWASP and CWE framework mapping.

    220 GitHub starsUsed in 2 repos~2.4k tokens
    SecurityAuto-check passed
  • Triage Codeql

    netdata/netdata

    Inspect, review or triage GitHub Code Scanning alerts, including CodeQL findings; apply verified dismissals when authorized.

    81k GitHub stars~1.8k tokensUpdated today
    SecurityAuto-check: notes

More from seqra/opentaint

All 16 skills in this repo
  • Analyze an OpenTaint scan's dropped external methods and decide which of them are propagators and optionally sinks.

    162 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Orchestrate Stage

    seqra/opentaint

    Run one stage of the OpenTaint pipeline by coordinating leaf subagents and deterministic joins.

    162 GitHub stars~806 tokensUpdated today
    Auto-check passed
  • Appsec Agent

    seqra/opentaint

    Run an end-to-end OpenTaint application-security analysis while owning the long project build and scans and delegating each other pipeline stage.

    162 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Build Project

    seqra/opentaint

    Build a target project into an opentaint project model. An agent skill from seqra/opentaint.

    162 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Model a method's taint propagation as a passThrough approximation.

    162 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Create Rule

    seqra/opentaint

    Author and verify an OpenTaint rule. An agent skill from seqra/opentaint.

    162 GitHub stars~2.6k tokensUpdated today
    Auto-check passed

Categories

Questions about Create Dataflow Approximation

What does Create Dataflow Approximation do?

Model a method's taint propagation as code-based dataflow approximation and refine it against a test project until the sample passes. Create Dataflow Approximation is an agent skill from seqra/opentaint. Model a method's taint propagation as code-based dataflow approximation and refine it against a test project until the sample passes.

When should I use Create Dataflow Approximation?

Create Dataflow Approximation fits situations like: A dropped method that requires code-based approximation; tasks that involve Static analysis and SAST.

How do I install Create Dataflow Approximation in Claude Code?

Run `npx skills add seqra/opentaint --skill create-dataflow-approximation -a claude-code`. Or copy the skill folder (skills/create-dataflow-approximation in seqra/opentaint) into .claude/skills/create-dataflow-approximation in your project. Claude Code loads it when a task matches its description.

How do I install Create Dataflow Approximation in Codex?

Run `npx skills add seqra/opentaint --skill create-dataflow-approximation -a codex`. Or copy the skill folder (skills/create-dataflow-approximation in seqra/opentaint) into .agents/skills/create-dataflow-approximation in your project. Codex loads it when a task matches its description.

Can I use Create Dataflow Approximation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add seqra/opentaint --skill create-dataflow-approximation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-dataflow-approximation, .gemini/skills/create-dataflow-approximation, .github/skills/create-dataflow-approximation and .opencode/skills/create-dataflow-approximation in your project.

What does Create Dataflow Approximation need to run?

Going by SKILL.md and its folder, Create Dataflow Approximation needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Create Dataflow Approximation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Create Dataflow Approximation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Create Dataflow Approximation use?

Create Dataflow Approximation is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Create Dataflow Approximation use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Create Dataflow Approximation?

Skills that share tags, products or a category with Create Dataflow Approximation: Semgrep (vigolium/piolium, 140 stars), C To Ast (Narwhal-Lab/MagicSkills, 316 stars), Semgrep Security Scan (trailofbits/skills, 7.4k stars) and LLM Sast Scanner (SunWeb3Sec/llm-sast-scanner, 286 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Create Dataflow Approximation?

seqra (a GitHub organization) maintains it in seqra/opentaint, which has 162 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 7, 2026.

Source: seqra/opentaint on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.