Agent skill

Conventions Ingestion

by stella in stella/stella

Apply when building or reviewing external ingestion, imports, connector polling, webhooks, extraction workers, sync cursors, checkpoints, or repair jobs.

Apache-2.0Auto-check passedBackend & APIs

Install Conventions Ingestion

skills CLI
$ npx skills add stella/stella --skill conventions-ingestion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install stella/stella conventions-ingestion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/stella/stella.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/conventions-ingestion .claude/skills/conventions-ingestion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
conventions-ingestion
GitHub stars
258
Token cost
~2.9k tokens
SKILL.md length
1,602 words
Files
2
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

Apply when building or reviewing external ingestion, imports, connector polling, webhooks, extraction workers, sync cursors, checkpoints, or repair jobs.

  • Works in 12 steps: Stable identity. Give every source item… → Idempotent persistence. Upsert, claim,… → Explicit outcomes. Distinguish terminal… → …
  • Tasks that involve Webhooks
  • SKILL.md covers Target Property, Required Design, Checkpoint Boundary and Source-Surface Census, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Conventions Ingestion is an agent skill from stella/stella. Apply when building or reviewing external ingestion, imports, connector polling, webhooks, extraction workers, sync cursors, checkpoints, or repair jobs. Enforces replay safety, idempotency, durable progress, and bounded recovery.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Backend & APIs, covering Webhooks. The repository describes itself as: Open-source legal workspace. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Webhooks

Example prompts

  • “/conventions-ingestion”

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Stable identity. Give every source item a stable, tenant-scoped identity.
  2. Idempotent persistence. Upsert, claim, or transition by stable identity.
  3. Explicit outcomes. Distinguish terminal outcomes (applied, unchanged,
  4. Checkpoint last. Advance a cursor, watermark, or checkpoint only after
  5. Atomic database batches. For database-only work, persist the items and
  6. External side effects. Object storage, search indexes, email, and remote
  7. Compare-and-set cursors. Capture the cursor loaded at run start and
  8. Stale-work protection. Mutable inputs need a source version or content
  9. Durable execution. Do not rely on detached promises or process memory
  10. Bounded recovery. Repair scans and list reads use cursor pagination and
  11. Owned schema. The vertical slice owns its source, item, attempt, and
  12. Bounded external calls. Every remote request has an explicit timeout,

What it can do on your machine

Read from SKILL.md and the folder at commit a9618db. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Conventions Ingestion loads about 2.9k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 1,602 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from stella/stella at commit a9618db, republished under its Apache-2.0 licence (© stella). 1,602 words, ~2,914 tokens.

Download SKILL.mdSave it as .claude/skills/conventions-ingestion/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
conventions-ingestion
description
Apply when building or reviewing external ingestion, imports, connector polling, webhooks, extraction workers, sync cursors, checkpoints, or repair jobs. Enforces replay safety, idempotency, durable progress, and bounded recovery.

Replay-Safe Ingestion Conventions

Apply to any workflow that turns external or asynchronous input into durable stella state: paginated imports, connector sync, webhooks, file extraction, queue workers, migrations, and repair scans.

Target Property

A retry, duplicate delivery, worker restart, or overlapping run must converge to the same durable state as one successful run. Idempotency is one ingredient; replay safety also requires correct checkpoint ordering, durable retries, and protection from stale work.

Required Design

  1. Stable identity. Give every source item a stable, tenant-scoped identity. Enforce it with a database unique constraint whose leading columns preserve the tenant or source boundary. Do not rely on a hash collision check alone.
  2. Idempotent persistence. Upsert, claim, or transition by stable identity. Reapplying the same input must not duplicate rows, counters, notifications, or other effects.
  3. Explicit outcomes. Distinguish terminal outcomes (applied, unchanged, deliberately rejected) from retryable failures. A skipped item is terminal only when losing it is intentional and auditable.
  4. Checkpoint last. Advance a cursor, watermark, or checkpoint only after every earlier item is terminal or has a durable retry record. On an ambiguous failure, hold the old checkpoint and replay.
  5. Atomic database batches. For database-only work, persist the items and checkpoint in one short transaction with commitReplaySafeIngestionBatch from apps/api/src/lib/replay-safe-ingestion.ts.
  6. External side effects. Object storage, search indexes, email, and remote APIs cannot join the database transaction. Make the database record the source identity/fingerprint and retry state first; use deterministic object keys or provider idempotency keys. Persist the cursor only after all page work is durable.
  7. Compare-and-set cursors. Capture the cursor loaded at run start and require it in the checkpoint update. A stale run must return the persisted winner, never overwrite newer progress. Public corpus ingestion uses advanceCorpusIngestionCheckpoint from apps/api/src/lib/corpus-ingestion-checkpoint.ts.
  8. Stale-work protection. Mutable inputs need a source version or content fingerprint. A late worker must compare the claimed version before overwriting newer state. AI outputs also include schema, model, prompt, and parser versions in their identity/provenance.
  9. Durable execution. Do not rely on detached promises or process memory for required work. Use a durable queue/outbox and deterministic job identity. Add a bounded repair scan when enqueue and commit cannot be atomic.
  10. Bounded recovery. Repair scans and list reads use cursor pagination and configured limits. Workers remain stateless and safe under concurrency.
  11. Owned schema. The vertical slice owns its source, item, attempt, and failure tables. Shared code provides transaction and identity primitives, not a cross-domain ingestion framework.
  12. Bounded external calls. Every remote request has an explicit timeout, bounded retry policy with jitter, provider-aware rate limiting, and a maximum concurrency. Persist retry state; do not hold a database transaction while waiting on the provider.
  13. Poison-item isolation. One malformed or permanently rejected item must not stall an entire source forever. Persist the item identity, classified terminal/retryable outcome, sanitized error context, and operator-visible repair path before allowing later progress.
  14. Accounted-for source fields. Every field a source states on a page the adapter already fetches is stored, or excluded with the reason. See below.
  15. A publisher is not a court. The deciding court comes from the record: the court code in its ECLI first, then the record's own court field, resolved through the shared resolver (apps/api/src/lib/case-law/cz-ecli-courts.ts for Czech sources). Never a per-adapter constant, even where the portal is believed to serve one court's decisions: portals republish other courts, and the court name is what authority weighting is read off. An inventory declaration must name the field the value is read from, and a value produced from a literal satisfies no declaration. no-literal-decision-court rejects a literal court in a case-law adapter.

Checkpoint Boundary

Direct Drizzle writes to syncCursor are banned by no-direct-ingestion-checkpoint-write. Public corpus cursors go through advanceCorpusIngestionCheckpoint; database-only batches keep the write inside the persistCheckpoint callback passed to commitReplaySafeIngestionBatch. The lint rule enforces the visible boundary; it does not prove that preceding external effects are durable.

Source-Surface Census

A page an adapter never fetched is invisible to any inventory of the pages it did. So the declaration comes in two steps, and the first is the surfaces: SourceAdapter.sourceSurfaces is required, and it is total over the addresses the publisher serves for one decision — SOURCE_SURFACES as a kebab-case list, then the map as const satisfies Record<<that union>, SourceSurfaceDisposition> as the surfaces of a literal written as const satisfies SourceSurfaceCensus.

Each surface is one of three things:

  • storedSourceSurface(part) — fetched with the decision and kept as that envelope part. The conformance suite reads the part back out of an envelope this repository can produce, so an adapter with no such capture cannot declare one;
  • excludedSourceSurface(reason) — the same constructor discipline as the field inventory: a blank reason does not compile, and "not fetched today" is not a reason. A rendering of a payload already kept, a corpus-wide index, a query-scoped export and a page the publisher's robots policy disallows are;
  • backlogSurface(adapter, reason) — it belongs in the row and is not there yet. Only an adapter already on source-surface-backlog-baseline.json can be named, each entry is listed there, and source-surface-census.test.ts fails both ways, so the set only shrinks. Recording a surface deletes its line.
Show full SKILL.md (751 more words)Show less

Source-Field Inventory

A field an adapter never noticed is indistinguishable from one it decided to leave. Case-law adapters therefore declare what their source states, and the declaration is part of the contract: SourceAdapter.sourceFields is required, so an adapter without one does not compile.

An inventory has three parts, in the adapter beside the readers it mirrors:

  • SOURCE_FIELDS, the list of names the source labels on the per-decision pages this adapter parses, as const;
  • a disposition map written as const satisfies Record<<that union>, SourceFieldDisposition>, so the map is total by type and a new name without a decision does not compile. Each entry is { disposition: "stored", target } (a metadata key, a result field, the parsed document, or the row's identity), or excludedSourceField(reason), where the reason says what the field is and why the row does not carry it. "Not read today" is not a reason; duplicate of a stored field, derived elsewhere, no field on the row, and data minimization are. The constructor is the only way to write an exclusion, and a blank reason does not compile;
  • listSourceFields(parts), which reads the stored envelope back and answers what the publisher labelled across it. The whole envelope, not one page: a source states fields on the listing row as well as on the detail payload, and a reader given one of them declares the others out of scope by accident.

source-field-inventory.test.ts drives every registered adapter from the registry: each adapter's fixture is built, its stored envelope goes through its own listSourceFields, every name that comes out must be in the map, and every field the map stores must be on the decision built from that fixture, at the target the disposition names. A field on the page that is in neither set fails with its name, and so does a field the map declares that the envelope never states. Its coverage map is total over the registry, so a source registered without a fixture does not compile.

Refresh a fixture from the live page when the source changes. The suite certifies the adapter against the page it is given, so a fixture that stopped matching the publisher certifies nothing.

Raw Holds Every Fetched Response

An inventory decides what is read; the stored raw decides what can still be read later. Where a source serves one decision across several responses (a detail page beside the document), store all of them, with encodeSourceRawEnvelope and SOURCE_RAW_ENVELOPE_CONTENT_TYPE, naming each part by its role. A raw that holds only the page the parser read makes a field captured later unrecoverable for every stored row: replay can only re-read what was kept.

The envelope holds text. A response the adapter keeps as bytes goes through sourceRawBytes, which the pipeline stores instead of sourceRaw, so an adapter that sets both loses the payload that names the decision.

reparseStoredRaw decodes with decodeSourceRawEnvelope and handles null, which is what a row stored before its adapter had an envelope reads as. The shapes those rows hold are registered per adapter in LEGACY_RAW_SHAPES, so a reader knows which payload it is looking at and a migrated adapter deletes its line. Bump the adapter's entry in PARSER_VERSIONS when the replay's output changes, and state in the pull request what a replay does and does not backfill: rows stored before the change hold what they held, and only a re-crawl adds to them.

Verification

Test the behavior that types and lint cannot prove:

  • replay the same batch and assert the same fixed point;
  • fail item persistence and assert the checkpoint does not advance;
  • fail checkpoint persistence and assert database item writes roll back;
  • deliver duplicates concurrently and assert one durable effect;
  • crash after a remote side effect and before acknowledgment, then replay;
  • finish stale work after a newer version and assert it cannot overwrite;
  • leave an enqueue gap and assert the bounded repair scan finds it;
  • exhaust the retry budget for one poison item and assert later items still reach a durable terminal state without silently dropping the failure;
  • overlap two workers at the provider concurrency limit and assert calls stay bounded and the persisted winner cannot be overwritten;
  • read a source fixture back through the adapter's own listSourceFields and assert every field it states is stored or excluded with a reason.

Prefer invariant and state-machine tests over one example retry.

Existing References

  • Upload finalization: apps/api/src/handlers/uploads/update.ts
  • Case-law ingestion: apps/api/src/handlers/case-law/ingestion/pipeline.ts and the modules under apps/api/src/handlers/case-law/ingestion/pipeline/
  • Legislation ingestion: apps/api/src/handlers/legislation/ingestion.ts
  • Hosted usage webhook deduplication: apps/api/src/lib/hosted-usage-provider/webhook-store.ts

These are examples, not blanket proof: audit each new side effect and checkpoint independently.

© stella, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/conventions-ingestion of stella/stella.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit a9618db

Compare with similar skills

Conventions Ingestion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Conventions Ingestion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Conventions Ingestion this skillstella/stella258—~2.9kAutomated safety check: PassApache-2.0
Novu Design Workflownovuhq/novu40k—~2.6kAutomated safety check: PassCustom licence
Golivemikehasa/golive-skill1.3k—~13kAutomated safety check: NotesMIT
Stripe Appsfossasia/eventyay1.7k1 repos~3.6kAutomated safety check: PassApache-2.0
Dingtalk Messageagentscope-ai/ReMe3.6k—~1.6kAutomated safety check: PassApache-2.0
PR Review Provideryansongda/pay5.4k—~2.4kAutomated safety check: PassMIT

Similar skills

  • Design notification workflows the Novu way — choose channels, set severity, decide when a workflow is critical, configure digests, and route based on subscriber state.

    40k GitHub stars~2.6k tokensUpdated today
    Backend & APIsAuto-check passed
  • Golive

    mikehasa/golive-skill

    Take an agent-written app from repo to live production on the user's OWN accounts, with providers they choose (hosting, database, auth, payments, email, domain/DNS).

    1.3k GitHub stars~13k tokensUpdated 6 days ago
    Backend & APIsAuto-check: notes
  • Stripe Apps

    fossasia/eventyay

    A skill your agent uses when building, modifying, or reviewing a Stripe App — or when the user describes something that implies one (e.g.

    1.7k GitHub starsUsed in 1 repo~3.6k tokens
    Backend & APIsAuto-check passed
  • Dingtalk Message

    agentscope-ai/ReMe

    钉钉消息发送技能。支持企业内部机器人(批量单聊/群聊)和 Webhook 自定义机器人两种接入方式,支持多机器人管理,支持文本、Markdown、链接、ActionCard、FeedCard等多种消息类型。

    3.6k GitHub stars~1.6k tokensUpdated today
    Backend & APIsAuto-check passed
  • PR Review Provider

    yansongda/pay

    A skill your agent uses when reviewing PRs that add or modify a payment Provider in yansongda/pay - covers plugin pipeline, multi-tenant safety, signature verification, docs, and naming conventions.

    5.4k GitHub stars~2.4k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Stripe Best Practices

    kanchengw/cnllm

    Guides Stripe integration decisions — API selection (Checkout Sessions vs PaymentIntents), Connect platform setup (Accounts v2, controller properties), billing/subscriptions, Treasury financial…

    173 GitHub starsUsed in 2 repos~925 tokens
    Backend & APIsAuto-check passed

More from stella/stella

All 24 skills in this repo
  • Plan

    stella/stella

    Create a concise, evidence-backed implementation plan in the repository planning area when the user explicitly asks for a plan.

    258 GitHub stars~917 tokensUpdated today
    Auto-check passed
  • Answer From Sources

    stella/stella

    Answers data-protection (GDPR) questions grounded in the regulation and supervisory guidance, with a citation for every claim.

    258 GitHub stars~735 tokensUpdated today
    Auto-check passed
  • Check Against Rules

    stella/stella

    Reviews a non-disclosure agreement against the firm's NDA checklist and reports findings with citations.

    258 GitHub stars~856 tokensUpdated today
    Auto-check passed
  • Intake To Draft

    stella/stella

    Collects the facts of an unpaid invoice, then drafts a payment demand letter.

    258 GitHub stars~537 tokensUpdated today
    Auto-check passed
  • Conventions Perf

    stella/stella

    Apply when a performance-guard check (network baseline, bundle baseline, DB query count, loader-prefetch lint, RC bailouts) fails or when touching a hot route/endpoint.

    258 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Apply when writing or reviewing React effects in apps/web. An agent skill from stella/stella.

    258 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Conventions Ingestion

What does Conventions Ingestion do?

Apply when building or reviewing external ingestion, imports, connector polling, webhooks, extraction workers, sync cursors, checkpoints, or repair jobs. Conventions Ingestion is an agent skill from stella/stella. Apply when building or reviewing external ingestion, imports, connector polling, webhooks, extraction workers, sync cursors, checkpoints, or repair jobs.

When should I use Conventions Ingestion?

Conventions Ingestion fits situations like: tasks that involve Webhooks.

How do I install Conventions Ingestion in Claude Code?

Run `npx skills add stella/stella --skill conventions-ingestion -a claude-code`. Or copy the skill folder (.agents/skills/conventions-ingestion in stella/stella) into .claude/skills/conventions-ingestion in your project. Claude Code loads it when a task matches its description.

How do I install Conventions Ingestion in Codex?

Run `npx skills add stella/stella --skill conventions-ingestion -a codex`. Or copy the skill folder (.agents/skills/conventions-ingestion in stella/stella) into .agents/skills/conventions-ingestion in your project. Codex loads it when a task matches its description.

Can I use Conventions Ingestion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add stella/stella --skill conventions-ingestion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/conventions-ingestion, .gemini/skills/conventions-ingestion, .github/skills/conventions-ingestion and .opencode/skills/conventions-ingestion in your project.

What does Conventions Ingestion need to run?

SKILL.md names no scripts, command-line tools or credentials: Conventions Ingestion is instructions for the agent only.

Does Conventions Ingestion access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Conventions Ingestion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Conventions Ingestion use?

Conventions Ingestion is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Conventions Ingestion use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Conventions Ingestion?

Skills that share tags, products or a category with Conventions Ingestion: Novu Design Workflow (novuhq/novu, 40k stars), Golive (mikehasa/golive-skill, 1.3k stars), Stripe Apps (fossasia/eventyay, 1.7k stars) and Dingtalk Message (agentscope-ai/ReMe, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Conventions Ingestion?

stella (a GitHub organization) maintains it in stella/stella, which has 258 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 10, 2026.

Source: stella/stella on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.