Agent skill

Backend Architecture

by overmind-core in overmind-core/overmind

Deeper backend map of overbae — module layout, celery queue topology, the span-only tracing model and its API surface, capabilities and toml sync, behaviour-keyed scoring, auth and guests, model…

AGPL-3.0Auto-check passedAI & LLM Engineering

Install Backend Architecture

skills CLI
$ npx skills add overmind-core/overmind --skill backend-architecture -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install overmind-core/overmind backend-architecture --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/overmind-core/overmind.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/backend-architecture .claude/skills/backend-architecture && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
backend-architecture
GitHub stars
597
Token cost
~10k tokens
SKILL.md length
5,153 words
Files
1
Skills in repo
20
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Deeper backend map of overbae — module layout, celery queue topology, the span-only tracing model and its API surface, capabilities and toml sync, behaviour-keyed scoring, auth and guests, model…

  • Navigating unfamiliar backend subsystems
  • SKILL.md covers Layout, Celery topology, Tracing and Evaluation, plus 6 more sections
  • Calls make and modal; needs CLERK_API_SECRET_KEY and VITE_CLERK_PUBLISHABLE_KEY
  • Wiring cross-subsystem behavior

What it does

Backend Architecture is an agent skill from overmind-core/overmind. Deeper backend map of overbae — module layout, celery queue topology, the span-only tracing model and its API surface, capabilities and toml sync, behaviour-keyed scoring, auth and guests, model serving and base weights. Use when navigating unfamiliar backend subsystems or wiring cross-subsystem behavior.

Its SKILL.md is about 10k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Backend development, Background jobs and LLM inference and serving. It works with Model Context Protocol. The repository describes itself as: The platform for continuously improving AI agents. The licence is AGPL-3.0.

When your agent uses it

  • Navigating unfamiliar backend subsystems
  • Wiring cross-subsystem behavior

Example prompts

  • “/backend-architecture”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 3dec73c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make
    • modal

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CLERK_API_SECRET_KEY
    • VITE_CLERK_PUBLISHABLE_KEY
    • STRIPE_SECRET_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Backend Architecture loads about 10k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 5,153 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from overmind-core/overmind at commit 3dec73c, republished under its AGPL-3.0 licence (© overmind-core). 5,153 words, ~10,261 tokens.

Download SKILL.mdSave it as .claude/skills/backend-architecture/SKILL.md (or your agent's skills folder).
name
backend-architecture
description
Deeper backend map of overbae — module layout, celery queue topology, the span-only tracing model and its API surface, capabilities and toml sync, behaviour-keyed scoring, auth and guests, model serving and base weights. Use when navigating unfamiliar backend subsystems or wiring cross-subsystem behavior.

overbae backend map

Single Django app overbae, project-scoped tenancy.

Layout

  • overbae/api/ — DRF views/serializers, one module per surface (eval_views.py, datasets.py, optimizer.py, otlp.py, billing.py, mcp.py, …). views.py and serializers.py are the legacy monoliths — new surfaces get their own module. Global exception handler returns {detail, code, error_id}.
  • overbae/models/ — split by domain (iam.py, capabilities.py, traces.py, evaluation.py, finetuning.py, datasets.py, optimizer.py, billing.py, …).
  • overbae/services/ — business logic; largest subtrees: eval/ (rubric/judging/cascade/runner), datasets/ (Parquet store, landing, cell runner, contract, alignment, diff, use), mcp/ (Streamable HTTP MCP server — procedure in the mcp skill), codebase/, scan/ (AI-surface extraction, footprint matcher, scan application), capabilities/ (identity, lifecycle, graph).
  • overbae/tasks/ — Celery tasks, roughly one module per feature; utils/task_lock.py for locking.
  • overbae/modal/ — Modal.com GPU workers (vLLM serving, SFT, PII NER).
  • overbae/management/commands/ — backfills and syncs.

Celery topology

Five workers, six queues. Workers are resource profiles; queues are fairness classes.

WorkerPoolConcQueuesHolds
controlthreads8controlorchestration, chord callbacks, FSM advances, beat
iothreads24io,io_tracesjudges, eval launch, live trace scoring, connector polling, rebinding
batchprefork6batchsample generation, connector chunks
landingprefork1landingsource import and split publication
interactiveprefork4interactiveworkshop cell runs and agent turns

Three constraints set the worker split. Only prefork enforces time_limit and revoke(terminate=True), so every time-limited task routes to batch, landing or interactive. A loaded prefork child costs hundreds of MB, so wide fan-out cannot be prefork. Orchestration holds its own lane because a chord callback stuck behind work never finalises its run. run_eval_run and preload_capability_eval_set route to io because Modal warm-up and judge authoring block for minutes; the import and workshop reconcilers and queue metrics run on control with short expiries and are dropped when it is full.

io_traces is a second queue on the io worker, not a second worker: one worker over both round-robins (kombu's redis default), so an unbounded trace-scoring burst cannot queue ahead of user-started eval scoring.

interactive sets --prefetch-multiplier=1. batch and landing also disable Redis prefetch so busy processes do not reserve waiting work. Landing workers consume only landing; raising bulk concurrency cannot protect import latency when every bulk slot is occupied.

Evaluation preparation uses EvalGenerationWork receipts and a global scheduler row lock. EVAL_MAX_IN_FLIGHT bounds admitted provider work (default 12), with EVAL_MAX_IN_FLIGHT_PER_RUN (default 2) and least-recently-admitted project/run fairness. Waiting samples remain in Postgres; completion and a 30-second control tick replenish available slots. Completion wakeups coalesce through a cache key, so at most one dispatcher waits on the lock. Failed/cancelled parents cannot claim or continue generation. Lost started attempts become explicit unknown outcomes and are not automatically repeated.

Chat HTTP clients enforce one cancellable elapsed deadline across request bodies, SSE heartbeats and retry waits. They close their connections on expiration; a socket read timeout alone is not a total time bound.

Hosted commands come from docker/worker-topology.json. Deployment runs migrations, verifies a healthy landing-only consumer on the new image, then updates workers, beat and API. Landing and interactive workers run with their queue as the Celery hostname, and overbae.worker_health <queue> checks that the consumer reads only its queue and registers that queue's task; a rollout of either waits for that health check. scripts/plan_landing_capacity.py generates the reviewed ECS, IAM, bounded scale-out and CloudWatch queue-age/newly-blocked-import/monitor-health plan. NewlyBlockedImports counts imports blocked in the last 5 minutes by a system cause (queue_timeout, worker_timeout, dispatch_failed, attempts_exhausted); a user-data failure does not raise it. Scale-in requires an explicit idle drain; automatic scale-in is disabled.

Routing lives in CELERY_TASK_ROUTES and must stay in sync with make worker (one process standing in for the whole fleet, so its -Q lists every queue) , docker-compose and the deployment command manifest. tests/test_celery_topology.py enforces it, and asserts no time-limited task lands on a threads lane. Workers hot-restart via watchmedo on .py changes.

Tracing

Span-only: there is no Trace table. A trace is the set of spans sharing a trace_id; the root span has parent_span_id IS NULL. span_id is the primary key; trace_id is an indexed 32-char hex CharField.

Span carries OTel-native fields plus platform linkage:

  • Identity: span_id, trace_id, parent_span_id
  • Classification: span_type (llm_call/tool_call/retrieval/workflow), operation
  • Timing: start_time_ns, end_time_ns, duration_ns — nanosecond ints, no datetime column
  • Status: status_code, status_message
  • Resource/scope: service_name, resource_attrs, scope_name, scope_version
  • Payload: attributes, events, links
  • Linkage, filled by process_span: capability (resolved through IdentityAlias, never created), job, iteration
  • Ingest time: received_at

Ingest: POST /api/v1/traces (with a /v1/traces compat alias) takes an OTLP protobuf export → parse ResourceSpans → ScopeSpans → Span → flatten → bulk upsert on span_id → process_span per span for linkage.

Span cost has one owner: services/span_pricing.stamp_span_cost runs on every OTLP and connector span before it is stored. A reported cost (any alias, coalesced to genai.cost) is kept; otherwise genai-prices sets genai.cost from model, tokens and cache reads at the span's start time. The span, trace totals and Capability.usage_stats all read that value, and the SDK never prices. Gateway billing is separate on purpose: model_catalog.estimate_cost charges what OpenRouter charged us.

Read surface, backed by SpanViewSet rather than a trace viewset:

  • GET /api/traces/ — one row per trace, its head span (Span.trace_heads): the root once it arrives, else the earliest span, so traces stream in as spans land. Each row carries trace_status (completed = root present, live = rootless and recent, interrupted = rootless and quiet past TRACE_SETTLE_SECONDS)
  • GET /api/traces/{trace_id}/ — every span sharing that trace_id, plus root summary and trace_status
  • GET /api/traces/services/ — distinct service_name; GET /api/traces/models/ — distinct models
  • No /api/spans/ route exists.

Query params are span-native (received_at__gte, trace_id, span_type, operation, service_name, has_error).

Evaluation

Two systems share the evaluator engine but never share a judge. ctx["eval_surface"] decides which, and base.judge_module(ctx) is the only place that chooses — nested callers inherit the surface through ctx rather than importing a judge directly.

SurfaceSet byJudgeScore comes from
generativetasks/eval.pyevaluators/gen_judge.py, ChecklistResultweighted fraction of checklist verdicts, computed in Python
trace_scoring (default)services/eval/trace_scoring.pyevaluators/judge.py, JudgeResultthe model's own score field

The default is trace_scoring. rubric_compiler mirrors the split: build_checklist_prompt for generative, build_judge_prompt for trace scoring.

services/eval/context_check.py supplies model-aware advisory estimates for candidate generation and generative judges through REST POST /api/eval-runs/context-check/, run detail and MCP readiness/resources. Readiness accepts the proposed variants; context warnings never change ready, gate launch or exclude training candidates. Unknown limits remain unverified. Estimates include reserved output; before generation, reference-answer size is only a proxy for judge input. Runtime checks use actual prompts and tool history. The shared judge funnel no longer clips the supplied prompt to a fixed character count. Output exhaustion may retry once with a larger absolute budget bounded by published context/output limits and 64k tokens; unknown capacity or a permanent context rejection does not repeat. _judge sub-scores retain attempt budgets, finish reasons, usage, request IDs and costs. Exhaustion/rejection is a technical error, not a quality zero.

services/eval/context_suggestions.py adds preflight-only alternatives from the configured registry, filtered by published context/output and tool/schema support. Judge suggestions avoid the evaluated provider family. Registry order is retained, not sorted by price. cost_basis scopes USD estimates to all checked inputs plus full reserved output for one pass per model/judge; unknown prices are null. Console shows suggestions during setup, never auto-selects, and uses the existing ranked training candidates and their recommended-configuration training estimates for training alternatives. Trained benchmark hosting costs must not borrow base-model token rates. MCP readiness exposes the same evaluation suggestion contract; training catalog and estimates remain available through its training tools.

Training setup puts the judge selector below the eval set. Blank FinetuningJob.eval_judge_model preserves each evaluator's saved choice (including mixed sets); an explicit choice overrides only generative judge/fallback snapshots for this job, never the library or the Jev decision backend. Before/after runs share those frozen snapshots, and sibling baseline reuse includes judge-selection identity. REST context preview and MCP readiness accept judge_model and return judge_models: all registry judge options, annotated with estimated fit across every grader and aggregate USD budget per evaluated model/dataset pass. Unknown or undersized options stay selectable. MCP start_finetune accepts the same immutable eval_judge_model. Console tints judge/benchmark selectors only for estimated limit overflows, mutes undersized options without disabling them, and uses a warning-coloured Start button to open a short risk summary with Start anyway. Unverified, unavailable and pending checks stay neutral and do not open the warning summary. Context issues never require a model change to launch.

Bounded decisions use core/decisions.py: registered Jev on OpenRouter's /systemone endpoint, structured choice questions, validated complete probability distributions, project/contract-scoped caching, a conservative UTF-8 context guard, and shared Redis request/token pacing. Missing configuration, capacity, invalid responses, oversized context or low confidence fall back through services/eval/decisions.py; confidence is not calibrated correctness. Generative question resolution uses closed, required per-question properties compatible with strict provider schemas, preserves original question IDs and the selected judge, and validates coverage and choices before conversion. No evidence is truncated for Jev. config.decision pins backend (jev or generative), model, confidence floor and adapter-policy version; new snapshots and rubric digests include it. The actual served revision and fallback/cost provenance are stored in _decision sub-scores.

Configurable evaluator policies default to generative. Explicit Jev policies cover checklist verdicts, extracted-claim support, per-turn dimensions, cascade step ratings and confident successful numeric behaviour steps; qualify them against representative labelled examples before relying on their scores. Fixed grounding, capability classification and workshop semantic services retain bounded Jev routing. Independent question batches retain accepted answers and generatively resolve only uncertainty or unfinished questions after a later transport failure. Generated choices carry reasoning, not fabricated Jev confidence. Claim verification reuses extracted claims on fallback. Failed/uncertain behaviour steps retain full generative causal diagnosis. Label-only categorical judges on both evaluation surfaces can opt into Jev Choice with declared labels and unchanged label-to-score mappings; uncertain or insufficient evidence falls back. Categorical checklists, behaviour contracts, panels and cascade routes retain generative review. Jev labels carry decision provenance, not generated explanations. Holistic outcomes, session ledgers, claim extraction, rubric authoring and cascade summaries remain generative. Grounding keeps insufficient evidence outside its support/contradiction denominator; lookup and persistence use one policy-aware identity. Combined dataset description, evaluator relevance and rubric authoring use one generative call. Cascade provenance is stored once per batch, unknown costs include retries and holistic stages, and decision latency measures the whole adapter call including failures. Jev is excluded from inference and chat-model pickers. Provider calls in tests are blocked unless explicitly stubbed.

Generated samples snapshot the runner's initial messages and tool schemas after system-prompt injection in trajectory.model_request; metadata.output_start separates context from newly generated messages. This is the initial runner request, not provider chat-template bytes or every replay turn. services/eval/sample_io.py supplies the REST detail and MCP resource with separate input, generated output, and grading reference. Historical samples fall back only to the verified pinned dataset row and label it input_source=dataset, never an exact captured request; absent sources are unavailable. The eval-run MCP resource exposes five bounded sample inspections with grader reasoning.

Trace scoring carves a trace into units in services/eval/units.py (explicit precedence: turn spans > entry_point invocations > key-segment shim > structural root; beside the lattice, a run boundary enclosing turn slices becomes a run-grain execution surface when a run-grain behaviour binds it, and a boundary-less single-function-span trace is an unscorable orphan fragment); services/eval/trace_scoring.py judges the units and services/behaviour/binder.py binds them.

Standalone REST creation and MCP run_evaluation accept the same optional judge_model used by context preview. Blank preserves mixed library choices; explicit overrides are frozen in EvalRun.judge_model and all generative judge/fallback snapshots, including replay judges attached after launch. Existing runs reject judge changes. Deterministic checks and Jev policy stay unchanged. Console New evaluation uses the shared judge selector, context estimates and advisory Start warning; run detail shows snapshot judges. The MCP run resource exposes the override and bounded run_evaluators judge identities.

Capabilities and sync

The agent is the project itself — one graph per project, no table. Capability rows are its nodes (UI: "Capability"; the sidebar's "Agent" is the product).

  • A scan never deletes: absence sets status=leftover and the row reactivates in place when the code returns; observed rows (telemetry- or hand-made) are exempt from leftover.
  • DELETE /api/capabilities/{id}/ is a soft delete (status=deleted): the row and its data stay, but it is invisible to lists, scans, and identity lookup, so a later scan of the same code mints a fresh row. There is no merge, split, retire, or history.
  • Ingest never creates a capability: services/capabilities/identity.lookup resolves ids/names/slugs through IdentityAlias (renames included), an unknown identity leaves the span unbound, and tasks/capability_rebind re-binds the backlog after scans and hand-made rows.
  • Wire identity is id-only: ingest binds spans by overmind.capability.id alone; overmind.agent.* attributes are not read; overmind.capability.name is a display label that never resolves. The overmind/<uuid> model alias never changes.
  • Discovery is local: /overmind setup runs overmind chassis (AST inventory + call graph), writes overmind_capabilities.json, convert_json_to_toml fills prompt spans, drops fabricated anchors and unverifiable provenance, then stamps trajectory_map[].verified, then overmind sync.
  • Toml sync is POST/GET /api/v1/sync (API-key auth). It maps onto Capability (current/leftover), stores system_prompt / eval_metrics / capability_card / eval_matrix in improvement_metadata (eval_matrix is intent metadata only — not materialized into graders), mints Behaviour rows from capability_card.trajectory_map, keeps repo_summary / trace_provider / toml version in Project.settings, writes leftovers back to toml with archived = true, and enqueues the async Default-set preload (Tier-0 card compiler + Tier-1 LLM judges).
  • Repository provenance: overmind chassis saves a local .overmind/scan.json receipt before inspection. Conversion verifies branch, commit and a SHA-256 fingerprint of tracked and unignored working-tree files (excluding generated scan files), then persists repository_snapshot in toml. A changed checkout requires a rescan. Sync never recaptures Git state: it stores provenance in Project.settings and stamps last_synced_at separately; a push without provenance clears the prior label. The agent graph and overmind://project/current expose both. Missing Git or a missing scan receipt leaves provenance unknown.

Scoring

Behaviour-keyed. The codebase scan mints Behaviour/BehaviourVersion contracts per capability (services/behaviour/registry.py, re-anchored across rescans, retired never deleted). Trace scoring carves units (services/eval/units.py), binds each as a TaskExecution (services/behaviour/binder.py — unbound units still materialize), and writes Verdict rows plus the feedback_score["trace_scoring"] block the console reads. The session score is a fold over the append-only ConversationEvent ask ledger. API: api/behaviours.py.

Data Workshop

Source → cells → derived versions, edited only through the dataset's own agent. The full model, runner sandbox, agent tools, landing and export contract: data-workshop skill.

Auth and tenancy

Clerk when CLERK_API_SECRET_KEY is set; blank secret is self-hosted local JWT (POST /api/auth/local/ — email + password, create on first use, no verification). Project-scoped tenancy (User ↔ Project via ProjectMembership), no org layer. Everything queryable is filtered by project.

  • A guest (User.is_guest, minted by POST /api/auth/guest/) holds one project and can read and claim. GuestJWTAuthentication (api/authentication.py) refuses every other write with guest_upgrade_required; a view opts in with guest_allowed = True — never a permission class or middleware, since a view's own permission_classes replaces the defaults and a guest identity exists only through that token. Guests get no free credits; /demo points at local setup. A claim moves the memberships to the Clerk account and deactivates the guest; tasks/guest_cleanup.py deletes inactive guests and unclaimed ones after 7 days. Guest claim requires Clerk.
  • Console: VITE_SELF_HOSTED=true and a blank VITE_CLERK_PUBLISHABLE_KEY skip ClerkProvider and show the local login form; otherwise Clerk.
  • Project invites: with Clerk, an invitation email is sent; without Clerk the ProjectInvite row is stored and claimed on first local (or Clerk) sign-in for that email.
  • Commercial billing (remaining-credit 402s, Free/Pro quotas, Stripe Checkout) injects when STRIPE_SECRET_KEY is set (overbae/services/billing_provider.py). Ledger charges always run. Empty key → uncapped OSS: spend is recorded and shown, gates and grants no-op.
  • API keys are either scope=account (every project the user belongs to) or scope=project with one resourceIds entry. Creating a key with project always mints project scope.

Training preparation

Workshop findings are advisory before model-specific work: incomplete task/evidence/answer/schema reviews, capability mismatches and train/eval overlap produce warnings, not launch blockers. Exact preprocessing checks technical compatibility and does not repair datasets. Fine-tuning validation is read-only, and successful creation freezes the selected train, validation and eval_cell versions atomically. The same pinned eval cell is reused for before/after runs; dataset versions and capability cannot be swapped after creation. Internal loss-validation splitting preserves duplicates and cannot silently clean selected data.

The workshop is model-independent. services/training_preparation.py caches exact Modal preprocessing by training/validation cell fingerprints, model, training type, context length, training stack and the SFT asset code fingerprint. REST POST /api/training-preparations/ and MCP prepare_training_data share the service; MCP returns a training_preparation job receipt readable through get_job and the jobs resource. Preparation uses CPU-only prepare_<train_function> in the matching Modal training image; it does not start GPU training.

Console setup does not submit or await preprocessing and has no per-model preparation cards or workshop review recommendations. run_finetuning requests or reuses the exact artifact after launch, reports preparation as job progress, and starts GPU training only when ready. Explicit REST/MCP preparation remains available.

Modal recommendations use row-length estimates to size training context without treating estimates as compatibility verdicts. Job preparation repeats sizing across the pinned training and enabled validation cells, preserving a larger requested context. If exact tokenization exceeds that context but fits the model's training-type limit, it requests a matching larger preparation and rechecks every row before GPU submission. Exact lengths need no estimate headroom; over-limit rows and other incompatibilities still block. No dataset is rewritten, truncated or filtered. A training-job retry can recover an undersized context through this same path.

TrainingPreparation persists queued/starting/running/ready/incompatible/failed, the remote call ID and a 20-minute deadline. A 15-second controller queues short io observers. Restarts poll the saved call; transport failures retain it. An unacknowledged submission fails closed. Explicit retry of a confirmed failure cancels its saved remote operation before requeueing; uncertain submissions cannot be retried automatically. REST /retry/ and MCP retry_failed use the same guard. Retrying a failed/cancelled training job also retries its cached confirmed-failed preparation before queuing training, unless GPU training already has a remote ID. Ready or in-flight preparations are reused; unsafe failures and incompatible data refuse the retry without resetting job or eval state.

sft_assets/preprocess.py uses the trainer's pretok path, reporting exact token counts, shifted supervised targets, decoded masked previews and incompatible source-row IDs. It rejects over-context rows and zero-supervision examples without truncation or silent deletion. A ready artifact stores token IDs/labels, checksum and tokenizer; upload_dataset only materializes rows present in that artifact. The GPU engine validates context/vocabulary and consumes those tokens without preprocessing them again. Changes to the model/context or data require a matching artifact. Other providers retain their existing provider-side preprocessing; do not label those configurations exact-preflight-ready.

Changes to the SFT assets or preparation functions require modal deploy overbae/modal/modal_sft_worker.py; recreating local Docker services does not deploy Modal code. A worker/code fingerprint mismatch blocks preparation until the matching worker is deployed. Preprocessing infrastructure failures return a safe, actionable error and remain failed, distinct from incompatible rows. Each preparation retains subprocess stdout/stderr at preparations/<id>/preprocess_stdout.log on the SFT volume, including failed attempts; raw diagnostics do not enter user-facing error messages.

Unsloth token accuracy is measured from the final decoder output through the actual output head, in bounded chunks over supervised next-token targets. Fused CE stays enabled; accuracy does not depend on returned logits or a batch/context threshold. Counts are token-weighted across microbatches, with separate training and validation windows; checkpoint recomputation does not count twice. A missing decoder capture fails explicitly rather than silently dropping the metric. A rolling Modal deploy leaves active calls on their existing code; it cannot add metrics to an already-running trainer or reconstruct historical accuracy.

Show full SKILL.md (2,007 more words)Show less

Serving and weights

services/model_activation.py owns making a ready deployment live. REST active_model writes and MCP set_active_model persist a per-capability ModelActivation; the 15s deployment reconciler advances checking, waking/verifying via the existing pre_warm handle, and atomic routing switch. Until verification succeeds the previous target serves. Claims expire after 45s, transport errors retain the remote handle, and activation has a 50-minute deadline. Failed activation is retried through the same selection operation; duplicate in-progress selections are idempotent, competing selections are rejected, and clearing routing invalidates the operation generation. The capability retains its previous selection for verified switch-back. First/last application request timestamps only advance on successful API-key calls to the capability alias, excluding optimiser traffic and requests begun before the current switch. Internal probes and pinned calls never confirm alias integration.

services/inference_metrics.py supplies REST and MCP deployment metrics and activity. Both surfaces share period (1h, 24h, 7d, 30d, all) and source (application, all) filters; API defaults remain all time/all traffic, while the Console defaults to application traffic in the last 24 hours. The same selection applies to totals, failure counts, cost coverage, latest request/failure and chart buckets. Empty time buckets have zero requests and null latency; successful end-to-end timings include cold requests. Percentiles use the most recent 5,000 measured successful calls (per bucket for activity). Failed calls count as requests but do not contaminate successful latency distributions. Historical end-to-end timings and traffic sources remain unknown. services/inference_live.py reads Modal stats from the parameter-bound worker method without invoking inference. Pool identity includes base identity, GPU/image, context and LoRA rank; LoRA deployments use the shared base name. Counts are cached for 5 seconds and remote reads have a 10-second deadline. Recent successful inference (including deployment/activation verification) and gateway warming stamps remain fresh on every read. Zero measured runners means asleep; unavailable measurements remain unknown. Class-wide counts cannot establish one deployment's warmth.

services/serving_context.py estimates the pinned evaluation workload (including the capability prompt and reference-output headroom) independently of training context_length. Model recommendations with an eval dataset and REST/MCP launch share the serving plan. Estimated context overflow warns without excluding a candidate or rejecting launch. The plan respects native context without assuming RoPE scaling and still validates single-GPU capacity. Deployment and base-model evaluation use the plan; all before/after calls reserve the same output budget. A shared baseline grows only from a settled deployment state. Existing ready fine-tunes are not silently reconfigured.

modal_shared/context_budget.py supplies a positive default output reservation and rejects prompt truncation. Both streaming and non-streaming serving enforce it; vLLM validates the exact templated input plus reserved output before generation. MCP run_inference uses the same default reservation (8,192 tokens) when max_tokens is omitted or null, passes explicit positive integers unchanged, and returns non-retryable context_length_exceeded errors for requests that do not fit. Core LLM calls preserve finish reasons and reject token-limited responses before parsers or tool executors consume them. Eval runners retain the partial answer and usage, mark the sample degraded, and write skipped judgments rather than quality zeros. Finalization synchronizes training-eval status explicitly; degraded runs cannot become successful training-quality comparisons. MCP inference returns finish reason, model truncation and response-content clipping separately.

The MCP deployment resource exposes current worker state and measurements through inference_live.worker_status, independently of deployment readiness, alias selection and historical metrics filters. It uses the Console's cached pool measurements and recent traffic/warming stamps without invoking inference. Unavailable counts stay null; state remains unknown unless recent successful traffic or a warming stamp supplies evidence.

Training jobs and eval sets may have no capability. Training still requires an eval dataset and an eval set with enabled generative evaluators; a job without a capability can also select a ready trained benchmark. Unassigned eval sets are project-scoped and cannot be activated for live trace scoring. eval.eval_set.create_with_evaluators atomically creates a set and its library members in their applicable roles; REST and MCP share that creation path.

Training has four persisted evaluation choices: eval_incumbent_before/after and eval_model_before/after. The training model's before target is its untouched base; after is its ready trained deployment. The training setup's Benchmark model selector chooses the codebase incumbent or a ready, project-scoped trained model for this run. REST and MCP start_finetune accept its serving model ID as baseline_model; the choice is immutable once the job is created and never changes active_model or serving. Omitting it uses the capability's separate benchmark_model default (set through MCP set_benchmark_model), or the codebase model when that default is null. The Console defaults to the codebase incumbent and sends its selection explicitly. New incumbent evaluations reject an unavailable selection. The capability Models tab only selects live routing. When OpenRouter is configured and its cached catalog has an exact model ID or Hugging Face checkpoint match, before evals use OpenRouter without a base deployment. Unavailable models or catalogs keep the existing provider/Modal fallback. A started eval or attached baseline deployment pins the route until an explicit retry; catalog refreshes do not switch live evals or provision duplicate baseline infrastructure. Defaults are untouched-base-before + trained-after; incumbent comparisons are separate opt-ins. The matched base is the preferred delta reference. All four can be off, but the eval dataset and eval set remain required. Baseline evals launch as independent jobs alongside GPU training; their completion and failure do not gate training. Retries retain EvalRuns but detach failed baseline-eval links. After evals start once training finishes. The same pinned eval cell and grader snapshots are reused across the job's evaluations. Generate variants snapshot the capability prompt in params, injecting it only when the row has no system turn. They evaluate the raw model with recorded context, not the live application. Every selected training evaluation uses all rows of the pinned version, with no row cap or sampling; shared group baselines must also cover all rows. Automatic intermediate-checkpoint judging is replaced by this explicit schedule; historical checkpoint results remain readable. MCP start_finetune accepts all four choices and the fine-tune resource returns evaluation_plan. The Console projects selected evaluations into waiting rows immediately and replaces each by kind when its run arrives. Loss-curve polling continues after deployment while evaluations remain pending. services.deployment persists each deployment's stage, remote call handle, generation, retry budget, deadline, and evaluation waiters. The single 15-second deployment controller covers training outputs and baselines; short IO tasks claim a row for 45 seconds, spawn or observe one remote call, then exit. Worker death reconnects to the saved call rather than restarting GPU work. Each attempt has a four-hour deadline and at most three confirmed-failure attempts with 30/60-second backoff. Transport errors retain the handle. Submission without a saved acknowledgement fails closed; never launch another call until remote state is resolved. Failed baseline deployments notify waiting evaluations even without a remote handle or while cancellation is pending. Baseline stage progress is mirrored to the training job while it runs; eval score updates merge into freshly locked progress. Base preparation always uses fetch_base_model before publishing/merging and booting. Cancelled jobs and superseded generations cannot become ready; terminal notifications and remote cancellation are durable work too. Explicit retries reset the generation and budget without losing checkpoint metadata. Training remains successful if only deployment fails; the dependent evaluation is marked failed. REST and MCP share the retry service and expose stage, attempt, retry time, deadline, and last error. The deployed-model list includes only deployments linked to training jobs, before pagination and filtering. Baseline eval infrastructure remains addressable by ID but is excluded from the Console's model lists and counts.

Training recommendations classify the selected capability's recorded codebase context (codebase.task_type), cached by context fingerprint. Task type drives benchmark skill weights; dataset stats still drive context/tool compatibility, hyperparameters and cost. With no selected capability, use the dataset's stored semantic task or profile heuristic, even if the dataset has a capability mapping. An unclassifiable selected capability stays unknown and ungraded; it does not fall back to the dataset. Console and MCP readiness use the same recommendation service; an uncached capability classification may call an LLM.

  • Serving splits by finetune shape in deployment.serves_as_adapter. A dense or explicitly supported MoE Modal-trained LoRA is served as an adapter on a shared BF16 base — publish_adapter copies the adapter, and every deployment on that base shares one container pool, so a second adapter deploys in seconds. Everything else (full finetunes, unsupported MoE, non-Modal providers) is merged and quantized into a private checkpoint. The shared base must stay BF16: an FP8 base measurably degrades adapter quality. MoE models ("moe": true) require verified inference.lora_supported in models.json. Qwen3-Coder uses native mixed-format MoE LoRA with the fused expert format declared at adapter load; weights remain separate from the BF16 base. Shared workers resolve the family from the sealed base manifest. inference.min_vram_gb enforces measured GPU capacity floors; Gemma 4 E2B and E4B require at least L40S. A deployment whose stored GPU is below that floor is raised before the next boot.
  • Base weights live in exactly one place: .base_models/{org--model} on the weights Volume. fetch_base_model is the only writer and a global mutex (max_containers=1), so concurrent callers cannot corrupt a shared dir or download twice. Celery waits on that Function (.remote()) before spawning a GPU train job, so a cold base downloads on CPU. Training then loads BASE_MODEL_PATH and fails if the snapshot is incomplete — no Hub fallback, no AWS for bases, no catalog prewarm.
  • Nothing on the overmind-sft Volume is read at serve time — GPU workers mount weights and vllm-cache, plus inference-artifacts for shared-base LoRA. tasks/cleanup_modal.py sweeps it daily, ages run dirs with no FinetuningJob row off the Volume's own mtimes, and drops runs/{run_id}/final/ once the deploy has landed (a READY deployment means the serving copy is on the weights Volume and a confirmed S3 archive covers a rebuild). The same beat drops spent /weights/.staging/{id}/ trees. stage_modal_checkpoint and publish_adapter fall back to download_checkpoint_from_s3 when the checkpoint is absent — the one place an S3 round trip is allowed, never the first deploy. Both delete the job's staging dir as soon as the serving copy is written.
  • Shared-base LoRA workers use a separate {gpu}_{image}_lora class for every existing GPU/image pair; the selector is unchanged. fetch_base_model hashes and seals .base-manifest.json once, then validates its file identities on reuse. The base digest participates in the container parameters. Snapshot capture initializes vLLM and CUDA graphs, exports adapter-free runtime-layout weights to immutable generations on overmind-inference-artifacts, caches parsed headers, and uses level-2 sleep to omit large weight backups. Artifact identity includes base, effective serving arguments, GPU, image, loader source and runtime versions; it is not portable across arbitrary engine profiles. A new LoRA on a compatible base/profile reuses these assets.
  • Restore refreshes the Volumes, validates identities, wakes weight allocations, streams 2 GiB shards in manifest order into native pinned buffers, copies into the original CUDA graph addresses, resets empty LoRA banks, then wakes KV memory and checks health. LoRA images pin Run:ai 0.16.1; one reader uses 16 native threads, eight Torch threads and an 8 GB buffer for bases up to 80 GB, otherwise 16 GB. The pin is required by the upstream allocation site. Snapshots contain no tenant adapters; adapters are attached after restore under the existing lock. Lifecycle failures are recorded and surfaced on the input, not raised from enter into an endless retry. Internal sleep/reload/adapter-control endpoints cannot be forwarded through public inference methods.
  • Full-checkpoint workers retain their non-snapshot startup and compile-cache persistence. Only those concrete classes own the normal post-snapshot startup hook: putting it on the common base causes Modal to run it alongside the LoRA restore hook. pre_warm still verifies a deployment before READY; for LoRA it also completes first-time artifact/snapshot preparation. CUDAGRAPH_CAPTURE_SIZES and GPU selection are unchanged. The Modal gateway and Django SSE write idle pings immediately and every 15s; Django non-stream completions wait up to 15s for a result before starting newline keepalives. Playground SSE retains its keepalive. The edge ALB drops idle connections after 60s. Early non-stream errors return 400/502; after a keepalive commits headers, Django aborts an incomplete JSON response so clients raise a transport error. SSE failures use an error frame. Requests carry X-Request-ID, and failures and end-to-end timing are recorded alongside engine timing. The Django client rejects upstream JSON errors and consumes SSE lines without byte-buffer accumulation. GPU scheduling remains outside the loader's control; no cold-start SLA is implied.

© overmind-core, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/backend-architecture of overmind-core/overmind.

Open the folder on GitHubat commit 3dec73c

Compare with similar skills

Backend Architecture next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Backend Architecture compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Backend Architecture this skillovermind-core/overmind597—~10kAutomated safety check: PassAGPL-3.0
Cookbook Aimldatabricks-solutions/databricks-apps-cookbook183—~1.7kAutomated safety check: PassCustom licence
Pgvector Semantic Searchtimescale/pg-aiguide1.9k1 repos~3.8kAutomated safety check: PassApache-2.0
Pixiu LLM Gatewayapache/dubbo-go-pixiu568—~2.5kAutomated safety check: PassApache-2.0
Reachai Onboardingw8123/EnterpriseAgentFramework836—~6.1kAutomated safety check: PassMIT
Perfupraullenchai/Rapid-MLX3.9k—~1.6kAutomated safety check: NotesCustom licence

Similar skills

  • Cookbook Aiml

    databricks-solutions/databricks-apps-cookbook

    Invoke ML models, run vector search, and connect to MCP servers from Databricks Apps.

    183 GitHub stars~1.7k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Pixiu LLM Gateway

    apache/dubbo-go-pixiu

    Creates and validates dubbo-go-pixiu LLM gateway conf.yaml. An agent skill from apache/dubbo-go-pixiu.

    568 GitHub stars~2.5k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Reachai Onboarding

    w8123/EnterpriseAgentFramework

    Integrate Java business systems with ReachAI SDK registration, SDK instance heartbeat, gateway/embed access, and optional API Management handoff.

    836 GitHub stars~6.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Perfup

    raullenchai/Rapid-MLX

    Autonomous performance optimization: research, PoC, benchmark, implement, review, PR

    3.9k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Integrate Model

    tryonlabs/opentryon

    Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner).

    551 GitHub stars~1.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from overmind-core/overmind

All 20 skills in this repo
  • API Endpoints

    overmind-core/overmind

    End-to-end workflow for adding or changing a backend API endpoint — which module the serializer and view belong in, URL registration, OpenAPI client regeneration, and typed consumption from the…

    597 GitHub stars~830 tokensUpdated today
    Auto-check: notes
  • Finetuning Model Onboarding

    overmind-core/overmind

    Rules for adding a new model or model family to the finetuning pipeline, or changing finetuning behavior for an existing one — engine-agnostic customization via family hooks instead of if/else in…

    597 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Frontend Design

    overmind-core/overmind

    Overmind Console design system — semantic tokens, shared primitives, geometry and icons, the border-contrast floor, the duplicated table implementations, and the verification scripts.

    597 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • MCP

    overmind-core/overmind

    End-to-end workflow for adding or changing Overmind MCP tools, resources, prompts, authentication, or result contracts — server layers, catalog registration, MCP-impact classification, and required…

    597 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • PR Etiquette

    overmind-core/overmind

    How to open a complete pull request on overmind-core/overmind — the CI gates, the cross-cutting surfaces a change must carry with it (MCP, blast radius, the docs repo), gh pr edit being broken here…

    597 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Seed Demo Data

    overmind-core/overmind

    Run or modify the seeddemo management command (the one-project Support Copilot demo) without breaking the beat-safety invariants that keep celery workers from re-driving seeded rows.

    597 GitHub stars~973 tokensUpdated today
    Auto-check passed

Questions about Backend Architecture

What does Backend Architecture do?

Deeper backend map of overbae — module layout, celery queue topology, the span-only tracing model and its API surface, capabilities and toml sync, behaviour-keyed scoring, auth and guests, model…. Backend Architecture is an agent skill from overmind-core/overmind. Deeper backend map of overbae — module layout, celery queue topology, the span-only tracing model and its API surface, capabilities and toml sync, behaviour-keyed scoring, auth and guests, model serving and base weights.

When should I use Backend Architecture?

Backend Architecture fits situations like: navigating unfamiliar backend subsystems; wiring cross-subsystem behavior.

How do I install Backend Architecture in Claude Code?

Run `npx skills add overmind-core/overmind --skill backend-architecture -a claude-code`. Or copy the skill folder (.agents/skills/backend-architecture in overmind-core/overmind) into .claude/skills/backend-architecture in your project. Claude Code loads it when a task matches its description.

How do I install Backend Architecture in Codex?

Run `npx skills add overmind-core/overmind --skill backend-architecture -a codex`. Or copy the skill folder (.agents/skills/backend-architecture in overmind-core/overmind) into .agents/skills/backend-architecture in your project. Codex loads it when a task matches its description.

Can I use Backend Architecture in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add overmind-core/overmind --skill backend-architecture -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/backend-architecture, .gemini/skills/backend-architecture, .github/skills/backend-architecture and .opencode/skills/backend-architecture in your project.

What does Backend Architecture need to run?

Going by SKILL.md and its folder, Backend Architecture needs the command-line tools its instructions call (make and modal) and credentials named CLERK_API_SECRET_KEY, VITE_CLERK_PUBLISHABLE_KEY and STRIPE_SECRET_KEY. Our summary lists: Python 3; Docker.

Does Backend Architecture access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Backend Architecture safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Backend Architecture use?

Backend Architecture is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Backend Architecture use?

About 10k tokens (SKILL.md is roughly 41k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Backend Architecture?

Skills that share tags, products or a category with Backend Architecture: Cookbook Aiml (databricks-solutions/databricks-apps-cookbook, 183 stars), Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars), Pixiu LLM Gateway (apache/dubbo-go-pixiu, 568 stars) and Reachai Onboarding (w8123/EnterpriseAgentFramework, 836 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Backend Architecture?

overmind-core (a GitHub organization) maintains it in overmind-core/overmind, which has 597 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 8, 2026.

Source: overmind-core/overmind on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.