Agent skill

Together Reference Architecture

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Design a production Together AI service with a typed provider boundary, policy-based model routing, serverless and dedicated lanes, batch workers, telemetry, budgets, and reversible degradation.

MITAuto-check passedBackend & APIs

Install Together Reference Architecture

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill together-reference-architecture -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace together-reference-architecture --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/together-reference-architecture .claude/skills/together-reference-architecture && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
together-reference-architecture
GitHub stars
2.8k
Token cost
~1.1k tokens
SKILL.md length
379 words
Files
2 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Design a production Together AI service with a typed provider boundary, policy-based model routing, serverless and dedicated lanes, batch workers, telemetry, budgets, and reversible degradation.

  • Works in 6 steps: Partition workloads into interactive,… → Define a typed provider adapter and… → Place bounded queues around bursty and… → …
  • Defining the integration topology
  • SKILL.md covers Overview, Prerequisites, Tool Discipline and Current Contract, plus 7 more sections
  • Needs TOGETHER_API_KEY

What it does

Together Reference Architecture is an agent skill from jeremylongshore/tons-of-skills-marketplace. Design a production Together AI service with a typed provider boundary, policy-based model routing, serverless and dedicated lanes, batch workers, telemetry, budgets, and reversible degradation. Use when defining the integration topology. Trigger with "Together architecture", "Together model gateway", or "design Together service".

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/official-docs.md`). Compatibility notes: Designed for Claude Code; implementation may require cloud, queue, secret-store, and Together AI access

It sits in Backend & APIs, covering Serverless and Model routing and gateways. It works with Together AI. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Defining the integration topology
  • With Together architecture
  • Together model gateway
  • Design Together service

Example prompts

  • “Together architecture”
  • “Together model gateway”
  • “design Together service”
  • “/together-reference-architecture”

Requirements

  • A credential in TOGETHER_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code; implementation may require cloud, queue, secret-store, and Together AI access
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, WebFetch, Write, Edit

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Partition workloads into interactive, offline batch, training, and reserved-capacity classes.
  2. Define a typed provider adapter and policy service for model, bounds, fallback, and deprecation state.
  3. Place bounded queues around bursty and asynchronous work with durable IDs and reconciliation.
  4. Add per-model request/token control, circuit breaking, usage/cost telemetry, and quality sampling.
  5. Define data redaction, retention, tenant isolation, and secret rotation at each trust boundary.
  6. Document degradation, model migration, batch recovery, dedicated scale-to-zero, and provider-exit paths.

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • WebFetch
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.together.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • TOGETHER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code; implementation may require cloud, queue, secret-store, and Together AI access

    From compatibility in the SKILL.md frontmatter.

Context cost

Together Reference Architecture loads about 1.1k tokens when it runs, and up to ~1.8k if it reads all its reference files. Until then it costs about 91 tokens; SKILL.md has 379 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 379 words, ~1,063 tokens.

Download SKILL.mdSave it as .claude/skills/together-reference-architecture/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
together-reference-architecture
description
Design a production Together AI service with a typed provider boundary, policy-based model routing, serverless and dedicated lanes, batch workers, telemetry, budgets, and reversible degradation. Use when defining the integration topology. Trigger with "Together architecture", "Together model gateway", or "design Together service".
allowed-tools
Read, Glob, Grep, WebFetch, Write, Edit
compatibility
Designed for Claude Code; implementation may require cloud, queue, secret-store, and Together AI access
argument-hint
[repository-path] [workload-profile]
version
1.9.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT
tags
saas, together-ai, architecture
model
inherit
effort
high

Together AI Reference Architecture

Overview

This skill turns workload requirements into explicit real-time, batch, and dedicated paths with one governed provider boundary and observable cost/quality behavior.

Prerequisites

  • Workload classes, modalities, volumes, latency objectives, and data classifications
  • Model-quality evaluations and fallback constraints
  • Availability, cost, retention, residency, and recovery objectives
  • Existing gateway, queue, telemetry, and secret-management topology

Tool Discipline

Use Read, Glob, and Grep to map callers, trust boundaries, queues, storage, and observability. Use WebFetch for current Together capabilities and limits. Use Write or Edit only for approved diagrams, ADRs, interfaces, or configuration.

Current Contract

  • Interactive serverless and dedicated models share inference request shapes, but capacity and billing differ.
  • Batch is an asynchronous file/job path with arbitrary result order and separate error artifacts.
  • Model IDs, prices, redirects, and availability are runtime policy inputs, not constants scattered through callers.
  • Credentials are project-scoped; isolate environments and inject keys at the gateway or worker boundary.

Authentication

Use separate TOGETHER_API_KEY references per environment and workload authority. Keep provider credentials server-side. Internal callers authenticate to the application gateway independently; they never receive the Together key.

Instructions

  1. Partition workloads into interactive, offline batch, training, and reserved-capacity classes.
  2. Define a typed provider adapter and policy service for model, bounds, fallback, and deprecation state.
  3. Place bounded queues around bursty and asynchronous work with durable IDs and reconciliation.
  4. Add per-model request/token control, circuit breaking, usage/cost telemetry, and quality sampling.
  5. Define data redaction, retention, tenant isolation, and secret rotation at each trust boundary.
  6. Document degradation, model migration, batch recovery, dedicated scale-to-zero, and provider-exit paths.
Show full SKILL.md (121 more words)Show less

Approval Boundaries

Do not introduce provider failover, cross-region data movement, dedicated capacity, or automatic model substitution without security, quality, reliability, and cost owners.

Output

Return component/flow topology, trust boundaries, provider interfaces, model policy, capacity lanes, observability, budgets, failure modes, rollback, and decision owners.

Error Handling

ConditionResponse
Requirements conflictRecord the tradeoff and seek the named decision owner.
Provider unavailableApply bounded circuit/degradation policy; do not retry indefinitely.
Model deprecatedRoute through evaluated migration policy, not an ad hoc replacement.
Batch partially failsReconcile by ID and retry only approved failed records.

Examples

The example below shows the minimum redacted evidence expected from a successful invocation of this operator workflow.

text
interactive=serverless-gateway; offline=batch-worker; reserved=dedicated-v2; secrets=per-environment

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/.curated/together-reference-architecture of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/official-docs.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Together Reference Architecture next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Together Reference Architecture compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Together Reference Architecture this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMIT
Together Fireworksericrisco/rsc-harness180—~3.3kAutomated safety check: PassMIT
Neonxai-org/plugin-marketplace288—~3.9kAutomated safety check: NotesNone
OmniRoute Combo Routingdiegosouzapw/OmniRoute75k—~2.1kAutomated safety check: PassMIT
ModalK-Dense-AI/scientific-agent-skills48k1 repos~4.5kAutomated safety check: NotesApache-2.0
Finetuningawslabs/agent-plugins916—~2.3kAutomated safety check: PassApache-2.0

Similar skills

  • Together Fireworks

    ericrisco/rsc-harness

    A skill your agent uses when calling open-weight LLMs on Together AI or Fireworks AI's OpenAI-compatible endpoints — baseurl plus namespaced model id, the cheapest model that clears the bar…

    180 GitHub stars~3.3k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Neon

    xai-org/plugin-marketplace

    Overview of the Neon platform for apps and agents, spanning Postgres, Auth, Data API, and the new services: Object Storage, Compute Functions, and AI Gateway.

    288 GitHub stars~3.9k tokensUpdated yesterday
    Backend & APIsAuto-check: notes
  • OmniRoute Combo Routing

    diegosouzapw/OmniRoute

    Manages OmniRoute routing combos through its REST API: create and update combos, choose from 19 strategies, set fallback chains, test outcomes and read metrics.

    75k GitHub stars~2.1k tokensUpdated today
    Backend & APIsAuto-check passed
  • Modal

    K-Dense-AI/scientific-agent-skills

    Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs.

    48k GitHub starsUsed in 1 repo~4.5k tokens
    Backend & APIsAuto-check: notes
  • Finetuning

    awslabs/agent-plugins

    Official

    Generates code that fine-tunes a base model using SageMaker serverless training jobs.

    916 GitHub stars~2.3k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • LLM Gateway

    sickn33/agentic-awesome-skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Backend & APIsAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Together Reference Architecture

What does Together Reference Architecture do?

Design a production Together AI service with a typed provider boundary, policy-based model routing, serverless and dedicated lanes, batch workers, telemetry, budgets, and reversible degradation. Together Reference Architecture is an agent skill from jeremylongshore/tons-of-skills-marketplace. Design a production Together AI service with a typed provider boundary, policy-based model routing, serverless and dedicated lanes, batch workers, telemetry, budgets, and reversible degradation.

When should I use Together Reference Architecture?

Together Reference Architecture fits situations like: defining the integration topology; with Together architecture; together model gateway; design Together service.

How do I install Together Reference Architecture in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill together-reference-architecture -a claude-code`. Or copy the skill folder (skills/.curated/together-reference-architecture in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/together-reference-architecture in your project. Claude Code loads it when a task matches its description.

How do I install Together Reference Architecture in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill together-reference-architecture -a codex`. Or copy the skill folder (skills/.curated/together-reference-architecture in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/together-reference-architecture in your project. Codex loads it when a task matches its description.

Can I use Together Reference Architecture in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill together-reference-architecture -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/together-reference-architecture, .gemini/skills/together-reference-architecture, .github/skills/together-reference-architecture and .opencode/skills/together-reference-architecture in your project.

What does Together Reference Architecture need to run?

Going by SKILL.md and its folder, Together Reference Architecture needs credentials named TOGETHER_API_KEY. Our summary lists: A credential in TOGETHER_API_KEY. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, Write, Edit. Compatibility (from SKILL.md): Designed for Claude Code; implementation may require cloud, queue, secret-store, and Together AI access.

Does Together Reference Architecture access the network?

SKILL.md names 1 domain. As links in the text: docs.together.ai. This is read from the text; nothing was executed.

Is Together Reference Architecture safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Together Reference Architecture use?

Together Reference Architecture is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Together Reference Architecture use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 758 tokens, read only when the agent opens those files.

What are the alternatives to Together Reference Architecture?

Skills that share tags, products or a category with Together Reference Architecture: Together Fireworks (ericrisco/rsc-harness, 180 stars), Neon (xai-org/plugin-marketplace, 288 stars), OmniRoute Combo Routing (diegosouzapw/OmniRoute, 75k stars) and Modal (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Together Reference Architecture?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.