Agentsop Framework Selection
agentsope/SkillAlchemy
Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-project SOPs (LangGraph…
LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.
$ npx skills add ancoleman/ai-design-components --skill model-serving -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ancoleman/ai-design-components model-serving --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/model-serving .claude/skills/model-serving && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "model-serving" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/model-serving into .claude/skills/model-serving/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-serving", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ancoleman/ai-design-components/tree/main/skills/model-servingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ancoleman/ai-design-components --skill model-serving -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ancoleman/ai-design-components model-serving --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/model-serving .agents/skills/model-serving && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "model-serving" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/model-serving into .agents/skills/model-serving/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-serving", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill model-serving -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ancoleman/ai-design-components model-serving --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/model-serving .cursor/skills/model-serving && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "model-serving" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/model-serving into .cursor/skills/model-serving/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-serving", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ancoleman/ai-design-components.git --path skills/model-serving--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ancoleman/ai-design-components --skill model-serving -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ancoleman/ai-design-components model-serving --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/model-serving .gemini/skills/model-serving && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "model-serving" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/model-serving into .gemini/skills/model-serving/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-serving", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ancoleman/ai-design-components model-servingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ancoleman/ai-design-components --skill model-serving -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/model-serving .github/skills/model-serving && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "model-serving" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/model-serving into .github/skills/model-serving/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-serving", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill model-serving -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ancoleman/ai-design-components model-serving --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/model-serving .opencode/skills/model-serving && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "model-serving" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/model-serving into .opencode/skills/model-serving/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-serving", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
model-servingLLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.
Model Serving is an agent skill from ancoleman/ai-design-components. LLM and ML model deployment for inference. Use when serving models in production, building AI APIs, or optimizing inference. Covers vLLM (LLM serving), TensorRT-LLM (GPU optimization), Ollama (local), BentoML (ML deployment), Triton (multi-model), LangChain (orchestration), LlamaIndex (RAG), and streaming patterns.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 28 other files, including scripts and reference files (for example `examples/k8s-vllm-deployment/README.md`, `examples/langchain-agents/README.md` and `examples/langchain-agents/main.py`).
It sits in AI & LLM Engineering, covering LLM inference and serving, Machine learning and Building AI agents. It works with vLLM, LangChain, NVIDIA AI Platform and LlamaIndex. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pippythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Model Serving loads about 3.4k tokens when it runs, and up to ~21k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 673 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 673 words, ~3,371 tokens.
.claude/skills/model-serving/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.Deploy LLM and ML models for production inference with optimized serving engines, streaming response patterns, and orchestration frameworks. Focuses on self-hosted model serving, GPU optimization, and integration with frontend applications.
vLLM (Recommended Primary)
TensorRT-LLM
Ollama
Decision Framework:
Self-hosted LLM deployment needed?
├─ Yes, need maximum throughput → vLLM
├─ Yes, need absolute max GPU efficiency → TensorRT-LLM
├─ Yes, local development only → Ollama
└─ No, use managed API (OpenAI, Anthropic) → No serving layer neededBentoML (Recommended)
Triton Inference Server
LangChain
LlamaIndex
# Install
pip install vllm
# Serve a model (OpenAI-compatible API)
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--dtype auto \
--max-model-len 4096 \
--gpu-memory-utilization 0.9 \
--port 8000Key Parameters:
--dtype: Model precision (auto, float16, bfloat16)--max-model-len: Context window size--gpu-memory-utilization: GPU memory fraction (0.8-0.95)--tensor-parallel-size: Number of GPUs for model parallelismBackend (FastAPI):
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from openai import OpenAI
import json
app = FastAPI()
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
@app.post("/chat/stream")
async def chat_stream(message: str):
async def generate():
stream = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": message}],
stream=True,
max_tokens=512
)
for chunk in stream:
if chunk.choices[0].delta.content:
token = chunk.choices[0].delta.content
yield f"data: {json.dumps({'token': token})}\n\n"
yield f"data: {json.dumps({'done': True})}\n\n"
return StreamingResponse(
generate(),
media_type="text/event-stream",
headers={"Cache-Control": "no-cache"}
)Frontend (React):
// Integration with ai-chat skill
const sendMessage = async (message: string) => {
const response = await fetch('/chat/stream', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ message })
})
const reader = response.body!.getReader()
const decoder = new TextDecoder()
while (true) {
const { done, value } = await reader.read()
if (done) break
const chunk = decoder.decode(value)
const lines = chunk.split('\n\n')
for (const line of lines) {
if (line.startsWith('data: ')) {
const data = JSON.parse(line.slice(6))
if (data.token) {
setResponse(prev => prev + data.token)
}
}
}
}
}import bentoml
from bentoml.io import JSON
import numpy as np
@bentoml.service(
resources={"cpu": "2", "memory": "4Gi"},
traffic={"timeout": 10}
)
class IrisClassifier:
model_ref = bentoml.models.get("iris_classifier:latest")
def __init__(self):
self.model = bentoml.sklearn.load_model(self.model_ref)
@bentoml.api(batchable=True, max_batch_size=32)
def classify(self, features: list[dict]) -> list[str]:
X = np.array([[f['sepal_length'], f['sepal_width'],
f['petal_length'], f['petal_width']] for f in features])
predictions = self.model.predict(X)
return ['setosa', 'versicolor', 'virginica'][predictions]from langchain_openai import OpenAIEmbeddings, ChatOpenAI
from langchain_community.vectorstores import Qdrant
from langchain.chains import RetrievalQA
from langchain.text_splitter import RecursiveCharacterTextSplitter
# Load and chunk documents
text_splitter = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=50)
chunks = text_splitter.split_documents(documents)
# Create vector store
embeddings = OpenAIEmbeddings()
vectorstore = Qdrant.from_documents(
chunks,
embeddings,
url="http://localhost:6333",
collection_name="docs"
)
# Create retrieval chain
llm = ChatOpenAI(model="gpt-4o")
qa_chain = RetrievalQA.from_chain_type(
llm=llm,
retriever=vectorstore.as_retriever(search_kwargs={"k": 3}),
return_source_documents=True
)
# Query
result = qa_chain({"query": "What is PagedAttention?"})Rule of thumb for LLMs:
GPU Memory (GB) = Model Parameters (B) × Precision (bytes) × 1.2Examples:
Quantization reduces memory:
# Enable quantization (AWQ for 4-bit)
vllm serve TheBloke/Llama-3.1-8B-AWQ \
--quantization awq \
--gpu-memory-utilization 0.9
# Multi-GPU deployment (tensor parallelism)
vllm serve meta-llama/Llama-3.1-70B-Instruct \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.9Continuous batching (vLLM default):
Adaptive batching (BentoML):
@bentoml.api(
batchable=True,
max_batch_size=32,
max_latency_ms=1000 # Wait max 1s to fill batch
)
def predict(self, inputs: list[np.ndarray]) -> list[float]:
# BentoML automatically batches requests
return self.model.predict(np.array(inputs))See examples/k8s-vllm-deployment/ for complete YAML manifests.
Key considerations:
nvidia.com/gpu: 1/health endpointFor production, add rate limiting, authentication, and monitoring:
Kong Configuration:
services:
- name: vllm-service
url: http://vllm-llama-8b:8000
plugins:
- name: rate-limiting
config:
minute: 60 # 60 requests per minute per API key
- name: key-auth
- name: prometheusEssential LLM metrics:
Prometheus instrumentation:
from prometheus_client import Counter, Histogram
requests_total = Counter('llm_requests_total', 'Total requests')
tokens_generated = Counter('llm_tokens_generated', 'Total tokens')
request_duration = Histogram('llm_request_duration_seconds', 'Request duration')
@app.post("/chat")
async def chat(request):
requests_total.inc()
start = time.time()
response = await generate(request)
tokens_generated.inc(len(response.tokens))
request_duration.observe(time.time() - start)
return responseThis skill provides the backend serving layer for the ai-chat skill.
Flow:
Frontend (React) → API Gateway → vLLM Server → GPU Inference
↑ ↓
└─────────── SSE Stream (tokens) ─────────────────┘See references/streaming-sse.md for complete implementation patterns.
Architecture:
User Query → LangChain
├─> Vector DB (Qdrant) for retrieval
├─> Combine context + query
└─> LLM (vLLM) for generationSee references/langchain-orchestration.md and examples/langchain-rag-qdrant/ for complete patterns.
For batch processing or non-real-time inference:
Client → API → Message Queue (Celery) → Workers (vLLM) → Results DBUseful for:
Use scripts/benchmark_inference.py to measure the deployment:
python scripts/benchmark_inference.py \
--endpoint http://localhost:8000/v1/chat/completions \
--model meta-llama/Llama-3.1-8B-Instruct \
--concurrency 32 \
--requests 1000Outputs:
Detailed Guides:
references/vllm.md - vLLM setup, PagedAttention, optimizationreferences/tgi.md - Text Generation Inference patternsreferences/bentoml.md - BentoML deployment patternsreferences/langchain-orchestration.md - LangChain RAG and agentsreferences/inference-optimization.md - Quantization, batching, GPU tuningWorking Examples:
examples/vllm-serving/ - Complete vLLM + FastAPI streaming setupexamples/ollama-local/ - Local development with Ollamaexamples/langchain-agents/ - LangChain agent patternsUtility Scripts:
scripts/benchmark_inference.py - Throughput and latency benchmarkingscripts/validate_model_config.py - Validate deployment configurationsvLLM provides OpenAI-compatible endpoints for easy migration:
# Before (OpenAI)
from openai import OpenAI
client = OpenAI(api_key="sk-...")
# After (vLLM)
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed"
)
# Same API calls work!
response = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "Hello"}]
)Route requests to different models based on task:
MODEL_ROUTING = {
"small": "meta-llama/Llama-3.1-8B-Instruct", # Fast, cheap
"large": "meta-llama/Llama-3.1-70B-Instruct", # Accurate, expensive
"code": "codellama/CodeLlama-34b-Instruct" # Code-specific
}
@app.post("/chat")
async def chat(message: str, task: str = "small"):
model = MODEL_ROUTING[task]
# Route to appropriate vLLM instanceTrack token usage:
import tiktoken
def estimate_cost(text: str, model: str, price_per_1k: float):
encoding = tiktoken.encoding_for_model(model)
tokens = len(encoding.encode(text))
return (tokens / 1000) * price_per_1k
# Compare costs
openai_cost = estimate_cost(text, "gpt-4o", 0.005) # $5 per 1M tokens
self_hosted_cost = 0 # Fixed GPU cost, unlimited tokensOut of GPU memory:
--max-model-len--gpu-memory-utilization (try 0.8)--quantization awq)Low throughput:
--gpu-memory-utilization (try 0.95)High latency:
scripts/benchmark_inference.pyexamples/ollama-local/ for GPU-free testingexamples/vllm-serving/examples/langchain-rag-qdrant/examples/k8s-vllm-deployment/© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 21 other files (scripts, references) in skills/model-serving of ancoleman/ai-design-components.
Open the folder on GitHubat commit 76551b7
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ancoleman/ai-design-components, which our catalogue first saw on October 7, 2026.
Model Serving next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Model Serving this skillancoleman/ai-design-components | 526 | 1 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Agentsop Framework Selectionagentsope/SkillAlchemy | 459 | — | ~5.8k | Automated safety check: Pass | MIT | |
| Jetson LLM BenchmarkNVIDIA/skills | 3.5k | 1 repos | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Agentsop LLM Engine Selectionagentsope/SkillAlchemy | 459 | — | ~6.1k | Automated safety check: Pass | MIT | |
| DGX Spark Memory and Thermal Opswshobson/agents | 40k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Mem0 Platform SDKmem0ai/mem0 | 67k | 2 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 |
agentsope/SkillAlchemy
Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-project SOPs (LangGraph…
NVIDIA/skills
Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.
agentsope/SkillAlchemy
Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.
wshobson/agents
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
mem0ai/mem0
Adds persistent memory to AI apps with the Mem0 Python and TypeScript SDKs: store, search, update and delete user memories, with framework integrations.
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
ancoleman/ai-design-components
Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.
ancoleman/ai-design-components
Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.
ancoleman/ai-design-components
Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.
ancoleman/ai-design-components
Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.
ancoleman/ai-design-components
Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.
ancoleman/ai-design-components
Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.
Categories
LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components. Model Serving is an agent skill from ancoleman/ai-design-components. LLM and ML model deployment for inference.
Model Serving fits situations like: serving models in production; building AI APIs; optimizing inference.
Run `npx skills add ancoleman/ai-design-components --skill model-serving -a claude-code`. Or copy the skill folder (skills/model-serving in ancoleman/ai-design-components) into .claude/skills/model-serving in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ancoleman/ai-design-components --skill model-serving -a codex`. Or copy the skill folder (skills/model-serving in ancoleman/ai-design-components) into .agents/skills/model-serving in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill model-serving -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-serving, .gemini/skills/model-serving, .github/skills/model-serving and .opencode/skills/model-serving in your project.
Going by SKILL.md and its folder, Model Serving needs Python for the scripts in its folder and the command-line tools its instructions call (pip and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Model Serving is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 18k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Model Serving: Agentsop Framework Selection (agentsope/SkillAlchemy, 459 stars), Jetson LLM Benchmark (NVIDIA/skills, 3.5k stars), Agentsop LLM Engine Selection (agentsope/SkillAlchemy, 459 stars) and DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.
Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.