Better Auth Security Best Practices
EpicenterHQ/epicenter
Better Auth security hardening: rate limits, secrets, CSRF, trusted origins, cookies, sessions, OAuth tokens, and audit logging.
Best practices for building an unauthenticated public Q&A chatbot widget.
$ npx skills add swyxio/skills --skill public-qa-chatbot -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install swyxio/skills public-qa-chatbot --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/public-qa-chatbot .claude/skills/public-qa-chatbot && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "public-qa-chatbot" agent skill from https://github.com/swyxio/skills/tree/main/public-qa-chatbot into .claude/skills/public-qa-chatbot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "public-qa-chatbot", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/swyxio/skills/tree/main/public-qa-chatbotType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add swyxio/skills --skill public-qa-chatbot -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install swyxio/skills public-qa-chatbot --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/public-qa-chatbot .agents/skills/public-qa-chatbot && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "public-qa-chatbot" agent skill from https://github.com/swyxio/skills/tree/main/public-qa-chatbot into .agents/skills/public-qa-chatbot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "public-qa-chatbot", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill public-qa-chatbot -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install swyxio/skills public-qa-chatbot --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/public-qa-chatbot .cursor/skills/public-qa-chatbot && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "public-qa-chatbot" agent skill from https://github.com/swyxio/skills/tree/main/public-qa-chatbot into .cursor/skills/public-qa-chatbot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "public-qa-chatbot", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/swyxio/skills.git --path public-qa-chatbot--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add swyxio/skills --skill public-qa-chatbot -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install swyxio/skills public-qa-chatbot --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/public-qa-chatbot .gemini/skills/public-qa-chatbot && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "public-qa-chatbot" agent skill from https://github.com/swyxio/skills/tree/main/public-qa-chatbot into .gemini/skills/public-qa-chatbot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "public-qa-chatbot", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install swyxio/skills public-qa-chatbotInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add swyxio/skills --skill public-qa-chatbot -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/public-qa-chatbot .github/skills/public-qa-chatbot && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "public-qa-chatbot" agent skill from https://github.com/swyxio/skills/tree/main/public-qa-chatbot into .github/skills/public-qa-chatbot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "public-qa-chatbot", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill public-qa-chatbot -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install swyxio/skills public-qa-chatbot --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/public-qa-chatbot .opencode/skills/public-qa-chatbot && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "public-qa-chatbot" agent skill from https://github.com/swyxio/skills/tree/main/public-qa-chatbot into .opencode/skills/public-qa-chatbot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "public-qa-chatbot", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
public-qa-chatbotBest practices for building an unauthenticated public Q&A chatbot widget.
Public QA Chatbot is an agent skill from swyxio/skills. Best practices for building an unauthenticated public Q&A chatbot widget. Covers rate limiting, security hardening, cost optimization, semantic caching, observability, UX patterns, chat scroll behavior, and architecture. Tech-agnostic with concrete examples from a production implementation.
Its SKILL.md is about 9.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including assets (for example `MINTLIFY_VIRTUAL_FILESYSTEM.md`, `agents/openai.yaml` and `assets/vite-react-tanstack-chat-demo/package-lock.json`).
It sits in Backend & APIs, covering Rate limiting, Security review and Caching. It works with TanStack. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (TypeScript), which the agent can run.
Shell commands in SKILL.md call:
npmpnpmFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
tanstack.commintlify.comgithub.comai.engineeraistudio.google.complatform.openai.comagentskills.iosdk.vercel.aiupstash.combraintrust.devlangfuse.comarcjet.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYREDIS_TOKENVECTOR_TOKENBRAINTRUST_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Public QA Chatbot loads about 9.8k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 4,134 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 4,134 words, ~9,767 tokens.
.claude/skills/public-qa-chatbot/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.A comprehensive skill for building unauthenticated, public-facing Q&A chatbot widgets on marketing sites, conference pages, documentation portals, and similar contexts where you need to serve anonymous visitors while controlling cost and abuse.
Use this as a design and review catalog, not a requirement to add every listed service or subsystem. Select controls for the public exposure, cost, privacy, and UX risks created by the requested change. Server-side secret isolation, server-authoritative enforcement of abuse limits, and safe handling of user content are invariants where those boundaries are touched. Caching, vector search, tracing, virtualization, voice, and multi-layer infrastructure are recommendations or opportunities unless the requested scale or failure mode makes one necessary. Stop when the named widget behavior is usable and its concrete public-abuse/cost risks are bounded; report adjacent improvements as follow-ups.
Distilled from a production implementation powering the AI Engineer Europe 2026 conference chatbot, with additional chat-scroll lessons from TanStack Virtual's chat guidance and agentic retrieval lessons from Mintlify's virtual filesystem assistant. See MINTLIFY_VIRTUAL_FILESYSTEM.md for a clean markdown reference version of the Mintlify pattern.
For a runnable React/TanStack Virtual demo of long chat scroll behavior plus an expanded bottom command shelf, use assets/vite-react-tanstack-chat-demo. The demo includes hover/double-click message controls, subtle token/latency stats, tool-call and multimodal examples, assistant response variants via left/right swipe, and a Realtime voice capture strip with live transcription and an audiogram.
Image: Public Q&A chatbot demo with live voice transcription and bottom command shelf
Run the demo locally:
cd assets/vite-react-tanstack-chat-demo
npm install
npm run dev -- --port 5179To exercise the Realtime voice path, start the dev server with OPENAI_API_KEY set. Keep the standard API key on the server side only; the browser should receive an ephemeral Realtime client secret.
This skill is written to be tech-agnostic. The reference implementation uses the stack below, but each component is swappable:
| Component | Reference choice | Alternatives |
|---|---|---|
| LLM provider | Gemini 3.1 Flash-Lite (via @ai-sdk/google) | OpenAI GPT-4o-mini, Anthropic Claude Haiku, Mistral, Llama via Groq/Together |
| AI SDK | Vercel AI SDK v6 (ai) | LangChain, LlamaIndex, direct provider SDKs |
| Hosting | Vercel (serverless functions) | Cloudflare Workers, AWS Lambda, Railway, Fly.io, Render |
| Rate limiting | Upstash Redis (@upstash/ratelimit) | Cloudflare Rate Limiting, AWS WAF, Redis (self-hosted), Arcjet |
| Semantic cache | Upstash Vector + Gemini Embeddings | Pinecone, Weaviate, Qdrant, pgvector, Cloudflare Vectorize |
| Agentic docs retrieval | Read-only virtual filesystem over indexed docs | Plain RAG, hosted search API, real sandbox only for async/developer tools |
| Embedding model | Gemini text-embedding-004 (128 dims) | OpenAI text-embedding-3-small, Cohere Embed v3, Voyage AI |
| Observability | Braintrust (wrapAISDK) | Langfuse, Helicone, LangSmith, OpenTelemetry, Datadog LLM Obs |
| Frontend | React (inline component) | Vue, Svelte, vanilla JS, Web Components |
| Long chat virtualization | TanStack Virtual chat support | Native scroll for short widgets, react-virtuoso, custom virtual list only when already proven |
Do not require virtualization for every public FAQ widget. A short, bounded chat can stay as a simple DOM list. Reach for a virtualized chat list when conversations can grow long, rows have dynamic heights, older history prepends, or streaming output makes scroll anchoring fragile. When using React and a virtualized list is justified, prefer TanStack Virtual's chat support over custom scroll math.
Do not require agentic document browsing for every public FAQ widget either. Plain RAG is sufficient for short, stable FAQs. Add a virtual docs filesystem when answers live across multiple pages, users ask for exact syntax, docs have a meaningful hierarchy, or top-k retrieval often misses the section an expert would grep for.
Apply limits at multiple granularities to prevent abuse:
// Example constants
const LIMITS = {
turnsPerSession: 9,
sessionsPerVisitorPerDay: 15,
globalSessionsPerDay: 3000,
};In-memory rate limiting resets on every serverless cold start and isn't shared across instances. Use a distributed store for production:
Upstash Redis (reference):
import { Ratelimit } from "@upstash/ratelimit";
import { Redis } from "@upstash/redis";
const redis = new Redis({ url: REDIS_URL, token: REDIS_TOKEN });
const limiter = new Ratelimit({
redis,
limiter: Ratelimit.slidingWindow(15, "1 d"), // 15 per day
prefix: "chatbot:visitor",
});
const { success } = await limiter.limit(clientIp);Alternatives:
ioredis + custom sliding window logicAlways keep an in-memory fallback for local development:
const useDistributed = !!redisUrl && !!redisToken;
if (!useDistributed) {
// Fall back to in-memory Map for local dev
}Never trust client-reported turn counts or session flags. The server must count turns from the messages array itself:
// Server counts turns - never trust client-reported values
const userTurnCount = messages.filter(m => m.role === "user").length;
const isNewSession = userTurnCount <= 1;Only increment the session counter after the server confirms a successful response, not when the user submits. This prevents phantom session counts from failed requests, network errors, or aborted streams:
// Client-side: count after first assistant response arrives
useEffect(() => {
const hasAssistantMessage = messages.some(m => m.role === "assistant");
if (hasAssistantMessage && !sessionCounted.current) {
sessionCounted.current = true;
incrementSessionCount();
}
}, [messages]);When a request is not a new session (i.e. a follow-up turn in an existing conversation), skip daily session counter increments entirely. Only the first turn of a conversation should count as a "session" for rate limiting purposes:
if (!isNewSession) {
return { allowed: true }; // Skip session counting for follow-up turns
}When rate-limited, let users input their own API key to continue chatting. This turns abuse into the user's own cost while preserving good UX:
// Skip rate limiting when user provides their own key
if (!userApiKey) {
const limit = await checkRateLimit(ip, turnCount, isNewSession);
if (!limit.allowed) {
return res.status(429).json({ error: limit.reason, rateLimited: true });
}
}
const apiKey = userApiKey || serverKey;Provide a direct link to obtain a key (e.g. https://aistudio.google.com/apikey for Gemini, https://platform.openai.com/api-keys for OpenAI).
Check the Origin or Referer header against an allowlist. This prevents cross-site request abuse where third parties embed scripts that burn your API quota:
const origin = req.headers.origin ?? req.headers.referer ?? "";
const allowedHosts = ["localhost", "yourdomain.com", "vercel.app"];
if (origin && !allowedHosts.some(h => origin.includes(h))) {
return res.status(403).json({ error: "Forbidden" });
}Note: Substring matching (
origin.includes(h)) is acceptable for v1 but could theoretically match crafted domains. For stricter validation, parse the URL and compare the hostname.
Cap both the number of messages and individual message length to prevent token-stuffing attacks that run up your LLM bill:
const MAX_MESSAGES = 10;
const MAX_MESSAGE_LENGTH = 2000;
const trimmedMessages = messages.slice(-MAX_MESSAGES).map(m => ({
...m,
parts: m.parts.map(p =>
p.type === "text" && typeof p.text === "string"
? { ...p, text: p.text.slice(0, MAX_MESSAGE_LENGTH) }
: p
),
}));Also limit model output: maxOutputTokens: 500 for short Q&A answers.
Never trust as casts for user-supplied values. Validate against a known set:
const VALID_PAGES = new Set(["europe", "home", "worldsfair"]);
if (!VALID_PAGES.has(page)) {
return res.status(400).json({ error: "Invalid page parameter." });
}For documentation-backed chatbots, access control must happen before retrieval, not after answer generation. If the bot exposes semantic search, exact search, or a virtual docs filesystem, apply the same visibility filter to every surface:
isPublic, groups, tenantId, docsVersion, or equivalent metadata with indexed chunks so filters are cheap and testable.} catch {
console.error("Chat API error");
return res.status(500).json({
error: "An error occurred processing your request. Please try again.",
});
}Use the platform's trusted headers. On Vercel: x-real-ip > x-vercel-forwarded-for > x-forwarded-for. The standard x-forwarded-for is spoofable by clients.
Alternatives:
CF-Connecting-IPX-Forwarded-For (first IP is trustworthy when set by AWS)Fastly-Client-IPIf you only need text responses, explicitly restrict the model:
const model = provider("gemini-3.1-flash-lite", {
responseModalities: ["TEXT"], // Gemini-specific
// For OpenAI: modalities: ["text"]
});Also state "text-only assistant" in the system prompt as a defense-in-depth measure.
Use vector similarity search to cache and reuse responses for semantically similar questions. Most effective for FAQ-style chatbots where users ask the same questions in different words.
Upstash Vector (reference):
import { Index } from "@upstash/vector";
const vectorIndex = new Index({ url: VECTOR_URL, token: VECTOR_TOKEN });
// Lookup: check cache before calling LLM
const results = await vectorIndex.query({
vector: await getEmbedding(question),
topK: 1,
includeMetadata: true,
filter: `page = '${page}'`,
});
if (results[0]?.score >= 0.92 && results[0]?.metadata?.answer) {
return results[0].metadata.answer; // Cache hit - skip LLM call
}
// Store: cache after LLM responds (fire-and-forget)
void vectorIndex.upsert({
id: `cache-${Date.now()}`,
vector: embedding,
metadata: { question, answer, page, cachedAt: Date.now() },
});Key decisions:
Alternatives:
Always store a cachedAt timestamp in cache entry metadata. On lookup, reject entries older than your TTL (e.g. 7 days). This prevents stale answers from persisting indefinitely, especially when FAQ content changes:
const CACHE_TTL_MS = 7 * 24 * 60 * 60 * 1000; // 7 days
if (Date.now() - result.metadata.cachedAt > CACHE_TTL_MS) {
// Stale - treat as cache miss
}When returning a cached response, use the same streaming protocol as live LLM responses. Don't switch to a different response format (e.g. manual Data Stream Protocol vs. UI Message Stream). Inconsistent formats cause client-side parsing errors and broken UX:
// BAD: different format for cache hits
res.write(`0:${JSON.stringify(cachedText)}\n`); // Manual Data Stream Protocol
// GOOD: same format for both paths
const stream = createUIMessageStream({ /* ... */ });
pipeUIMessageStreamToResponse(stream, res);For docs assistants that expose grep-style tools, avoid scanning every page or chunk over the network. Use a two-stage exact-search path:
$contains, full-text indexes, trigram search, or metadata filters by section/path.page and chunk_index.{ path, docsVersion } so repeated grep/cat workflows do not hit the database twice.Log candidate count and final hit count. If the coarse filter returns too many pages, ask the model to narrow the query instead of silently running an expensive full-corpus scan.
Offer a browsable FAQ list alongside the chat interface. This serves users who have common questions without making any LLM calls at all:
// Structured FAQ data for UI rendering
export const FAQ_QUESTIONS: Array<{
category: string;
question: string;
answer: string;
}> = [
{ category: "Ticketing", question: "Can I get a refund?", answer: "Yes, per our refund policy..." },
// ...
];Organize by category with expandable sections. Clicking a question can either show the pre-written answer directly or send it to the chat for a more detailed LLM response.
For a constrained Q&A chatbot, you rarely need the most powerful model:
| Model | Input cost | Output cost | Best for |
|---|---|---|---|
| Gemini 3.1 Flash-Lite | $0.25/1M | $1.50/1M | Cheapest, good for FAQ |
| GPT-4o-mini | $0.15/1M | $0.60/1M | Good balance of cost/quality |
| Claude Haiku | $0.25/1M | $1.25/1M | Fast, good at following instructions |
| Llama 3.3 70B (via Groq) | Free tier available | Free tier available | Cost-sensitive prototypes |
Set maxOutputTokens to the minimum needed (e.g. 500 tokens for 2-4 sentence answers). This caps cost per request and keeps responses concise.
Pre-build and cache the system prompt context at module level. This avoids re-computing expensive string concatenations on every request:
let cachedContext: Record<string, string> | null = null;
function buildContext(): Record<string, string> {
if (cachedContext) return cachedContext;
// ... expensive computation ...
cachedContext = result;
return cachedContext;
}Instrument all LLM calls with input/output, latency, token usage, and cost. This is essential for monitoring abuse, debugging, and cost tracking.
Braintrust (reference):
import { initLogger, wrapAISDK } from "braintrust";
initLogger({ projectName: "my-chatbot", apiKey: BRAINTRUST_API_KEY });
const { streamText } = wrapAISDK(ai); // Auto-traces all callsAlternatives:
Track cache hit rates to understand cost savings and tune the similarity threshold. A cache hit is a "free" response that saved an LLM call.
Trace retrieval tools separately from LLM calls. For semantic search, exact search, and virtual filesystem tools, log:
grep/cat/ls exploration.This tells you whether agentic retrieval is improving answer quality or just adding cost and latency.
Avoid logging full error objects, API keys, or user PII. Log just enough to debug (error type, status codes, IP hashes).
Enable markdown in chat responses and instruct the model to use it via the system prompt:
You may use markdown formatting in your responses when appropriate:
- Use **bold** for emphasis on key information like dates, prices, or venue names
- Use [links](url) when referencing websites
- Use bullet points for lists of speakers, sessions, or options
- Keep formatting light and readableReact: react-markdown + remark-gfm
Vue: vue-markdown-render
Vanilla JS: marked or markdown-it
Let users reposition and resize the chat window. Persist geometry to localStorage so it survives page reloads. Clamp positions to viewport bounds:
const newX = Math.max(0, Math.min(
e.clientX - dragOffset.x,
window.innerWidth - geometry.width
));Always stream responses for perceived speed. Use your SDK's streaming API rather than waiting for the full response. The first token appearing quickly matters more than total latency.
For the UI, update one in-progress assistant message as tokens arrive. Do not append a new message row per token. Token-level rows are expensive, break transcript semantics, and make scroll anchoring harder.
Public Q&A widgets often start short, so a plain scroll container is fine until there is evidence it is not. Add virtualization when the widget can hold long histories, rich markdown, tool results, images, code blocks, history pagination, or token-streaming messages that grow in height.
When virtualization is warranted, recommend TanStack Virtual's chat support for React implementations, but keep it optional and swappable. The important lessons are the scroll contracts:
flex-direction: column-reverse, inverted transforms, and hand-maintained scrollTop += delta bookkeeping.setMessages((current) => [...olderMessages, ...current]).80px, rather than exact-bottom checks that are brittle across browsers and dynamic heights.hasMoreHistory, loading flags, and request dedupe in app state. The virtualizer should receive the current ordered message array, not own data fetching.TanStack Virtual maps these lessons to anchorTo: 'end', followOnAppend, scrollEndThreshold, stable getItemKey, measureElement, isAtEnd(), getDistanceFromEnd(), and scrollToEnd(). These APIs are useful defaults, not a hard dependency.
Do not treat the composer as only a textbox. For AI applications, the bottom of the screen is valuable thumb-reachable space for the actions users need while forming a prompt: attach, tools, model, voice, send, mode, reasoning depth, runtime context, and tool launchers.
Use a progressive bottom command shelf when the app has enough controls to justify it:
Plan / Build, effort level, device/project/branch, and budget or usage.Every optional service should have a fallback:
| Service | If unavailable... |
|---|---|
| Redis (rate limiting) | Fall back to in-memory counters |
| Vector DB (cache) | Skip semantic caching, always call LLM |
| Observability (tracing) | Skip tracing, log locally |
| Server API key | Prompt user for BYOK |
| Virtualized chat list | Fall back to a bounded native scroll list with transcript limits |
// Pattern: optional service with graceful fallback
const vectorIndex = vectorUrl && vectorToken
? new Index({ url: vectorUrl, token: vectorToken })
: null; // null = skip caching
if (vectorIndex) { /* try cache */ }
// Always falls through to LLM callShow top FAQ questions on hover over the chat bubble. This gives users an immediate sense of what the chatbot can help with and reduces "what do I ask?" friction.
When embedding a chatbot widget on a page that supports dark/light mode, make the chatbot colors contrast with the page background:
Accept the page's theme state (e.g. isDark prop) and derive all colors from a single theme palette function. Use useMemo to avoid recalculating on every render:
const theme = useMemo(() => getTheme(isDark), [isDark]);
// getTheme returns 40+ color tokens: bg, text, borders, buttons, surfaces, shadowsDefine comprehensive color tokens so every UI element adapts. This avoids hardcoded colors scattered throughout the component and makes the entire widget respond to theme changes in one place.
Design the chatbot as a single component that accepts props so it can be dropped into any page with different branding/context:
<Chatbot
page="europe"
accentColor="#7C3AED"
title="AI Engineer Europe Assistant"
/>Instead of stuffing all data into the system prompt, expose tools that the model can call on-demand. This keeps the context window smaller and responses more accurate:
tools: {
search_speakers: tool({
description: "Search for speakers by name, company, or role",
inputSchema: jsonSchema<{ search?: string }>({ ... }),
execute: async (args) => searchSpeakers(args),
}),
search_sessions: tool({
description: "Search sessions by title, speaker, day, type, or track",
inputSchema: jsonSchema<{ search?: string; day?: string }>({ ... }),
execute: async (args) => searchSessions(args),
}),
}Top-k RAG works for simple FAQ questions, but it breaks down when the answer spans several pages, the user needs exact syntax, or the correct page does not land in the nearest embedding results. For documentation-backed chatbots, consider exposing the knowledge base as a read-only virtual filesystem so the model can explore with familiar tools such as ls, cat, find, and grep.
The important idea is to give the model the filesystem workflow, not necessarily a real filesystem. Mintlify's ChromaFs pattern maps shell commands onto an existing docs index instead of booting a sandbox for every visitor. That matters for public chatbot latency and cost: their article reports p90 session creation dropping from about 46s with sandbox/repo setup to about 100ms with a virtual filesystem over Chroma.
Recommended shape:
Set<path> plus Map<directory, children> so ls, cd, and basic find do not need network calls.cat /path/page.mdx by fetching all chunks for that page, sorting by chunk_index, and reassembling the full page. Cache page reads during the session so repeated inspection is cheap.ls, but fetch content only when the model runs cat.EROFS-style error so the assistant can explore freely without state cleanup or cross-user mutation risk.grep as a two-stage search: use the vector/document database as a coarse filter to identify candidate pages, then run exact string or regex matching in memory over the fetched candidates. This gives exact-match behavior without scanning every file over the network.Expose the virtual filesystem as narrow tools rather than a general shell when possible:
tools: {
list_docs: tool({
description: "List child paths under a documentation directory.",
inputSchema: jsonSchema<{ path: string }>({ ... }),
execute: async ({ path }) => docsFs.ls(path),
}),
read_doc: tool({
description: "Read a full documentation page by path.",
inputSchema: jsonSchema<{ path: string }>({ ... }),
execute: async ({ path }) => docsFs.cat(path),
}),
search_docs_exact: tool({
description: "Search docs by exact string or regex and return matching paths/snippets.",
inputSchema: jsonSchema<{ pattern: string; regex?: boolean }>({ ... }),
execute: async ({ pattern, regex }) => docsFs.grep(pattern, { regex }),
}),
}Use this pattern when the chatbot needs to behave like a docs expert. Keep normal semantic search as the first-pass tool for broad questions, then let the model escalate to grep/cat/ls when it needs exact wording, syntax, cross-page synthesis, or source-grounded citations.
Structure the system prompt with these sections in order:
For chatbot endpoints that need streaming + external service calls (Redis, Vector DB, observability), use a standard API route / serverless function rather than edge functions. Edge functions have stricter size/dependency limits and cold start characteristics that can cause issues with multiple SDK imports.
For large documentation sites, generate a docs manifest alongside the chunk index:
type DocsPath = {
path: string; // "/auth/oauth.mdx"
title: string; // "OAuth"
isPublic: boolean;
groups: string[];
updatedAt: string;
sourceId: string;
docsVersion: string;
};At runtime, load the access-pruned manifest into memory:
Set<string> for valid file paths.Map<string, string[]> for directory-to-children lookup.find_docs behavior.This makes list_docs, find_docs, and path validation memory-only. Rebuild or invalidate the tree when docsVersion changes.
Maintain two representations of FAQ data:
question, answer, category fields for rendering the FAQ list view// System prompt context (flat text)
export const FAQ_KNOWLEDGE_BASE = `
## TICKETING & PRICING
Q: Can I get a refund?
A: Yes, per our refund policy...
`;
// UI list view (structured)
export const FAQ_QUESTIONS = [
{ category: "Ticketing", question: "Can I get a refund?", answer: "Yes..." },
];Chunked vector results are good for discovery, but they are often too lossy for final answers. For read_doc(path) or citation verification, fetch the whole page:
const chunks = await vectorIndex.query({
topK: 200,
includeMetadata: true,
filter: `path = '${path}' AND docsVersion = '${docsVersion}'`,
});
return chunks
.sort((a, b) => a.metadata.chunk_index - b.metadata.chunk_index)
.map(chunk => chunk.metadata.text)
.join("\n\n");Cache full-page reads by { path, docsVersion }. This lets the model answer exact syntax, multi-section, and "compare these pages" questions with the same source material a human docs reader would inspect.
Always include practical information (venue name, address, dates, ticket URLs) directly in the context. These are the most common questions and should never require a tool call.
Libraries like html2canvas that clone and manipulate the DOM can interfere with React's virtual DOM reconciliation, causing page reloads, lost state, or broken event handlers. If you need page screenshots, use native browser APIs (navigator.mediaDevices.getDisplayMedia) or capture at the server level instead.
The failure mode for long chatbot widgets is usually scattered scroll math: column-reverse, inverted transforms, manual offset deltas, unconditional scrollToBottom, and index-based keys. These hacks often pass short manual tests and then fail when older history loads, the assistant streams a long markdown answer, or the user reads history while new output arrives.
Prefer a single scroll contract:
LLM model IDs change frequently and may require suffixes like -preview. A wrong model ID can return a 200 OK response with an empty or errored stream body, making it look like a frontend bug. Always verify the exact model ID against the provider's docs and test with a real API call before deploying.
Never skip pnpm build / npm run build before pushing to a branch. TypeScript errors, import issues, and other compilation failures caught locally are much faster to fix than waiting for CI. This is especially important when multiple people are editing the same files.
Use this checklist when building a new public Q&A chatbot:
ls/cat/grep-style exploration without per-user sandboxessrc/pages/api/chat.ts and src/components/Chatbot.tsx)© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 13 other files (assets) in public-qa-chatbot of swyxio/skills.
Open the folder on GitHubat commit 038ef34
Public QA Chatbot next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Public QA Chatbot this skillswyxio/skills | 176 | — | ~9.8k | Automated safety check: Pass | MIT | |
| Better Auth Security Best PracticesEpicenterHQ/epicenter | 4.8k | — | ~896 | Automated safety check: Pass | Custom licence | |
| Tanstack Query Best PracticesDeckardGer/tanstack-agent-skills | 222 | 1 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Trpc PatternsFranciscoMoretti/chat-js | 1.2k | — | ~288 | Automated safety check: Pass | Apache-2.0 | |
| Tanstack Integration Best PracticesDeckardGer/tanstack-agent-skills | 222 | — | ~644 | Automated safety check: Pass | MIT | |
| API IntegrationHack23/cia | 239 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 |
EpicenterHQ/epicenter
Better Auth security hardening: rate limits, secrets, CSRF, trusted origins, cookies, sessions, OAuth tokens, and audit logging.
DeckardGer/tanstack-agent-skills
TanStack Query (React Query) best practices for data fetching, caching, mutations, and server state management.
FranciscoMoretti/chat-js
Implement ChatJS tRPC routers, TanStack Query hooks, mutations, and cache invalidation.
DeckardGer/tanstack-agent-skills
Best practices for integrating TanStack Query with TanStack Router and TanStack Start.
Hack23/cia
External API integration patterns, retry logic, circuit breakers, caching, rate limiting for government data APIs
agentfront/frontmcp
A skill your agent uses when configuring a FrontMCP server through frontmcp.config or the @FrontMcp options.
swyxio/skills
Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.
swyxio/skills
Design, implement, audit, or refresh protected username and handle namespaces for public products.
swyxio/skills
Fully automated new Mac setup for fullstack web developers and AI engineers.
swyxio/skills
Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…
swyxio/skills
Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.
swyxio/skills
Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.
Works with
Categories
Best practices for building an unauthenticated public Q&A chatbot widget. Public QA Chatbot is an agent skill from swyxio/skills. Best practices for building an unauthenticated public Q&A chatbot widget.
Public QA Chatbot fits situations like: tasks that involve Rate limiting; tasks that involve Security review; tasks that involve Caching.
Run `npx skills add swyxio/skills --skill public-qa-chatbot -a claude-code`. Or copy the skill folder (public-qa-chatbot in swyxio/skills) into .claude/skills/public-qa-chatbot in your project. Claude Code loads it when a task matches its description.
Run `npx skills add swyxio/skills --skill public-qa-chatbot -a codex`. Or copy the skill folder (public-qa-chatbot in swyxio/skills) into .agents/skills/public-qa-chatbot in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill public-qa-chatbot -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/public-qa-chatbot, .gemini/skills/public-qa-chatbot, .github/skills/public-qa-chatbot and .opencode/skills/public-qa-chatbot in your project.
Going by SKILL.md and its folder, Public QA Chatbot needs TypeScript for the scripts in its folder, the command-line tools its instructions call (npm and pnpm) and credentials named OPENAI_API_KEY, REDIS_TOKEN, VECTOR_TOKEN and BRAINTRUST_API_KEY. Our summary lists: Node.js; A credential in OPENAI_API_KEY; A credential in REDIS_TOKEN.
SKILL.md names 12 domains. As links in the text: tanstack.com, mintlify.com, github.com, ai.engineer, aistudio.google.com, platform.openai.com, agentskills.io, sdk.vercel.ai, upstash.com, braintrust.dev, langfuse.com and arcjet.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Public QA Chatbot is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 9.8k tokens (SKILL.md is roughly 39k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Public QA Chatbot: Better Auth Security Best Practices (EpicenterHQ/epicenter, 4.8k stars), Tanstack Query Best Practices (DeckardGer/tanstack-agent-skills, 222 stars), Trpc Patterns (FranciscoMoretti/chat-js, 1.2k stars) and Tanstack Integration Best Practices (DeckardGer/tanstack-agent-skills, 222 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
swyxio (a GitHub user) maintains it in swyxio/skills, which has 176 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.
Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.