Caveman Gateway Setup
JuliusBrussee/caveman
Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.
Answer questions about LLM and agentic-application behavior from data already ingested into Elastic: latency and error rate, token and cost utilization, response quality and guardrail events, and…
$ npx skills add elastic/agent-skills --skill observability-llm-obs -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install elastic/agent-skills observability-llm-obs --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/elastic/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/observability/llm-obs .claude/skills/observability-llm-obs && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "observability-llm-obs" agent skill from https://github.com/elastic/agent-skills/tree/main/skills/observability/llm-obs into .claude/skills/observability-llm-obs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-llm-obs", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/elastic/agent-skills/tree/main/skills/observability/llm-obsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add elastic/agent-skills --skill observability-llm-obs -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install elastic/agent-skills observability-llm-obs --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/elastic/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/observability/llm-obs .agents/skills/observability-llm-obs && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "observability-llm-obs" agent skill from https://github.com/elastic/agent-skills/tree/main/skills/observability/llm-obs into .agents/skills/observability-llm-obs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-llm-obs", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add elastic/agent-skills --skill observability-llm-obs -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install elastic/agent-skills observability-llm-obs --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/elastic/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/observability/llm-obs .cursor/skills/observability-llm-obs && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "observability-llm-obs" agent skill from https://github.com/elastic/agent-skills/tree/main/skills/observability/llm-obs into .cursor/skills/observability-llm-obs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-llm-obs", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/elastic/agent-skills.git --path skills/observability/llm-obs--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add elastic/agent-skills --skill observability-llm-obs -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install elastic/agent-skills observability-llm-obs --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/elastic/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/observability/llm-obs .gemini/skills/observability-llm-obs && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "observability-llm-obs" agent skill from https://github.com/elastic/agent-skills/tree/main/skills/observability/llm-obs into .gemini/skills/observability-llm-obs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-llm-obs", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install elastic/agent-skills observability-llm-obsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add elastic/agent-skills --skill observability-llm-obs -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/elastic/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/observability/llm-obs .github/skills/observability-llm-obs && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "observability-llm-obs" agent skill from https://github.com/elastic/agent-skills/tree/main/skills/observability/llm-obs into .github/skills/observability-llm-obs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-llm-obs", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add elastic/agent-skills --skill observability-llm-obs -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install elastic/agent-skills observability-llm-obs --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/elastic/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/observability/llm-obs .opencode/skills/observability-llm-obs && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "observability-llm-obs" agent skill from https://github.com/elastic/agent-skills/tree/main/skills/observability/llm-obs into .opencode/skills/observability-llm-obs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "observability-llm-obs", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
observability-llm-obsAnswer questions about LLM and agentic-application behavior from data already ingested into Elastic: latency and error rate, token and cost utilization, response quality and guardrail events, and…
Observability LLM Obs is an agent skill from elastic/agent-skills, published by the product's own GitHub organization. Answer questions about LLM and agentic-application behavior from data already ingested into Elastic: latency and error rate, token and cost utilization, response quality and guardrail events, and agentic call-chain orchestration. Use when the user asks about LLM monitoring, GenAI observability, token spend or AI cost, model latency, prompt or guardrail failures, or how an agent's tool-call chain executed.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/esql-recipes.md`, `references/genai-fields.md` and `references/integrations.md`). Compatibility notes: Requires the elastic CLI (= 0.2) with es and kb support, and an Elasticsearch deployment holding LLM telemetry from APM/OTLP traces or an Elastic LLM…
It sits in DevOps & Cloud, covering Observability and LLM observability. It works with Elasticsearch. The repository describes itself as: Official Elastic Skills. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit baa5111. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are esql).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
elastic.cogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires the `elastic` CLI (>= 0.2) with `es` and `kb` support, and an Elasticsearch deployment holding LLM telemetry from APM/OTLP traces or an Elastic LLM integration. Base floor is Elasticsearch 8.11+ or Serverless. The time-series path for integration metrics needs `TS` (Stack GA 9.4) and `TRANGE` (Stack GA 9.3); both are GA on Serverless. Below Stack 9.4 fall back to `FROM` with `BUCKET`. SLO and alerting lookups require Kibana. No Kibana UI is required.
From compatibility in the SKILL.md frontmatter.
Observability LLM Obs loads about 4.9k tokens when it runs, and up to ~9.9k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 2,206 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from elastic/agent-skills at commit baa5111, republished under its Apache-2.0 licence (© elastic). 2,206 words, ~4,921 tokens.
.claude/skills/observability-llm-obs/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Answer questions about monitoring LLMs and agentic components using data actually ingested into Elastic — nothing else. The four questions this skill answers are LLM performance, cost and token utilization, response quality, and call chaining or agentic workflow orchestration.
A given deployment typically uses one or more ingestion paths: APM/OTLP traces, and/or integration metrics and logs. Which one exists is a discovery result, not an assumption — never assume both are present. For ES|QL syntax, commands, and query patterns, use the elasticsearch-esql skill. For service-level latency and error triage that is not LLM-specific, use the observability-sre-triage skill.
<!-- begin-partial: preamble -->
This skill executes Elasticsearch operations through the elastic CLI. If the
elastic CLI is not installed, tell the user what it is needed for. Do
not guess credentials, call the HTTP API directly, or attempt other workarounds.
This skill references operations in HTTP-shorthand form (e.g., GET /, GET /_cat/indices, GET /{index}/_mapping,
GET /{index}/_settings/index.mode, POST /_query). The Operations table at the end of this document
maps each shorthand to the equivalent elastic CLI command — always use the CLI rather than calling the HTTP API
directly.
<!-- end-partial: preamble -->
The CLI check above gates querying the cluster — it does not gate analysis. When the user has already supplied the evidence in their question (metric values, counts, status reasons, log lines, alert payloads, configuration), reason from that evidence and deliver the conclusion.
When you genuinely do need data the user has not provided, still say what you would check and how — name the specific query, index, and field that would settle the question — and then ask for CLI setup. An answer that names the check is useful without a cluster; one that only asks for setup is not.
Applies to every response produced under this skill.
| Ingestion path | Index patterns | What it can answer |
|---|---|---|
| OTel / EDOT traces via OTLP | traces-*.otel-*, or the generic traces-* | Per-request latency, tokens, models, finish reasons, and call chains |
| Elastic APM agent traces | traces-apm*, or the generic traces-* | Per-request latency and outcome; GenAI attributes if the SDK adds them |
| APM/OTel metrics | metrics-apm*, metrics-*.otel-* | Aggregated service-level rates and latencies |
| Elastic LLM integration metrics | metrics-<integration>.* (for example metrics-aws_bedrock_agentcore.*) | Aggregate token counts, invocations, latency, sometimes cost |
| Elastic LLM integration logs | logs-<integration>.* | Prompt/response records, guardrail and content-filter events |
Use the generic traces-* pattern to find trace data regardless of whether an Elastic APM agent or OpenTelemetry
collected it. Instrumentation can come from EDOT, OpenLLMetry, OpenLIT, or Langtrace exporting to OTLP — all of them
land LLM and agent spans in the trace data streams.
Only traces can reconstruct a call chain. trace.id, span.id, and parent.id are the only way to rebuild an
agentic call chain. Integration metrics are pre-aggregated and cannot do it — if the question is about orchestration and
only integration metrics exist, say the chain cannot be reconstructed from the available data.
Background reading: LLM and agentic AI observability, EDOT LLM use cases, Observability Labs — LLM Observability.
Run these in order. Do not skip to step 5.
Verify the connection and detect the version. Call GET /. The decision this drives is which query language
surface is available: build_flavor: "serverless" means all ES|QL features including TS, TRANGE, and TBUCKET
are available. Otherwise use version.number: on Stack, TS is preview in 9.2 and GA in 9.4, and TRANGE is GA in
9.3.0, so treat Stack 9.4 as the floor for the whole time-series path and fall back to FROM with
BUCKET(@timestamp, ...) below it. Getting this wrong produces a query that fails to parse.
Determine which ingestion paths exist. The decision is whether this deployment has traces, integration data, or
both — and it changes which questions are answerable at all. List data streams with GET /_data_stream/<name> (pass
traces-*, metrics-*, logs-*) or resolve patterns with GET /_resolve/index/<pattern>. Look for trace data
streams, metrics-apm*, and any metrics-* or logs-* matching a known LLM integration dataset. If neither path
has LLM data, say so and stop — do not answer from general knowledge about the model provider.
Discover the real LLM field names. The decision is which exact field paths to write into ES|QL, and it cannot be
guessed. Naming differs across instrumentations: gen_ai.* versus llm.* versus integration-specific names, and the
semantic conventions themselves have shifted (older instrumentation emits gen_ai.system, newer emits
gen_ai.provider.name). Use GET /_field_caps with a field pattern such as *gen_ai*, *llm*, *token*,
*cost*, or read GET /<index>/_mapping, then sample a document to confirm the values are populated. Attribute
nesting differs by ingestion path — OTel-native trace data streams expose span attributes as attributes.<name>
(and often as a bare passthrough <name>), not as span.attributes.<name>. Confirm which form resolves before
writing a query. See references/genai-fields.md for the attribute catalog and the
resolution rules.
Decide whether a cost field exists. Cost is not part of the OpenTelemetry GenAI specification. Some
instrumentations add a custom attribute such as llm.response.cost.usd_estimate, and some integrations expose a cost
metric, but many deployments have neither. Look for it explicitly in step 3. If it is absent, the answer to a cost
question is that cost is not instrumented — report token counts instead and name the gap.
Choose ONE consistent source per question. When both APM traces and integration metrics exist, pick one and use it for the whole answer. Mixing them double-counts and produces two different numbers for the same quantity, because the integration polls the provider's own accounting while the traces record what the client observed. Route by question type: traces for per-request analysis, call chains, and anything needing a trace hierarchy; integration metrics for aggregate token and cost totals over long windows. State which source the answer came from.
Check alerts and SLOs when relevance is plausible. The decision is whether a degradation is already known and
tracked. Find rules with GET kbn:/api/alerting/rules/_find and SLOs with GET kbn:/api/observability/slos, then
filter to those targeting LLM-related services or integration metrics — the field names from step 3 tell you which
rules are related. Firing alerts, or SLOs in violated or degrading status, are evidence of degraded performance. Note
that the SLO API's sli.kql.custom indicator takes KQL rather than ES|QL; that is an API contract, not a
recommendation to use KQL elsewhere.
Write queries with POST /_query. Always bound the time range, add service.name when present, and LIMIT results.
Use coarse buckets when only a trend is needed rather than scanning a wide window at fine granularity.
| Question | Traces path | Integration path |
|---|---|---|
| Latency, throughput, error rate | Filter on the GenAI operation or model attribute; COUNT(*) per bucket, AVG(span.duration.us), and failures via event.outcome == "failure" | Request-rate, latency, and error metrics by model dimension |
| Tokens and cost | SUM the input and output token attributes by time, model, or service; add a cost attribute only if one exists | Token and cost metrics aggregated by time and model |
| Response quality and safety | event.outcome, error.type, and the finish-reason attribute; prompts and responses only if captured and not redacted | Guardrail blocks, content-filter events, and policy violations |
| Call chaining and orchestration | Traces only — group by trace.id, walk parent.id to span.id, aggregate by span name or GenAI operation | Not answerable — metrics are pre-aggregated |
On the trace path, per-span latency comes from duration (nanoseconds, populated on every OTel-native span) or
span.duration.us (microseconds, an APM-compatibility field that is null on many spans — sorting on it can silently
drop the slowest step). Confirm which is populated before using it. The slowest child span is the bottleneck.
Aggregating by span name or the GenAI operation attribute shows the distribution of step types across a workflow —
retrieval, LLM call, tool use.
For integration-specific data streams and field names — OpenAI, Azure OpenAI, Azure AI Foundry, Amazon Bedrock, Bedrock AgentCore, GCP Vertex AI — see references/integrations.md. For longer worked queries including the trace-hierarchy walk and the time-series integration pattern, see references/esql-recipes.md.
The field paths below are illustrative. Confirm the real paths from step 3 before running anything — this skill's own method is to discover field names first, and the correct nesting depends on the ingestion path.
"How many tokens are we burning per model?" — confirm the token attributes exist and resolve, then sum them by model and time bucket. Report the totals returned:
FROM traces-*
| WHERE @timestamp > NOW() - 24 hours AND attributes.gen_ai.request.model IS NOT NULL
| EVAL in_tok = TO_LONG(attributes.gen_ai.usage.input_tokens),
out_tok = TO_LONG(attributes.gen_ai.usage.output_tokens)
| STATS input_tokens = SUM(in_tok), output_tokens = SUM(out_tok)
BY hour = BUCKET(@timestamp, 1 hour), attributes.gen_ai.request.model
| SORT hour
| LIMIT 500"What is our LLM spend?" — look for a cost field with GET /_field_caps on *cost* before querying. If none
exists, report token utilization and state plainly that cost is not instrumented in this deployment; adding a cost
attribute at the instrumentation layer, or enabling an integration that reports cost, is the fix. Do not multiply tokens
by a price you assumed.
"Which model is slowest, and is it failing?" — latency and error rate in one pass, grouped by model:
FROM traces-*
| WHERE @timestamp > NOW() - 24 hours AND attributes.gen_ai.request.model IS NOT NULL
| STATS request_count = COUNT(*),
failures = COUNT(*) WHERE event.outcome == "failure",
avg_duration_us = AVG(span.duration.us)
BY attributes.gen_ai.request.model
| EVAL error_rate = failures::double / request_count
| SORT avg_duration_us DESC
| LIMIT 100"Why is our agent slow?" — this needs the trace hierarchy, so it requires trace data. Find the traces containing
more than one LLM or tool span, ranked by time spent in those spans, then walk into the worst trace by trace.id to
find the bottleneck span. Because the WHERE restricts to LLM spans, the sum below is LLM time, not end-to-end trace
duration — name the column so it cannot be misread as wall-clock latency:
FROM traces-*
| WHERE @timestamp > NOW() - 3 hours AND attributes.gen_ai.operation.name IS NOT NULL
| STATS llm_span_count = COUNT(*), llm_duration_us = SUM(span.duration.us) BY trace.id
| WHERE llm_span_count > 1
| SORT llm_duration_us DESC
| LIMIT 50"Are prompts getting blocked?" — check the finish-reason attribute and error.type on the trace path, or the
integration's guardrail log events. A finish reason such as a content filter is a quality signal, not a transport error.
"Is there anything already alerting on this?" — GET kbn:/api/alerting/rules/_find and
GET kbn:/api/observability/slos, filtered to the LLM-related services or integration metrics identified in step 2.
GET /_field_caps, GET /<index>/_mapping, or a sample document. Never guess an attribute path.TS with TRANGE and TBUCKET on Serverless or Stack 9.4+; fall
back to FROM with BUCKET(@timestamp, ...) below that. Alias the bucket (BY bucket = TBUCKET(1 hour)) and sort on
the alias.integer in one
backing index and long in another after a rollover, which makes ES|QL refuse the field. Cast with TO_LONG(...) in
an EVAL before aggregating.TS command reference applies on Stack 9.4+
and Serverless, and the
FROM command reference applies
elsewhere.| HTTP API (shorthand) | elastic CLI command |
|---|---|
GET / | elastic es info |
GET /_data_stream/<name> | elastic es indices get-data-stream --name '<name>' |
GET /_resolve/index/<pattern> | elastic es indices resolve-index --name '<pattern>' |
GET /<index>/_mapping | elastic es indices get-mapping --index '<index>' |
GET /_field_caps | elastic es field-caps --index '<index>' --fields '<fields>' |
POST /_query | elastic es esql query --format tsv --query '<esql>' |
GET kbn:/api/alerting/rules/_find | elastic kb alerting get-alerting-rules-find --filter '<filter>' |
GET kbn:/api/observability/slos | elastic kb slo find-slos-op --space-id '<space>' --kql-query '<kql>' |
© elastic, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/observability/llm-obs of elastic/agent-skills.
Open the folder on GitHubat commit baa5111
Observability LLM Obs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Observability LLM Obs this skillelastic/agent-skills | 592 | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| Caveman Gateway SetupJuliusBrussee/caveman | 110k | 1 repos | ~2.6k | Automated safety check: Warn | Apache-2.0 | |
| UModel Root Cause Analysisalibaba/UnifiedModel | 412 | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Google Agents CLI Observabilitypifferologo/cloud-agents-cli | 129 | 1 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Agent Platform Alert Configurationgoogle/skills | 21k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Agent Observability Experiment Bootstrapdatadog-labs/agent-skills | 177 | — | ~2.3k | Automated safety check: Pass | MIT |
JuliusBrussee/caveman
Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
pifferologo/cloud-agents-cli
This skill should be used when the user wants to "set up tracing", "monitor my ADK agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring…
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
datadog-labs/agent-skills
Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK.
elastic/kibana
Implement and quality-check OpenTelemetry metric instrumentation in Kibana code that uses @kbn/metrics.
elastic/agent-skills
Triage Elastic Security alerts — gather context, classify threats, create cases, and acknowledge.
elastic/agent-skills
Create, search, update, and manage SOC cases via the Kibana Cases API.
elastic/agent-skills
Create, tune, and manage Elastic Security detection rules (SIEM and Endpoint).
elastic/agent-skills
Create and manage Kibana Dashboards and Lens visualizations.
elastic/agent-skills
Generate sample security events, attack scenarios, and synthetic alerts for Elastic Security.
elastic/agent-skills
Onboard an Elastic Cloud organization: configure the elastic CLI's Cloud context and API key, establish a default region, then invite users, assign predefined or custom Serverless project roles, and…
Works with
Categories
Answer questions about LLM and agentic-application behavior from data already ingested into Elastic: latency and error rate, token and cost utilization, response quality and guardrail events, and…. Observability LLM Obs is an agent skill from elastic/agent-skills, published by the product's own GitHub organization. Answer questions about LLM and agentic-application behavior from data already ingested into Elastic: latency and error rate, token and cost utilization, response quality and guardrail events, and agentic call-chain orchestration.
Observability LLM Obs fits situations like: the user asks about LLM monitoring; genAI observability; guardrail failures; how an agents tool-call chain executed.
Run `npx skills add elastic/agent-skills --skill observability-llm-obs -a claude-code`. Or copy the skill folder (skills/observability/llm-obs in elastic/agent-skills) into .claude/skills/observability-llm-obs in your project. Claude Code loads it when a task matches its description.
Run `npx skills add elastic/agent-skills --skill observability-llm-obs -a codex`. Or copy the skill folder (skills/observability/llm-obs in elastic/agent-skills) into .agents/skills/observability-llm-obs in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add elastic/agent-skills --skill observability-llm-obs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/observability-llm-obs, .gemini/skills/observability-llm-obs, .github/skills/observability-llm-obs and .opencode/skills/observability-llm-obs in your project.
SKILL.md names no scripts, command-line tools or credentials: Observability LLM Obs is instructions for the agent only. Compatibility (from SKILL.md): Requires the `elastic` CLI (>= 0.2) with `es` and `kb` support, and an Elasticsearch deployment holding LLM telemetry from APM/OTLP traces or an Elastic LLM integration. Base floor is Elasticsearch 8.11+ or Serverless. The time-series path for integration metrics needs `TS` (Stack GA 9.4) and `TRANGE` (Stack GA 9.3); both are GA on Serverless. Below Stack 9.4 fall back to `FROM` with `BUCKET`. SLO and alerting lookups require Kibana. No Kibana UI is required. .
SKILL.md names 2 domains. As links in the text: elastic.co and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Observability LLM Obs is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Observability LLM Obs: Caveman Gateway Setup (JuliusBrussee/caveman, 110k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 412 stars), Google Agents CLI Observability (pifferologo/cloud-agents-cli, 129 stars) and Agent Platform Alert Configuration (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
elastic (a GitHub organization, an official publisher) maintains it in elastic/agent-skills, which has 592 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on October 7, 2026.
Source: elastic/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.