Agent skill

Langfuse

by davila7 in davila7/claude-code-templates

Expert in Langfuse - the open-source LLM observability platform.

MITAuto-check passedAI & LLM Engineering

Install Langfuse

skills CLI
$ npx skills add davila7/claude-code-templates --skill langfuse -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates langfuse --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/ai-research/langfuse .claude/skills/langfuse && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langfuse
GitHub stars
32k
Used in
5 other repos
Token cost
~1.4k tokens
SKILL.md length
235 words
Files
1
Skills in repo
477
Repo updated
First seen
Licence
MIT

At a glance

Expert in Langfuse - the open-source LLM observability platform.

  • Llm observability
  • SKILL.md covers Capabilities, Requirements, Patterns and Anti-Patterns, plus 2 more sections
  • Reaches cloud.langfuse.com
  • Prompt management

What it does

Langfuse is an agent skill from davila7/claude-code-templates. Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production. Use when: langfuse, llm observability, llm tracing, prompt management, llm evaluation.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM observability and Building AI agents. It works with Langfuse, OpenAI, LangChain and LlamaIndex. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.

When your agent uses it

  • Llm observability
  • Prompt management

Example prompts

  • “/langfuse”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4c82aba. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • cloud.langfuse.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Langfuse loads about 1.4k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 235 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 4c82aba, republished under its MIT licence (© davila7). 235 words, ~1,449 tokens.

Download SKILL.mdSave it as .claude/skills/langfuse/SKILL.md (or your agent's skills folder).
name
langfuse
description
Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production. Use when: langfuse, llm observability, llm tracing, prompt management, llm evaluation.
source
vibeship-spawner-skills (Apache 2.0)

Langfuse

Role: LLM Observability Architect

You are an expert in LLM observability and evaluation. You think in terms of traces, spans, and metrics. You know that LLM applications need monitoring just like traditional software - but with different dimensions (cost, quality, latency). You use data to drive prompt improvements and catch regressions.

Capabilities

  • LLM tracing and observability
  • Prompt management and versioning
  • Evaluation and scoring
  • Dataset management
  • Cost tracking
  • Performance monitoring
  • A/B testing prompts

Requirements

  • Python or TypeScript/JavaScript
  • Langfuse account (cloud or self-hosted)
  • LLM API keys

Patterns

Basic Tracing Setup

Instrument LLM calls with Langfuse

When to use: Any LLM application

python
from langfuse import Langfuse

# Initialize client
langfuse = Langfuse(
    public_key="pk-...",
    secret_key="sk-...",
    host="https://cloud.langfuse.com"  # or self-hosted URL
)

# Create a trace for a user request
trace = langfuse.trace(
    name="chat-completion",
    user_id="user-123",
    session_id="session-456",  # Groups related traces
    metadata={"feature": "customer-support"},
    tags=["production", "v2"]
)

# Log a generation (LLM call)
generation = trace.generation(
    name="gpt-4o-response",
    model="gpt-4o",
    model_parameters={"temperature": 0.7},
    input={"messages": [{"role": "user", "content": "Hello"}]},
    metadata={"attempt": 1}
)

# Make actual LLM call
response = openai.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello"}]
)

# Complete the generation with output
generation.end(
    output=response.choices[0].message.content,
    usage={
        "input": response.usage.prompt_tokens,
        "output": response.usage.completion_tokens
    }
)

# Score the trace
trace.score(
    name="user-feedback",
    value=1,  # 1 = positive, 0 = negative
    comment="User clicked helpful"
)

# Flush before exit (important in serverless)
langfuse.flush()
OpenAI Integration

Automatic tracing with OpenAI SDK

When to use: OpenAI-based applications

python
from langfuse.openai import openai

# Drop-in replacement for OpenAI client
# All calls automatically traced

response = openai.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello"}],
    # Langfuse-specific parameters
    name="greeting",  # Trace name
    session_id="session-123",
    user_id="user-456",
    tags=["test"],
    metadata={"feature": "chat"}
)

# Works with streaming
stream = openai.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Tell me a story"}],
    stream=True,
    name="story-generation"
)

for chunk in stream:
    print(chunk.choices[0].delta.content, end="")

# Works with async
import asyncio
from langfuse.openai import AsyncOpenAI

async_client = AsyncOpenAI()

async def main():
    response = await async_client.chat.completions.create(
        model="gpt-4o",
        messages=[{"role": "user", "content": "Hello"}],
        name="async-greeting"
    )
LangChain Integration

Trace LangChain applications

When to use: LangChain-based applications

python
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langfuse.callback import CallbackHandler

# Create Langfuse callback handler
langfuse_handler = CallbackHandler(
    public_key="pk-...",
    secret_key="sk-...",
    host="https://cloud.langfuse.com",
    session_id="session-123",
    user_id="user-456"
)

# Use with any LangChain component
llm = ChatOpenAI(model="gpt-4o")

prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a helpful assistant."),
    ("user", "{input}")
])

chain = prompt | llm

# Pass handler to invoke
response = chain.invoke(
    {"input": "Hello"},
    config={"callbacks": [langfuse_handler]}
)

# Or set as default
import langchain
langchain.callbacks.manager.set_handler(langfuse_handler)

# Then all calls are traced
response = chain.invoke({"input": "Hello"})

# Works with agents, retrievers, etc.
from langchain.agents import create_openai_tools_agent

agent = create_openai_tools_agent(llm, tools, prompt)
agent_executor = AgentExecutor(agent=agent, tools=tools)

result = agent_executor.invoke(
    {"input": "What's the weather?"},
    config={"callbacks": [langfuse_handler]}
)

Anti-Patterns

❌ Not Flushing in Serverless

Why bad: Traces are batched. Serverless may exit before flush. Data is lost.

Instead: Always call langfuse.flush() at end. Use context managers where available. Consider sync mode for critical traces.

❌ Tracing Everything

Why bad: Noisy traces. Performance overhead. Hard to find important info.

Instead: Focus on: LLM calls, key logic, user actions. Group related operations. Use meaningful span names.

❌ No User/Session IDs

Why bad: Can't debug specific users. Can't track sessions. Analytics limited.

Instead: Always pass user_id and session_id. Use consistent identifiers. Add relevant metadata.

Limitations

  • Self-hosted requires infrastructure
  • High-volume may need optimization
  • Real-time dashboard has latency
  • Evaluation requires setup

Works well with: langgraph, crewai, structured-output, autonomous-agents

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in cli-tool/components/skills/ai-research/langfuse of davila7/claude-code-templates.

Open the folder on GitHubat commit 4c82aba

Used in 5 other repositories

We found 14 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Langfuse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langfuse compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langfuse this skilldavila7/claude-code-templates32k5 repos~1.4kAutomated safety check: PassMIT
Langfusemajiayu000/claude-skill-registry6663 repos~3.1kAutomated safety check: PassMIT
Upgrade Stripekanchengw/cnllm1754 repos~1.4kAutomated safety check: PassApache-2.0
Agent Prompt Engineeringagentailor/fullstack-langgraph-nextjs-agent132—~3.6kAutomated safety check: PassMIT
Phoenix LLM ObservabilityOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Agent Eval Casesagentailor/fullstack-langgraph-nextjs-agent132—~5.3kAutomated safety check: PassMIT

Similar skills

  • Langfuse

    majiayu000/claude-skill-registry

    Expert in Langfuse - the open-source LLM observability platform.

    666 GitHub starsUsed in 3 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Upgrade Stripe

    kanchengw/cnllm

    Guide for upgrading Stripe API versions and SDKs. An agent skill from kanchengw/cnllm.

    175 GitHub starsUsed in 4 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Prompt Engineering

    agentailor/fullstack-langgraph-nextjs-agent

    Comprehensive guide for designing, refining, and auditing system prompts for autonomous AI agents based on Anthropic's production practices.

    132 GitHub stars~3.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Phoenix LLM Observability

    Orchestra-Research/AI-Research-SKILLs

    Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Cases

    agentailor/fullstack-langgraph-nextjs-agent

    Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

    132 GitHub stars~5.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agentsop Observability Setup

    agentsope/SkillAlchemy

    Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.

    457 GitHub stars~4.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from davila7/claude-code-templates

All 477 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 10 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 2 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Questions about Langfuse

What does Langfuse do?

Expert in Langfuse - the open-source LLM observability platform. Langfuse is an agent skill from davila7/claude-code-templates. Expert in Langfuse - the open-source LLM observability platform.

When should I use Langfuse?

Langfuse fits situations like: llm observability; prompt management.

How do I install Langfuse in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill langfuse -a claude-code`. Or copy the skill folder (cli-tool/components/skills/ai-research/langfuse in davila7/claude-code-templates) into .claude/skills/langfuse in your project. Claude Code loads it when a task matches its description.

How do I install Langfuse in Codex?

Run `npx skills add davila7/claude-code-templates --skill langfuse -a codex`. Or copy the skill folder (cli-tool/components/skills/ai-research/langfuse in davila7/claude-code-templates) into .agents/skills/langfuse in your project. Codex loads it when a task matches its description.

Can I use Langfuse in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill langfuse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langfuse, .gemini/skills/langfuse, .github/skills/langfuse and .opencode/skills/langfuse in your project.

What does Langfuse need to run?

SKILL.md names no scripts, command-line tools or credentials: Langfuse is instructions for the agent only. Our summary lists: Python 3.

Does Langfuse access the network?

SKILL.md names 1 domain. In commands or code: cloud.langfuse.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Langfuse safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langfuse use?

Langfuse is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langfuse use?

About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Langfuse?

Skills that share tags, products or a category with Langfuse: Langfuse (majiayu000/claude-skill-registry, 666 stars), Upgrade Stripe (kanchengw/cnllm, 175 stars), Agent Prompt Engineering (agentailor/fullstack-langgraph-nextjs-agent, 132 stars) and Phoenix LLM Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langfuse?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,432 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 7, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.