Expert in Langfuse - the open-source LLM observability platform.

MITAuto-check passedAI & LLM Engineering

Install Langfuse

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill langfuse -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry langfuse --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-llm/langfuse .claude/skills/langfuse && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langfuse
GitHub stars
666
Used in
3 other repos
Token cost
~3.1k tokens
SKILL.md length
1,028 words
Files
2
Skills in repo
971
Repo updated
First seen
Licence
MIT

At a glance

Expert in Langfuse - the open-source LLM observability platform.

  • Tasks that involve LLM observability
  • SKILL.md covers Capabilities, Prerequisites, Scope and Ecosystem, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Building AI agents

What it does

Langfuse is an agent skill from majiayu000/claude-skill-registry. Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in AI & LLM Engineering, covering LLM observability and Building AI agents. It works with Langfuse, OpenAI, LangChain and LlamaIndex. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM observability
  • Tasks that involve Building AI agents

Example prompts

  • “/langfuse”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 000116a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • cloud.langfuse.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Langfuse loads about 3.1k tokens when it runs. Until then it costs about 67 tokens; SKILL.md has 1,028 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 000116a, republished under its MIT licence (© majiayu000). 1,028 words, ~3,056 tokens.

Download SKILL.mdSave it as .claude/skills/langfuse/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
langfuse
description
Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production.
risk
unknown
source
vibeship-spawner-skills (Apache 2.0)
date_added
2026-02-27

Langfuse

Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production.

Role: LLM Observability Architect

You are an expert in LLM observability and evaluation. You think in terms of traces, spans, and metrics. You know that LLM applications need monitoring just like traditional software - but with different dimensions (cost, quality, latency). You use data to drive prompt improvements and catch regressions.

Expertise
  • Tracing architecture
  • Prompt versioning
  • Evaluation strategies
  • Cost optimization
  • Quality monitoring

Capabilities

  • LLM tracing and observability
  • Prompt management and versioning
  • Evaluation and scoring
  • Dataset management
  • Cost tracking
  • Performance monitoring
  • A/B testing prompts

Prerequisites

  • 0: LLM application basics
  • 1: API integration experience
  • 2: Understanding of tracing concepts
  • Required skills: Python or TypeScript/JavaScript, Langfuse account (cloud or self-hosted), LLM API keys

Scope

  • 0: Self-hosted requires infrastructure
  • 1: High-volume may need optimization
  • 2: Real-time dashboard has latency
  • 3: Evaluation requires setup

Ecosystem

Primary
  • Langfuse Cloud
  • Langfuse Self-hosted
  • Python SDK
  • JS/TS SDK
Common_integrations
  • LangChain
  • LlamaIndex
  • OpenAI SDK
  • Anthropic SDK
  • Vercel AI SDK
Platforms
  • Any Python/JS backend
  • Serverless functions
  • Jupyter notebooks

Patterns

Basic Tracing Setup

Instrument LLM calls with Langfuse

When to use: Any LLM application

from langfuse import Langfuse

Initialize client

langfuse = Langfuse( public_key="pk-...", secret_key="sk-...", host="https://cloud.langfuse.com" # or self-hosted URL )

Create a trace for a user request

trace = langfuse.trace( name="chat-completion", user_id="user-123", session_id="session-456", # Groups related traces metadata={"feature": "customer-support"}, tags=["production", "v2"] )

Log a generation (LLM call)

generation = trace.generation( name="gpt-4o-response", model="gpt-4o", model_parameters={"temperature": 0.7}, input={"messages": [{"role": "user", "content": "Hello"}]}, metadata={"attempt": 1} )

Make actual LLM call

response = openai.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "Hello"}] )

Complete the generation with output

generation.end( output=response.choices[0].message.content, usage={ "input": response.usage.prompt_tokens, "output": response.usage.completion_tokens } )

Score the trace

trace.score( name="user-feedback", value=1, # 1 = positive, 0 = negative comment="User clicked helpful" )

Flush before exit (important in serverless)

langfuse.flush()

OpenAI Integration

Automatic tracing with OpenAI SDK

When to use: OpenAI-based applications

from langfuse.openai import openai

Drop-in replacement for OpenAI client

All calls automatically traced

response = openai.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "Hello"}], # Langfuse-specific parameters name="greeting", # Trace name session_id="session-123", user_id="user-456", tags=["test"], metadata={"feature": "chat"} )

Works with streaming

stream = openai.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "Tell me a story"}], stream=True, name="story-generation" )

for chunk in stream: print(chunk.choices[0].delta.content, end="")

Works with async

import asyncio from langfuse.openai import AsyncOpenAI

async_client = AsyncOpenAI()

async def main(): response = await async_client.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "Hello"}], name="async-greeting" )

LangChain Integration

Trace LangChain applications

When to use: LangChain-based applications

from langchain_openai import ChatOpenAI from langchain_core.prompts import ChatPromptTemplate from langfuse.callback import CallbackHandler

Create Langfuse callback handler

langfuse_handler = CallbackHandler( public_key="pk-...", secret_key="sk-...", host="https://cloud.langfuse.com", session_id="session-123", user_id="user-456" )

Use with any LangChain component

llm = ChatOpenAI(model="gpt-4o")

prompt = ChatPromptTemplate.from_messages([ ("system", "You are a helpful assistant."), ("user", "{input}") ])

chain = prompt | llm

Pass handler to invoke

response = chain.invoke( {"input": "Hello"}, config={"callbacks": [langfuse_handler]} )

Or set as default

import langchain langchain.callbacks.manager.set_handler(langfuse_handler)

Then all calls are traced

response = chain.invoke({"input": "Hello"})

Works with agents, retrievers, etc.

from langchain.agents import create_openai_tools_agent

agent = create_openai_tools_agent(llm, tools, prompt) agent_executor = AgentExecutor(agent=agent, tools=tools)

result = agent_executor.invoke( {"input": "What's the weather?"}, config={"callbacks": [langfuse_handler]} )

Prompt Management

Version and deploy prompts

When to use: Managing prompts across environments

from langfuse import Langfuse

langfuse = Langfuse()

Fetch prompt from Langfuse

(Create in UI or via API first)

prompt = langfuse.get_prompt("customer-support-v2")

Get compiled prompt with variables

compiled = prompt.compile( customer_name="John", issue="billing question" )

Use with OpenAI

response = openai.chat.completions.create( model=prompt.config.get("model", "gpt-4o"), messages=compiled, temperature=prompt.config.get("temperature", 0.7) )

trace = langfuse.trace(name="support-chat") generation = trace.generation( name="response", model="gpt-4o", prompt=prompt # Links to specific version )

Create/update prompts via API

langfuse.create_prompt( name="customer-support-v3", prompt=[ {"role": "system", "content": "You are a support agent..."}, {"role": "user", "content": "{{user_message}}"} ], config={ "model": "gpt-4o", "temperature": 0.7 }, labels=["production"] # or ["staging", "development"] )

Show full SKILL.md (383 more words)Show less

Fetch specific label

prompt = langfuse.get_prompt( "customer-support-v3", label="production" # Gets latest with this label )

Evaluation and Scoring

Evaluate LLM outputs systematically

When to use: Quality assurance and improvement

from langfuse import Langfuse

langfuse = Langfuse()

Manual scoring in code

trace = langfuse.trace(name="qa-flow")

After getting response

trace.score( name="relevance", value=0.85, # 0-1 scale comment="Response addressed the question" )

trace.score( name="correctness", value=1, # Binary: 0 or 1 data_type="BOOLEAN" )

LLM-as-judge evaluation

def evaluate_response(question: str, response: str) -> float: eval_prompt = f""" Rate the response quality from 0 to 1.

Question: {question}
Response: {response}

Output only a number between 0 and 1.
"""

result = openai.chat.completions.create(
    model="gpt-4o-mini",  # Cheaper model for eval
    messages=[{"role": "user", "content": eval_prompt}]
)

return float(result.choices[0].message.content.strip())

Score asynchronously

score = evaluate_response(question, response) trace.score( name="quality-llm-judge", value=score )

Create evaluation dataset

dataset = langfuse.create_dataset(name="support-qa-v1")

Add items to dataset

langfuse.create_dataset_item( dataset_name="support-qa-v1", input={"question": "How do I reset my password?"}, expected_output="Go to settings > security > reset password" )

Run evaluation on dataset

dataset = langfuse.get_dataset("support-qa-v1")

for item in dataset.items: # Generate response response = generate_response(item.input["question"])

# Link to dataset item
trace = langfuse.trace(name="eval-run")
trace.generation(
    name="response",
    input=item.input,
    output=response
)

# Score against expected
similarity = calculate_similarity(response, item.expected_output)
trace.score(name="similarity", value=similarity)

# Link trace to dataset item
item.link(trace, "eval-run-1")
Decorator Pattern

Clean instrumentation with decorators

When to use: Function-based applications

from langfuse.decorators import observe, langfuse_context

@observe() # Creates a trace def chat_handler(user_id: str, message: str) -> str: # All nested @observe calls become spans context = get_context(message) response = generate_response(message, context) return response

@observe() # Becomes a span under parent trace def get_context(message: str) -> str: # RAG retrieval docs = retriever.get_relevant_documents(message) return "\n".join([d.page_content for d in docs])

@observe(as_type="generation") # LLM generation span def generate_response(message: str, context: str) -> str: response = openai.chat.completions.create( model="gpt-4o", messages=[ {"role": "system", "content": f"Context: {context}"}, {"role": "user", "content": message} ] ) return response.choices[0].message.content

Add metadata and scores

@observe() def main_flow(user_input: str): # Update current trace langfuse_context.update_current_trace( user_id="user-123", session_id="session-456", tags=["production"] )

result = process(user_input)

# Score the trace
langfuse_context.score_current_trace(
    name="success",
    value=1 if result else 0
)

return result

Works with async

@observe() async def async_handler(message: str): result = await async_generate(message) return result

Collaboration

Delegation Triggers
  • agent|langgraph|graph -> langgraph (Need to build agent to monitor)
  • crewai|multi-agent|crew -> crewai (Need to build crew to monitor)
  • structured output|extraction -> structured-output (Need to build extraction to monitor)
Observable LangGraph Agent

Skills: langfuse, langgraph

Workflow:

1. Build agent with LangGraph
2. Add Langfuse callback handler
3. Trace all LLM calls and tool uses
4. Score outputs for quality
5. Monitor and iterate
Monitored RAG Pipeline

Skills: langfuse, structured-output

Workflow:

1. Build RAG with retrieval and generation
2. Trace retrieval and LLM calls
3. Score relevance and accuracy
4. Track costs and latency
5. Optimize based on data
Evaluated Agent System

Skills: langfuse, langgraph, structured-output

Workflow:

1. Build agent with structured outputs
2. Create evaluation dataset
3. Run evaluations with traces
4. Compare prompt versions
5. Deploy best performers

Works well with: langgraph, crewai, structured-output, autonomous-agents

When to Use

  • User mentions or implies: langfuse
  • User mentions or implies: llm observability
  • User mentions or implies: llm tracing
  • User mentions or implies: prompt management
  • User mentions or implies: llm evaluation
  • User mentions or implies: monitor llm
  • User mentions or implies: debug llm

Limitations

  • Use this skill only when the task clearly matches the scope described above.
  • Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
  • Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ai-llm/langfuse of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 000116a

Used in 3 other repositories

We found 18 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Langfuse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langfuse compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langfuse this skillmajiayu000/claude-skill-registry6663 repos~3.1kAutomated safety check: PassMIT
Langfusedavila7/claude-code-templates32k5 repos~1.4kAutomated safety check: PassMIT
Upgrade Stripekanchengw/cnllm1754 repos~1.4kAutomated safety check: PassApache-2.0
Phoenix Integration SnippetsArize-ai/phoenix12k—~1.4kAutomated safety check: PassApache-2.0
Routerbase API Integrationaiskillstore/marketplace430—~964Automated safety check: PassNone
Stripe Best Practiceskanchengw/cnllm1753 repos~925Automated safety check: PassApache-2.0

Similar skills

  • Langfuse

    davila7/claude-code-templates

    Expert in Langfuse - the open-source LLM observability platform.

    32k GitHub starsUsed in 5 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Upgrade Stripe

    kanchengw/cnllm

    Guide for upgrading Stripe API versions and SDKs. An agent skill from kanchengw/cnllm.

    175 GitHub starsUsed in 4 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Generates onboarding code snippets for Phoenix tracing integrations and wires them into the project onboarding UI.

    12k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Routerbase API Integration

    aiskillstore/marketplace

    Integrate applications with RouterBase, the OpenAI-compatible model gateway at https://routerbase.com/v1.

    430 GitHub stars~964 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Stripe Best Practices

    kanchengw/cnllm

    Guides Stripe integration decisions — API selection (Checkout Sessions vs PaymentIntents), Connect platform setup (Accounts v2, controller properties), billing/subscriptions, Treasury financial…

    175 GitHub starsUsed in 3 repos~925 tokens
    Backend & APIsAuto-check passed
  • Add Example Agent

    GetBindu/Bindu

    Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.

    10k GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes

More from majiayu000/claude-skill-registry

All 971 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed
  • Open Notebook

    majiayu000/claude-skill-registry

    Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis.

    666 GitHub starsUsed in 4 repos~2.4k tokens
    Auto-check passed

Questions about Langfuse

What does Langfuse do?

Expert in Langfuse - the open-source LLM observability platform. Langfuse is an agent skill from majiayu000/claude-skill-registry. Expert in Langfuse - the open-source LLM observability platform.

When should I use Langfuse?

Langfuse fits situations like: tasks that involve LLM observability; tasks that involve Building AI agents.

How do I install Langfuse in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill langfuse -a claude-code`. Or copy the skill folder (skills/ai-llm/langfuse in majiayu000/claude-skill-registry) into .claude/skills/langfuse in your project. Claude Code loads it when a task matches its description.

How do I install Langfuse in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill langfuse -a codex`. Or copy the skill folder (skills/ai-llm/langfuse in majiayu000/claude-skill-registry) into .agents/skills/langfuse in your project. Codex loads it when a task matches its description.

Can I use Langfuse in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill langfuse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langfuse, .gemini/skills/langfuse, .github/skills/langfuse and .opencode/skills/langfuse in your project.

What does Langfuse need to run?

SKILL.md names no scripts, command-line tools or credentials: Langfuse is instructions for the agent only. Our summary lists: Python 3.

Does Langfuse access the network?

SKILL.md names 1 domain. As links in the text: cloud.langfuse.com. This is read from the text; nothing was executed.

Is Langfuse safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langfuse use?

Langfuse is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langfuse use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Langfuse?

Skills that share tags, products or a category with Langfuse: Langfuse (davila7/claude-code-templates, 32k stars), Upgrade Stripe (kanchengw/cnllm, 175 stars), Phoenix Integration Snippets (Arize-ai/phoenix, 12k stars) and Routerbase API Integration (aiskillstore/marketplace, 430 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langfuse?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 971 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.