Agent skill

Langfuse Core Workflow B

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Execute Langfuse secondary workflow: Evaluation, scoring, and datasets.

MITAuto-check passedAI & LLM Engineering

Install Langfuse Core Workflow B

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langfuse-core-workflow-b -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace langfuse-core-workflow-b --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/saas-packs/langfuse-pack/skills/langfuse-core-workflow-b .claude/skills/langfuse-core-workflow-b && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langfuse-core-workflow-b
GitHub stars
2.8k
Token cost
~2.1k tokens
SKILL.md length
256 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Execute Langfuse secondary workflow: Evaluation, scoring, and datasets.

  • Works in 6 steps: Score Traces via SDK → User Feedback Collection → Prompt Management → …
  • Implementing LLM evaluation
  • SKILL.md covers Overview, Prerequisites, Instructions and Error Handling, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Langfuse Core Workflow B is an agent skill from jeremylongshore/tons-of-skills-marketplace. Execute Langfuse secondary workflow: Evaluation, scoring, and datasets. Use when implementing LLM evaluation, adding user feedback, or setting up automated quality scoring and experiment datasets. Trigger with phrases like "langfuse evaluation", "langfuse scoring", "rate llm outputs", "langfuse feedback", "langfuse datasets", "langfuse experiments".

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering LLM observability. It works with Langfuse. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Implementing LLM evaluation
  • Adding user feedback
  • Setting up automated quality scoring and experiment datasets
  • With phrases like langfuse evaluation

Example prompts

  • “langfuse evaluation”
  • “langfuse scoring”
  • “rate llm outputs”
  • “/langfuse-core-workflow-b”

Requirements

  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(npm:*), Grep

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Score Traces via SDK
  2. User Feedback Collection
  3. Prompt Management
  4. Create and Populate Datasets
  5. Run Experiments with the Experiment Runner
  6. LLM-as-a-Judge Evaluation

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(npm:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • langfuse.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Langfuse Core Workflow B loads about 2.1k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 256 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 256 words, ~2,130 tokens.

Download SKILL.mdSave it as .claude/skills/langfuse-core-workflow-b/SKILL.md (or your agent's skills folder).
name
langfuse-core-workflow-b
description
Execute Langfuse secondary workflow: Evaluation, scoring, and datasets. Use when implementing LLM evaluation, adding user feedback, or setting up automated quality scoring and experiment datasets. Trigger with phrases like "langfuse evaluation", "langfuse scoring", "rate llm outputs", "langfuse feedback", "langfuse datasets", "langfuse experiments".
allowed-tools
Read, Write, Edit, Bash(npm:*), Grep
compatibility
Designed for Claude Code
version
1.17.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, langfuse, llm, workflow, evaluation

Langfuse Core Workflow B: Evaluation, Scoring & Datasets

Overview

Implement LLM output evaluation using Langfuse scores (numeric, categorical, boolean), the experiment runner SDK for dataset-driven benchmarks, prompt management with versioned prompts, and LLM-as-a-Judge evaluation patterns.

Prerequisites

  • Langfuse SDK configured with API keys
  • Traces already being collected (see langfuse-core-workflow-a)
  • For v4+: @langfuse/client installed

Instructions

Step 1: Score Traces via SDK

Langfuse supports three score data types: Numeric, Categorical, and Boolean.

typescript
import { LangfuseClient } from "@langfuse/client";

const langfuse = new LangfuseClient();

// Numeric score (e.g., 0-1 quality rating)
await langfuse.score.create({
  traceId: "trace-abc-123",
  name: "relevance",
  value: 0.92,
  dataType: "NUMERIC",
  comment: "Highly relevant answer with good context usage",
});

// Categorical score (e.g., pass/fail classification)
await langfuse.score.create({
  traceId: "trace-abc-123",
  observationId: "gen-xyz-456", // Optional: score a specific generation
  name: "quality-tier",
  value: "excellent",
  dataType: "CATEGORICAL",
});

// Boolean score (e.g., thumbs up/down)
await langfuse.score.create({
  traceId: "trace-abc-123",
  name: "user-approved",
  value: 1, // 1 = true, 0 = false
  dataType: "BOOLEAN",
  comment: "User clicked thumbs up",
});
Step 2: User Feedback Collection
typescript
// API endpoint for frontend feedback widget
app.post("/api/feedback", async (req, res) => {
  const { traceId, rating, comment } = req.body;

  // Thumbs up/down
  await langfuse.score.create({
    traceId,
    name: "user-feedback",
    value: rating === "positive" ? 1 : 0,
    dataType: "BOOLEAN",
    comment,
  });

  // Granular star rating (1-5)
  if (req.body.stars) {
    await langfuse.score.create({
      traceId,
      name: "star-rating",
      value: req.body.stars,
      dataType: "NUMERIC",
      comment: `${req.body.stars}/5 stars`,
    });
  }

  res.json({ success: true });
});
Step 3: Prompt Management
typescript
// Fetch a versioned prompt from Langfuse
const textPrompt = await langfuse.prompt.get("summarize-article", {
  type: "text",
  label: "production", // or "latest", "staging"
});

// Compile with variables -- replaces {{variable}} placeholders
const compiled = textPrompt.compile({
  maxLength: "100 words",
  tone: "professional",
});

// Chat prompts return message arrays
const chatPrompt = await langfuse.prompt.get("customer-support", {
  type: "chat",
});

const messages = chatPrompt.compile({
  customerName: "Alice",
  issue: "billing question",
});
// messages = [{ role: "system", content: "..." }, { role: "user", content: "..." }]
Step 4: Create and Populate Datasets
typescript
// Create a dataset for evaluation
await langfuse.api.datasets.create({
  name: "customer-support-v1",
  description: "Test cases for customer support chatbot",
  metadata: { version: "1.0", domain: "support" },
});

// Add test items
const testCases = [
  {
    input: { query: "How do I cancel my subscription?" },
    expectedOutput: { intent: "cancellation", sentiment: "neutral" },
    metadata: { category: "billing" },
  },
  {
    input: { query: "Your product is amazing!" },
    expectedOutput: { intent: "feedback", sentiment: "positive" },
    metadata: { category: "feedback" },
  },
];

for (const testCase of testCases) {
  await langfuse.api.datasetItems.create({
    datasetName: "customer-support-v1",
    input: testCase.input,
    expectedOutput: testCase.expectedOutput,
    metadata: testCase.metadata,
  });
}
Step 5: Run Experiments with the Experiment Runner
typescript
import { LangfuseClient } from "@langfuse/client";

const langfuse = new LangfuseClient();

// Define the task function -- your LLM application logic
async function classifyIntent(input: { query: string }): Promise<string> {
  const response = await openai.chat.completions.create({
    model: "gpt-4o-mini",
    messages: [
      { role: "system", content: "Classify the user intent. Return one word." },
      { role: "user", content: input.query },
    ],
    temperature: 0,
  });
  return response.choices[0].message.content?.trim() || "";
}

// Define evaluator functions
function exactMatch({ output, expectedOutput }: {
  output: string;
  expectedOutput: { intent: string };
}) {
  return {
    name: "exact-match",
    value: output.toLowerCase() === expectedOutput.intent.toLowerCase() ? 1 : 0,
    dataType: "BOOLEAN" as const,
  };
}

// Run the experiment
const result = await langfuse.runExperiment({
  datasetName: "customer-support-v1",
  runName: "gpt-4o-mini-classifier-v1",
  runDescription: "Testing intent classification with gpt-4o-mini",
  task: classifyIntent,
  evaluators: [exactMatch],
});

console.log(`Experiment complete. ${result.runs.length} items evaluated.`);
// View results in Langfuse UI: Datasets > customer-support-v1 > Runs
Step 6: LLM-as-a-Judge Evaluation
typescript
async function llmJudge({ output, input, expectedOutput }: {
  output: string;
  input: { query: string };
  expectedOutput: { intent: string; sentiment: string };
}) {
  const judgment = await openai.chat.completions.create({
    model: "gpt-4o",
    temperature: 0,
    messages: [
      {
        role: "system",
        content: `You are an AI evaluator. Score the response 0-10 on accuracy and helpfulness.
Return JSON: {"score": <number>, "reasoning": "<explanation>"}`,
      },
      {
        role: "user",
        content: `Query: ${input.query}\nExpected: ${JSON.stringify(expectedOutput)}\nActual: ${output}`,
      },
    ],
    response_format: { type: "json_object" },
  });

  const result = JSON.parse(judgment.choices[0].message.content || "{}");

  return {
    name: "llm-judge-quality",
    value: result.score / 10, // Normalize to 0-1
    dataType: "NUMERIC" as const,
    comment: result.reasoning,
  };
}

// Use as an evaluator in experiments
await langfuse.runExperiment({
  datasetName: "customer-support-v1",
  runName: "judge-evaluation-v1",
  task: classifyIntent,
  evaluators: [exactMatch, llmJudge],
});

Error Handling

IssueCauseSolution
Scores not appearingAPI call failed silentlyAwait score.create() and check for errors
Score validation errorWrong data typeMatch value type to dataType (number/string/0-1)
LLM judge inconsistentHigh temperatureSet temperature: 0 for evaluation calls
Dataset item missingWrong dataset nameVerify exact name match (case-sensitive)
Experiment not in UIRun not flushedCheck runExperiment completed without errors

Output

Produce a versioned dataset or prompt reference, an experiment run identifier, and per-item plus aggregate scores. Summarize the threshold, sample size, and failed cases so a release decision is reproducible.

Examples

Create a small customer-support-v1 dataset, run the exact-match evaluator at temperature zero, and inspect failed items before changing the prompt. Add the LLM-as-a-judge only as a second score; retain deterministic exact-match or rubric evidence as the release gate.

Resources

Next Steps

For common error debugging, see langfuse-common-errors. For CI/CD integration of evaluations, see langfuse-ci-integration.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/saas-packs/langfuse-pack/skills/langfuse-core-workflow-b of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Langfuse Core Workflow B next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langfuse Core Workflow B compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langfuse Core Workflow B this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.1kAutomated safety check: PassMIT
Langfuse Codebase Navigatorlangfuse/langfuse36k—~1.4kAutomated safety check: PassCustom licence
Langfuse Integration Pagelangfuse/langfuse-docs246—~3.7kAutomated safety check: PassMIT
Langfuselangfuse/skills301—~2.1kAutomated safety check: NotesMIT
Add Yourself To Team Langfuselangfuse/langfuse-docs246—~548Automated safety check: PassMIT
Weekly Production Reviewlangfuse/langfuse36k—~4.1kAutomated safety check: PassCustom licence

Similar skills

  • Navigate Langfuse repositories, code areas, and agent skills.

    36k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Langfuse Integration Page

    langfuse/langfuse-docs

    Create a new Langfuse integration page in the langfuse-docs repo.

    246 GitHub stars~3.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Langfuse

    langfuse/skills

    Interact with Langfuse and access its documentation: tracing, monitoring, creating datasets, running experiments, and evaluating AI applications.

    301 GitHub stars~2.1k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check: notes
  • Add Yourself To Team Langfuse

    langfuse/langfuse-docs

    Add a new team member to Langfuse's canonical team data and shared team table.

    246 GitHub stars~548 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Weekly Production Review

    langfuse/langfuse

    Prepare Langfuse weekly production reviews covering failures, fixes, open issues, and tracking gaps.

    36k GitHub stars~4.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Agent Setup Maintenance

    langfuse/langfuse

    Shared workflow for editing Langfuse's repo-owned agent setup under .agents/.

    36k GitHub stars~799 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Langfuse Core Workflow B

What does Langfuse Core Workflow B do?

Execute Langfuse secondary workflow: Evaluation, scoring, and datasets. Langfuse Core Workflow B is an agent skill from jeremylongshore/tons-of-skills-marketplace. Execute Langfuse secondary workflow: Evaluation, scoring, and datasets.

When should I use Langfuse Core Workflow B?

Langfuse Core Workflow B fits situations like: implementing LLM evaluation; adding user feedback; setting up automated quality scoring and experiment datasets; with phrases like langfuse evaluation.

How do I install Langfuse Core Workflow B in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langfuse-core-workflow-b -a claude-code`. Or copy the skill folder (plugins/saas-packs/langfuse-pack/skills/langfuse-core-workflow-b in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langfuse-core-workflow-b in your project. Claude Code loads it when a task matches its description.

How do I install Langfuse Core Workflow B in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langfuse-core-workflow-b -a codex`. Or copy the skill folder (plugins/saas-packs/langfuse-pack/skills/langfuse-core-workflow-b in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langfuse-core-workflow-b in your project. Codex loads it when a task matches its description.

Can I use Langfuse Core Workflow B in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langfuse-core-workflow-b -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langfuse-core-workflow-b, .gemini/skills/langfuse-core-workflow-b, .github/skills/langfuse-core-workflow-b and .opencode/skills/langfuse-core-workflow-b in your project.

What does Langfuse Core Workflow B need to run?

SKILL.md names no scripts, command-line tools or credentials: Langfuse Core Workflow B is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(npm:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Langfuse Core Workflow B access the network?

SKILL.md names 1 domain. As links in the text: langfuse.com. This is read from the text; nothing was executed.

Is Langfuse Core Workflow B safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langfuse Core Workflow B use?

Langfuse Core Workflow B is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langfuse Core Workflow B use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Langfuse Core Workflow B?

Skills that share tags, products or a category with Langfuse Core Workflow B: Langfuse Codebase Navigator (langfuse/langfuse, 36k stars), Langfuse Integration Page (langfuse/langfuse-docs, 246 stars), Langfuse (langfuse/skills, 301 stars) and Add Yourself To Team Langfuse (langfuse/langfuse-docs, 246 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langfuse Core Workflow B?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.