Agent skill

Langfuse CI Integration

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Configure Langfuse CI/CD integration with GitHub Actions and automated testing.

MITAuto-check passedDevOps & Cloud

Install Langfuse CI Integration

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langfuse-ci-integration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace langfuse-ci-integration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langfuse-ci-integration .claude/skills/langfuse-ci-integration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langfuse-ci-integration
GitHub stars
2.8k
Token cost
~2.4k tokens
SKILL.md length
267 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Configure Langfuse CI/CD integration with GitHub Actions and automated testing.

  • Works in 5 steps: GitHub Actions Workflow for AI Quality… → Prompt Regression Tests → Experiment-Driven Quality Gates → …
  • Setting up automated testing
  • SKILL.md covers Overview, Prerequisites, Instructions and CI Best Practices, plus 4 more sections
  • Calls npx; reaches cloud.langfuse.com; needs LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY

What it does

Langfuse CI Integration is an agent skill from jeremylongshore/tons-of-skills-marketplace. Configure Langfuse CI/CD integration with GitHub Actions and automated testing. Use when setting up automated testing, configuring CI pipelines, or integrating Langfuse tests into your build process. Trigger with phrases like "langfuse CI", "langfuse GitHub Actions", "langfuse automated tests", "CI langfuse", "langfuse pipeline".

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in DevOps & Cloud, covering LLM observability and CI/CD. It works with Langfuse and GitHub Actions. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Setting up automated testing
  • Configuring CI pipelines
  • Integrating Langfuse tests into your build process
  • With phrases like langfuse CI

Example prompts

  • “langfuse CI”
  • “langfuse GitHub Actions”
  • “langfuse automated tests”
  • “/langfuse-ci-integration”

Requirements

  • Node.js
  • A credential in LANGFUSE_PUBLIC_KEY
  • A credential in LANGFUSE_SECRET_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(gh:*)

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. GitHub Actions Workflow for AI Quality Tests
  2. Prompt Regression Tests
  3. Experiment-Driven Quality Gates
  4. Automated Prompt Deployment
  5. Score Regression Monitoring

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(gh:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • cloud.langfuse.com

    Also links to:

    • langfuse.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LANGFUSE_PUBLIC_KEY
    • LANGFUSE_SECRET_KEY
    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Langfuse CI Integration loads about 2.4k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 267 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 267 words, ~2,432 tokens.

Download SKILL.mdSave it as .claude/skills/langfuse-ci-integration/SKILL.md (or your agent's skills folder).
name
langfuse-ci-integration
description
Configure Langfuse CI/CD integration with GitHub Actions and automated testing. Use when setting up automated testing, configuring CI pipelines, or integrating Langfuse tests into your build process. Trigger with phrases like "langfuse CI", "langfuse GitHub Actions", "langfuse automated tests", "CI langfuse", "langfuse pipeline".
allowed-tools
Read, Write, Edit, Bash(gh:*)
compatibility
Designed for Claude Code
version
1.17.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, langfuse, testing, ci-cd

Langfuse CI Integration

Overview

Integrate Langfuse into CI/CD pipelines: trace validation tests, prompt regression testing, experiment-driven quality gates, automated prompt deployment from version control, and score monitoring.

Prerequisites

  • Langfuse API keys stored as GitHub secrets (LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY)
  • Test framework (Vitest or Jest)
  • OpenAI API key for LLM tests

Instructions

Step 1: GitHub Actions Workflow for AI Quality Tests
yaml
# .github/workflows/langfuse-tests.yml
name: AI Quality Tests

on:
  pull_request:
    paths: ["src/ai/**", "src/prompts/**", "tests/ai/**"]

jobs:
  ai-quality:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: "20", cache: "npm" }
      - run: npm ci

      - name: Run AI quality tests with tracing
        env:
          LANGFUSE_PUBLIC_KEY: ${{ secrets.LANGFUSE_PUBLIC_KEY }}
          LANGFUSE_SECRET_KEY: ${{ secrets.LANGFUSE_SECRET_KEY }}
          LANGFUSE_BASE_URL: ${{ vars.LANGFUSE_BASE_URL || 'https://cloud.langfuse.com' }}
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
        run: npx vitest run tests/ai/ --reporter=verbose

      - name: Langfuse connectivity check
        env:
          LANGFUSE_PUBLIC_KEY: ${{ secrets.LANGFUSE_PUBLIC_KEY }}
          LANGFUSE_SECRET_KEY: ${{ secrets.LANGFUSE_SECRET_KEY }}
        run: |
          node -e "
            const { LangfuseClient } = require('@langfuse/client');
            const lf = new LangfuseClient();
            lf.prompt.get('__ci-health__').catch(() => {});
            console.log('Langfuse SDK initialized OK');
          "
Step 2: Prompt Regression Tests
typescript
// tests/ai/prompt-quality.test.ts
import { describe, it, expect, afterAll } from "vitest";
import { LangfuseClient } from "@langfuse/client";
import { startActiveObservation, updateActiveObservation } from "@langfuse/tracing";
import OpenAI from "openai";

const langfuse = new LangfuseClient();
const openai = new OpenAI();

describe("Prompt Quality Regression", () => {
  it("summarization prompt produces valid output", async () => {
    const prompt = await langfuse.prompt.get("summarize-article", { type: "text" });
    const compiled = prompt.compile({ maxLength: "100 words" });

    const result = await startActiveObservation(
      { name: "ci-test-summarize", asType: "generation" },
      async () => {
        updateActiveObservation({ model: "gpt-4o-mini", input: compiled });

        const response = await openai.chat.completions.create({
          model: "gpt-4o-mini",
          messages: [{ role: "user", content: compiled }],
          temperature: 0,
        });

        const output = response.choices[0].message.content || "";
        updateActiveObservation({
          output,
          usage: {
            promptTokens: response.usage?.prompt_tokens,
            completionTokens: response.usage?.completion_tokens,
          },
        });
        return output;
      }
    );

    expect(result.length).toBeGreaterThan(20);
    expect(result.length).toBeLessThan(600);
  });

  it("classification prompt returns valid intent", async () => {
    const prompt = await langfuse.prompt.get("classify-intent", { type: "text" });
    const compiled = prompt.compile({ userMessage: "I want to cancel my subscription" });

    const response = await openai.chat.completions.create({
      model: "gpt-4o-mini",
      messages: [{ role: "user", content: compiled }],
      temperature: 0,
    });

    const intent = response.choices[0].message.content?.trim().toLowerCase() || "";
    const validIntents = ["billing", "cancellation", "support", "feedback"];
    expect(validIntents).toContain(intent);
  });
});
Step 3: Experiment-Driven Quality Gates
typescript
// tests/ai/experiment-gate.test.ts
import { describe, it, expect } from "vitest";
import { LangfuseClient } from "@langfuse/client";
import OpenAI from "openai";

const langfuse = new LangfuseClient();
const openai = new OpenAI();

describe("Quality Gate: Intent Classification", () => {
  it("scores above 80% accuracy on test dataset", async () => {
    async function classifyIntent(input: { query: string }) {
      const response = await openai.chat.completions.create({
        model: "gpt-4o-mini",
        messages: [
          { role: "system", content: "Classify intent. Return one word." },
          { role: "user", content: input.query },
        ],
        temperature: 0,
      });
      return response.choices[0].message.content?.trim() || "";
    }

    const result = await langfuse.runExperiment({
      datasetName: "intent-classification-test",
      runName: `ci-${process.env.GITHUB_SHA?.slice(0, 7) || "local"}`,
      task: classifyIntent,
      evaluators: [
        ({ output, expectedOutput }) => ({
          name: "exact-match",
          value: output.toLowerCase() === expectedOutput.intent.toLowerCase() ? 1 : 0,
          dataType: "BOOLEAN" as const,
        }),
      ],
    });

    // Calculate accuracy
    const scores = result.runs.flatMap((r) => r.scores || []);
    const accuracy = scores.filter((s) => s.value === 1).length / scores.length;

    console.log(`Accuracy: ${(accuracy * 100).toFixed(1)}%`);
    expect(accuracy).toBeGreaterThanOrEqual(0.8);
  });
});
Step 4: Automated Prompt Deployment
yaml
# .github/workflows/deploy-prompts.yml
name: Deploy Prompts to Langfuse

on:
  push:
    branches: [main]
    paths: ["src/prompts/**"]

jobs:
  deploy:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: "20", cache: "npm" }
      - run: npm ci

      - name: Deploy prompts
        env:
          LANGFUSE_PUBLIC_KEY: ${{ secrets.LANGFUSE_PUBLIC_KEY }}
          LANGFUSE_SECRET_KEY: ${{ secrets.LANGFUSE_SECRET_KEY }}
        run: node scripts/deploy-prompts.mjs
typescript
// scripts/deploy-prompts.mjs
import { LangfuseClient } from "@langfuse/client";
import { readdirSync, readFileSync } from "fs";
import { join } from "path";

const langfuse = new LangfuseClient();
const promptDir = join(process.cwd(), "src/prompts");

for (const file of readdirSync(promptDir).filter((f) => f.endsWith(".json"))) {
  const config = JSON.parse(readFileSync(join(promptDir, file), "utf-8"));

  await langfuse.api.prompts.create({
    name: config.name,
    prompt: config.template,
    type: config.type || "text",
    config: config.config || {},
    labels: ["production", `deploy-${new Date().toISOString().split("T")[0]}`],
  });

  console.log(`Deployed: ${config.name}`);
}
Step 5: Score Regression Monitoring
typescript
// scripts/check-quality-regression.ts
import { LangfuseClient } from "@langfuse/client";

const langfuse = new LangfuseClient();

async function checkRegression() {
  const scores = await langfuse.api.scores.list({
    name: "quality",
    limit: 100,
  });

  const values = scores.data.map((s) => s.value).filter((v): v is number => v !== null);
  const avg = values.reduce((a, b) => a + b, 0) / values.length;

  console.log(`Average quality score: ${avg.toFixed(3)} (n=${values.length})`);

  if (avg < 0.7) {
    console.error("QUALITY REGRESSION: Score below 0.7 threshold");
    process.exit(1);
  }
}

checkRegression();

CI Best Practices

PracticeWhy
Use temperature: 0 in CI testsDeterministic outputs, fewer false failures
Separate CI API keysIsolate test traces from production
Run experiments on dataset changesCatch regressions before deploy
Assert on ranges, not exact stringsLLM output varies even at temp 0
Flush/shutdown in afterAllEnsure all traces reach Langfuse

Error Handling

IssueCauseSolution
Traces not in dashboardNo flush in CIAdd sdk.shutdown() or afterAll flush
Flaky quality testsNon-deterministic LLMUse temperature: 0, assert on ranges
Prompt not foundNot yet deployedDeploy prompts before running tests
Missing secrets in CINot configuredAdd to GitHub Settings > Secrets > Actions

Output

The pipeline reports the test suite result, experiment/run identifier, and aggregate quality score. A regression exits non-zero so the pull request is blocked; a pass leaves a traceable Langfuse run linked to the commit SHA.

Examples

Run the focused quality suite locally before opening a pull request:

bash
LANGFUSE_PUBLIC_KEY=pk-lf-... LANGFUSE_SECRET_KEY=sk-lf-... \
  npx vitest run tests/ai/ --reporter=verbose

For an intentionally changed prompt, first update the expected dataset result, then run the experiment gate and inspect the resulting run in Langfuse before merging.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/langfuse-ci-integration of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Langfuse CI Integration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langfuse CI Integration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langfuse CI Integration this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.4kAutomated safety check: PassMIT
Create Repo Agentlangfuse/langfuse36k—~730Automated safety check: PassCustom licence
Infra Scalinglangfuse/langfuse36k—~4.4kAutomated safety check: PassCustom licence
Ak Dev New Tracing Provideryaalalabs/agent-kernel192—~3.5kAutomated safety check: PassApache-2.0
Qv Devops Runner Queuetetherto/qvac685—~413Automated safety check: PassApache-2.0
Analyze GitHub Action Logswithastro/astro63k1 repos~1.3kAutomated safety check: PassCustom licence

Similar skills

  • Create Repo Agent

    langfuse/langfuse

    Design, implement, review, or harden Langfuse repo-owned autonomous agents.

    36k GitHub stars~730 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Infra Scaling

    langfuse/langfuse

    Tune and review Langfuse autoscaling for web, web-iso, and web-ingestion.

    36k GitHub stars~4.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Ak Dev New Tracing Provider

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new observability/tracing provider to Agent Kernel.

    192 GitHub stars~3.5k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Inspect GitHub Actions self-hosted runner queues when a developer provides a run URL, job URL, runner label, or reports a blocked CI job, or invokes /qv-devops-runner-queue.

    685 GitHub stars~413 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Official

    Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.

    63k GitHub starsUsed in 1 repo~1.3k tokens
    DevOps & CloudAuto-check passed
  • GitHub Actions Templates

    bartstc/vite-ts-react-template

    Create production-ready GitHub Actions workflows for automated testing, building, and deploying applications.

    122 GitHub starsUsed in 14 repos~1.9k tokens
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Langfuse CI Integration

What does Langfuse CI Integration do?

Configure Langfuse CI/CD integration with GitHub Actions and automated testing. Langfuse CI Integration is an agent skill from jeremylongshore/tons-of-skills-marketplace. Configure Langfuse CI/CD integration with GitHub Actions and automated testing.

When should I use Langfuse CI Integration?

Langfuse CI Integration fits situations like: setting up automated testing; configuring CI pipelines; integrating Langfuse tests into your build process; with phrases like langfuse CI.

How do I install Langfuse CI Integration in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langfuse-ci-integration -a claude-code`. Or copy the skill folder (skills/.curated/langfuse-ci-integration in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langfuse-ci-integration in your project. Claude Code loads it when a task matches its description.

How do I install Langfuse CI Integration in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langfuse-ci-integration -a codex`. Or copy the skill folder (skills/.curated/langfuse-ci-integration in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langfuse-ci-integration in your project. Codex loads it when a task matches its description.

Can I use Langfuse CI Integration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langfuse-ci-integration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langfuse-ci-integration, .gemini/skills/langfuse-ci-integration, .github/skills/langfuse-ci-integration and .opencode/skills/langfuse-ci-integration in your project.

What does Langfuse CI Integration need to run?

Going by SKILL.md and its folder, Langfuse CI Integration needs the command-line tools its instructions call (npx) and credentials named LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY and OPENAI_API_KEY. Our summary lists: Node.js; A credential in LANGFUSE_PUBLIC_KEY; A credential in LANGFUSE_SECRET_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(gh:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Langfuse CI Integration access the network?

SKILL.md names 2 domains. In commands or code: cloud.langfuse.com; the agent is likely to contact it when it follows the instructions. As links in the text: langfuse.com. This is read from the text; nothing was executed.

Is Langfuse CI Integration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langfuse CI Integration use?

Langfuse CI Integration is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langfuse CI Integration use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Langfuse CI Integration?

Skills that share tags, products or a category with Langfuse CI Integration: Create Repo Agent (langfuse/langfuse, 36k stars), Infra Scaling (langfuse/langfuse, 36k stars), Ak Dev New Tracing Provider (yaalalabs/agent-kernel, 192 stars) and Qv Devops Runner Queue (tetherto/qvac, 685 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langfuse CI Integration?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.