Agent skill

Deepeval

by sammcj in sammcj/agentic-coding

A skill your agent uses when discussing or working with DeepEval (the python AI evaluation framework)

Apache-2.0Auto-check: notesTesting & QA

Install Deepeval

skills CLI
$ npx skills add sammcj/agentic-coding --skill deepeval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sammcj/agentic-coding deepeval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Skills_disabled/deepeval .claude/skills/deepeval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deepeval
GitHub stars
162
Token cost
~3.5k tokens
SKILL.md length
596 words
Files
5 (incl. references)
Skills in repo
64
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when discussing or working with DeepEval (the python AI evaluation framework)

  • Working with DeepEval (the python AI evaluation framework)
  • SKILL.md covers Overview, Installation, Quick Start and Core Workflows, plus 7 more sections
  • Calls pip; needs OPENAI_API_KEY

What it does

Deepeval is an agent skill from sammcj/agentic-coding. Use when discussing or working with DeepEval (the python AI evaluation framework)

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/async_performance.md`, `references/custom_metrics.md` and `references/dataset_management.md`).

It sits in Testing & QA. It works with Python and pytest. The repository describes itself as: Agentic Coding Rules, Templates etc... The licence is Apache-2.0.

When your agent uses it

  • Working with DeepEval (the python AI evaluation framework)

Example prompts

  • “/deepeval”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 2f25ced. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • deepeval.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deepeval loads about 3.5k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 23 tokens; SKILL.md has 596 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~25k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:46
    DeepEval automatically loads `.env.local` then `.env`:
  • NoteMentions a .env fileSKILL.md:49
    # .env

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sammcj/agentic-coding at commit 2f25ced, republished under its Apache-2.0 licence (© sammcj). 596 words, ~3,545 tokens.

Download SKILL.mdSave it as .claude/skills/deepeval/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
deepeval
description
Use when discussing or working with DeepEval (the python AI evaluation framework)

DeepEval

Overview

DeepEval is a pytest-based framework for testing LLM applications. It provides 50+ evaluation metrics covering RAG pipelines, conversational AI, agents, safety, and custom criteria. DeepEval integrates into development workflows through pytest, supports multiple LLM providers, and includes component-level tracing with the @observe decorator.

Repository: https://github.com/confident-ai/deepeval Documentation: https://deepeval.com

Installation

bash
pip install -U deepeval

Requires Python 3.9+.

Quick Start

Basic pytest test
python
import pytest
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric

def test_chatbot():
    metric = AnswerRelevancyMetric(threshold=0.7, model="athropic-claude-sonnet-4-5")
    test_case = LLMTestCase(
        input="What if these shoes don't fit?",
        actual_output="You have 30 days for full refund"
    )
    assert_test(test_case, [metric])

Run with: deepeval test run test_chatbot.py

Environment setup

DeepEval automatically loads .env.local then .env:

bash
# .env
OPENAI_API_KEY="sk-..."

Core Workflows

RAG Evaluation

Evaluate both retrieval and generation phases:

python
from deepeval.metrics import (
    ContextualPrecisionMetric,
    ContextualRecallMetric,
    ContextualRelevancyMetric,
    AnswerRelevancyMetric,
    FaithfulnessMetric
)

# Retrieval metrics
contextual_precision = ContextualPrecisionMetric(threshold=0.7)
contextual_recall = ContextualRecallMetric(threshold=0.7)
contextual_relevancy = ContextualRelevancyMetric(threshold=0.7)

# Generation metrics
answer_relevancy = AnswerRelevancyMetric(threshold=0.7)
faithfulness = FaithfulnessMetric(threshold=0.8)

test_case = LLMTestCase(
    input="What are the side effects of aspirin?",
    actual_output="Common side effects include stomach upset and nausea.",
    expected_output="Aspirin side effects include gastrointestinal issues.",
    retrieval_context=[
        "Aspirin common side effects: stomach upset, nausea, vomiting.",
        "Serious aspirin side effects: gastrointestinal bleeding.",
    ]
)

evaluate(test_cases=[test_case], metrics=[
    contextual_precision, contextual_recall, contextual_relevancy,
    answer_relevancy, faithfulness
])

Component-level tracing:

python
from deepeval.tracing import observe, update_current_span

@observe(metrics=[contextual_relevancy])
def retriever(query: str):
    chunks = your_vector_db.search(query)
    update_current_span(
        test_case=LLMTestCase(input=query, retrieval_context=chunks)
    )
    return chunks

@observe(metrics=[answer_relevancy, faithfulness])
def generator(query: str, chunks: list):
    response = your_llm.generate(query, chunks)
    update_current_span(
        test_case=LLMTestCase(
            input=query,
            actual_output=response,
            retrieval_context=chunks
        )
    )
    return response

@observe
def rag_pipeline(query: str):
    chunks = retriever(query)
    return generator(query, chunks)
Conversational AI Evaluation

Test multi-turn dialogues:

python
from deepeval.test_case import Turn, ConversationalTestCase
from deepeval.metrics import (
    RoleAdherenceMetric,
    KnowledgeRetentionMetric,
    ConversationCompletenessMetric,
    TurnRelevancyMetric
)

convo_test_case = ConversationalTestCase(
    chatbot_role="professional, empathetic medical assistant",
    turns=[
        Turn(role="user", content="I have a persistent cough"),
        Turn(role="assistant", content="How long have you had this cough?"),
        Turn(role="user", content="About a week now"),
        Turn(role="assistant", content="A week-long cough should be evaluated.")
    ]
)

metrics = [
    RoleAdherenceMetric(threshold=0.7),
    KnowledgeRetentionMetric(threshold=0.7),
    ConversationCompletenessMetric(threshold=0.6),
    TurnRelevancyMetric(threshold=0.7)
]

evaluate(test_cases=[convo_test_case], metrics=metrics)
Agent Evaluation

Test tool usage and task completion:

python
from deepeval.test_case import ToolCall
from deepeval.metrics import (
    TaskCompletionMetric,
    ToolUseMetric,
    ArgumentCorrectnessMetric
)

agent_test_case = ConversationalTestCase(
    turns=[
        Turn(role="user", content="When did Trump first raise tariffs?"),
        Turn(
            role="assistant",
            content="Let me search for that information.",
            tools_called=[
                ToolCall(
                    name="WebSearch",
                    arguments={"query": "Trump first raised tariffs year"}
                )
            ]
        ),
        Turn(role="assistant", content="Trump first raised tariffs in 2018.")
    ]
)

evaluate(
    test_cases=[agent_test_case],
    metrics=[
        TaskCompletionMetric(threshold=0.7),
        ToolUseMetric(threshold=0.7),
        ArgumentCorrectnessMetric(threshold=0.7)
    ]
)
Safety Evaluation

Check for harmful content:

python
from deepeval.metrics import (
    ToxicityMetric,
    BiasMetric,
    PIILeakageMetric,
    HallucinationMetric
)

def safety_gate(output: str, input: str) -> tuple[bool, list]:
    """Returns (passed, reasons) tuple"""
    test_case = LLMTestCase(input=input, actual_output=output)

    safety_metrics = [
        ToxicityMetric(threshold=0.5),
        BiasMetric(threshold=0.5),
        PIILeakageMetric(threshold=0.5)
    ]

    failures = []
    for metric in safety_metrics:
        metric.measure(test_case)
        if not metric.is_successful():
            failures.append(f"{metric.name}: {metric.reason}")

    return len(failures) == 0, failures

Metric Selection Guide

RAG Metrics

Retrieval Phase:

  • ContextualPrecisionMetric - Relevant chunks ranked higher than irrelevant ones
  • ContextualRecallMetric - All necessary information retrieved
  • ContextualRelevancyMetric - Retrieved chunks relevant to input

Generation Phase:

  • AnswerRelevancyMetric - Output addresses the input query
  • FaithfulnessMetric - Output grounded in retrieval context
Conversational Metrics
  • TurnRelevancyMetric - Each turn relevant to conversation
  • KnowledgeRetentionMetric - Information retained across turns
  • ConversationCompletenessMetric - All aspects addressed
  • RoleAdherenceMetric - Chatbot maintains assigned role
  • TopicAdherenceMetric - Conversation stays on topic
Agent Metrics
  • TaskCompletionMetric - Task successfully completed
  • ToolUseMetric - Correct tools selected
  • ArgumentCorrectnessMetric - Tool arguments correct
  • MCPUseMetric - MCP correctly used
Safety Metrics
  • ToxicityMetric - Harmful content detection
  • BiasMetric - Biased outputs identification
  • HallucinationMetric - Fabricated information
  • PIILeakageMetric - Personal information leakage
Custom Metrics

G-Eval (LLM-based):

python
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCaseParams

custom_metric = GEval(
    name="Professional Tone",
    criteria="Determine if response maintains professional, empathetic tone",
    evaluation_params=[LLMTestCaseParams.ACTUAL_OUTPUT],
    threshold=0.7,
    model="anthropic-claude-sonnet-4-5"
)

BaseMetric subclass:

See references/custom_metrics.md for complete guide on creating custom metrics with BaseMetric subclassing and deterministic scorers (ROUGE, BLEU, BERTScore).

Configuration

LLM Provider Setup

DeepEval supports OpenAI, Anthropic Claude, Google Gemini, AWS Bedrock, and 100+ providers via LiteLLM. Anthropic models are preferred.

CLI configuration (global):

bash
deepeval set-azure-openai --openai-endpoint=... --openai-api-key=... --deployment-name=...
deepeval set-ollama deepseek-r1:1.5b

Python configuration (per-metric):

python
from deepeval.models import AnthropicModel, OllamaModel

anthropic_model = AnthropicModel(
    model_id=settings.anthropic_model_id,
    client_args={"api_key": settings.anthropic_api_key},
    temperature=settings.agent_temperature
)

metric = AnswerRelevancyMetric(model=anthropic_model)

See references/model_providers.md for complete provider configuration guide.

Performance Optimisation

Async mode is enabled by default. Configure with AsyncConfig and CacheConfig:

python
from deepeval import evaluate, AsyncConfig, CacheConfig

evaluate(
    test_cases=[...],
    metrics=[...],
    async_config=AsyncConfig(
        run_async=True,
        max_concurrent=20,    # Reduce if rate limited
        throttle_value=0      # Delay between test cases (seconds)
    ),
    cache_config=CacheConfig(
        use_cache=True,       # Read from cache
        write_cache=True      # Write to cache
    )
)

CLI parallelisation:

bash
deepeval test run -n 4 -c -i  # 4 processes, cached, ignore errors

Best practices:

  • Limit to 5 metrics maximum (2-3 generic + 1-2 custom)
  • Use the latest available Anthropic Claude Sonnet or Haiku models
  • Reduce max_concurrent to 5 if hitting rate limits
  • Use evaluate() function over individual measure() calls

See references/async_performance.md for detailed performance optimisation guide.

Dataset Management

Loading datasets
python
from deepeval.dataset import EvaluationDataset, Golden

dataset = EvaluationDataset()

# From CSV
dataset.add_goldens_from_csv_file(
    file_path="./test_data.csv",
    input_col_name="question",
    expected_output_col_name="answer",
    context_col_name="context",
    context_col_delimiter="|"
)

# From JSON
dataset.add_goldens_from_json_file(
    file_path="./test_data.json",
    input_key_name="query",
    expected_output_key_name="response"
)
Synthetic generation
python
from deepeval.synthesizer import Synthesizer

synthesizer = Synthesizer()

# From documents
goldens = synthesizer.generate_goldens_from_docs(
    document_paths=["./docs/knowledge_base.pdf"],
    max_goldens_per_document=10,
    evolution_types=["REASONING", "MULTICONTEXT", "COMPARATIVE"]
)

# From scratch
goldens = synthesizer.generate_goldens_from_scratch(
    subject="customer support for SaaS product",
    task="answer user questions about billing",
    max_goldens=20
)

Evolution types: REASONING, MULTICONTEXT, CONCRETISING, CONSTRAINED, COMPARATIVE, HYPOTHETICAL, IN_BREADTH

See references/dataset_management.md for complete dataset guide including versioning and cloud integration.

Test Case Types

Single-turn (LLMTestCase)
python
from deepeval.test_case import LLMTestCase

test_case = LLMTestCase(
    input="What if these shoes don't fit?",
    actual_output="You have 30 days for full refund",
    expected_output="We offer 30-day full refund",
    retrieval_context=["All customers eligible for 30 day refund"],
    tools_called=[ToolCall(name="...", arguments={"...": "..."})]
)
Multi-turn (ConversationalTestCase)
python
from deepeval.test_case import Turn, ConversationalTestCase

convo_test_case = ConversationalTestCase(
    chatbot_role="helpful customer service agent",
    turns=[
        Turn(role="user", content="I need help with my order"),
        Turn(role="assistant", content="I'd be happy to help"),
        Turn(role="user", content="It hasn't arrived yet")
    ]
)
Multimodal (MLLMTestCase)
python
from deepeval.test_case import MLLMTestCase, MLLMImage

m_test_case = MLLMTestCase(
    input=["Describe this image", MLLMImage(url="./photo.png", local=True)],
    actual_output=["A red bicycle leaning against a wall"]
)
Show full SKILL.md (238 more words)Show less

CI/CD Integration

yaml
# .github/workflows/test.yml
name: LLM Tests
on: [push, pull_request]

jobs:
  evaluate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5
      - name: Install dependencies
        run: pip install deepeval
      - name: Run evaluations
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
        run: deepeval test run tests/

References

Detailed implementation guides:

  • references/model_providers.md - Complete guide for configuring OpenAI, Anthropic, Gemini, Bedrock, and local models. Includes provider-specific considerations, cost analysis, and troubleshooting.

  • references/custom_metrics.md - Complete guide for creating custom metrics by subclassing BaseMetric. Includes deterministic scorers (ROUGE, BLEU, BERTScore) and LLM-based evaluation patterns.

  • references/async_performance.md - Complete guide for optimising evaluation performance with async mode, caching, concurrency tuning, and rate limit handling.

  • references/dataset_management.md - Complete guide for dataset loading, saving, synthetic generation, versioning, and cloud integration with Confident AI.

Best Practices

Metric Selection
  • Match metrics to use case (RAG systems need retrieval + generation metrics)
  • Start with 2-3 essential metrics, expand as needed
  • Use appropriate thresholds (0.7-0.8 for production, 0.5-0.6 for development)
  • Combine complementary metrics (answer relevancy + faithfulness)
Test Case Design
  • Create representative examples covering common queries and edge cases
  • Include context when needed (retrieval_context for RAG, expected_output for G-Eval)
  • Use datasets for scale testing
  • Version test cases over time
Evaluation Workflow
  • Component-level first - Use @observe for individual parts
  • End-to-end validation before deployment
  • Automate in CI/CD with deepeval test run
  • Track results over time with Confident AI cloud
Testing Anti-Patterns

Avoid:

  • Testing only happy paths
  • Using unrealistic inputs
  • Ignoring metric reasons
  • Setting thresholds too high initially
  • Running full test suite on every change

Do:

  • Test edge cases and failure modes
  • Use real user queries as test inputs
  • Read and analyse metric reasons
  • Adjust thresholds based on empirical results
  • Use component-level tests during development
  • Separate config and eval content from code

© sammcj, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in Skills_disabled/deepeval of sammcj/agentic-coding.

  • SKILL.md
  • references/async_performance.md
  • references/custom_metrics.md
  • references/dataset_management.md
  • references/model_providers.md

Open the folder on GitHubat commit 2f25ced

Compare with similar skills

Deepeval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deepeval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deepeval this skillsammcj/agentic-coding162—~3.5kAutomated safety check: NotesApache-2.0
Running Testsbrendanhasz/probflow175—~657Automated safety check: PassMIT
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence
ONNX Runtime Test Runnermicrosoft/onnxruntime22k—~1.8kAutomated safety check: PassMIT
Simple Modern Uvjlevy/simple-modern-uv301—~1.9kAutomated safety check: PassMIT

Similar skills

  • Running Tests

    brendanhasz/probflow

    Run Python unit test suites strictly using the uv package manager and pytest.

    175 GitHub stars~657 tokensUpdated 13 days ago
    Testing & QAAuto-check passed
  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • ONNX Runtime Test Runner

    microsoft/onnxruntime

    Official

    Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.

    22k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Simple Modern Uv

    jlevy/simple-modern-uv

    Start, selectively modernize, fully migrate, or update Python projects using simple-modern-uv practices: uv, ruff, BasedPyright, pytest, GitHub Actions CI, and tag-driven PyPI publishing.

    301 GitHub stars~1.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Test Coverage Review

    areed1192/finance-news-aggregator

    Audit, plan, write, and verify unit tests for Python projects using pytest.

    149 GitHub stars~2.6k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed

More from sammcj/agentic-coding

All 64 skills in this repo
  • Yue2 Music

    sammcj/agentic-coding

    A skill your agent uses when generating songs with YuE2, covering a recording via SheetSage2 audio-to-ABC, editing a score or lyrics with melody preservation, or building a reproducible listening…

    162 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Bento Slides

    sammcj/agentic-coding

    A skill your agent uses when creating or editing Bento (.bento.html) slide decks, including any request for a single-file HTML slide deck.

    162 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Idrive Backup

    sammcj/agentic-coding

    A skill your agent uses whenever the user wants you to manage, discuss or diagnose iDrive Backup configuration on macOS

    162 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check: notes
  • Piper Tts Training

    sammcj/agentic-coding

    Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches.

    162 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • PPTX To Md

    sammcj/agentic-coding

    Convert a PPTX slide deck into per-slide markdown that preserves both the verbatim text and the meaning of embedded screenshots, diagrams and charts in their original layout positions.

    162 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Skill Creator Primer

    sammcj/agentic-coding

    You MUST load this skill before the skill-creator skill AND before making ANY change to, or conducting a review of ANY Agent Skill.

    162 GitHub stars~9.8k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Deepeval

What does Deepeval do?

A skill your agent uses when discussing or working with DeepEval (the python AI evaluation framework). Deepeval is an agent skill from sammcj/agentic-coding.

When should I use Deepeval?

Deepeval fits situations like: working with DeepEval (the python AI evaluation framework).

How do I install Deepeval in Claude Code?

Run `npx skills add sammcj/agentic-coding --skill deepeval -a claude-code`. Or copy the skill folder (Skills_disabled/deepeval in sammcj/agentic-coding) into .claude/skills/deepeval in your project. Claude Code loads it when a task matches its description.

How do I install Deepeval in Codex?

Run `npx skills add sammcj/agentic-coding --skill deepeval -a codex`. Or copy the skill folder (Skills_disabled/deepeval in sammcj/agentic-coding) into .agents/skills/deepeval in your project. Codex loads it when a task matches its description.

Can I use Deepeval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sammcj/agentic-coding --skill deepeval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepeval, .gemini/skills/deepeval, .github/skills/deepeval and .opencode/skills/deepeval in your project.

What does Deepeval need to run?

Going by SKILL.md and its folder, Deepeval needs the command-line tools its instructions call (pip) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Deepeval access the network?

SKILL.md names 2 domains. As links in the text: github.com and deepeval.com. This is read from the text; nothing was executed.

Is Deepeval safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Deepeval use?

Deepeval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deepeval use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.

What are the alternatives to Deepeval?

Skills that share tags, products or a category with Deepeval: Running Tests (brendanhasz/probflow, 175 stars), Adk Verify Snippets (google/adk-python, 22k stars), Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars) and ONNX Runtime Test Runner (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deepeval?

sammcj (a GitHub user) maintains it in sammcj/agentic-coding, which has 162 GitHub stars. The repository holds 64 skills in this directory. The repository was last updated on October 9, 2026.

Source: sammcj/agentic-coding on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.