Running Tests
brendanhasz/probflow
Run Python unit test suites strictly using the uv package manager and pytest.
A skill your agent uses when discussing or working with DeepEval (the python AI evaluation framework)
$ npx skills add sammcj/agentic-coding --skill deepeval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sammcj/agentic-coding deepeval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Skills_disabled/deepeval .claude/skills/deepeval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "deepeval" agent skill from https://github.com/sammcj/agentic-coding/tree/main/Skills_disabled/deepeval into .claude/skills/deepeval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepeval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sammcj/agentic-coding/tree/main/Skills_disabled/deepevalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sammcj/agentic-coding --skill deepeval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sammcj/agentic-coding deepeval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .agents/skills && cp -r skills-src/Skills_disabled/deepeval .agents/skills/deepeval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "deepeval" agent skill from https://github.com/sammcj/agentic-coding/tree/main/Skills_disabled/deepeval into .agents/skills/deepeval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepeval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sammcj/agentic-coding --skill deepeval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sammcj/agentic-coding deepeval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/Skills_disabled/deepeval .cursor/skills/deepeval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "deepeval" agent skill from https://github.com/sammcj/agentic-coding/tree/main/Skills_disabled/deepeval into .cursor/skills/deepeval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepeval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sammcj/agentic-coding.git --path Skills_disabled/deepeval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sammcj/agentic-coding --skill deepeval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sammcj/agentic-coding deepeval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/Skills_disabled/deepeval .gemini/skills/deepeval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "deepeval" agent skill from https://github.com/sammcj/agentic-coding/tree/main/Skills_disabled/deepeval into .gemini/skills/deepeval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepeval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sammcj/agentic-coding deepevalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sammcj/agentic-coding --skill deepeval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .github/skills && cp -r skills-src/Skills_disabled/deepeval .github/skills/deepeval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "deepeval" agent skill from https://github.com/sammcj/agentic-coding/tree/main/Skills_disabled/deepeval into .github/skills/deepeval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepeval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sammcj/agentic-coding --skill deepeval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sammcj/agentic-coding deepeval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/Skills_disabled/deepeval .opencode/skills/deepeval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "deepeval" agent skill from https://github.com/sammcj/agentic-coding/tree/main/Skills_disabled/deepeval into .opencode/skills/deepeval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepeval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
deepevalA skill your agent uses when discussing or working with DeepEval (the python AI evaluation framework)
Deepeval is an agent skill from sammcj/agentic-coding. Use when discussing or working with DeepEval (the python AI evaluation framework)
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/async_performance.md`, `references/custom_metrics.md` and `references/dataset_management.md`).
It sits in Testing & QA. It works with Python and pytest. The repository describes itself as: Agentic Coding Rules, Templates etc... The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 2f25ced. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comdeepeval.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Deepeval loads about 3.5k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 23 tokens; SKILL.md has 596 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
DeepEval automatically loads `.env.local` then `.env`:# .envAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sammcj/agentic-coding at commit 2f25ced, republished under its Apache-2.0 licence (© sammcj). 596 words, ~3,545 tokens.
.claude/skills/deepeval/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.DeepEval is a pytest-based framework for testing LLM applications. It provides 50+ evaluation metrics covering RAG pipelines, conversational AI, agents, safety, and custom criteria. DeepEval integrates into development workflows through pytest, supports multiple LLM providers, and includes component-level tracing with the @observe decorator.
Repository: https://github.com/confident-ai/deepeval Documentation: https://deepeval.com
pip install -U deepevalRequires Python 3.9+.
import pytest
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric
def test_chatbot():
metric = AnswerRelevancyMetric(threshold=0.7, model="athropic-claude-sonnet-4-5")
test_case = LLMTestCase(
input="What if these shoes don't fit?",
actual_output="You have 30 days for full refund"
)
assert_test(test_case, [metric])Run with: deepeval test run test_chatbot.py
DeepEval automatically loads .env.local then .env:
# .env
OPENAI_API_KEY="sk-..."Evaluate both retrieval and generation phases:
from deepeval.metrics import (
ContextualPrecisionMetric,
ContextualRecallMetric,
ContextualRelevancyMetric,
AnswerRelevancyMetric,
FaithfulnessMetric
)
# Retrieval metrics
contextual_precision = ContextualPrecisionMetric(threshold=0.7)
contextual_recall = ContextualRecallMetric(threshold=0.7)
contextual_relevancy = ContextualRelevancyMetric(threshold=0.7)
# Generation metrics
answer_relevancy = AnswerRelevancyMetric(threshold=0.7)
faithfulness = FaithfulnessMetric(threshold=0.8)
test_case = LLMTestCase(
input="What are the side effects of aspirin?",
actual_output="Common side effects include stomach upset and nausea.",
expected_output="Aspirin side effects include gastrointestinal issues.",
retrieval_context=[
"Aspirin common side effects: stomach upset, nausea, vomiting.",
"Serious aspirin side effects: gastrointestinal bleeding.",
]
)
evaluate(test_cases=[test_case], metrics=[
contextual_precision, contextual_recall, contextual_relevancy,
answer_relevancy, faithfulness
])Component-level tracing:
from deepeval.tracing import observe, update_current_span
@observe(metrics=[contextual_relevancy])
def retriever(query: str):
chunks = your_vector_db.search(query)
update_current_span(
test_case=LLMTestCase(input=query, retrieval_context=chunks)
)
return chunks
@observe(metrics=[answer_relevancy, faithfulness])
def generator(query: str, chunks: list):
response = your_llm.generate(query, chunks)
update_current_span(
test_case=LLMTestCase(
input=query,
actual_output=response,
retrieval_context=chunks
)
)
return response
@observe
def rag_pipeline(query: str):
chunks = retriever(query)
return generator(query, chunks)Test multi-turn dialogues:
from deepeval.test_case import Turn, ConversationalTestCase
from deepeval.metrics import (
RoleAdherenceMetric,
KnowledgeRetentionMetric,
ConversationCompletenessMetric,
TurnRelevancyMetric
)
convo_test_case = ConversationalTestCase(
chatbot_role="professional, empathetic medical assistant",
turns=[
Turn(role="user", content="I have a persistent cough"),
Turn(role="assistant", content="How long have you had this cough?"),
Turn(role="user", content="About a week now"),
Turn(role="assistant", content="A week-long cough should be evaluated.")
]
)
metrics = [
RoleAdherenceMetric(threshold=0.7),
KnowledgeRetentionMetric(threshold=0.7),
ConversationCompletenessMetric(threshold=0.6),
TurnRelevancyMetric(threshold=0.7)
]
evaluate(test_cases=[convo_test_case], metrics=metrics)Test tool usage and task completion:
from deepeval.test_case import ToolCall
from deepeval.metrics import (
TaskCompletionMetric,
ToolUseMetric,
ArgumentCorrectnessMetric
)
agent_test_case = ConversationalTestCase(
turns=[
Turn(role="user", content="When did Trump first raise tariffs?"),
Turn(
role="assistant",
content="Let me search for that information.",
tools_called=[
ToolCall(
name="WebSearch",
arguments={"query": "Trump first raised tariffs year"}
)
]
),
Turn(role="assistant", content="Trump first raised tariffs in 2018.")
]
)
evaluate(
test_cases=[agent_test_case],
metrics=[
TaskCompletionMetric(threshold=0.7),
ToolUseMetric(threshold=0.7),
ArgumentCorrectnessMetric(threshold=0.7)
]
)Check for harmful content:
from deepeval.metrics import (
ToxicityMetric,
BiasMetric,
PIILeakageMetric,
HallucinationMetric
)
def safety_gate(output: str, input: str) -> tuple[bool, list]:
"""Returns (passed, reasons) tuple"""
test_case = LLMTestCase(input=input, actual_output=output)
safety_metrics = [
ToxicityMetric(threshold=0.5),
BiasMetric(threshold=0.5),
PIILeakageMetric(threshold=0.5)
]
failures = []
for metric in safety_metrics:
metric.measure(test_case)
if not metric.is_successful():
failures.append(f"{metric.name}: {metric.reason}")
return len(failures) == 0, failuresRetrieval Phase:
ContextualPrecisionMetric - Relevant chunks ranked higher than irrelevant onesContextualRecallMetric - All necessary information retrievedContextualRelevancyMetric - Retrieved chunks relevant to inputGeneration Phase:
AnswerRelevancyMetric - Output addresses the input queryFaithfulnessMetric - Output grounded in retrieval contextTurnRelevancyMetric - Each turn relevant to conversationKnowledgeRetentionMetric - Information retained across turnsConversationCompletenessMetric - All aspects addressedRoleAdherenceMetric - Chatbot maintains assigned roleTopicAdherenceMetric - Conversation stays on topicTaskCompletionMetric - Task successfully completedToolUseMetric - Correct tools selectedArgumentCorrectnessMetric - Tool arguments correctMCPUseMetric - MCP correctly usedToxicityMetric - Harmful content detectionBiasMetric - Biased outputs identificationHallucinationMetric - Fabricated informationPIILeakageMetric - Personal information leakageG-Eval (LLM-based):
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCaseParams
custom_metric = GEval(
name="Professional Tone",
criteria="Determine if response maintains professional, empathetic tone",
evaluation_params=[LLMTestCaseParams.ACTUAL_OUTPUT],
threshold=0.7,
model="anthropic-claude-sonnet-4-5"
)BaseMetric subclass:
See references/custom_metrics.md for complete guide on creating custom metrics with BaseMetric subclassing and deterministic scorers (ROUGE, BLEU, BERTScore).
DeepEval supports OpenAI, Anthropic Claude, Google Gemini, AWS Bedrock, and 100+ providers via LiteLLM. Anthropic models are preferred.
CLI configuration (global):
deepeval set-azure-openai --openai-endpoint=... --openai-api-key=... --deployment-name=...
deepeval set-ollama deepseek-r1:1.5bPython configuration (per-metric):
from deepeval.models import AnthropicModel, OllamaModel
anthropic_model = AnthropicModel(
model_id=settings.anthropic_model_id,
client_args={"api_key": settings.anthropic_api_key},
temperature=settings.agent_temperature
)
metric = AnswerRelevancyMetric(model=anthropic_model)See references/model_providers.md for complete provider configuration guide.
Async mode is enabled by default. Configure with AsyncConfig and CacheConfig:
from deepeval import evaluate, AsyncConfig, CacheConfig
evaluate(
test_cases=[...],
metrics=[...],
async_config=AsyncConfig(
run_async=True,
max_concurrent=20, # Reduce if rate limited
throttle_value=0 # Delay between test cases (seconds)
),
cache_config=CacheConfig(
use_cache=True, # Read from cache
write_cache=True # Write to cache
)
)CLI parallelisation:
deepeval test run -n 4 -c -i # 4 processes, cached, ignore errorsBest practices:
max_concurrent to 5 if hitting rate limitsevaluate() function over individual measure() callsSee references/async_performance.md for detailed performance optimisation guide.
from deepeval.dataset import EvaluationDataset, Golden
dataset = EvaluationDataset()
# From CSV
dataset.add_goldens_from_csv_file(
file_path="./test_data.csv",
input_col_name="question",
expected_output_col_name="answer",
context_col_name="context",
context_col_delimiter="|"
)
# From JSON
dataset.add_goldens_from_json_file(
file_path="./test_data.json",
input_key_name="query",
expected_output_key_name="response"
)from deepeval.synthesizer import Synthesizer
synthesizer = Synthesizer()
# From documents
goldens = synthesizer.generate_goldens_from_docs(
document_paths=["./docs/knowledge_base.pdf"],
max_goldens_per_document=10,
evolution_types=["REASONING", "MULTICONTEXT", "COMPARATIVE"]
)
# From scratch
goldens = synthesizer.generate_goldens_from_scratch(
subject="customer support for SaaS product",
task="answer user questions about billing",
max_goldens=20
)Evolution types: REASONING, MULTICONTEXT, CONCRETISING, CONSTRAINED, COMPARATIVE, HYPOTHETICAL, IN_BREADTH
See references/dataset_management.md for complete dataset guide including versioning and cloud integration.
from deepeval.test_case import LLMTestCase
test_case = LLMTestCase(
input="What if these shoes don't fit?",
actual_output="You have 30 days for full refund",
expected_output="We offer 30-day full refund",
retrieval_context=["All customers eligible for 30 day refund"],
tools_called=[ToolCall(name="...", arguments={"...": "..."})]
)from deepeval.test_case import Turn, ConversationalTestCase
convo_test_case = ConversationalTestCase(
chatbot_role="helpful customer service agent",
turns=[
Turn(role="user", content="I need help with my order"),
Turn(role="assistant", content="I'd be happy to help"),
Turn(role="user", content="It hasn't arrived yet")
]
)from deepeval.test_case import MLLMTestCase, MLLMImage
m_test_case = MLLMTestCase(
input=["Describe this image", MLLMImage(url="./photo.png", local=True)],
actual_output=["A red bicycle leaning against a wall"]
)# .github/workflows/test.yml
name: LLM Tests
on: [push, pull_request]
jobs:
evaluate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- name: Install dependencies
run: pip install deepeval
- name: Run evaluations
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
run: deepeval test run tests/Detailed implementation guides:
references/model_providers.md - Complete guide for configuring OpenAI, Anthropic, Gemini, Bedrock, and local models. Includes provider-specific considerations, cost analysis, and troubleshooting.
references/custom_metrics.md - Complete guide for creating custom metrics by subclassing BaseMetric. Includes deterministic scorers (ROUGE, BLEU, BERTScore) and LLM-based evaluation patterns.
references/async_performance.md - Complete guide for optimising evaluation performance with async mode, caching, concurrency tuning, and rate limit handling.
references/dataset_management.md - Complete guide for dataset loading, saving, synthetic generation, versioning, and cloud integration with Confident AI.
retrieval_context for RAG, expected_output for G-Eval)@observe for individual partsdeepeval test runAvoid:
Do:
© sammcj, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in Skills_disabled/deepeval of sammcj/agentic-coding.
Open the folder on GitHubat commit 2f25ced
Deepeval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Deepeval this skillsammcj/agentic-coding | 162 | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | |
| Running Testsbrendanhasz/probflow | 175 | — | ~657 | Automated safety check: Pass | MIT | |
| Adk Verify Snippetsgoogle/adk-python | 22k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Hermetic Python Unit TestsdimensionalOS/dimos | 4.6k | — | ~1.4k | Automated safety check: Pass | Custom licence | |
| ONNX Runtime Test Runnermicrosoft/onnxruntime | 22k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Simple Modern Uvjlevy/simple-modern-uv | 301 | — | ~1.9k | Automated safety check: Pass | MIT |
brendanhasz/probflow
Run Python unit test suites strictly using the uv package manager and pytest.
google/adk-python
Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…
dimensionalOS/dimos
Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.
microsoft/onnxruntime
Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.
jlevy/simple-modern-uv
Start, selectively modernize, fully migrate, or update Python projects using simple-modern-uv practices: uv, ruff, BasedPyright, pytest, GitHub Actions CI, and tag-driven PyPI publishing.
areed1192/finance-news-aggregator
Audit, plan, write, and verify unit tests for Python projects using pytest.
sammcj/agentic-coding
A skill your agent uses when generating songs with YuE2, covering a recording via SheetSage2 audio-to-ABC, editing a score or lyrics with melody preservation, or building a reproducible listening…
sammcj/agentic-coding
A skill your agent uses when creating or editing Bento (.bento.html) slide decks, including any request for a single-file HTML slide deck.
sammcj/agentic-coding
A skill your agent uses whenever the user wants you to manage, discuss or diagnose iDrive Backup configuration on macOS
sammcj/agentic-coding
Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches.
sammcj/agentic-coding
Convert a PPTX slide deck into per-slide markdown that preserves both the verbatim text and the meaning of embedded screenshots, diagrams and charts in their original layout positions.
sammcj/agentic-coding
You MUST load this skill before the skill-creator skill AND before making ANY change to, or conducting a review of ANY Agent Skill.
Categories
A skill your agent uses when discussing or working with DeepEval (the python AI evaluation framework). Deepeval is an agent skill from sammcj/agentic-coding.
Deepeval fits situations like: working with DeepEval (the python AI evaluation framework).
Run `npx skills add sammcj/agentic-coding --skill deepeval -a claude-code`. Or copy the skill folder (Skills_disabled/deepeval in sammcj/agentic-coding) into .claude/skills/deepeval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sammcj/agentic-coding --skill deepeval -a codex`. Or copy the skill folder (Skills_disabled/deepeval in sammcj/agentic-coding) into .agents/skills/deepeval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sammcj/agentic-coding --skill deepeval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepeval, .gemini/skills/deepeval, .github/skills/deepeval and .opencode/skills/deepeval in your project.
Going by SKILL.md and its folder, Deepeval needs the command-line tools its instructions call (pip) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.
SKILL.md names 2 domains. As links in the text: github.com and deepeval.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Deepeval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Deepeval: Running Tests (brendanhasz/probflow, 175 stars), Adk Verify Snippets (google/adk-python, 22k stars), Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars) and ONNX Runtime Test Runner (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sammcj (a GitHub user) maintains it in sammcj/agentic-coding, which has 162 GitHub stars. The repository holds 64 skills in this directory. The repository was last updated on October 9, 2026.
Source: sammcj/agentic-coding on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.