Agent Eval Cases
agentailor/fullstack-langgraph-nextjs-agent
Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.
Set up compliance exports, drift detection, evaluations, scoring, and learning analytics
$ npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ucsandman/DashClaw compliance-drift-evals --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ucsandman/DashClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/dashclaw-agent/compliance-drift-evals .claude/skills/compliance-drift-evals && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "compliance-drift-evals" agent skill from https://github.com/ucsandman/DashClaw/tree/main/.claude/skills/dashclaw-agent/compliance-drift-evals into .claude/skills/compliance-drift-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compliance-drift-evals", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ucsandman/DashClaw/tree/main/.claude/skills/dashclaw-agent/compliance-drift-evalsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ucsandman/DashClaw compliance-drift-evals --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ucsandman/DashClaw.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/dashclaw-agent/compliance-drift-evals .agents/skills/compliance-drift-evals && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "compliance-drift-evals" agent skill from https://github.com/ucsandman/DashClaw/tree/main/.claude/skills/dashclaw-agent/compliance-drift-evals into .agents/skills/compliance-drift-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compliance-drift-evals", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ucsandman/DashClaw compliance-drift-evals --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ucsandman/DashClaw.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/dashclaw-agent/compliance-drift-evals .cursor/skills/compliance-drift-evals && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "compliance-drift-evals" agent skill from https://github.com/ucsandman/DashClaw/tree/main/.claude/skills/dashclaw-agent/compliance-drift-evals into .cursor/skills/compliance-drift-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compliance-drift-evals", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ucsandman/DashClaw.git --path .claude/skills/dashclaw-agent/compliance-drift-evals--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ucsandman/DashClaw compliance-drift-evals --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ucsandman/DashClaw.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/dashclaw-agent/compliance-drift-evals .gemini/skills/compliance-drift-evals && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "compliance-drift-evals" agent skill from https://github.com/ucsandman/DashClaw/tree/main/.claude/skills/dashclaw-agent/compliance-drift-evals into .gemini/skills/compliance-drift-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compliance-drift-evals", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ucsandman/DashClaw compliance-drift-evalsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ucsandman/DashClaw.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/dashclaw-agent/compliance-drift-evals .github/skills/compliance-drift-evals && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "compliance-drift-evals" agent skill from https://github.com/ucsandman/DashClaw/tree/main/.claude/skills/dashclaw-agent/compliance-drift-evals into .github/skills/compliance-drift-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compliance-drift-evals", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ucsandman/DashClaw compliance-drift-evals --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ucsandman/DashClaw.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/dashclaw-agent/compliance-drift-evals .opencode/skills/compliance-drift-evals && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "compliance-drift-evals" agent skill from https://github.com/ucsandman/DashClaw/tree/main/.claude/skills/dashclaw-agent/compliance-drift-evals into .opencode/skills/compliance-drift-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compliance-drift-evals", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
compliance-drift-evalsSet up compliance exports, drift detection, evaluations, scoring, and learning analytics
Compliance Drift Evals is an agent skill from ucsandman/DashClaw. Set up compliance exports, drift detection, evaluations, scoring, and learning analytics
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation, GitOps and AI governance. It works with Model Context Protocol. The repository describes itself as: Remote approvals, policy checks, and execution evidence for unattended AI agents. The licence is MIT.
Read from SKILL.md and the folder at commit 704824d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DASHCLAW_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Compliance Drift Evals loads about 1.8k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 263 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ucsandman/DashClaw at commit 704824d, republished under its MIT licence (© ucsandman). 263 words, ~1,838 tokens.
.claude/skills/compliance-drift-evals/SKILL.md (or your agent's skills folder).DashClaw's analytical capabilities for governance evidence, behavioral monitoring, and agent quality tracking.
Generate audit-ready evidence bundles for regulatory frameworks.
| Framework | ID | Description |
|---|---|---|
| SOC 2 | soc2 | Service Organization Control |
| NIST AI RMF | nist-ai-rmf | AI Risk Management Framework |
| EU AI Act | eu-ai-act | European AI regulation |
| ISO 42001 | iso42001 | AI Management System |
// V1 SDK
const exp = await claw.createComplianceExport({
name: 'Q1 2026 SOC 2 Audit',
frameworks: ['soc2'],
format: 'json', // or 'md'
window_days: 90,
include_evidence: true,
include_remediation: true,
include_trends: true
});# API
curl -X POST "$DASHCLAW_BASE_URL/api/compliance/exports" \
-H "x-api-key: $DASHCLAW_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name":"Q1 Audit","frameworks":["soc2"],"window_days":90}'await claw.createComplianceSchedule({
name: 'Weekly SOC 2',
frameworks: ['soc2'],
cron_expression: '0 9 * * 1', // Every Monday at 9am
window_days: 7,
include_evidence: true
});const gaps = await claw.analyzeGaps('soc2');
// Returns: missing controls, partial coverage, recommendationsconst trends = await claw.getComplianceTrends({ framework: 'soc2', limit: 12 });
// Monthly coverage scores over timeStatistical behavioral drift detection using z-scores. Pure math — no LLM required.
| Metric | What It Measures |
|---|---|
risk_score | Are actions getting riskier? |
confidence | Is agent confidence dropping? |
duration_ms | Are actions taking longer? |
cost_estimate | Are costs increasing? |
tokens_total | Is token usage growing? |
learning_score | Is the agent learning? |
| z-score | Severity | Meaning |
|---|---|---|
| ≥ 1.5 | info | Notable deviation |
| ≥ 2.0 | warning | Significant drift |
| ≥ 3.0 | critical | Severe anomaly |
// Establish baseline from last 30 days
await claw.computeDriftBaselines({
agent_id: 'my-agent',
lookback_days: 30
});const drift = await claw.detectDrift({
agent_id: 'my-agent',
window_days: 7
});
// drift.alerts: [{ metric, z_score, severity, current_value, baseline_mean }]await claw.acknowledgeDriftAlert(alertId);const stats = await claw.getDriftStats({ agent_id: 'my-agent' });
// { total_alerts, unacknowledged, by_severity, by_metric }Score agent outputs using 5 built-in scorer types.
| Type | LLM Required | Description |
|---|---|---|
regex | No | Pattern matching against output |
contains | No | Keyword/phrase detection |
numeric_range | No | Value within expected range |
custom_function | No | Arbitrary JavaScript logic |
llm_judge | Yes (optional) | LLM-based quality assessment |
// Regex scorer — check for PII
await claw.createScorer({
name: 'no-pii-in-output',
scorerType: 'regex',
config: {
pattern: '\\b\\d{3}-\\d{2}-\\d{4}\\b', // SSN pattern
invert: true // Score 1 if NOT found (good)
},
description: 'Ensures no SSN patterns in output'
});
// Numeric range scorer
await claw.createScorer({
name: 'response-time-check',
scorerType: 'numeric_range',
config: {
field: 'duration_ms',
min: 0,
max: 5000
}
});await claw.createScore({
actionId: 'ar_abc123',
scorerName: 'no-pii-in-output',
score: 1.0, // 0-1 scale
label: 'pass',
reasoning: 'No PII patterns detected'
});const run = await claw.createEvalRun({
name: 'Weekly quality check',
scorerId: 'sc_abc123',
actionFilters: { days: 7 }
});
// Scores all matching actions from the last 7 daysMulti-dimensional risk and quality scoring with auto-calibration.
await claw.createScoringProfile({
name: 'deploy-quality',
description: 'Quality scoring for deployment actions',
composite_method: 'weighted_average', // or: minimum, geometric_mean
dimensions: [
{
name: 'risk',
weight: 0.4,
source: 'risk_score',
scale: [
{ min: 0, max: 40, label: 'low', score: 1.0 },
{ min: 40, max: 70, label: 'medium', score: 0.6 },
{ min: 70, max: 100, label: 'high', score: 0.2 }
]
},
{
name: 'speed',
weight: 0.3,
source: 'duration_ms',
scale: [
{ min: 0, max: 5000, label: 'fast', score: 1.0 },
{ min: 5000, max: 30000, label: 'normal', score: 0.7 },
{ min: 30000, max: null, label: 'slow', score: 0.3 }
]
},
{
name: 'cost',
weight: 0.3,
source: 'cost_estimate',
scale: [
{ min: 0, max: 1, label: 'cheap', score: 1.0 },
{ min: 1, max: 10, label: 'moderate', score: 0.6 },
{ min: 10, max: null, label: 'expensive', score: 0.2 }
]
}
]
});const suggestions = await claw.autoCalibrate({
lookback_days: 30
});
// Returns percentile-based scale suggestions from historical dataReplace hardcoded risk scores with rule-based computation:
await claw.createRiskTemplate({
name: 'deploy-risk',
base_risk: 50,
rules: [
{ field: 'systems_touched', operator: 'contains', value: 'production', add: 30 },
{ field: 'reversible', operator: '==', value: false, add: 20 },
{ field: 'metadata.has_rollback', operator: '==', value: true, add: -15 }
]
});Track agent improvement over time. DashClaw's unique moat.
| Level | Episodes | Success Rate | Avg Score |
|---|---|---|---|
| Novice | 0+ | any | any |
| Developing | 10+ | 40%+ | 40+ |
| Competent | 50+ | 60%+ | 55+ |
| Proficient | 150+ | 75%+ | 65+ |
| Expert | 500+ | 85%+ | 75+ |
| Master | 1000+ | 92%+ | 85+ |
const velocity = await claw.computeLearningVelocity({
agent_id: 'my-agent',
lookback_days: 90,
period: 'weekly'
});
// Linear regression slope of performance over timeconst curves = await claw.computeLearningCurves({
agent_id: 'my-agent',
lookback_days: 180
});
// Per-action-type learning curves showing improvement trajectoryconst summary = await claw.getLearningAnalyticsSummary({
agent_id: 'my-agent'
});
// { maturity_level, velocity, total_episodes, success_rate, avg_score }© ucsandman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/dashclaw-agent/compliance-drift-evals of ucsandman/DashClaw.
Open the folder on GitHubat commit 704824d
Compliance Drift Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Compliance Drift Evals this skillucsandman/DashClaw | 311 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Agent Eval Casesagentailor/fullstack-langgraph-nextjs-agent | 132 | — | ~5.3k | Automated safety check: Pass | MIT | |
| Agent Harness DesignAnastasiyaW/codex-claude-code-config | 154 | — | ~764 | Automated safety check: Pass | MIT | |
| LLM Eval Pipeline Auditai-evals-course/evals-skills | 1.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| China AI Compliance AuditjnMetaCode/shellward | 140 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Kit SDKmark3labs/kit | 141 | — | ~2.6k | Automated safety check: Pass | MIT |
agentailor/fullstack-langgraph-nextjs-agent
Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.
AnastasiyaW/codex-claude-code-config
Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against…
ai-evals-course/evals-skills
Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.
jnMetaCode/shellward
按中国法规(网安法 / PIPL / 等保2.0 / 数据出境 / AI生成内容标识)审计一个 AI 项目的代码仓库,产出每条都带 文件:行 取证、经独立复核、经脚本校验的合规报告。当用户问「这个项目上线合不合规」「调用了 OpenAI/Claude 算不算数据出境」「要不要做 AI 标识」「帮我做合规自查/等保/PIPL 检查」时使用。Audit an AI project's…
mark3labs/kit
Guide for building Go applications with the Kit SDK. An agent skill from mark3labs/kit.
ZIZKA-AI-SL/ZizkaDB
Set up and start the local ZizkaDB development stack. An agent skill from ZIZKA-AI-SL/ZizkaDB.
ucsandman/DashClaw
Governance behavior for AI agents governed by DashClaw. An agent skill from ucsandman/DashClaw.
ucsandman/DashClaw
The single command that gets a DashClaw change ON MAIN AND LIVE — it resolves everything blocking production, never defers, and never hands back a checklist.
ucsandman/DashClaw
Turn a bug symptom into a structured, reproducible bug report — summary, environment, exact repro steps, actual vs expected, and evidence (logs, error text, failing route/test) — and then optionally…
ucsandman/DashClaw
Governance behavior for Muse agents governed by DashClaw. An agent skill from ucsandman/DashClaw.
ucsandman/DashClaw
Contribute to the DashClaw codebase — architecture, scaffolding, tests, CI
ucsandman/DashClaw
Create and test DashClaw guard policies for agent governance
Works with
Categories
Set up compliance exports, drift detection, evaluations, scoring, and learning analytics. Compliance Drift Evals is an agent skill from ucsandman/DashClaw.
Compliance Drift Evals fits situations like: tasks that involve LLM evaluation; tasks that involve GitOps; tasks that involve AI governance.
Run `npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a claude-code`. Or copy the skill folder (.claude/skills/dashclaw-agent/compliance-drift-evals in ucsandman/DashClaw) into .claude/skills/compliance-drift-evals in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a codex`. Or copy the skill folder (.claude/skills/dashclaw-agent/compliance-drift-evals in ucsandman/DashClaw) into .agents/skills/compliance-drift-evals in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compliance-drift-evals, .gemini/skills/compliance-drift-evals, .github/skills/compliance-drift-evals and .opencode/skills/compliance-drift-evals in your project.
Going by SKILL.md and its folder, Compliance Drift Evals needs the command-line tools its instructions call (curl) and credentials named DASHCLAW_API_KEY. Our summary lists: A credential in DASHCLAW_API_KEY.
SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Compliance Drift Evals is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Compliance Drift Evals: Agent Eval Cases (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Agent Harness Design (AnastasiyaW/codex-claude-code-config, 154 stars), LLM Eval Pipeline Audit (ai-evals-course/evals-skills, 1.5k stars) and China AI Compliance Audit (jnMetaCode/shellward, 140 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ucsandman (a GitHub user) maintains it in ucsandman/DashClaw, which has 311 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 10, 2026.
Source: ucsandman/DashClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.