Agent skill

Compliance Drift Evals

by ucsandman in ucsandman/DashClaw

Set up compliance exports, drift detection, evaluations, scoring, and learning analytics

MITAuto-check passedAI & LLM Engineering

Install Compliance Drift Evals

skills CLI
$ npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ucsandman/DashClaw compliance-drift-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ucsandman/DashClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/dashclaw-agent/compliance-drift-evals .claude/skills/compliance-drift-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
compliance-drift-evals
GitHub stars
311
Token cost
~1.8k tokens
SKILL.md length
263 words
Files
1
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Set up compliance exports, drift detection, evaluations, scoring, and learning analytics

  • Tasks that involve LLM evaluation
  • SKILL.md covers Compliance Exports, Drift Detection, Evaluations and Scoring Profiles, plus 1 more section
  • Calls curl; needs DASHCLAW_API_KEY
  • Tasks that involve GitOps

What it does

Compliance Drift Evals is an agent skill from ucsandman/DashClaw. Set up compliance exports, drift detection, evaluations, scoring, and learning analytics

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation, GitOps and AI governance. It works with Model Context Protocol. The repository describes itself as: Remote approvals, policy checks, and execution evidence for unattended AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve GitOps
  • Tasks that involve AI governance

Example prompts

  • “/compliance-drift-evals”

Requirements

  • A credential in DASHCLAW_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 704824d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DASHCLAW_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Compliance Drift Evals loads about 1.8k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 263 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ucsandman/DashClaw at commit 704824d, republished under its MIT licence (© ucsandman). 263 words, ~1,838 tokens.

Download SKILL.mdSave it as .claude/skills/compliance-drift-evals/SKILL.md (or your agent's skills folder).
name
compliance-drift-evals
description
Set up compliance exports, drift detection, evaluations, scoring, and learning analytics
license
MIT
metadata.author
ucsandman
metadata.version
1.0.0
metadata.category
analytics

Compliance, Drift, Evaluations & Learning

DashClaw's analytical capabilities for governance evidence, behavioral monitoring, and agent quality tracking.


Compliance Exports

Generate audit-ready evidence bundles for regulatory frameworks.

Supported Frameworks
FrameworkIDDescription
SOC 2soc2Service Organization Control
NIST AI RMFnist-ai-rmfAI Risk Management Framework
EU AI Acteu-ai-actEuropean AI regulation
ISO 42001iso42001AI Management System
Create an Export
javascript
// V1 SDK
const exp = await claw.createComplianceExport({
  name: 'Q1 2026 SOC 2 Audit',
  frameworks: ['soc2'],
  format: 'json',        // or 'md'
  window_days: 90,
  include_evidence: true,
  include_remediation: true,
  include_trends: true
});
bash
# API
curl -X POST "$DASHCLAW_BASE_URL/api/compliance/exports" \
  -H "x-api-key: $DASHCLAW_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name":"Q1 Audit","frameworks":["soc2"],"window_days":90}'
Scheduled Exports
javascript
await claw.createComplianceSchedule({
  name: 'Weekly SOC 2',
  frameworks: ['soc2'],
  cron_expression: '0 9 * * 1',  // Every Monday at 9am
  window_days: 7,
  include_evidence: true
});
Gap Analysis
javascript
const gaps = await claw.analyzeGaps('soc2');
// Returns: missing controls, partial coverage, recommendations
javascript
const trends = await claw.getComplianceTrends({ framework: 'soc2', limit: 12 });
// Monthly coverage scores over time

Drift Detection

Statistical behavioral drift detection using z-scores. Pure math — no LLM required.

6 Tracked Metrics
MetricWhat It Measures
risk_scoreAre actions getting riskier?
confidenceIs agent confidence dropping?
duration_msAre actions taking longer?
cost_estimateAre costs increasing?
tokens_totalIs token usage growing?
learning_scoreIs the agent learning?
Severity Thresholds
z-scoreSeverityMeaning
≥ 1.5infoNotable deviation
≥ 2.0warningSignificant drift
≥ 3.0criticalSevere anomaly
Compute Baselines
javascript
// Establish baseline from last 30 days
await claw.computeDriftBaselines({
  agent_id: 'my-agent',
  lookback_days: 30
});
Detect Drift
javascript
const drift = await claw.detectDrift({
  agent_id: 'my-agent',
  window_days: 7
});

// drift.alerts: [{ metric, z_score, severity, current_value, baseline_mean }]
Acknowledge Alerts
javascript
await claw.acknowledgeDriftAlert(alertId);
Monitor Drift Stats
javascript
const stats = await claw.getDriftStats({ agent_id: 'my-agent' });
// { total_alerts, unacknowledged, by_severity, by_metric }

Evaluations

Score agent outputs using 5 built-in scorer types.

Scorer Types
TypeLLM RequiredDescription
regexNoPattern matching against output
containsNoKeyword/phrase detection
numeric_rangeNoValue within expected range
custom_functionNoArbitrary JavaScript logic
llm_judgeYes (optional)LLM-based quality assessment
Create a Scorer
javascript
// Regex scorer — check for PII
await claw.createScorer({
  name: 'no-pii-in-output',
  scorerType: 'regex',
  config: {
    pattern: '\\b\\d{3}-\\d{2}-\\d{4}\\b',  // SSN pattern
    invert: true  // Score 1 if NOT found (good)
  },
  description: 'Ensures no SSN patterns in output'
});

// Numeric range scorer
await claw.createScorer({
  name: 'response-time-check',
  scorerType: 'numeric_range',
  config: {
    field: 'duration_ms',
    min: 0,
    max: 5000
  }
});
Score an Action
javascript
await claw.createScore({
  actionId: 'ar_abc123',
  scorerName: 'no-pii-in-output',
  score: 1.0,        // 0-1 scale
  label: 'pass',
  reasoning: 'No PII patterns detected'
});
Batch Evaluation Run
javascript
const run = await claw.createEvalRun({
  name: 'Weekly quality check',
  scorerId: 'sc_abc123',
  actionFilters: { days: 7 }
});
// Scores all matching actions from the last 7 days

Scoring Profiles

Multi-dimensional risk and quality scoring with auto-calibration.

Create a Profile
javascript
await claw.createScoringProfile({
  name: 'deploy-quality',
  description: 'Quality scoring for deployment actions',
  composite_method: 'weighted_average',  // or: minimum, geometric_mean
  dimensions: [
    {
      name: 'risk',
      weight: 0.4,
      source: 'risk_score',
      scale: [
        { min: 0, max: 40, label: 'low', score: 1.0 },
        { min: 40, max: 70, label: 'medium', score: 0.6 },
        { min: 70, max: 100, label: 'high', score: 0.2 }
      ]
    },
    {
      name: 'speed',
      weight: 0.3,
      source: 'duration_ms',
      scale: [
        { min: 0, max: 5000, label: 'fast', score: 1.0 },
        { min: 5000, max: 30000, label: 'normal', score: 0.7 },
        { min: 30000, max: null, label: 'slow', score: 0.3 }
      ]
    },
    {
      name: 'cost',
      weight: 0.3,
      source: 'cost_estimate',
      scale: [
        { min: 0, max: 1, label: 'cheap', score: 1.0 },
        { min: 1, max: 10, label: 'moderate', score: 0.6 },
        { min: 10, max: null, label: 'expensive', score: 0.2 }
      ]
    }
  ]
});
Auto-Calibration
javascript
const suggestions = await claw.autoCalibrate({
  lookback_days: 30
});
// Returns percentile-based scale suggestions from historical data
Risk Templates

Replace hardcoded risk scores with rule-based computation:

javascript
await claw.createRiskTemplate({
  name: 'deploy-risk',
  base_risk: 50,
  rules: [
    { field: 'systems_touched', operator: 'contains', value: 'production', add: 30 },
    { field: 'reversible', operator: '==', value: false, add: 20 },
    { field: 'metadata.has_rollback', operator: '==', value: true, add: -15 }
  ]
});

Learning Analytics

Track agent improvement over time. DashClaw's unique moat.

Maturity Model
LevelEpisodesSuccess RateAvg Score
Novice0+anyany
Developing10+40%+40+
Competent50+60%+55+
Proficient150+75%+65+
Expert500+85%+75+
Master1000+92%+85+
Compute Learning Velocity
javascript
const velocity = await claw.computeLearningVelocity({
  agent_id: 'my-agent',
  lookback_days: 90,
  period: 'weekly'
});
// Linear regression slope of performance over time
Learning Curves
javascript
const curves = await claw.computeLearningCurves({
  agent_id: 'my-agent',
  lookback_days: 180
});
// Per-action-type learning curves showing improvement trajectory
Analytics Summary
javascript
const summary = await claw.getLearningAnalyticsSummary({
  agent_id: 'my-agent'
});
// { maturity_level, velocity, total_episodes, success_rate, avg_score }

© ucsandman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/dashclaw-agent/compliance-drift-evals of ucsandman/DashClaw.

Open the folder on GitHubat commit 704824d

Compare with similar skills

Compliance Drift Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Compliance Drift Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Compliance Drift Evals this skillucsandman/DashClaw311—~1.8kAutomated safety check: PassMIT
Agent Eval Casesagentailor/fullstack-langgraph-nextjs-agent132—~5.3kAutomated safety check: PassMIT
Agent Harness DesignAnastasiyaW/codex-claude-code-config154—~764Automated safety check: PassMIT
LLM Eval Pipeline Auditai-evals-course/evals-skills1.5k—~2.5kAutomated safety check: PassApache-2.0
China AI Compliance AuditjnMetaCode/shellward140—~1.1kAutomated safety check: PassApache-2.0
Kit SDKmark3labs/kit141—~2.6kAutomated safety check: PassMIT

Similar skills

  • Agent Eval Cases

    agentailor/fullstack-langgraph-nextjs-agent

    Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

    132 GitHub stars~5.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agent Harness Design

    AnastasiyaW/codex-claude-code-config

    Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against…

    154 GitHub stars~764 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Eval Pipeline Audit

    ai-evals-course/evals-skills

    Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

    1.5k GitHub stars~2.5k tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed
  • China AI Compliance Audit

    jnMetaCode/shellward

    按中国法规(网安法 / PIPL / 等保2.0 / 数据出境 / AI生成内容标识)审计一个 AI 项目的代码仓库,产出每条都带 文件:行 取证、经独立复核、经脚本校验的合规报告。当用户问「这个项目上线合不合规」「调用了 OpenAI/Claude 算不算数据出境」「要不要做 AI 标识」「帮我做合规自查/等保/PIPL 检查」时使用。Audit an AI project's…

    140 GitHub stars~1.1k tokensUpdated 12 days ago
    SecurityAuto-check passed
  • Kit SDK

    mark3labs/kit

    Guide for building Go applications with the Kit SDK. An agent skill from mark3labs/kit.

    141 GitHub stars~2.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Zizkadb Dev Setup

    ZIZKA-AI-SL/ZizkaDB

    Set up and start the local ZizkaDB development stack. An agent skill from ZIZKA-AI-SL/ZizkaDB.

    129 GitHub stars~535 tokensUpdated today
    DevelopmentAuto-check: notes

More from ucsandman/DashClaw

All 13 skills in this repo
  • Dashclaw Governance

    ucsandman/DashClaw

    Governance behavior for AI agents governed by DashClaw. An agent skill from ucsandman/DashClaw.

    311 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Dashclaw Ship

    ucsandman/DashClaw

    The single command that gets a DashClaw change ON MAIN AND LIVE — it resolves everything blocking production, never defers, and never hands back a checklist.

    311 GitHub stars~7.2k tokensUpdated yesterday
    Auto-check passed
  • Repro

    ucsandman/DashClaw

    Turn a bug symptom into a structured, reproducible bug report — summary, environment, exact repro steps, actual vs expected, and evidence (logs, error text, failing route/test) — and then optionally…

    311 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Muse Governance

    ucsandman/DashClaw

    Governance behavior for Muse agents governed by DashClaw. An agent skill from ucsandman/DashClaw.

    311 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Build Dashclaw

    ucsandman/DashClaw

    Contribute to the DashClaw codebase — architecture, scaffolding, tests, CI

    311 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Create Policies

    ucsandman/DashClaw

    Create and test DashClaw guard policies for agent governance

    311 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed

Questions about Compliance Drift Evals

What does Compliance Drift Evals do?

Set up compliance exports, drift detection, evaluations, scoring, and learning analytics. Compliance Drift Evals is an agent skill from ucsandman/DashClaw.

When should I use Compliance Drift Evals?

Compliance Drift Evals fits situations like: tasks that involve LLM evaluation; tasks that involve GitOps; tasks that involve AI governance.

How do I install Compliance Drift Evals in Claude Code?

Run `npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a claude-code`. Or copy the skill folder (.claude/skills/dashclaw-agent/compliance-drift-evals in ucsandman/DashClaw) into .claude/skills/compliance-drift-evals in your project. Claude Code loads it when a task matches its description.

How do I install Compliance Drift Evals in Codex?

Run `npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a codex`. Or copy the skill folder (.claude/skills/dashclaw-agent/compliance-drift-evals in ucsandman/DashClaw) into .agents/skills/compliance-drift-evals in your project. Codex loads it when a task matches its description.

Can I use Compliance Drift Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ucsandman/DashClaw --skill compliance-drift-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compliance-drift-evals, .gemini/skills/compliance-drift-evals, .github/skills/compliance-drift-evals and .opencode/skills/compliance-drift-evals in your project.

What does Compliance Drift Evals need to run?

Going by SKILL.md and its folder, Compliance Drift Evals needs the command-line tools its instructions call (curl) and credentials named DASHCLAW_API_KEY. Our summary lists: A credential in DASHCLAW_API_KEY.

Does Compliance Drift Evals access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Compliance Drift Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Compliance Drift Evals use?

Compliance Drift Evals is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Compliance Drift Evals use?

About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Compliance Drift Evals?

Skills that share tags, products or a category with Compliance Drift Evals: Agent Eval Cases (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Agent Harness Design (AnastasiyaW/codex-claude-code-config, 154 stars), LLM Eval Pipeline Audit (ai-evals-course/evals-skills, 1.5k stars) and China AI Compliance Audit (jnMetaCode/shellward, 140 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Compliance Drift Evals?

ucsandman (a GitHub user) maintains it in ucsandman/DashClaw, which has 311 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 10, 2026.

Source: ucsandman/DashClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.