Phoenix Pxi Playwright
Arize-ai/phoenix
Write, extend, and debug PXI Playwright E2E tests for Phoenix.
This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target.
$ npx skills add jacob-dietle/context-os --skill eval-loop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jacob-dietle/context-os eval-loop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jacob-dietle/context-os.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/eval-loop .claude/skills/eval-loop && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-loop" agent skill from https://github.com/jacob-dietle/context-os/tree/main/.claude/skills/eval-loop into .claude/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jacob-dietle/context-os/tree/main/.claude/skills/eval-loopType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jacob-dietle/context-os --skill eval-loop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jacob-dietle/context-os eval-loop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jacob-dietle/context-os.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/eval-loop .agents/skills/eval-loop && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-loop" agent skill from https://github.com/jacob-dietle/context-os/tree/main/.claude/skills/eval-loop into .agents/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jacob-dietle/context-os --skill eval-loop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jacob-dietle/context-os eval-loop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jacob-dietle/context-os.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/eval-loop .cursor/skills/eval-loop && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-loop" agent skill from https://github.com/jacob-dietle/context-os/tree/main/.claude/skills/eval-loop into .cursor/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jacob-dietle/context-os.git --path .claude/skills/eval-loop--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jacob-dietle/context-os --skill eval-loop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jacob-dietle/context-os eval-loop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jacob-dietle/context-os.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/eval-loop .gemini/skills/eval-loop && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-loop" agent skill from https://github.com/jacob-dietle/context-os/tree/main/.claude/skills/eval-loop into .gemini/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jacob-dietle/context-os eval-loopInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jacob-dietle/context-os --skill eval-loop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jacob-dietle/context-os.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/eval-loop .github/skills/eval-loop && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-loop" agent skill from https://github.com/jacob-dietle/context-os/tree/main/.claude/skills/eval-loop into .github/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jacob-dietle/context-os --skill eval-loop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jacob-dietle/context-os eval-loop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jacob-dietle/context-os.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/eval-loop .opencode/skills/eval-loop && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-loop" agent skill from https://github.com/jacob-dietle/context-os/tree/main/.claude/skills/eval-loop into .opencode/skills/eval-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-loop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-loopThis skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target.
Eval Loop is an agent skill from jacob-dietle/context-os. This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target. Traces symptoms to root causes, sets measurable targets with automated backpressure (unit tests, Playwright, LLM-as-judge, or rubric scoring), and iterates until targets pass. Use when user reports a quality gap ("this is a 3/10"), when shipping a feature that needs a quality bar ("what would 10/10 look like?"), or when a class of problems keeps…
Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `eval/statistical-validity-checks.md`, `references/backpressure-patterns.md` and `references/llm-judge-rubrics.md`).
It sits in Testing & QA, covering Root cause analysis, LLM evaluation and Quizzes and assessments. It works with Playwright. The licence is MIT.
Read from SKILL.md and the folder at commit 1027e3f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown, jsonl and typescript).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Loop loads about 5.2k tokens when it runs, and up to ~7.3k if it reads all its reference files. Until then it costs about 172 tokens; SKILL.md has 1,668 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jacob-dietle/context-os at commit 1027e3f, republished under its MIT licence (© jacob-dietle). 1,668 words, ~5,197 tokens.
.claude/skills/eval-loop/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Generalized quality iteration loop for any dimension — UX, data architecture, code quality, feature completeness, content tone. Traces specific symptoms to structural root causes, defines measurable targets, and iterates with automated backpressure until targets pass.
Meta-Principle: "Don't fix the symptom, fix the class — but verify by measuring the symptom."
Scope-Principle (added): "This loop verifies presence, shape, and content quality. It does NOT verify predictive validity. Scoring/classification problems require statistical gates this skill does not provide — route them to eval-driven-scoring."
Apply this skill when:
Do NOT use for:
/specification-driven-development)/eval-driven-scoring. This is a HARD route, not a suggestion.Relationship to eval-driven-scoring: This skill is the generalized eval loop for deterministic quality (presence, shape, tone, content). eval-driven-scoring is the specialized sibling for predictive quality (classification, ranking). They are NOT interchangeable — predictive problems need holdouts, base rates, and discriminative ratios that this skill does not enforce.
Before Step 1, classify the quality problem. Routing wrong here wastes work and can produce confidently wrong scoring models.
| Type | Example prompts | Route |
|---|---|---|
| Presence / shape | "field X doesn't appear", "link is missing", "button 404s", "schema missing column" | Continue with this skill |
| Content quality | "email tone is off", "copy is unclear", "card title is generic" | Continue with this skill |
| Predictive / scoring | "rule identifies good leads", "model predicts churn", "score ranks ICP accurately", "lead scoring misclassifies" | STOP. Route to /eval-driven-scoring and read references/statistical-validity-checks.md. Do not proceed with this skill. |
| Mixed | Data-shape targets + scoring targets together | Split targets: presence/shape continue here; scoring/predictive targets route to eval-driven-scoring separately |
Backpressure in this skill is assertion-based — "does this link exist", "does this card render", "is the post_url field populated". Those verifications pass when the fix is applied correctly.
Backpressure for predictive problems requires:
None of these are in this skill. If you use this skill's backpressure on a scoring problem, you will produce a model that perfectly "passes" on known positives and fails on reality. This is not hypothetical — it has happened and produced client-visible damage.
| User says | Type | Action |
|---|---|---|
| "Signal cards don't link to HubSpot" | Presence | Continue Step 1 |
| "Cold emails sound robotic" | Content quality | Continue Step 1 |
| "Our lead scoring is wrong on these deals" | Predictive | STOP. Route out. |
| "Build a model to identify good-fit contacts" | Predictive | STOP. Route out. |
| "Scoring model labels are right but UI doesn't show why" | Mixed | Split: UI target continues here; scoring target routes out |
SYMPTOM (specific, observed)
"Lisa can't click the contact to see who they are"
│
▼
GENERALIZE (what class of problem is this?)
"Signal cards lack provenance — user can't trace to source"
│
▼
ROOT CAUSES (first-principle gaps that create this class)
1. Agent doesn't include post URLs in evidence
2. contact_linkedin not mapped to UI component
3. reasoning_trace field exists but UI doesn't render it
4. No signal dedup across runs
│
▼
TARGETS (measurable, with verification method)
T1: "Every signal card has a clickable LinkedIn link" → Playwright assertion
T2: "Reasoning trace visible on expand" → Playwright assertion
T3: "95% of signals include source post URL" → Unit test on pipeline output
T4: "No duplicate accounts across consecutive runs" → Unit test on D1 query
│
▼
FIX (one root cause at a time)
│
▼
VERIFY (run the target's backpressure test)
Pass → mark target complete, move to next
Fail → diagnose, adjust fix, re-verify
│
▼
LOOP until all targets passCollect the specific complaints. Not interpretations — the actual observable problems.
## Symptoms (observed)
1. "I can't click on the contact and see who they are"
2. "There's no way to know if this was previously rejected"
3. "I can't see the LinkedIn post that triggered this signal"
4. "The evidence section just says 'LinkedIn' — that's paper thin"Ask: "Is this the complete list, or are there more?" Surface ALL symptoms before proceeding. Fixing one while ignoring others wastes iterations.
Group symptoms by the structural gap they share.
One symptom can point to multiple classes. Multiple symptoms often share a root.
## Problem Classes
A. PROVENANCE — User cannot trace a signal to its source data
Symptoms: 1, 3, 4
B. CONTACT INTELLIGENCE — User cannot evaluate who the contact is
Symptoms: 1, 2
C. SIGNAL DEDUP — System doesn't track cross-run history
Symptoms: 2Why generalize? Fixing the class fixes current AND future symptoms in that class. Fixing symptoms one-by-one is whack-a-mole.
For each problem class, trace to concrete technical/architectural gaps.
Root causes are things that can be fixed in code, data, prompts, or architecture. They are NOT symptoms restated.
## Root Causes
### Class A: PROVENANCE
A1. Agent prompt doesn't instruct to include LinkedIn post URL in evidence_json
A2. evidence_json schema has no `post_url` field
A3. SignalCard component doesn't render source links
A4. Supabase intelligence_posts VIEW has post_url but pipeline doesn't query it
### Class B: CONTACT INTELLIGENCE
B1. contact_linkedin field exists in D1 but not mapped to UI
B2. No contact card component — just name + email as text
B3. No mailto: link on email
B4. No HubSpot link for the contact (only company)Test each root cause: "If I fixed ONLY this, would it improve the symptom?" If yes, it's a real root cause. If not, dig deeper.
Every root cause fix gets a measurable target. Every target gets a verification method.
The verification method is the backpressure — the thing that tells you the fix actually worked, and catches regressions.
Choose the lightest verification that catches the problem:
| Verification Type | When to Use | Example |
|---|---|---|
| Unit test | Logic, data shape, contracts | expect(signal.evidence.post_url).toBeTruthy() |
| Integration test | API responses, data flow | fetch('/signals/1').then(r => expect(r.contact_linkedin).toMatch(...)) |
| Playwright assertion | UI presence, clickability, layout | expect(page.locator('.contact-linkedin-link')).toBeVisible() |
| Playwright + screenshot | Visual quality, design intent | Screenshot → human review or LLM-as-judge |
| LLM-as-judge | Tone, quality, rubric compliance | "Rate this email 1-10 on the Shakespeare rubric" |
| Rubric scoring | Content quality, brand voice | Structured rubric with weighted dimensions |
| Manual checklist | Last resort for subjective quality | "Does this feel right?" (avoid if possible) |
Rule: If you can write an automated test, write one. If the quality is inherently subjective, use LLM-as-judge with a rubric. Manual checklists are a smell — they don't enable iteration.
Rule (added): Backpressure here verifies this specific instance matches criteria. It does NOT verify the rule generalizes to unseen data. If the target is "rule correctly classifies", Step 0 routing should have sent you elsewhere.
## Targets
### T1: Signal cards link to contact LinkedIn profile
- Root cause: B1
- Backpressure: Playwright
- Test: `expect(page.locator('a[href*="linkedin.com/in/"]').first()).toBeVisible()`
- Pass criteria: Link visible, href matches contact_linkedin from API
### T2: 95% of signals include source post URL
- Root cause: A1, A2
- Backpressure: Unit test
- Test: Run pipeline on 20 accounts → `signals.filter(s => s.evidence.post_url).length / signals.length >= 0.95`
- Pass criteria: ≥95% of generated signals have non-null post_url
### T3: Cold emails sound like Shakespeare
- Root cause: Prompt doesn't reference style guide
- Backpressure: LLM-as-judge
- Rubric: references/shakespeare-rubric.md
- Test: Score 10 emails → average ≥ 8/10
- Pass criteria: Mean score ≥ 8, no individual score < 6Fix one root cause at a time. Run its backpressure test. Keep or revert.
Fix A1 (add post_url to agent prompt)
→ Run T2 (unit test: 95% have post_url)
→ Result: 60% have post_url (FAIL)
→ Diagnose: agent finds post but doesn't always extract URL
→ Adjust: add explicit instruction "always include post_url field"
→ Re-run T2: 98% (PASS) → KEEP
Fix B1 (map contact_linkedin to SignalCard)
→ Run T1 (Playwright: LinkedIn link visible)
→ Result: PASS
→ KEEPIteration Rules:
Following the autoresearch pattern, state lives in two files so a fresh agent can continue:
| File | Purpose |
|---|---|
eval-session.md | Living document: symptoms, problem classes, root causes, targets, iteration log |
eval-results.jsonl | Append-only log: one line per target verification run |
{"timestamp":"2026-03-31T17:00:00Z","target":"T1","method":"playwright","result":"pass","details":"LinkedIn link visible for all 3 signals"}
{"timestamp":"2026-03-31T17:05:00Z","target":"T2","method":"unit_test","result":"fail","details":"12/20 signals have post_url (60%)","iteration":1}
{"timestamp":"2026-03-31T17:30:00Z","target":"T2","method":"unit_test","result":"pass","details":"19/20 signals have post_url (95%)","iteration":2}User: "This is a 3/10, the signal cards are paper thin"
→ Step 0: Classify (presence/content — continue)
→ Step 1: Surface all symptoms (interview user)
→ Step 2-4: Generalize → root causes → targets
→ Step 5: Iterate until targets pass
→ Outcome: "Here's what changed + verification results"User: "We're about to ship the signal dashboard, what would 10/10 look like?"
→ Step 0: Classify (presence/content — continue)
→ Step 1: Define ideal experience (user stories or persona walkthrough)
→ Step 2: Identify gaps between current and ideal
→ Step 3-4: Root causes → targets
→ Step 5: Iterate until targets pass
→ Outcome: "Ship confidence: all N targets passing"User: "Our lead scoring model is wrong on these deals"
→ Step 0: Classify (predictive — STOP)
→ Response: "This is a predictive problem. Routing to eval-driven-scoring.
That skill requires: ground truth provenance audit, holdout split,
discriminative ratio check. Do not attempt here — backpressure in this
skill would produce a model that passes on known positives and fails
on reality."When symptoms span layers (data architecture + frontend + agent prompts), group targets by layer and fix bottom-up:
1. Data layer first (agent prompt, schema, pipeline)
→ Unit test backpressure
2. API layer next (routes, response shapes)
→ Integration test backpressure
3. UI layer last (components, pages)
→ Playwright backpressureWhy bottom-up? UI fixes on broken data are wasted work. Fix the data, then the UI will have something real to show.
Some quality gaps span UX + data + architecture simultaneously. The signal card problem is a good example:
| Dimension | Root Cause | Target | Backpressure |
|---|---|---|---|
| Data | Agent doesn't include post URL | 95% of signals have post_url | Unit test |
| Data | No signal dedup across runs | 0 duplicate accounts in consecutive runs | Unit test |
| API | contact_linkedin not in API response | /signals/:id returns contact_linkedin | Integration test |
| UX | No LinkedIn link on contact name | Link visible and clickable | Playwright |
| UX | No reasoning trace display | "Why this signal?" expandable visible | Playwright |
| UX | No mailto: on email | Email is mailto: link | Playwright |
| Content | Email body too long (88 words vs 50-75 target) | Mean word count ≤ 75 | Unit test on pipeline output |
Fix order: Data → API → UX → Content (bottom-up).
Note: None of the above are predictive targets. They verify shape, presence, and bounded content quality. If the list included "scoring rule correctly identifies ICP," Step 0 would have routed that target to eval-driven-scoring.
Individual symptoms are reactive — you fix what's broken. But product standards are proactive — they define what "done" means for EVERY page before it ships. A product standard is a set of rules that apply uniformly, not per-symptom.
When to apply: After fixing specific symptoms, step back and ask: "Does every page in this product meet the same bar?" If not, the eval loop isn't done — you fixed the symptom but not the class.
A product standard is a checklist derived from the generalized problem classes. Each item is a rule that applies to every page/component, not to a specific instance.
## [Product Name] Quality Standard
### Entity Linking
- Every company name → links to CRM (HubSpot, Salesforce, etc.)
- Every company name → links to internal account detail page
- Every contact name → links to LinkedIn profile (when URL available)
- Every contact name → links to CRM contact page (when ID available)
- Every email → mailto: link
- Every external URL → opens in new tab with rel="noopener"
### Data Provenance
- Every signal/insight → traceable to source data in ≤2 clicks
- Every AI-generated content → "Why?" expandable showing reasoning
- Every score/metric → breakdown visible on hover or expand
### Actionability
- No buttons that 404 or error — hide what's not built yet
- No empty states without explanation ("No data" → "No signals generated yet. Next run: 9am UTC")
- Every card/row has a clear primary action (approve, view detail, open in CRM)
### Consistency
- Same entity type links to the same destination across all pages
- Same visual pattern for the same data type (scores, dates, status badges)
- Nav reflects actual available pages — no links to removed routesStep 0 (before Step 1 of the eval loop): Check if a product standard exists for this product. If yes, run the checklist against every page FIRST. Any failures become automatic targets.
After Step 5 (after fixing specific symptoms): Re-run the product standard checklist against ALL pages. Specific symptoms may be fixed but the standard may reveal the same class of problem on other pages you didn't check.
Backpressure for product standards: Each checklist item maps to a Playwright assertion that can run against every route:
// Product standard: every company name links to HubSpot
for (const route of ["/", "/accounts", "/linkedin", "/jobs"]) {
test(`${route}: company names link to HubSpot`, async ({ page }) => {
await page.goto(baseUrl + route);
const companyLinks = page.locator('a[href*="app.hubspot.com"]');
const count = await companyLinks.count();
if (count > 0) {
const href = await companyLinks.first().getAttribute("href");
expect(href).toMatch(/app\.hubspot\.com\/contacts\/\d+\/company\/\d+/);
}
});
}BAD:
User: "Signal cards don't link to HubSpot"
→ Fix signal cards → ship
→ Accounts list has same problem → unfixed
→ User finds it next week → trust erodes
GOOD:
User: "Signal cards don't link to HubSpot"
→ Generalize: "every entity should link to CRM"
→ Define product standard: entity linking rules
→ Apply to ALL pages before shipping
→ No surprise gaps on other pagesreferences/backpressure-patterns.md — Detailed patterns for each verification typereferences/llm-judge-rubrics.md — How to write effective rubrics for LLM-as-judgeeval-driven-scoring/references/statistical-validity-checks.md — What scoring problems need (you shouldn't — Step 0 routes you out)This skill generalizes patterns from:
eval-driven-scoring — LLM classification eval loops (autoresearch + RLM patterns)© jacob-dietle, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in .claude/skills/eval-loop of jacob-dietle/context-os.
Open the folder on GitHubat commit 1027e3f
Eval Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Loop this skilljacob-dietle/context-os | 111 | — | ~5.2k | Automated safety check: Pass | MIT | |
| Phoenix Pxi PlaywrightArize-ai/phoenix | 12k | — | ~2.6k | Automated safety check: Pass | Custom licence | |
| Fix Failing Playwright Specappsmithorg/appsmith | 41k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Debug Playwright Prowquay/quay | 2.8k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Control UI E2Eopenclaw/openclaw | 392k | — | ~2.9k | Automated safety check: Pass | MIT | |
| Quay Prow Triagequay/quay | 2.8k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 |
Arize-ai/phoenix
Write, extend, and debug PXI Playwright E2E tests for Phoenix.
appsmithorg/appsmith
Fixes failing Playwright specs by reading the error, classifying the cause in the test code and applying corrections that follow project conventions.
quay/quay
Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…
openclaw/openclaw
A skill your agent uses when designing, testing, fixing, or extending the OpenClaw Control UI GUI, including UI stress-test galleries with feedback inputs, Vitest + Playwright end-to-end checks…
quay/quay
Diagnose any Quay Prow job failure end to end: prowjob.json - top-level build log - JUnit - resolved failing step - Playwright results.json when the failing step is Playwright, continuing through…
TriliumNext/Trilium
Testing CKEditor 5 plugins in the Trilium monorepo. An agent skill from TriliumNext/Trilium.
jacob-dietle/context-os
A skill your agent uses when deciding what to work on next, when progress is stuck, or when the reflex is to build or automate before proving the current bottleneck.
jacob-dietle/context-os
This skill should be used to periodically defragment a multi-app/multi-service codebase — both CODE (duplicate deploy targets, colliding bindings, stale forks) and CONTEXT (parallel spec…
jacob-dietle/context-os
This skill should be used when producing content (newsletter posts, blog posts, LinkedIn posts) from existing corpus material.
jacob-dietle/context-os
This skill should be used when users ask about their work context, what they're working on, recent activity, file relationships, or knowledge graph structure.
jacob-dietle/context-os
This skill should be used when making architectural decisions, writing specs, or reviewing decisions that contain "future work", "v2", "simpler for now", "out of scope", or complexity claims.
jacob-dietle/context-os
This skill should be used when decomposing a spec into a multi-agent implementation plan with dependency ordering, parallelism decisions, contract testing, and verification strategy.
Works with
This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target. Eval Loop is an agent skill from jacob-dietle/context-os. This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target.
Eval Loop fits situations like: user reports a quality gap (this is a 3/10); shipping a feature that needs a quality bar (what would 10/10 look like?); A class of problems keeps recurring.
Run `npx skills add jacob-dietle/context-os --skill eval-loop -a claude-code`. Or copy the skill folder (.claude/skills/eval-loop in jacob-dietle/context-os) into .claude/skills/eval-loop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jacob-dietle/context-os --skill eval-loop -a codex`. Or copy the skill folder (.claude/skills/eval-loop in jacob-dietle/context-os) into .agents/skills/eval-loop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jacob-dietle/context-os --skill eval-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-loop, .gemini/skills/eval-loop, .github/skills/eval-loop and .opencode/skills/eval-loop in your project.
SKILL.md names no scripts, command-line tools or credentials: Eval Loop is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Eval Loop: Phoenix Pxi Playwright (Arize-ai/phoenix, 12k stars), Fix Failing Playwright Spec (appsmithorg/appsmith, 41k stars), Debug Playwright Prow (quay/quay, 2.8k stars) and Control UI E2E (openclaw/openclaw, 392k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jacob-dietle (a GitHub user) maintains it in jacob-dietle/context-os, which has 111 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on August 13, 2026.
Source: jacob-dietle/context-os on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.