Analyzing Offline Evaluations
PostHog/posthog
Inspect and compare offline AI evaluation experiments, diagnose case-level regressions, and follow scorer history across application, model, or prompt changes.
Downloads and analyzes the latest three distinct nightly evaluation runs (Claude Code, Codex CLI, and Jetski CLI) from the GCS remote dashboard to identify and flag unhealthy or low-performing tasks…
$ npx skills add GoogleChrome/modern-web-guidance-src --skill nightly-eval-investigation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GoogleChrome/modern-web-guidance-src nightly-eval-investigation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GoogleChrome/modern-web-guidance-src.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/nightly-eval-investigation .claude/skills/nightly-eval-investigation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nightly-eval-investigation" agent skill from https://github.com/GoogleChrome/modern-web-guidance-src/tree/main/.agents/skills/nightly-eval-investigation into .claude/skills/nightly-eval-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nightly-eval-investigation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GoogleChrome/modern-web-guidance-src/tree/main/.agents/skills/nightly-eval-investigationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GoogleChrome/modern-web-guidance-src --skill nightly-eval-investigation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GoogleChrome/modern-web-guidance-src nightly-eval-investigation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GoogleChrome/modern-web-guidance-src.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/nightly-eval-investigation .agents/skills/nightly-eval-investigation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nightly-eval-investigation" agent skill from https://github.com/GoogleChrome/modern-web-guidance-src/tree/main/.agents/skills/nightly-eval-investigation into .agents/skills/nightly-eval-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nightly-eval-investigation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GoogleChrome/modern-web-guidance-src --skill nightly-eval-investigation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GoogleChrome/modern-web-guidance-src nightly-eval-investigation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GoogleChrome/modern-web-guidance-src.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/nightly-eval-investigation .cursor/skills/nightly-eval-investigation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nightly-eval-investigation" agent skill from https://github.com/GoogleChrome/modern-web-guidance-src/tree/main/.agents/skills/nightly-eval-investigation into .cursor/skills/nightly-eval-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nightly-eval-investigation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GoogleChrome/modern-web-guidance-src.git --path .agents/skills/nightly-eval-investigation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GoogleChrome/modern-web-guidance-src --skill nightly-eval-investigation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GoogleChrome/modern-web-guidance-src nightly-eval-investigation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GoogleChrome/modern-web-guidance-src.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/nightly-eval-investigation .gemini/skills/nightly-eval-investigation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nightly-eval-investigation" agent skill from https://github.com/GoogleChrome/modern-web-guidance-src/tree/main/.agents/skills/nightly-eval-investigation into .gemini/skills/nightly-eval-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nightly-eval-investigation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GoogleChrome/modern-web-guidance-src nightly-eval-investigationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GoogleChrome/modern-web-guidance-src --skill nightly-eval-investigation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GoogleChrome/modern-web-guidance-src.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/nightly-eval-investigation .github/skills/nightly-eval-investigation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nightly-eval-investigation" agent skill from https://github.com/GoogleChrome/modern-web-guidance-src/tree/main/.agents/skills/nightly-eval-investigation into .github/skills/nightly-eval-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nightly-eval-investigation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GoogleChrome/modern-web-guidance-src --skill nightly-eval-investigation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GoogleChrome/modern-web-guidance-src nightly-eval-investigation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GoogleChrome/modern-web-guidance-src.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/nightly-eval-investigation .opencode/skills/nightly-eval-investigation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nightly-eval-investigation" agent skill from https://github.com/GoogleChrome/modern-web-guidance-src/tree/main/.agents/skills/nightly-eval-investigation into .opencode/skills/nightly-eval-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nightly-eval-investigation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nightly-eval-investigationDownloads and analyzes the latest three distinct nightly evaluation runs (Claude Code, Codex CLI, and Jetski CLI) from the GCS remote dashboard to identify and flag unhealthy or low-performing tasks…
Nightly Eval Investigation is an agent skill from GoogleChrome/modern-web-guidance-src. Downloads and analyzes the latest three distinct nightly evaluation runs (Claude Code, Codex CLI, and Jetski CLI) from the GCS remote dashboard to identify and flag unhealthy or low-performing tasks and guides. Use this skill whenever you need to run a bulk investigation on remote nightly runs, track agent health, or identify over-prescribed/brittle guides.
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/investigate.ts` and `scripts/publish_report.ts`).
The licence is Apache-2.0.
10 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0d002f0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (TypeScript), which the agent can run.
Shell commands in SKILL.md call:
gcloudnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gcloud, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nightly Eval Investigation loads about 3.2k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 1,443 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from GoogleChrome/modern-web-guidance-src at commit 0d002f0, republished under its Apache-2.0 licence (© GoogleChrome). 1,443 words, ~3,239 tokens.
.claude/skills/nightly-eval-investigation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.This skill automates the retrieval and multi-agent comparison of remote nightly evaluation runs from Google Cloud Storage (GCS) to diagnose system-wide guidance health, task drift, and over-prescribed guides.
[!IMPORTANT] CRITICAL CONSTRAINT - READ-ONLY INVESTIGATION ONLY This skill is strictly for diagnostics, investigation, and suggesting recommendations. The agent MUST NOT under any circumstances edit, modify, create, delete, or touch any files under the
guides/directory (including task files, guides, graders, expectations, or demos). All suggestions for fixes must be documented exclusively in the generated investigation report markdown file under the "Actionable Recommendations" section. No automated or manual code remediation should be performed.
All nightly evaluation suites are automatically uploaded to Google Cloud Storage:
gs://guidance-evals/nightly-YYYY-MM-DD_HH-MM-SS-[agent_type]/ (note: legacy folders may contain an optional trailing -[ldap] suffix)We track three core agent implementations in our periodic evaluations:
claude_code or claude)codex_cli or codex)jetski_cli or agy or jetski)For every task evaluated in the runs, the investigation tracks four specific metrics across the 3 runs:
MISSING_EXPECTED_GUIDE)guidesUsed array.tasks/task.md is failing to force discovery/retrieval of the correct reference guide.TOO_MANY_GUIDES_CONSUMED)LOW_GUIDED_PASS_RATE)HIGH_UNGUIDED_PASS_RATE)To perform a nightly evaluation investigation, follow these steps:
gcloud storage ls gs://guidance-evals/nightly-.claude_code or claudecodex_cli or codexjetski_cli, agy, or jetskiresults/suites/<folder_name>trace.zip and screenshots) to save bandwidth and disk space:gcloud storage rsync gs://guidance-evals/<folder_name> results/suites/<folder_name> --recursive --exclude ".*\.zip$|.*\.png$"Run the automated TypeScript analysis script:
node .agents/skills/nightly-eval-investigation/scripts/investigate.tsThis script will automatically cross-examine the results, flag unhealthy tasks, and output the report and context helper artifacts to the output directory:
out/nightly-investigation/nightly_investigation_report.mdout/nightly-investigation/flagged_tasks_context.json (contains extracted prompts, guide descriptions, and test headers for all flagged tasks to assist in diagnostics)For each task listed in the Flagged Tasks Table of nightly_investigation_report.md, you must perform a qualitative deep-dive to locate the root cause and recommend actionable fixes.
[!IMPORTANT] CRITICAL COMPLIANCE RULES:
- Do Not Modify Source Files: This is a passive/diagnostic investigation. Do NOT touch, edit, or modify any files under
guides/(such as task prompts, guide markdown files, or grader TypeScript files). All diagnostic findings and recommendations must be written only inside the markdown report.- Complete All Flagged Tasks: You MUST qualitatively investigate and fully populate the summary for every single flagged task. Leaving placeholder text or skipping any flagged task is unacceptable.
- Strict Heading Omission: Under
Diagnostic Details, only include headings/bullet points for the specific flags that were triggered for the task. Omit any non-triggered flags completely from the markdown file. Do not write "N/A" or "No issues".- Long-Running, Thorough Investigation: The qualitative investigation phase is a long-running, intensive task. You MUST thoroughly examine each flagged guide/test one by one, inspecting the target prompt (
task.md), reference guide (guide.md), and grading code (grader.ts) to build a high-quality, task-specific diagnostic summary and actionable recommendation. Do not rush, skip steps, or bundle tasks.
Follow this linear playbook for each flagged task:
Locate Source Files
Find the task's directory under the local workspace: guides/[category]/[use-case]/
tasks/task.mdguide.mdgrader.ts and expectations.mdReview Extracted Failed Assertions
The script investigate.ts automatically extracts and intersects the failed assertions across all three agents, writing them directly into the report.
LOW_GUIDED_PASS_RATE flag. For all other flags, it is completely omitted.Map Flags to Diagnostic Explanations Analyze the source files based on the specific flags that were triggered, and write a diagnostic summary explaining why the task was flagged under each active flag.
[!IMPORTANT] OMISSION RULE: Only include bullets for the specific flags that were triggered for the task. Omit any flags that were not triggered (do not write "N/A" or "No issues").
HIGH_UNGUIDED_PASS_RATE: Explain why unguided runs passed. Inspect tasks/task.md to see if the prompt is overly prescriptive (explicitly mentioning CSS attributes or API details that act as giveaways). Inspect grader.ts to check for loose/vacuous assertions.LOW_GUIDED_PASS_RATE: Explain why guided runs failed. Inspect guide.md for bugs, outdated modules, or incorrect syntax. Inspect grader.ts for calibration drift or rigid checks. Check for browser/emulation issues in the sandbox.MISSING_EXPECTED_GUIDE: Explain why the expected guide was missed. Inspect the task prompt in tasks/task.md to see if it lacks search keywords, or the guide metadata/synonyms in guide.md.TOO_MANY_GUIDES_CONSUMED: Explain why multiple guides (3 or more) were consumed. Check for overly broad prompts that cause query sprawl, or overlapping guide descriptions.Draft Specific, Actionable Recommendations Determine which recommendations to include based on the diagnostics mapped in previous steps. All recommendations must be directly justified by the findings from your investigation of the prompts, guides, and graders. Vague recommendations like "Fix prompt" or "Fix guide" are not acceptable; you must propose exact, concrete changes.
[!IMPORTANT] OMISSION RULE: Only include recommendation lines for components that actually require changes. If a component does not require updates (e.g., the grader is correct as-is), you MUST OMIT the corresponding recommendation line completely from the markdown file (do not write "Keep as is" or "No changes").
Synthesize the Markdown Report
Write the summary under the task's section in out/nightly-investigation/nightly_investigation_report.md following the template below.
Once the qualitative diagnostics and recommendations have been fully written and saved to the markdown report, present the final report to the user and inform them that they can run the automated publisher script to create the GitHub parent issue and the task-level engineering/devrel subissues.
[!WARNING] DO NOT EXECUTE THE PUBLISHER SCRIPT YOURSELF: The agent must never run the publisher script (
publish_report.ts) or create/publish any GitHub issues itself. It must only inform the user of the command so they can verify the report first and execute it manually.
node .agents/skills/nightly-eval-investigation/scripts/publish_report.tsPresent the final report to the user, providing a clickable link, and remind them of the publishing command:
All investigation reports MUST strictly follow these templates to ensure downstream compatibility:
out/nightly-investigation/nightly_investigation_report.md)# Nightly Evaluation Investigation Report
**Generated:** YYYY-MM-DD
**Suites Investigated:**
- **claude_code**: `[folder_name_1]`
- **codex_cli**: `[folder_name_2]`
- **jetski_cli**: `[folder_name_3]`
---
## Summary of Flagged Tasks
- **Total distinct tasks analyzed:** [total_count]
- **Total flagged tasks:** [flagged_count]
### Flagged Tasks Table
| Task / Guide | Flags | Details |
| :--- | :--- | :--- |
| `[task_name]` | `[FLAG_1]`, `[FLAG_2]` | [Detail explanation of flags] |
## Flagged Tasks Details
### `[task_name]`
#### Flags:
- **`[FLAG_1]`**: [Detail explanation]
#### Run Details:
| Agent | Guided Pass Rate | Unguided Pass Rate | Guides Consumed |
| :--- | :---: | :---: | :--- |
| **claude_code** | [guided_rate]% | [unguided_rate]% | `[guide_1]`, `[guide_2]` |
| **codex_cli** | [guided_rate]% | [unguided_rate]% | `[guide_1]` |
| **jetski_cli** | [guided_rate]% | [unguided_rate]% | *None* |
#### Qualitative Diagnostic Summary:
- **Diagnostic Details**:
- *Failed Assertions*:
- "[Exact assertion error message string]" (Test ID: `[test_id]`)
- **[FLAG_1]**: TODO
- **[FLAG_2]**: TODO
- **Actionable Recommendations**:
- [ ] **Prompt**: TODO
- [ ] **Guide**: TODO
- [ ] **Grader**: TODO
---© GoogleChrome, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in .agents/skills/nightly-eval-investigation of GoogleChrome/modern-web-guidance-src.
Open the folder on GitHubat commit 0d002f0
Nightly Eval Investigation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nightly Eval Investigation this skillGoogleChrome/modern-web-guidance-src | 1.1k | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Analyzing Offline EvaluationsPostHog/posthog | 40k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Eval-Driven Development Harnessaffaan-m/ECC | 276k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Evalalirezarezvani/claude-skills | 28k | — | ~618 | Automated safety check: Pass | MIT | |
| Eval Harnessaffaan-m/ECC | 276k | — | ~2.2k | Automated safety check: Pass | MIT | |
| Analyzing Email Headers For Phishing Investigationmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~3.2k | Automated safety check: Pass | Apache-2.0 |
PostHog/posthog
Inspect and compare offline AI evaluation experiments, diagnose case-level regressions, and follow scorer history across application, model, or prompt changes.
affaan-m/ECC
Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.
alirezarezvani/claude-skills
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
affaan-m/ECC
Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…
mukul975/Anthropic-Cybersecurity-Skills
Parse and analyze email headers (Received chain, Return-Path, Message-ID) to trace the true origin of a phishing email and validate SPF, DKIM, and DMARC results to confirm or rule out sender spoofing.
affaan-m/ECC
Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi
GoogleChrome/modern-web-guidance-src
Run a document coherence, link integrity, and git repository status audit across repository markdown files using a dedicated subagent.
GoogleChrome/modern-web-guidance-src
Build and publish Chrome Extensions using Manifest V3 best practices.
GoogleChrome/modern-web-guidance-src
Coding style, architectural conventions, and PR review standards for the modern-web-guidance-src (guidance) repository.
GoogleChrome/modern-web-guidance-src
Workflow for refactoring discipline-level guides (e.g., JavaScript, CSS) to remove "Common Knowledge" by generating and comparing against model-specific "Knowledge Mirrors".
GoogleChrome/modern-web-guidance-src
Best practices for creating expectations and grader files to evaluate guidance quality.
GoogleChrome/modern-web-guidance-src
Best practices for authoring guidance. An agent skill from GoogleChrome/modern-web-guidance-src.
Downloads and analyzes the latest three distinct nightly evaluation runs (Claude Code, Codex CLI, and Jetski CLI) from the GCS remote dashboard to identify and flag unhealthy or low-performing tasks…. Nightly Eval Investigation is an agent skill from GoogleChrome/modern-web-guidance-src. Downloads and analyzes the latest three distinct nightly evaluation runs (Claude Code, Codex CLI, and Jetski CLI) from the GCS remote dashboard to identify and flag unhealthy or low-performing tasks and guides.
Nightly Eval Investigation fits situations like: you need to run a bulk investigation on remote nightly runs; track agent health; identify over-prescribed/brittle guides.
Run `npx skills add GoogleChrome/modern-web-guidance-src --skill nightly-eval-investigation -a claude-code`. Or copy the skill folder (.agents/skills/nightly-eval-investigation in GoogleChrome/modern-web-guidance-src) into .claude/skills/nightly-eval-investigation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GoogleChrome/modern-web-guidance-src --skill nightly-eval-investigation -a codex`. Or copy the skill folder (.agents/skills/nightly-eval-investigation in GoogleChrome/modern-web-guidance-src) into .agents/skills/nightly-eval-investigation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GoogleChrome/modern-web-guidance-src --skill nightly-eval-investigation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nightly-eval-investigation, .gemini/skills/nightly-eval-investigation, .github/skills/nightly-eval-investigation and .opencode/skills/nightly-eval-investigation in your project.
Going by SKILL.md and its folder, Nightly Eval Investigation needs TypeScript for the scripts in its folder and the command-line tools its instructions call (gcloud and node). Our summary lists: Node.js.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Nightly Eval Investigation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nightly Eval Investigation: Analyzing Offline Evaluations (PostHog/posthog, 40k stars), Eval-Driven Development Harness (affaan-m/ECC, 276k stars), Eval (alirezarezvani/claude-skills, 28k stars) and Eval Harness (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GoogleChrome (a GitHub organization) maintains it in GoogleChrome/modern-web-guidance-src, which has 1,143 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 10, 2026.
Source: GoogleChrome/modern-web-guidance-src on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.