Exploring LLM Evaluations
PostHog/posthog
Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment).
A skill your agent uses when exploring the ai-agent-skills catalog to find, compare, and evaluate skills before installing.
$ npx skills add MoizIbnYousaf/ai-agent-skills --skill browse-and-evaluate -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install MoizIbnYousaf/ai-agent-skills browse-and-evaluate --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/MoizIbnYousaf/ai-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/browse-and-evaluate .claude/skills/browse-and-evaluate && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "browse-and-evaluate" agent skill from https://github.com/MoizIbnYousaf/ai-agent-skills/tree/main/skills/browse-and-evaluate into .claude/skills/browse-and-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browse-and-evaluate", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/MoizIbnYousaf/ai-agent-skills/tree/main/skills/browse-and-evaluateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add MoizIbnYousaf/ai-agent-skills --skill browse-and-evaluate -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install MoizIbnYousaf/ai-agent-skills browse-and-evaluate --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/MoizIbnYousaf/ai-agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/browse-and-evaluate .agents/skills/browse-and-evaluate && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "browse-and-evaluate" agent skill from https://github.com/MoizIbnYousaf/ai-agent-skills/tree/main/skills/browse-and-evaluate into .agents/skills/browse-and-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browse-and-evaluate", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add MoizIbnYousaf/ai-agent-skills --skill browse-and-evaluate -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install MoizIbnYousaf/ai-agent-skills browse-and-evaluate --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/MoizIbnYousaf/ai-agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/browse-and-evaluate .cursor/skills/browse-and-evaluate && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "browse-and-evaluate" agent skill from https://github.com/MoizIbnYousaf/ai-agent-skills/tree/main/skills/browse-and-evaluate into .cursor/skills/browse-and-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browse-and-evaluate", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/MoizIbnYousaf/ai-agent-skills.git --path skills/browse-and-evaluate--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add MoizIbnYousaf/ai-agent-skills --skill browse-and-evaluate -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install MoizIbnYousaf/ai-agent-skills browse-and-evaluate --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/MoizIbnYousaf/ai-agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/browse-and-evaluate .gemini/skills/browse-and-evaluate && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "browse-and-evaluate" agent skill from https://github.com/MoizIbnYousaf/ai-agent-skills/tree/main/skills/browse-and-evaluate into .gemini/skills/browse-and-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browse-and-evaluate", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install MoizIbnYousaf/ai-agent-skills browse-and-evaluateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add MoizIbnYousaf/ai-agent-skills --skill browse-and-evaluate -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/MoizIbnYousaf/ai-agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/browse-and-evaluate .github/skills/browse-and-evaluate && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "browse-and-evaluate" agent skill from https://github.com/MoizIbnYousaf/ai-agent-skills/tree/main/skills/browse-and-evaluate into .github/skills/browse-and-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browse-and-evaluate", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add MoizIbnYousaf/ai-agent-skills --skill browse-and-evaluate -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install MoizIbnYousaf/ai-agent-skills browse-and-evaluate --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/MoizIbnYousaf/ai-agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/browse-and-evaluate .opencode/skills/browse-and-evaluate && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "browse-and-evaluate" agent skill from https://github.com/MoizIbnYousaf/ai-agent-skills/tree/main/skills/browse-and-evaluate into .opencode/skills/browse-and-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "browse-and-evaluate", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
browse-and-evaluateA skill your agent uses when exploring the ai-agent-skills catalog to find, compare, and evaluate skills before installing.
Browse And Evaluate is an agent skill from MoizIbnYousaf/ai-agent-skills. Use when exploring the ai-agent-skills catalog to find, compare, and evaluate skills before installing. Always use --fields to limit output size and --dry-run before committing to an install.
Its SKILL.md is about 450 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Universal skill installer and package manager for AI coding agents. One command, 12+ runtimes. npx ai-agent-skills. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 6d95c78. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Browse And Evaluate loads about 453 tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 160 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from MoizIbnYousaf/ai-agent-skills at commit 6d95c78, republished under its MIT licence (© MoizIbnYousaf). 160 words, ~453 tokens.
.claude/skills/browse-and-evaluate/SKILL.md (or your agent's skills folder).Find the right skill for a task without flooding the context window or installing blindly.
--fields on list/search/info to keep output small. Default: --fields name,tier,workArea,description.--dry-run before installing anything.--format json in non-interactive pipelines. The CLI defaults to JSON when stdout is not a TTY.--limit when browsing large catalogs. Start with --limit 10.npx ai-agent-skills search <query> --fields name,tier,workArea,description --limit 10npx ai-agent-skills info <skill-name> --fields name,description,tags,collections,installCommandsnpx ai-agent-skills preview <skill-name>npx ai-agent-skills install <skill-name> --dry-runnpx ai-agent-skills install <skill-name>preview command sanitizes skill content to strip prompt injection patterns. If content looks truncated, check if suspicious patterns were removed.--list or --dry-run a collection before installing.--dry-run to verify the source is reachable.© MoizIbnYousaf, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/browse-and-evaluate of MoizIbnYousaf/ai-agent-skills.
Open the folder on GitHubat commit 6d95c78
Browse And Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Browse And Evaluate this skillMoizIbnYousaf/ai-agent-skills | 1.1k | — | ~453 | Automated safety check: Pass | MIT | |
| Exploring LLM EvaluationsPostHog/posthog | 40k | — | ~5.7k | Automated safety check: Pass | Custom licence | |
| Explorersupabase/supabase | 111k | — | ~845 | Automated safety check: Pass | Apache-2.0 | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 1 repos | ~8.1k | Automated safety check: Notes | MIT | |
| LLM Evaluationdavila7/claude-code-templates | 33k | 12 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Agent Evaluation Reportingsickn33/agentic-awesome-skills | 47k | 1 repos | ~2.1k | Automated safety check: Pass | MIT |
PostHog/posthog
Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment).
supabase/supabase
Build and modify Studio Explorer surfaces, including notebooks, chats, SQL snippets, query cells, and their shared toolbar patterns.
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
sickn33/agentic-awesome-skills
A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.
sickn33/agentic-awesome-skills
Evaluate agent behavior with versioned cases and explicit verifiers.
MoizIbnYousaf/ai-agent-skills
Database schema design, optimization, and migration patterns for PostgreSQL, MySQL, and NoSQL databases.
MoizIbnYousaf/ai-agent-skills
Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration.
MoizIbnYousaf/ai-agent-skills
A skill your agent uses when checking the overall health of a skills library.
MoizIbnYousaf/ai-agent-skills
Backend API design, database architecture, microservices patterns, and test-driven development.
MoizIbnYousaf/ai-agent-skills
Writing effective code documentation - API docs, README files, inline comments, and technical guides.
MoizIbnYousaf/ai-agent-skills
A skill your agent uses when building a managed team skills library for a real stack.
A skill your agent uses when exploring the ai-agent-skills catalog to find, compare, and evaluate skills before installing. Browse And Evaluate is an agent skill from MoizIbnYousaf/ai-agent-skills. Use when exploring the ai-agent-skills catalog to find, compare, and evaluate skills before installing.
Browse And Evaluate fits situations like: exploring the ai-agent-skills catalog to find; evaluate skills before installing.
Run `npx skills add MoizIbnYousaf/ai-agent-skills --skill browse-and-evaluate -a claude-code`. Or copy the skill folder (skills/browse-and-evaluate in MoizIbnYousaf/ai-agent-skills) into .claude/skills/browse-and-evaluate in your project. Claude Code loads it when a task matches its description.
Run `npx skills add MoizIbnYousaf/ai-agent-skills --skill browse-and-evaluate -a codex`. Or copy the skill folder (skills/browse-and-evaluate in MoizIbnYousaf/ai-agent-skills) into .agents/skills/browse-and-evaluate in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MoizIbnYousaf/ai-agent-skills --skill browse-and-evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/browse-and-evaluate, .gemini/skills/browse-and-evaluate, .github/skills/browse-and-evaluate and .opencode/skills/browse-and-evaluate in your project.
Going by SKILL.md and its folder, Browse And Evaluate needs the command-line tools its instructions call (npx). Our summary lists: Node.js.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Browse And Evaluate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 453 tokens (SKILL.md is roughly 1.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Browse And Evaluate: Exploring LLM Evaluations (PostHog/posthog, 40k stars), Explorer (supabase/supabase, 111k stars), Arize Evaluator (github/awesome-copilot, 40k stars) and LLM Evaluation (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
MoizIbnYousaf (a GitHub user) maintains it in MoizIbnYousaf/ai-agent-skills, which has 1,149 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on September 21, 2026.
Source: MoizIbnYousaf/ai-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.