LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Keeps the engagement record for client work. An agent skill from suboss87/FDEOps.
$ npx skills add suboss87/FDEOps --skill fde -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install suboss87/FDEOps fde --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/suboss87/FDEOps.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/fde .claude/skills/fde && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "fde" agent skill from https://github.com/suboss87/FDEOps/tree/Main/skills/fde into .claude/skills/fde/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fde", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/suboss87/FDEOps/tree/Main/skills/fdeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add suboss87/FDEOps --skill fde -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install suboss87/FDEOps fde --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/suboss87/FDEOps.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/fde .agents/skills/fde && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "fde" agent skill from https://github.com/suboss87/FDEOps/tree/Main/skills/fde into .agents/skills/fde/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fde", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add suboss87/FDEOps --skill fde -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install suboss87/FDEOps fde --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/suboss87/FDEOps.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/fde .cursor/skills/fde && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "fde" agent skill from https://github.com/suboss87/FDEOps/tree/Main/skills/fde into .cursor/skills/fde/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fde", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/suboss87/FDEOps.git --path skills/fde--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add suboss87/FDEOps --skill fde -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install suboss87/FDEOps fde --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/suboss87/FDEOps.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/fde .gemini/skills/fde && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "fde" agent skill from https://github.com/suboss87/FDEOps/tree/Main/skills/fde into .gemini/skills/fde/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fde", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install suboss87/FDEOps fdeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add suboss87/FDEOps --skill fde -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/suboss87/FDEOps.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/fde .github/skills/fde && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "fde" agent skill from https://github.com/suboss87/FDEOps/tree/Main/skills/fde into .github/skills/fde/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fde", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add suboss87/FDEOps --skill fde -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install suboss87/FDEOps fde --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/suboss87/FDEOps.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/fde .opencode/skills/fde && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "fde" agent skill from https://github.com/suboss87/FDEOps/tree/Main/skills/fde into .opencode/skills/fde/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fde", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
fdeKeeps the engagement record for client work. An agent skill from suboss87/FDEOps.
Fde is an agent skill from suboss87/FDEOps. Keeps the engagement record for client work. Use when they name a client or stakeholder. Use when they debrief a meeting or paste notes. Use when they ask what was agreed. Use when they run a POC, change the client's codebase, prove it on their staging, go live, or need evals before a model acts. Use when they prep a readout, when trust shifts, or they say @fde. Route and run the local fde CLI (or npx --yes fdeops). Never ask them to type commands. Not for ordinary code edits in an unbound repo.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 48 other files, including reference files (for example `agents/openai.yaml`, `references/ai.md` and `references/artifacts.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Forward deployed engineering skills for AI coding agents. The licence is MIT.
Read from SKILL.md and the folder at commit 64d1d8a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Fde loads about 2.9k tokens when it runs, and up to ~74k if it reads all its reference files. Until then it costs about 126 tokens; SKILL.md has 1,404 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from suboss87/FDEOps at commit 64d1d8a, republished under its MIT licence (© suboss87). 1,404 words, ~2,908 tokens.
.claude/skills/fde/SKILL.md (or your agent's skills folder). This skill also uses 46 other files; get the full folder from GitHub.Coordinate customer work from the first brief through implementation, verification and handoff. Choose the relevant task; never make the user pick a phase. All task skills also work individually. Reuse the customer's tools, decisions and operating process.
Read references/task-context.md for authority, data boundaries, CLI availability and evidence rules. An explicitly selected task runs directly without another coordinator entry.
.fde/, preferences or records just to draft, analyze or change code. Ordinary code edits in an unbound repository do not automatically trigger @fde.fde resume --init <client-name>. Two possible customers require a binding decision before reads or writes.The human asks in ordinary language or invokes a task skill. You run the required CLI commands. Never tell the FDE to type commands; never ask them to run the CLI. Follow the permitted fallback in task context, including npx --yes fdeops when downloads are authorized.
Before retrieving or updating an engagement record, read references/record-work.md for privacy-safe session entry, record commands, the memory contract and session digest. Load it once for the current task; standalone work does not require it. Confirm consequential record changes under task-context rules. A saved claim is not customer acceptance.
Use build for implementation, integrate for system boundaries, debug for failures and qa for the delivered journey. Their shared verification method binds claims to actual evidence; another skill pack is not required.
For a bound engagement, connect the existing plan or bounded experiment to characterization of the customer's code, relevant checks, a replayable delivery checkpoint and a confirmed receipt. Reuse their tests and runner. Where a model acts or judges, follow references/eval-pack.md and references/ai.md before release. Use ship for release authority, recovery and operating evidence.
Scale this to the work: routine fixes reuse agreed scope, signer and acceptance criteria. They do not require a new sponsor decision per edit. Keep implemented, verified, deployed, measured and accepted separate. A local pass is not a customer outcome.
Be direct, use the customer's terms, and act after a short playback. Ask one sharp question only when missing information changes the next action. State uncertainty rather than guessing. Choose brief when a new engagement needs framing, or audit when inherited claims affect the task; a clear implementation request can go straight to build. Engagement size changes depth, not the available skills.
The groups below are a navigation map, not required stages. Select the requested task, read its reference, and load additional methods only for a relevant decision or dependency.
| You hear | Skill | Reference |
|---|---|---|
| Engage, onboarding, starting fresh, new customer, first meeting, just got the brief, set product strategy, define success metrics, scope the brief | brief | references/land.md |
| Taking over, previous consultant left, joining mid-project | audit | references/audit.md |
| Need to understand who matters, who decides, map decision rights, who blocks quietly | who-decides | references/who-decides.md |
| Need to earn access, navigate AI policy, build credibility | earn-trust | references/earn-trust.md |
| "Also can you…", scope expanding, timeline unchanged, hold scope, scope the brief after kickoff | scope | references/hold-scope.md |
| You hear | Skill | Reference |
|---|---|---|
| Diagnose, don't know the real problem, brief feels wrong, shadow processes, frame discovery, understand the problem space, data not ready, data estate, catalog the data, parts of the problem, decompose | discover | references/discover.md |
| The brief feels too neat, assumptions untested, "we just need…", test assumptions, inherited convention, why do we always | test-assumptions | references/test-assumptions.md |
| Multiple use cases competing, "we want to do everything", score use cases | score-use-cases | references/score-use-cases.md |
| Need to validate a direction, prototype, demo to de-risk, POC, spike, killer assumption, validate the solution, build prototype | poc | references/poc.md |
| You hear | Skill | Reference |
|---|---|---|
| Align, break this down, what order, sequence the delivery, align the plan | plan | references/plan.md |
| Sponsor needs justification, need to defend budget or timeline, build the business case | business-case | references/business-case.md |
| Significant decision, multiple approaches, "what should we do?", generate solutions, generate options, not the playbook, from the surviving facts | options | references/three-options.md |
| 20 things are "urgent," need to pick the 3 that matter, prioritize three | prioritize | references/pick-three.md |
| You hear | Skill | Reference |
|---|---|---|
| What could go wrong, touching shared infrastructure, need to assess impact, assess impact, provision, IaC, shared infra | what-breaks | references/what-breaks.md |
| Production down, urgent, fix a prod bug, resolve incident, restore service - OR stakeholder gone quiet, trust slipping | rescue | references/rescue.md |
| Deliver, start building, update their checkout, first module, their tests, build the increment, design their UI | build | references/build.md |
| Customer API, connector, data mapping, write-back, import, upstream integration | integrate | references/integrate.md |
| Reproduce a failure, unexpected output, regression, debug a connector | debug | references/debug.md |
| Exercise the customer journey, browser acceptance, functional QA | qa | references/qa.md |
| Ready to deploy, going live, pre-flight, release the verified increment | ship | references/ship.md |
| Review this change, review the pull request, is it safe, does it match what we agreed | review | references/review.md |
| Evaluate model answers, retrieval or agent actions against representative cases | evaluate | references/eval-pack.md |
| Diff grew / scope creep in the PR / "did we only build what we said" / KEEP JUSTIFY SPLIT DROP | review (+ ship if going live) | references/review.md Stage 1 · references/ship.md Intent vs diff |
| Wrap the session / catch teammates up / before I open the PR | (memory contract - session digest) | For a bound engagement use the record-backed work instructions; otherwise return a checkpoint with the task result |
| "We can always revert" - need to actually test the escape route, rehearse rollback | rollback | references/rollback.md |
| You hear | Skill | Reference |
|---|---|---|
| Realize, weekly update due, "need to send the sponsor something", report the outcome | readout | references/readout.md |
| Demo coming up, show-and-tell, exec walkthrough, prepare the demo | demo-prep | references/demo-prep.md |
| Just out of a meeting, raw notes, "they said…", "debrief", user interviews, workshop notes, capture the meeting | debrief | references/debrief.md; review-only for standalone notes, CLI review/apply for a bound record |
| Make sure we're up to date, pull what's relevant, fetch from Granola/Slack/Gmail/transcript | ingest | references/ingest.md (capability check → stage → propose → confirm → apply) |
| Connect a new MCP / connect Granola Slack or Notion / what can you pull | connect | references/connect.md (+ references/source-setup.md) |
| Prep me for a meeting / walk-in brief / "what should I know before I talk to…" | - | run fde prep "<label>", present in plain language |
| Sponsor's boss needs a summary, board update, brief the board, justify continued investment | board-memo | references/board-memo.md |
| Status across all my customers, view the portfolio | dashboard | references/dashboard.md |
| You hear | Skill | Reference |
|---|---|---|
| Juggling 2+ customers, losing track, context-switching, switch engagements | switch-clients | references/switch-clients.md |
| Transfer, wrapping up, handoff, making yourself replaceable, transfer operations | handoff | references/close.md |
| Engagement ending, team needs to operate without you, write the runbook | runbook | references/runbook.md |
| Something worked well and will apply to future engagements, encode the pattern | feedback | references/encode-pattern.md |
| "Red-team this," "stress-test my plan," poke holes, challenge the plan, what am I missing | red-team | references/red-team.md |
| "Did we agree to X?", "Who decided?", "Why did we choose X?", scope dispute | - | run fde receipts <term>; return the decision, person, date, source and later changes; use targeted recall for missing context |
Overlays - activate alongside any skill on signal, don't wait to be told:
| Signal | Overlay |
|---|---|
| AI, ML, LLM, model, embeddings, RAG, agents, fine-tuning, inference, drift, train the model | references/ai.md |
| Golden set, eval suite, eval pack, pass/fail before AI ship, HITL gate for model, POC the model | references/eval-pack.md (+ ai.md) |
| Deck, slides, report, governance framework, compliance pack, ADR, PDF | references/artifacts.md |
| Patient data, PHI, HIPAA, EHR, clinical | references/healthcare.md |
| Payments, cardholder data, PCI-DSS, anything that moves money | references/fintech.md |
| Government agency, FedRAMP, ATO, CUI, classified | references/gov.md |
Ready to build at any point in an engagement: check that the supplied facts establish the outcome, constraints and verification path, then use build for the working slice. Use discover or plan only for a material gap or a real sequencing decision; do not run every stage or create records to demonstrate activity. On a takeover, audit inherited claims that affect the task. Two customers in one message: confirm which folder.
© suboss87, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 46 other files (references) in skills/fde of suboss87/FDEOps.
Open the folder on GitHubat commit 64d1d8a
Fde next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Fde this skillsuboss87/FDEOps | 954 | — | ~2.9k | Automated safety check: Pass | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
suboss87/FDEOps
Audit an inherited engagement or implementation against its evidence.
suboss87/FDEOps
Draft a board or executive summary of an engagement using outcomes, risks and investment decisions.
suboss87/FDEOps
Clarify a new customer brief, desired outcome, constraints and evidence gaps.
suboss87/FDEOps
Develop a business case for an initiative using costs, benefits, risks and evidence.
suboss87/FDEOps
Configure or diagnose access to a requested source using available host tools.
suboss87/FDEOps
Review the portfolio across existing engagement records. An agent skill from suboss87/FDEOps.
Categories
Keeps the engagement record for client work. An agent skill from suboss87/FDEOps. Fde is an agent skill from suboss87/FDEOps. Keeps the engagement record for client work.
Fde fits situations like: they name a client; they debrief a meeting; they ask what was agreed; change the clients codebase.
Run `npx skills add suboss87/FDEOps --skill fde -a claude-code`. Or copy the skill folder (skills/fde in suboss87/FDEOps) into .claude/skills/fde in your project. Claude Code loads it when a task matches its description.
Run `npx skills add suboss87/FDEOps --skill fde -a codex`. Or copy the skill folder (skills/fde in suboss87/FDEOps) into .agents/skills/fde in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add suboss87/FDEOps --skill fde -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fde, .gemini/skills/fde, .github/skills/fde and .opencode/skills/fde in your project.
Going by SKILL.md and its folder, Fde needs the command-line tools its instructions call (npx). Our summary lists: Node.js.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Fde is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 71k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Fde: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
suboss87 (a GitHub user) maintains it in suboss87/FDEOps, which has 954 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 8, 2026.
Source: suboss87/FDEOps on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.