Eval Harness
affaan-m/ECC
Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…
When the user needs to choose between technologies, frameworks, or tools — or says "which framework should I use", "compare X vs Y", "should we migrate from X to Y", "what database should I use"…
$ npx skills add shawnpang/startup-founder-skills --skill tech-stack-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install shawnpang/startup-founder-skills tech-stack-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/shawnpang/startup-founder-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tech-stack-eval .claude/skills/tech-stack-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tech-stack-eval" agent skill from https://github.com/shawnpang/startup-founder-skills/tree/main/skills/tech-stack-eval into .claude/skills/tech-stack-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tech-stack-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/shawnpang/startup-founder-skills/tree/main/skills/tech-stack-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add shawnpang/startup-founder-skills --skill tech-stack-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install shawnpang/startup-founder-skills tech-stack-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/shawnpang/startup-founder-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tech-stack-eval .agents/skills/tech-stack-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tech-stack-eval" agent skill from https://github.com/shawnpang/startup-founder-skills/tree/main/skills/tech-stack-eval into .agents/skills/tech-stack-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tech-stack-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add shawnpang/startup-founder-skills --skill tech-stack-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install shawnpang/startup-founder-skills tech-stack-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/shawnpang/startup-founder-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tech-stack-eval .cursor/skills/tech-stack-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tech-stack-eval" agent skill from https://github.com/shawnpang/startup-founder-skills/tree/main/skills/tech-stack-eval into .cursor/skills/tech-stack-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tech-stack-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/shawnpang/startup-founder-skills.git --path skills/tech-stack-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add shawnpang/startup-founder-skills --skill tech-stack-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install shawnpang/startup-founder-skills tech-stack-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/shawnpang/startup-founder-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tech-stack-eval .gemini/skills/tech-stack-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tech-stack-eval" agent skill from https://github.com/shawnpang/startup-founder-skills/tree/main/skills/tech-stack-eval into .gemini/skills/tech-stack-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tech-stack-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install shawnpang/startup-founder-skills tech-stack-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add shawnpang/startup-founder-skills --skill tech-stack-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/shawnpang/startup-founder-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tech-stack-eval .github/skills/tech-stack-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tech-stack-eval" agent skill from https://github.com/shawnpang/startup-founder-skills/tree/main/skills/tech-stack-eval into .github/skills/tech-stack-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tech-stack-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add shawnpang/startup-founder-skills --skill tech-stack-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install shawnpang/startup-founder-skills tech-stack-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/shawnpang/startup-founder-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tech-stack-eval .opencode/skills/tech-stack-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tech-stack-eval" agent skill from https://github.com/shawnpang/startup-founder-skills/tree/main/skills/tech-stack-eval into .opencode/skills/tech-stack-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tech-stack-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tech-stack-evalWhen the user needs to choose between technologies, frameworks, or tools — or says "which framework should I use", "compare X vs Y", "should we migrate from X to Y", "what database should I use"…
Tech Stack Eval is an agent skill from shawnpang/startup-founder-skills. When the user needs to choose between technologies, frameworks, or tools — or says "which framework should I use", "compare X vs Y", "should we migrate from X to Y", "what database should I use", "calculate TCO".
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: AI agent skills for tech startup founders — fundraising, sales, product, recruiting, engineering, legal, ops, and growth. Works with Claude Code, Cursor, Codex, and any Agent… The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4ad31b4. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
herokuawsFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tech Stack Eval loads about 2k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 706 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from shawnpang/startup-founder-skills at commit 4ad31b4, republished under its MIT licence (© shawnpang). 706 words, ~2,005 tokens.
.claude/skills/tech-stack-eval/SKILL.md (or your agent's skills folder).Do NOT use when: the decision is trivial (use team preference), the technology is already mandated, or this is an emergency production issue.
From startup-context: product type, team skills, tech stack, stage, scale, budget. Also ask:
# Tech Stack Evaluation: [Decision Title]
## Decision Context — what we are choosing and why it matters
## Candidates — table: technology, version, license, one-liner
## Evaluation Criteria — table: criterion, weight, why it matters
## Scoring Matrix — table: criterion (weight), scores per option, weighted total
## Ecosystem Health — table: GitHub stars, weekly downloads, last release, open issues, major users
## TCO Estimate — table: cost category by option over 12 months or 5 years
## Security & Compliance — vulnerability history, compliance readiness (SOC 2, GDPR)
## Recommendation — clear winner, rationale, confidence level, caveats
## Migration Path (if applicable) — phased plan with timeline and rollback strategySelect 6-8 and assign weights (total = 100%):
| Level | Criteria |
|---|---|
| Thriving | Regular releases (< 3 months), growing adoption, multiple corporate sponsors, active community |
| Stable | Regular releases (< 6 months), steady adoption, established community, no decline signs |
| At Risk | Infrequent releases (> 12 months), declining downloads, key maintainers leaving, few contributors |
Project over 12 months minimum (5 years for infrastructure decisions): compute, storage, bandwidth, licensing, engineering time (setup + ongoing maintenance x loaded cost), operational overhead (monitoring, on-call), and hidden costs (training, migration tooling, dual-running).
Engineering time is usually the largest cost for startups. A technology saving $200/month on hosting but costing 40 extra engineering hours to operate is a net loss.
| Risk Level | Criteria |
|---|---|
| Low | Additive change, no data migration, can run in parallel, < 2 weeks |
| Medium | Requires data migration or API changes, 2-8 weeks, can be phased |
| High | Core system replacement, > 8 weeks, requires downtime or big-bang cutover |
Use the strangler fig pattern: route new traffic to the new system, migrate old incrementally. Always maintain rollback capability. Set a concrete cut-off date -- half-migrated systems are the worst outcome.
| Level | Score | Interpretation |
|---|---|---|
| High | 80-100% | Clear winner, strong data, wide margin |
| Medium | 50-79% | Trade-offs present, recommendation holds but with caveats |
| Low | < 50% | Close call, limited data, suggest a proof-of-concept before committing |
architecture-design — chain when the tech stack decision feeds into a broader system designcicd-setup — chain to configure CI/CD for the chosen technologyExample prompt: "Compare React vs Vue for a SaaS dashboard. Priorities: developer productivity (40%), ecosystem (30%), performance (30%)."
Good output snippet:
## Scoring Matrix
| Criterion (weight) | React | Vue |
|--------------------------|-------|------|
| Developer Productivity (40%) | 4/5 | 4/5 |
| Ecosystem (30%) | 5/5 | 4/5 |
| Performance (30%) | 4/5 | 5/5 |
| **Weighted Total** | **4.3** | **4.3** |
Confidence: Medium (55%). Scores are nearly identical. Recommendation: React,
but only because your team has 2 years of React experience (not captured in
the matrix). If the team were greenfield, Vue's developer experience gives it
a slight edge. This is close enough to warrant team preference as the tiebreaker.Example prompt: "We're on Heroku at $2,400/mo. Should we migrate to AWS?"
Good output snippet:
## TCO Estimate (12 months)
| Category | Heroku | AWS |
|----------------------|-----------|-------------------|
| Compute | $1,200/mo | $480/mo (ECS) |
| Database | $800/mo | $350/mo (RDS) |
| Add-ons | $400/mo | $120/mo |
| Engineering (setup) | $0 | $12,000 one-time |
| Engineering (ongoing)| 2 hrs/mo | 8 hrs/mo |
| **Annual Total** | **$28,800** | **$18,000** |
Break-even at month 14. At Series A with a team of 6, wait until Heroku hits
$4,000/mo — engineering hours are better spent on product right now.© shawnpang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/tech-stack-eval of shawnpang/startup-founder-skills.
Open the folder on GitHubat commit 4ad31b4
Tech Stack Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tech Stack Eval this skillshawnpang/startup-founder-skills | 341 | — | ~2k | Automated safety check: Pass | MIT | |
| Eval Harnessaffaan-m/ECC | 275k | — | ~2.2k | Automated safety check: Pass | MIT | |
| Evalalirezarezvani/claude-skills | 28k | 1 repos | ~618 | Automated safety check: Pass | MIT | |
| Eval Harnessaffaan-m/ECC | 275k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Eval-Driven Development Harnessaffaan-m/ECC | 275k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Paperclip Evalspaperclipai/paperclip | 99k | — | ~839 | Automated safety check: Pass | MIT |
affaan-m/ECC
Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…
alirezarezvani/claude-skills
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
affaan-m/ECC
Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi
affaan-m/ECC
Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.
paperclipai/paperclip
Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing scripts, client components, bundles, or runtime behavior related to Never use eval() or unsafe dynamic code execution.
shawnpang/startup-founder-skills
When the user needs to review an existing contract, assess risk in proposed terms, or evaluate a contract before signing.
shawnpang/startup-founder-skills
When the user wants to apply to startup accelerators, incubators, or fellowship programs.
shawnpang/startup-founder-skills
When the user needs to design or evaluate system architecture — service boundaries, data models, API contracts, infrastructure topology, database selection, or dependency analysis.
shawnpang/startup-founder-skills
When the user needs to write a monthly or quarterly investor update, prepare a board deck, or communicate company progress to stakeholders.
shawnpang/startup-founder-skills
When the user needs to identify at-risk accounts, understand why customers are leaving, reduce churn rate, build health scores, design save plays, or create win-back campaigns.
shawnpang/startup-founder-skills
When the user needs to set up or improve CI/CD pipelines — GitHub Actions, GitLab CI, deployment automation, or says "set up CI", "automate deployment", "add tests to pipeline", "fix my build".
When the user needs to choose between technologies, frameworks, or tools — or says "which framework should I use", "compare X vs Y", "should we migrate from X to Y", "what database should I use"…. Tech Stack Eval is an agent skill from shawnpang/startup-founder-skills. When the user needs to choose between technologies, frameworks, or tools — or says "which framework should I use", "compare X vs Y", "should we migrate from X to Y", "what database should I use", "calculate TCO".
Tech Stack Eval fits situations like: needs to choose between technologies; says which framework should I use; should we migrate from X to Y; what database should I use.
Run `npx skills add shawnpang/startup-founder-skills --skill tech-stack-eval -a claude-code`. Or copy the skill folder (skills/tech-stack-eval in shawnpang/startup-founder-skills) into .claude/skills/tech-stack-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add shawnpang/startup-founder-skills --skill tech-stack-eval -a codex`. Or copy the skill folder (skills/tech-stack-eval in shawnpang/startup-founder-skills) into .agents/skills/tech-stack-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shawnpang/startup-founder-skills --skill tech-stack-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tech-stack-eval, .gemini/skills/tech-stack-eval, .github/skills/tech-stack-eval and .opencode/skills/tech-stack-eval in your project.
Going by SKILL.md and its folder, Tech Stack Eval needs the command-line tools its instructions call (heroku and aws).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tech Stack Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Tech Stack Eval: Eval Harness (affaan-m/ECC, 275k stars), Eval (alirezarezvani/claude-skills, 28k stars), Eval Harness (affaan-m/ECC, 275k stars) and Eval-Driven Development Harness (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
shawnpang (a GitHub user) maintains it in shawnpang/startup-founder-skills, which has 341 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on March 16, 2026.
Source: shawnpang/startup-founder-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.