Search
LangGraph · LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs. | FailproofAI/ | 5.3k | — | ~6k | Automated safety check: Pass | Unknown | 4 days ago |
| 2 | Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic. | agentailor/ | 132 | — | ~5.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 3 | INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith. | langchain-ai/ | 1.3k | — | ~8.7k | Automated safety check: Notes | MIT | 2 days ago |
| 4 | Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory… | jeremylongshore/ | 2.8k | — | ~3.7k | Automated safety check: Pass | MIT | yesterday |
| 5 | Build type-safe AI agents and graph-based workflows with PydanticAI and PydanticGraph. | magnus919/ | 115 | — | ~4.3k | Automated safety check: Pass | MIT | yesterday |