Search
LangChain · LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs. | FailproofAI/ | 5.3k | — | ~6k | Automated safety check: Pass | Unknown | 4 days ago |
| 2 | Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 3 | Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic. | agentailor/ | 132 | — | ~5.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 4 | Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step. | Jeffallan/ | 12k | — | ~2k | Automated safety check: Pass | MIT | 8 days ago |
| 5 | Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory… | jeremylongshore/ | 2.8k | — | ~3.7k | Automated safety check: Pass | MIT | yesterday |