Search

LangChain · LLM evaluation

5 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

FailproofAI/failproofai5.3k—~6kAutomated safety check: PassUnknown4 days ago
2

Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
3

Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

agentailor/fullstack-langgraph-nextjs-agent132—~5.3kAutomated safety check: PassMIT1 mo ago
4

Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

Jeffallan/claude-skills12k—~2kAutomated safety check: PassMIT8 days ago
5

Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory…

jeremylongshore/tons-of-skills-marketplace2.8k—~3.7kAutomated safety check: PassMITyesterday