Search

OpenTelemetry · LLM evaluation

8 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Evaluate and score agent behavior against a golden reference.

agentevals-dev/agentevals163—~904Automated safety check: PassApache-2.02 days ago
2

Propose an improved version of a prompt registered in a self-hosted AgentX (AgentX-trace-eval) instance, using real low-rated evaluation results as evidence, then publish it as a new version once…

AgentX-ai/AgentX-Trace-Eval106—~2kAutomated safety check: PassUnknown4 days ago
3

Inspect and debug live streaming agent sessions to understand what the agent did.

agentevals-dev/agentevals163—~534Automated safety check: PassApache-2.02 days ago
4

Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
5
5.Logfire EvalsOfficial

Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire.

pydantic/skills140—~3.6kAutomated safety check: PassMIT10 days ago
6

Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer.

Arize-ai/phoenix12k—~708Automated safety check: PassApache-2.0yesterday
7

Maintain the bundled TypeScript package docs that ship inside Phoenix npm packages.

Arize-ai/phoenix12k—~2.2kAutomated safety check: PassApache-2.0yesterday
8

Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.

Dynatrace/dynatrace-for-ai163—~4.5kAutomated safety check: PassApache-2.010 days ago