Search
PostHog · LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Teaches how to write and run evals on the products/posthogai/evalharness/ harness — sandboxed agent suites that execute the real coding agent in a Docker or Modal sandbox against a seeded Hedgebox… | PostHog/ | 40k | — | ~4k | Automated safety check: Notes | Unknown | yesterday |
| 2 | Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified. | PostHog/ | 40k | — | ~6.7k | Automated safety check: Pass | Unknown | yesterday |
| 3 | Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded… | PostHog/ | 40k | — | ~1.5k | Automated safety check: Pass | Unknown | yesterday |
| 4 | Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment). | PostHog/ | 40k | — | ~5.7k | Automated safety check: Pass | Unknown | yesterday |
| 5 | Set up an LLM-judge evaluation that extracts canonical use cases for a PostHog feature at scale and streams the results to a Slack channel as a live feed. | PostHog/ | 40k | — | ~7.6k | Automated safety check: Pass | Unknown | yesterday |