Search

PostHog · LLM evaluation

5 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1
1.Writing EvalsOfficial

Teaches how to write and run evals on the products/posthogai/evalharness/ harness — sandboxed agent suites that execute the real coding agent in a Docker or Modal sandbox against a seeded Hedgebox…

PostHog/posthog40k—~4kAutomated safety check: NotesUnknownyesterday
2

Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.

PostHog/posthog40k—~6.7kAutomated safety check: PassUnknownyesterday
3

Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded…

PostHog/posthog40k—~1.5kAutomated safety check: PassUnknownyesterday
4

Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment).

PostHog/posthog40k—~5.7kAutomated safety check: PassUnknownyesterday
5

Set up an LLM-judge evaluation that extracts canonical use cases for a PostHog feature at scale and streams the results to a Slack channel as a live feed.

PostHog/posthog40k—~7.6kAutomated safety check: PassUnknownyesterday