Agent skill

Verify Algorithms

by av1155 in av1155/houndarr

Verify probabilistic, distributional, or random behaviour empirically before changing search-engine code.

AGPL-3.0Auto-check passedDevelopment

Install Verify Algorithms

skills CLI
$ npx skills add av1155/houndarr --skill verify-algorithms -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install av1155/houndarr verify-algorithms --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/av1155/houndarr.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/verify-algorithms .claude/skills/verify-algorithms && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verify-algorithms
GitHub stars
292
Token cost
~1.6k tokens
SKILL.md length
832 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Verify probabilistic, distributional, or random behaviour empirically before changing search-engine code.

  • Works in 4 steps: Reproduce the algorithm in isolation… → Derive analytically what each page,… → Run hundreds of cycles through… → …
  • Another AI surfaces a claim about bias
  • SKILL.md covers When this rule fires, Required workflow, Tooling to use and What not to do, plus 2 more sections
  • Calls just and python

What it does

Verify Algorithms is an agent skill from av1155/houndarr. Verify probabilistic, distributional, or random behaviour empirically before changing search-engine code. Loads when reading or editing src/houndarr/engine/. Use when a user, code review, or another AI surfaces a claim about bias, ordering, page selection, randomness, or "we are searching the same things over and over".

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development. The repository describes itself as: Self-hosted arr companion for controlled missing, cutoff, and upgrade searches. The licence is AGPL-3.0.

When your agent uses it

  • Another AI surfaces a claim about bias
  • We are searching the same things over and over

Example prompts

  • “we are searching the same things over and over”
  • “/verify-algorithms”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Reproduce the algorithm in isolation against tests/mock_arr/, not
  2. Derive analytically what each page, item, or branch's probability
  3. Run hundreds of cycles through tests/mock_arr/probe_distribution.py
  4. Decide on evidence. If the empirical result agrees with the

What it can do on your machine

Read from SKILL.md and the folder at commit 9b39fdb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • just
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verify Algorithms loads about 1.6k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 832 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from av1155/houndarr at commit 9b39fdb, republished under its AGPL-3.0 licence (© av1155). 832 words, ~1,573 tokens.

Download SKILL.mdSave it as .claude/skills/verify-algorithms/SKILL.md (or your agent's skills folder).
name
verify-algorithms
description
Verify probabilistic, distributional, or random behaviour empirically before changing search-engine code. Loads when reading or editing src/houndarr/engine/. Use when a user, code review, or another AI surfaces a claim about bias, ordering, page selection, randomness, or "we are searching the same things over and over".
paths
src/houndarr/engine/**

Verifying claims about algorithms

Before modifying search-engine logic, scheduling, randomisation, ordering, distribution, or any code where probability or stateful iteration governs behaviour, verify the claim empirically and analytically first. Most reported "bugs" in this class turn out to be sample noise, observation bias, or misreadings of timing-dependent state, and shipping a fix for a non-bug introduces real risk for no real gain.

When this rule fires

Apply this workflow whenever a user, a code review, or another AI surfaces a claim along the lines of:

  • "X picks the wrong page / item / branch"
  • "Y is biased / unfair / skewed toward Z"
  • "Random does not feel random"
  • "The cycle order is broken"
  • "We are searching the same things over and over"

It does not apply to clear logic bugs, typos, or behaviour-change requests. The trigger is specifically: claims about probabilistic or distribution-shaped behaviour where the right answer is a measured histogram, not a code reading.

Required workflow

  1. Reproduce the algorithm in isolation against tests/mock_arr/, not against the live test instances or short-window log dumps. The live test *arrs hold tens of records, which is far below the sample size needed to distinguish bias from variance, and live state (cooldowns, hourly caps, *arr-side sort orders) confounds the measurement.
  2. Derive analytically what each page, item, or branch's probability should be under the current code. Read the loop, write the math down, and predict the distribution shape before running anything. "I think it should be uniform" is not a prediction; "uniform with chi-square below 16.92 at df=9" is.
  3. Run hundreds of cycles through tests/mock_arr/probe_distribution.py or a similar probe modelled on it. Compute chi-square, max/min ratio, and per-bucket standard deviation. Compare against the analytical prediction and against the 5% chi-square critical value at df = N - 1.
  4. Decide on evidence. If the empirical result agrees with the prediction and the chi-square lands below the critical value, the claim is wrong. Document the finding, reference the probe output, and close the investigation. If the result confirms real bias, scope the fix to the smallest change that closes the measured gap, then re-run the probe to prove the gap is gone.

Tooling to use

  • just mock-arr port=PORT items=N seed=S launches the seeded multi-app mock server with configurable item counts and a deterministic seed; identical seeds produce byte-identical responses.
  • .venv/bin/python -m tests.mock_arr.probe_distribution boots the mock in-process, drives the production run_instance_search for many cycles across a sweep of library sizes, and reports per-cycle start-page distributions plus full visit histograms. Use it as the template for any new programmatic probe.
  • The mock exposes GET /__page_log__/{app} and GET /__commands__/{app} for ground-truth request and dispatch records, plus POST /__reset__/{app} to clear them between configurations.
  • For statistical-power-bound questions (100k+ trials), a short pure-Python simulation of just the algorithm beats running through HTTP. Use it when the measurement is about the math, not the integration.
Show full SKILL.md (360 more words)Show less

What not to do

  • Do not treat a short-window dev-DB histogram (a few hours, dozens of cycles, a handful of items) as evidence of algorithmic bias. Cooldown phase, *arr-side sort order, and small-sample variance dominate that signal. The math you owe is a many-cycle distribution against a predicted shape.
  • Do not adopt an external diagnostic write-up without re-deriving the math yourself. Direction (page 1 vs page N) and magnitude (1.5x vs 5x) routinely invert in second-hand summaries of probabilistic algorithms, and shipping a fix for an inverted claim ships a regression.
  • Do not start coding because the claim is plausible. Plausibility is not evidence. The bar is a reproducible measurement that disagrees with the predicted distribution by more than chance.

Closing the loop

When measurement contradicts the claim, the writeup is the engineering contribution. Reference the probe output, state the measured statistics, explain what the original observation was actually picking up (cooldown saturation, recency effects, sort-order interaction, sample noise), and close the discussion. A correct "no change required" is a successful task, not a non-result.

Known emergent behaviours (already measured)

These are real but minor effects that have been verified by probe and deliberately left alone. Do not re-investigate them unless the operating point changes or a user reports a concrete regression.

  • Partial-last-page over-selection on missing/cutoff under random search order. When the engine's page_size does not divide totalRecords evenly, items on the (short) last page are drained every visit because the engine dispatches up to batch_size items per page. Measured at most 2x attention skew for the 1-9 items on the last page at default settings (batch=1, pageSize=10) and 4x in contrived configurations (batch=5, pageSize=20). Affects a small slice of the backlog; the only clean fix is a virtual flat-index draw which is a substantial redesign of _run_search_pass. Probe: tests/mock_arr/probe_cooldown.py.
  • Sonarr / Whisparr v2 windowed-rotation coverage time. The upgrade pass visits 5 series per cycle; full-library coverage takes approximately ceil(eligible_episodes * H / batch) cycles where H is the harmonic-coverage factor. Measured 91% theoretical and 85-89% empirical coverage at 60 cycles with batch=5 on 50 series. This is the intentional trade-off versus hammering one series with a single huge *arr fetch. Probe: tests/mock_arr/probe_upgrade_coverage.py.

© av1155, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/verify-algorithms of av1155/houndarr.

Open the folder on GitHubat commit 9b39fdb

Compare with similar skills

Verify Algorithms next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verify Algorithms compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verify Algorithms this skillav1155/houndarr292—~1.6kAutomated safety check: PassAGPL-3.0
Backend Code Reviewlanggenius/dify158k—~676Automated safety check: PassCustom licence
Native Data FetchingCherryHQ/cherry-studio-app4k6 repos~2.9kAutomated safety check: NotesMIT
Twenty App Entity Developmenttwentyhq/twenty58k—~1.8kAutomated safety check: PassCustom licence
Go Pedantrychromedp/chromedp13k—~3.7kAutomated safety check: PassMIT
Gumroad Prod Consoleantiwork/gumroad9.8k—~2.9kAutomated safety check: NotesMIT

Similar skills

  • Backend Code Review

    langgenius/dify

    Reviews backend code under api/ for concrete, reproducible defects, routes to rule packs for architecture, schema, repositories and SQLAlchemy, and ranks findings from P0 to P3.

    158k GitHub stars~676 tokensUpdated today
    DevelopmentAuto-check passed
  • Native Data Fetching

    CherryHQ/cherry-studio-app

    A skill your agent uses when implementing or debugging ANY network request, API call, or data fetching.

    4k GitHub starsUsed in 6 repos~2.9k tokens
    DevelopmentAuto-check: notes
  • Guides changes to an existing Twenty app: adding or editing objects, layouts, logic functions and front components, with a plan stated before multi-entity edits.

    58k GitHub stars~1.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Go Pedantry

    chromedp/chromedp

    This skill should be used when the user is writing Go code and needs guidance on Go-specific pedantry: error wrapping with fmt.Errorf and %w, interface design (accept interfaces return structs)…

    13k GitHub stars~3.7k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Gumroad Prod Console

    antiwork/gumroad

    Execute read-only Ruby/Rails commands against Gumroad's production database for debugging and investigation.

    9.8k GitHub stars~2.9k tokensUpdated today
    DevelopmentAuto-check: notes
  • LangBot Plugin Development

    langbot-app/LangBot

    Guides building, debugging and testing LangBot plugins: components, SDK calls, README and locale rules, SDK pitfalls and WebSocket-based testing.

    18k GitHub stars~3.9k tokensUpdated today
    DevelopmentAuto-check passed

More from av1155/houndarr

All 9 skills in this repo
  • Bump

    av1155/houndarr

    Bump Houndarr version and prepare a release PR. An agent skill from av1155/houndarr.

    292 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Check

    av1155/houndarr

    Run Houndarr's full quality gate (ruff lint, ruff format check, mypy, bandit, pytest) and report results in a single table.

    292 GitHub stars~366 tokensUpdated 2 days ago
    Auto-check passed
  • Houndarr Architecture

    av1155/houndarr

    Houndarr's source layout and architectural patterns at file granularity.

    292 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Houndarr Changelog

    av1155/houndarr

    Houndarr's CHANGELOG.md style guide and entry rules. An agent skill from av1155/houndarr.

    292 GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed
  • Houndarr Testing

    av1155/houndarr

    Houndarr's pytest patterns. An agent skill from av1155/houndarr.

    292 GitHub stars~1k tokensUpdated 2 days ago
    Auto-check passed
  • Houndarr CI

    av1155/houndarr

    Houndarr's CI workflow reference and branch protection. An agent skill from av1155/houndarr.

    292 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed

Questions about Verify Algorithms

What does Verify Algorithms do?

Verify probabilistic, distributional, or random behaviour empirically before changing search-engine code. Verify Algorithms is an agent skill from av1155/houndarr. Verify probabilistic, distributional, or random behaviour empirically before changing search-engine code.

When should I use Verify Algorithms?

Verify Algorithms fits situations like: another AI surfaces a claim about bias; we are searching the same things over and over.

How do I install Verify Algorithms in Claude Code?

Run `npx skills add av1155/houndarr --skill verify-algorithms -a claude-code`. Or copy the skill folder (.agents/skills/verify-algorithms in av1155/houndarr) into .claude/skills/verify-algorithms in your project. Claude Code loads it when a task matches its description.

How do I install Verify Algorithms in Codex?

Run `npx skills add av1155/houndarr --skill verify-algorithms -a codex`. Or copy the skill folder (.agents/skills/verify-algorithms in av1155/houndarr) into .agents/skills/verify-algorithms in your project. Codex loads it when a task matches its description.

Can I use Verify Algorithms in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add av1155/houndarr --skill verify-algorithms -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify-algorithms, .gemini/skills/verify-algorithms, .github/skills/verify-algorithms and .opencode/skills/verify-algorithms in your project.

What does Verify Algorithms need to run?

Going by SKILL.md and its folder, Verify Algorithms needs the command-line tools its instructions call (just and python). Our summary lists: Python 3.

Does Verify Algorithms access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Verify Algorithms safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verify Algorithms use?

Verify Algorithms is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verify Algorithms use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verify Algorithms?

Skills that share tags, products or a category with Verify Algorithms: Backend Code Review (langgenius/dify, 158k stars), Native Data Fetching (CherryHQ/cherry-studio-app, 4k stars), Twenty App Entity Development (twentyhq/twenty, 58k stars) and Go Pedantry (chromedp/chromedp, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verify Algorithms?

av1155 (a GitHub user) maintains it in av1155/houndarr, which has 292 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 5, 2026.

Source: av1155/houndarr on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.