A skill your agent uses when the user asks to "scout for new models", "check for model updates", "find better ASR models", "check sherpa-onnx releases", "look for new Whisper models", "search…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Model Scout

skills CLI
$ npx skills add RisorseArtificiali/anti-vocale --skill model-scout -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RisorseArtificiali/anti-vocale model-scout --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RisorseArtificiali/anti-vocale.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/model-scout .claude/skills/model-scout && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-scout
GitHub stars
118
Token cost
~2.9k tokens
SKILL.md length
1,373 words
Files
4 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user asks to "scout for new models", "check for model updates", "find better ASR models", "check sherpa-onnx releases", "look for new Whisper models", "search…

  • Works in 5 steps: Verify Baseline → Launch Parallel Research → Filter and Score → …
  • The user asks to scout for new models
  • SKILL.md covers Purpose, Current Baseline, Scope Handling and Evaluation Criteria, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Model Scout is an agent skill from RisorseArtificiali/anti-vocale. Use this skill when the user asks to "scout for new models", "check for model updates", "find better ASR models", "check sherpa-onnx releases", "look for new Whisper models", "search HuggingFace for ASR models", "run model scout", "any new Parakeet models", "check framework updates", "/model-scout", "find ASR models for <language", or when discussing improvements to on-device transcription quality for the Anti-Vocale app. Also triggers on periodic research requests about the ASR/LLM landscape for mobile speech…

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/model-inventory.md`, `references/report-template.md` and `references/search-queries.md`).

It sits in AI & LLM Engineering, covering Speech recognition and synthesis and Transcription. It works with ONNX, Hugging Face, Android and Kotlin. The repository describes itself as: Android app for transcribing voice messages locally on-device, with no internet required. The licence is Apache-2.0.

When your agent uses it

  • The user asks to scout for new models
  • Check for model updates
  • Find better ASR models
  • Check sherpa-onnx releases

Example prompts

  • “scout for new models”
  • “check for model updates”
  • “find better ASR models”
  • “/model-scout”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Verify Baseline
  2. Launch Parallel Research
  3. Filter and Score
  4. Synthesize Report
  5. Save Report

What it can do on your machine

Read from SKILL.md and the folder at commit d2d157e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • huggingface.co
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Scout loads about 2.9k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 152 tokens; SKILL.md has 1,373 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~152
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from RisorseArtificiali/anti-vocale at commit d2d157e, republished under its Apache-2.0 licence (© RisorseArtificiali). 1,373 words, ~2,893 tokens.

Download SKILL.mdSave it as .claude/skills/model-scout/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
model-scout
description
Use this skill when the user asks to "scout for new models", "check for model updates", "find better ASR models", "check sherpa-onnx releases", "look for new Whisper models", "search HuggingFace for ASR models", "run model scout", "any new Parakeet models", "check framework updates", "/model-scout", "find ASR models for <language>", or when discussing improvements to on-device transcription quality for the Anti-Vocale app. Also triggers on periodic research requests about the ASR/LLM landscape for mobile speech recognition, and on questions about what languages the app could serve better.
version
0.2.0

Model Scout Skill

Periodic reconnaissance skill for discovering new ASR models, framework releases, and on-device transcription techniques relevant to Anti-Vocale (Android voice message transcription app, global user base).

Purpose

Identify new or updated models and frameworks that could improve transcription quality in any language our users transcribe (Italian is the maintainer's reference benchmark, not the only axis), reduce model size, speed up inference on Android ARM64 devices, or extend language coverage. Coverage extends through TWO channels: the built-in catalog (app release required) and the external-model platform (community catalog plus user imports for the Transducer/Whisper/CTC/SenseVoice/Canary/Moonshine/Dolphin families (ExternalModels.kt declares seven); a catalog entry reaches users TODAY with no app release).

Current Baseline

Verify against source files before each run:

  • Framework versions: Read app/build.gradle.kts for sherpa-onnx and LiteRT-LM versions; read .sherpa-version at the repo root for the pinned sherpa-onnx tag
  • Model inventory: Read app/src/main/java/com/antivocale/app/transcription/*ModelManager.kt and *Downloader.kt plus data/ModelDownloader.kt (the Gemma inventory), and references/model-inventory.md
  • Community catalog: Read app/src/main/assets/ + the CURRENT versioned index (TASK-643: index-<versionName>.json, derive the name from the highest-versioned index-*.json present; the unsuffixed index.json is the frozen legacy channel for <=1.13.x and never receives new entries) for the cataloged entries and their languages (dedupe and coverage judgments run against this, not against memory)
  • External families: Read app/src/main/java/com/antivocale/app/transcription/ModelFamilySupport.kt for the import families and their constraints (featureDim, chunk caps)
  • Previous reports: Check docs/scout-reports/ for historical context

Known baseline (verify during run):

  • ASR built-in: Parakeet TDT 0.6b v3 (640MB stock-int8, 862MB smoothquant default, 25 EU langs), Whisper Small/Turbo/Medium (101 langs), Whisper Distil Large V3 IT (938MB), Qwen3-ASR 0.6b (938MB, 59 langs), Nemotron 3.5 (streaming, OnlineRecognizer, 45 langs), GigaAM v3 (Russian)
  • External families: Transducer / Whisper / CTC / SenseVoice / Canary / Moonshine / Dolphin (user-imported plus community catalog; seven families)
  • LLM: Gemma 4 via .litertlm through LlmTranscriptionBackend (the working audio path; Gemma 4 E2B and E4B variants)
  • Frameworks: sherpa-onnx (pinned via .sherpa-version), LiteRT-LM

Important context for LLM/GGUF findings: GGUF has NO path in the app at all (the llama-bro backend was removed entirely, TASK-639 2026-09-23: fork deleted upstream and the exports are text-only anyway); .litertlm via LiteRT-LM is the only Gemma runtime. When reporting Gemma GGUF variants, always note: "no GGUF runtime in the app; relevant only if a .litertlm conversion appears."

Scope Handling

Parse the user's request to determine scope. Default to full if unspecified:

ScopeFocus Area
fullAll areas below, plus the Italian reference sweep
asrASR/transcription models only
llmLLM models (litertlm, GGUF, multimodal)
frameworkssherpa-onnx, ONNX Runtime, LiteRT-LM releases
parakeetNVIDIA Parakeet family only
whisperWhisper family only
qwenQwen ASR family only
languages:X,YPer-language sweep for the named languages (e.g. languages:fa,uk,he): find the best models per language, built-in or community-catalog, and report coverage gaps

Evaluation Criteria

Score every finding 1-5 on each, then compute weighted composite:

  1. Target-language quality impact (weight: 3x): Will this improve WER/CER in a language our users transcribe? Name the languages the evidence covers and cite the benchmark or community report. Italian has our eval harness (eval/); for other languages rely on published WER, model cards, and community reports, and label the confidence honestly.
  2. Size efficiency (weight: 2x): Under 1GB for ASR / 5GB for LLM with good quality-per-MB?
  3. On-device feasibility (weight: 2x): Runs on Android ARM64 via sherpa-onnx or LiteRT-LM?
  4. Language coverage breadth (weight: 1x): Which languages does the app serve today (built-in catalog sets, the streaming entry, and the community catalog) and which gaps does this fill? A model for a language with NO current good option scores high here.
  5. Community-catalog importability (weight: 1x): Does it fit an external family (one of the seven import families: Transducer/Whisper/CTC/SenseVoice/Canary/Moonshine/Dolphin, sherpa-exportable)? A catalog candidate reaches users without an app release; pipeline = the community-model-conversion skill (proven on Whisper fine-tunes; for the other families flag conversion effort as unproven); intake threads = GH issues #69/#70.
  6. Maintenance health (weight: 1x): Actively maintained with recent commits?

Composite = weighted sum / 5, reported as X/10 (weights sum to 10, criteria score 1-5).

Execution Process

Step 1: Verify Baseline

Read current source files to confirm framework versions and model inventory. Check docs/scout-reports/ for previous findings to avoid repeating.

Step 2: Launch Parallel Research

Run all applicable research areas in parallel. These are logical areas run as parallel bash groups in THIS session, not Agent-tool dispatches: do NOT dispatch the model-scout agent recursively (it executes the whole flow itself and would race on the report path). See references/search-queries.md for the exact queries; transport is the crw CLI (crw search, crw scrape) with curl as fallback (the mcp__crw__ MCP tools may not be registered).

Area A: HuggingFace ASR Models: language-agnostic firehose first (csukuangfj author feed with the ASR pipeline-tag filter), then family queries, then any per-language sweeps the scope requested (docs inventory page only on full).

Area B: HuggingFace LLM Models: search for small multilingual litertlm/GGUF models, Gemma variants, and multimodal audio models.

Area C: Framework Releases: curl the GitHub API endpoints for sherpa-onnx and ONNX Runtime releases; WebFetch the LiteRT-LM releases page (HTML is its primary source).

Area D: Landscape Research: crw search for broad discovery and crw scrape for HuggingFace blog posts and recent ASR developments.

Show full SKILL.md (533 more words)Show less
Step 3: Filter and Score

Apply exclusion filters (too large, no quantization, no quality evidence for any user-relevant language, abandoned) and inclusion signals (sherpa-onnx pre-converted, published WER in a covered or requested language, fits an external family). Score remaining findings against evaluation criteria.

Step 4: Synthesize Report

Structure the report as specified in references/report-template.md. Include per-model cards with scores, the import path for each finding, prioritized recommendations, a language coverage section, and a watch list.

Step 5: Save Report

Save to docs/scout-reports/YYYY-MM-DD.md. Create directory if needed.

Behavioral Rules

  1. Evidence over speculation: every finding must have a verifiable source URL. Never fabricate model names, sizes, or benchmarks.
  2. Honest confidence: if quality in the claimed language cannot be verified, state so explicitly. Unverified quality is the norm for long-tail languages; say it, do not soften it.
  3. No marketing language: use neutral, technical assessments. No "groundbreaking" or "state-of-the-art" without citing benchmarks.
  4. Size is contextual: our ASR models range from ~300MB to 988MB built-in; community-catalog models may reasonably go larger if the family supports it. Don't arbitrarily label models "too large" without comparing to what we already ship or catalog.
  5. Integration realism: account for the full path. For built-in candidates: model download, ONNX conversion, sherpa-onnx integration, JNI bridge, ProGuard rules, UI changes, testing. For catalog candidates: the community-model-conversion pipeline (export, validation, the REQUIRED device-validation pass, catalog entry, docs, GH close-out).
  6. Parallel execution: always run CRW scrapes and searches concurrently to minimize wall-clock time.
  7. Historical awareness: reference previous scout reports when available.
  8. No padding: "No significant changes since last scout on YYYY-MM-DD" is a valid finding.

Tracked Items

Items the scout should check for updates on each run. For each item, search for new commits, comments, or status changes and report progress.

ItemURLWhyAdded
VibeVoice TTS support in sherpa-onnxhttps://github.com/k2-fsa/sherpa-onnx/issues/3106ONNX-exported VibeVoice models exist (FluffyBunnies/vibevoice-onnx-v2). If sherpa-onnx adds native support, could enable TTS features. Check for new comments, labels, or linked PRs.2026-04-30
Cohere Transcribe int8 in sherpa-onnxhttps://huggingface.co/CohereLabs/cohere-transcribe-03-2026First non-Whisper/non-Parakeet ASR in the sherpa-onnx ecosystem (2B params, 14 langs, Apache 2.0, 1.6GB int8). Check for: distilled/smaller variants, WER benchmarks per language, community quality reports.2026-05-02
LiteRT-community TFLite ASR modelshttps://huggingface.co/litert-communityQwen3-ASR 0.6B and Parakeet CTC 0.6b in TFLite with Qualcomm NPU builds. Check for: new model additions, TDT (not just CTC) conversions, broader Qualcomm SoC support, community benchmarks.2026-05-02

Key Sources to Monitor

Account/OrgPlatformPublishes
csukuangfjHuggingFacesherpa-onnx pre-converted ASR models across ALL languages (the firehose; sort by lastModified)
k2-fsaGitHubsherpa-onnx releases with new model support
mrfakenameHuggingFaceWhisper ONNX conversions
nvidiaHuggingFaceParakeet model family
litert-communityHuggingFaceTFLite ASR models (NPU builds)
OBLITERATUSHuggingFaceQuantized GGUF models
microsoftGitHubONNX Runtime releases

Error Handling

  • HuggingFace rate limit: wait 2s and retry
  • GitHub rate limit: WebFetch the HTML releases page instead
  • crw or curl failure on one source: log and continue, never block the entire report
  • No new findings: state clearly, do not pad the report

Additional Resources

Reference Files
  • references/search-queries.md: CRW tool calls, search queries, and scope-specific search strategies
  • references/report-template.md: full report structure with per-model card format and recommendation table layout
  • references/model-inventory.md: detailed current model inventory with download URLs, file formats, and language coverage

© RisorseArtificiali, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .claude/skills/model-scout of RisorseArtificiali/anti-vocale.

  • SKILL.md
  • references/model-inventory.md
  • references/report-template.md
  • references/search-queries.md

Open the folder on GitHubat commit d2d157e

Compare with similar skills

Model Scout next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Scout compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Scout this skillRisorseArtificiali/anti-vocale118—~2.9kAutomated safety check: PassApache-2.0
Xybrid Initxybrid-ai/xybrid469—~3kAutomated safety check: PassApache-2.0
Test Modelxybrid-ai/xybrid469—~1.3kAutomated safety check: PassApache-2.0
Transcribe Anythingswyxio/skills176—~8.5kAutomated safety check: PassMIT
Parakeet Sttsundial-org/awesome-openclaw-skills663—~771Automated safety check: PassNone
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT

Similar skills

  • Xybrid Init

    xybrid-ai/xybrid

    Generate model metadata for an ML model so it works with xybrid.

    469 GitHub stars~3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Test Model

    xybrid-ai/xybrid

    Test a model end-to-end using the xybrid execution system. An agent skill from xybrid-ai/xybrid.

    469 GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Transcribe Anything

    swyxio/skills

    Transcribes audio and video files to text using pluggable ASR backends.

    176 GitHub stars~8.5k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Parakeet Stt

    sundial-org/awesome-openclaw-skills

    Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU).

    663 GitHub stars~771 tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Volcengine Asr

    ysyecust/lecture-to-notes

    Transcribe local audio or video with Volcengine Doubao file ASR, including BigASR 1.0 Turbo direct upload and asynchronous 1.0 standard, 1.0 idle, or 2.0 standard jobs through TOS.

    273 GitHub stars~783 tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed

More from RisorseArtificiali/anti-vocale

  • Community Model Conversion

    RisorseArtificiali/anti-vocale

    Convert a HuggingFace ASR fine-tune into a sherpa-onnx external model, publish it, and add it to the Anti-Vocale community catalog.

    118 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed

Questions about Model Scout

What does Model Scout do?

A skill your agent uses when the user asks to "scout for new models", "check for model updates", "find better ASR models", "check sherpa-onnx releases", "look for new Whisper models", "search…. Model Scout is an agent skill from RisorseArtificiali/anti-vocale. Use this skill when the user asks to "scout for new models", "check for model updates", "find better ASR models", "check sherpa-onnx releases", "look for new Whisper models", "search HuggingFace for ASR models", "run model scout", "any new Parakeet models", "check framework updates", "/model-scout", "find ASR models for <language", or when discussing improvements to on-device transcription quality for the Anti-Vocale app.

When should I use Model Scout?

Model Scout fits situations like: the user asks to scout for new models; check for model updates; find better ASR models; check sherpa-onnx releases.

How do I install Model Scout in Claude Code?

Run `npx skills add RisorseArtificiali/anti-vocale --skill model-scout -a claude-code`. Or copy the skill folder (.claude/skills/model-scout in RisorseArtificiali/anti-vocale) into .claude/skills/model-scout in your project. Claude Code loads it when a task matches its description.

How do I install Model Scout in Codex?

Run `npx skills add RisorseArtificiali/anti-vocale --skill model-scout -a codex`. Or copy the skill folder (.claude/skills/model-scout in RisorseArtificiali/anti-vocale) into .agents/skills/model-scout in your project. Codex loads it when a task matches its description.

Can I use Model Scout in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RisorseArtificiali/anti-vocale --skill model-scout -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-scout, .gemini/skills/model-scout, .github/skills/model-scout and .opencode/skills/model-scout in your project.

What does Model Scout need to run?

SKILL.md names no scripts, command-line tools or credentials: Model Scout is instructions for the agent only.

Does Model Scout access the network?

SKILL.md names 2 domains. As links in the text: huggingface.co and github.com. This is read from the text; nothing was executed.

Is Model Scout safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Scout use?

Model Scout is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Scout use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.

What are the alternatives to Model Scout?

Skills that share tags, products or a category with Model Scout: Xybrid Init (xybrid-ai/xybrid, 469 stars), Test Model (xybrid-ai/xybrid, 469 stars), Transcribe Anything (swyxio/skills, 176 stars) and Parakeet Stt (sundial-org/awesome-openclaw-skills, 663 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Scout?

RisorseArtificiali (a GitHub organization) maintains it in RisorseArtificiali/anti-vocale, which has 118 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 9, 2026.

Source: RisorseArtificiali/anti-vocale on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.