Parakeet Stt
sundial-org/awesome-openclaw-skills
Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU).
Stage 2 of the Clinical ASR Flywheel. An agent skill from NVIDIA/skills.
$ npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills digital-health-clinical-asr-build --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/digital-health-clinical-asr-build .claude/skills/digital-health-clinical-asr-build && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "digital-health-clinical-asr-build" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-build into .claude/skills/digital-health-clinical-asr-build/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "digital-health-clinical-asr-build", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-buildType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills digital-health-clinical-asr-build --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/digital-health-clinical-asr-build .agents/skills/digital-health-clinical-asr-build && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "digital-health-clinical-asr-build" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-build into .agents/skills/digital-health-clinical-asr-build/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "digital-health-clinical-asr-build", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills digital-health-clinical-asr-build --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/digital-health-clinical-asr-build .cursor/skills/digital-health-clinical-asr-build && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "digital-health-clinical-asr-build" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-build into .cursor/skills/digital-health-clinical-asr-build/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "digital-health-clinical-asr-build", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/digital-health-clinical-asr-build--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills digital-health-clinical-asr-build --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/digital-health-clinical-asr-build .gemini/skills/digital-health-clinical-asr-build && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "digital-health-clinical-asr-build" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-build into .gemini/skills/digital-health-clinical-asr-build/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "digital-health-clinical-asr-build", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills digital-health-clinical-asr-buildInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/digital-health-clinical-asr-build .github/skills/digital-health-clinical-asr-build && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "digital-health-clinical-asr-build" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-build into .github/skills/digital-health-clinical-asr-build/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "digital-health-clinical-asr-build", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills digital-health-clinical-asr-build --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/digital-health-clinical-asr-build .opencode/skills/digital-health-clinical-asr-build && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "digital-health-clinical-asr-build" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-build into .opencode/skills/digital-health-clinical-asr-build/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "digital-health-clinical-asr-build", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
digital-health-clinical-asr-buildStage 2 of the Clinical ASR Flywheel. An agent skill from NVIDIA/skills.
Digital Health Clinical Asr Build is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Stage 2 of the Clinical ASR Flywheel. Use when curating clinical terms, tagging IPA, and synthesizing a NeMo manifest. NOT for scoring (use /digital-health-clinical-asr-eval).
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/manifest-schema.md`). Compatibility notes: NVIDIAAPIKEY (required) for hosted Magpie TTS via NVCF. DICTIONARYAPIKEY (optional) for Merriam-Webster Medical Dictionary lookup. Stage 1…
It sits in AI & LLM Engineering, covering Speech recognition and synthesis. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are csv).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DICTIONARY_API_KEYNVIDIA_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
NVIDIA_API_KEY (required) for hosted Magpie TTS via NVCF. DICTIONARY_API_KEY (optional) for Merriam-Webster Medical Dictionary lookup. Stage 1 (/digital-health-clinical-asr-setup) must have been completed first. All TTS, IPA, and synthesis recipes are inlined — no sibling agent skill required.
From compatibility in the SKILL.md frontmatter.
Digital Health Clinical Asr Build loads about 4.9k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 2,356 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 2,356 words, ~4,906 tokens.
.claude/skills/digital-health-clinical-asr-build/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.<!--
SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: Apache-2.0
-->
⚠ Agent: read this entire SKILL.md before answering. This stage is conversational and gated. Specifically: ask the user 1–2 specialty-aware clarifying questions before proposing terms (Step 2a), walk them through the two-tier IPA pipeline (override → merriam-webster → magpie_g2p) in Step 2c, hit the explicit QA-mode audition gate in Step 2d before full Cartesian synthesis, and name KER as the headline metric they'll see in Stage 3. Skipping any of these defeats the methodology.
You are the curate-and-synthesize stage. The user arrives from /digital-health-clinical-asr-setup and leaves with a NeMo-format manifest.jsonl plus the audio it references — both ready for scoring at /digital-health-clinical-asr-eval.
Be conversational. This is the warmest, most domain-aware step in the flywheel: you're asking a clinician (or someone who works with them) which terms hurt today and shaping a benchmark around their reality. Ask short, focused questions. Show the user what's being added. Don't lecture.
This stage transmits user-curated content to two external services. Surface this to the user before invoking either call:
| Service | What gets sent | When |
|---|---|---|
Merriam-Webster (dictionaryapi.com API or merriam-webster.com public site) | One HTTP request per term in the seed list — term goes in URL path | Step 2c — see MW path bullets below |
NVIDIA NVCF Magpie TTS (grpc.nvcf.nvidia.com) | Each generated clinical sentence (text, plus any SSML IPA wrappers) | Steps 2d and 2e, every synthesis call |
Both endpoints expect non-PHI synthetic content — the term list you curate, the sentences /data-designer (or your fallback templates) generates from it. Do not pass real patient records, real ASR transcripts, or any PHI through this skill. If the term list itself is sensitive (proprietary drug codenames, unreleased product names, customer-confidential indications), confirm with the user that external-API transmission is acceptable under their organization's data-governance policy before proceeding.
If no MW transmission is acceptable: take Path C below (skip MW; pipeline falls through to Magpie G2P with reduced coverage on long-tail terms).
Curate a clinical-specialty term list, generate eval audio for it through Magpie TTS with a two-tier IPA pipeline, and write a NeMo-format manifest tagged with the clinical-extension fields (term, entity_category, ipa_source, voice_id, noise_level, context_type). The output is the input to Stage 3.
By the end the user has:
$EVAL_DIR/cycle<N>/
├── audio/<slug>.wav synthesized clips
├── manifest.jsonl NeMo format + clinical extension
├── term_seed.csv the curated input
└── pronunciation_overrides.csv appendable across cycles($EVAL_DIR is the user's own choice — this skill does not impose a layout. The structure above is a recommendation, not a requirement.)
Activate on user phrases like:
Do not activate when (also: if the message mentions auth, API key, gRPC, streaming, riva-build, NIM deploy, NGC, or Docker, route per the bullets below and stop):
/digital-health-clinical-asr-eval/digital-health-clinical-asr-finetune/read-aloud (or /riva-tts)/riva-tts or /riva-asrriva-build / riva-deploy flags → /riva-asr-custom or /riva-tts-custom/riva-nim-setup/data-designer/digital-health-clinical-asr-setup completed — NVIDIA_API_KEY exported, Python deps installed, the six upstream skills confirmed./read-aloud (or /riva-tts) reachable. Hosted Magpie via NVCF is the default. Self-hosted Magpie NIM works but adds /riva-nim-setup to the prerequisite chain./data-designer reachable. Template fallback is acceptable for a first cycle if /data-designer is unavailable, but tag those rows so future cycles can re-generate.$EVAL_DIR/cycle<N>/ but does not enforce it.term_seed.csvAsk one question at a time. The goal is to surface 4–10 candidate terms with the right entity_category, not to write a textbook.
Questions, in order:
Propose 4–10 candidate terms with entity_category. Confirm with the user before writing. Then write term_seed.csv:
term,entity_category
cefazolin,drug
acetabular reamer,procedure
tibial plateau,anatomy
femoroacetabular impingement,condition
hemoglobin a1c,lab
respiratory therapist,roleThe category vocabulary is fixed. KER keys off it. Allowed values:
drug | procedure | anatomy | condition | lab | roleIf the user proposes a new category, push back: either it maps to one of the six, or the methodology needs a deliberate extension (which is a future cycle's job, not a one-off ad-hoc add).
/data-designerBrief /data-designer with:
For each row in
term_seed.csv, generate one or more natural English sentences embeddingtermin a way that fits the row'sentity_category. Output schema:{term, entity_category, sentence, context_type}. Generate 3–5context_typevariants per term. Initialcontext_typevocabulary:dictation,handoff,chart_note,history. Sentence length 10–30 words.
The output of this step is a per-term sentence variants file. Any filename is fine — pick one and use it consistently across the cycle directory.
Template fallback. If /data-designer is unavailable, use a 4-template fallback (one per context_type) and substitute term mechanically. Tag those rows in the manifest (context_type is set, the sentence is just less natural) so a future cycle can regenerate.
Every term passes through a 3-tier pipeline, in order:
pronunciation_overrides.csv carries verified IPA the team has audited. If term matches a row here, the override wins.merriam-webster.magpie_g2p.Every manifest row carries the ipa_source tag (override | merriam-webster | magpie_g2p). The delta between merriam-webster and magpie_g2p rows in the Stage 3 leaderboard is the proof the pronunciation strategy is working — call it out explicitly when you produce the leaderboard.
Three MW lookup choices — all tag merriam-webster. A: dictionaryapi.com JSON API + DICTIONARY_API_KEY (free at dictionaryapi.com) — recommended for standalone use. B: HTML scrape of merriam-webster.com — no key, brittle to site HTML changes; recipe inlined in references/pronunciation-pipeline.md. C: skip MW, fall through to Magpie G2P with weaker long-tail coverage. Both recipes + the full respelling→IPA table live in references/pronunciation-pipeline.md. The Path A function takes api_key as an arg (never reads os.environ); pass None to skip MW.
pronunciation_overrides.csv schema:
term,ipa,verified_by,verified_at,notes
cefazolin,sɛfəˈzoʊlɪn,brandoing,2026-05-13,confirmed against MW respelling + ear testAppend-only across cycles. Re-running the build later picks up new entries automatically.
Before running the full Cartesian product, synthesize one wav per term with: first voice, clean noise, default context. Audition each clip with the user.
For every term tagged magpie_g2p, propose an IPA candidate using clinical suffix patterns and validate against Magpie's en-US phoneme set before suggesting:
| Suffix | Stress pattern (example) |
|---|---|
-mycin | …ˈmaɪsɪn (vancomycin, gentamicin) |
-prazole | …ˈpreɪzoʊl (esomeprazole, omeprazole) |
-statin | …ˈstætɪn (atorvastatin, rosuvastatin) |
-sartan | …ˈsɑːrtən (losartan, valsartan) |
-azole | …ˈeɪzoʊl (fluconazole, ketoconazole) |
-cillin | …ˈsɪlɪn (amoxicillin, piperacillin) |
-parin | …ˈpɛərɪn (enoxaparin, heparin) |
Phoneme-validation pattern — live-probe Magpie's en-US neural G2P with a candidate IPA. If Magpie accepts the SSML, the IPA is in its inventory. Use the suffix patterns above as a pre-filter (cheap heuristic) and the live probe to confirm before committing to an override. The magpie_validates_ipa(ipa, api_key, voice_id) recipe — a minimal NVCF gRPC synthesis call that returns True/False fail-closed — is in references/pronunciation-pipeline.md.
Call it once per candidate IPA before showing it to the user. On user approval, append the verified IPA to pronunciation_overrides.csv. The row's ipa_source flips from magpie_g2p to override on the next manifest generation.
HITL audition gate before Step 2e — fail-closed. Do not synthesize the full Cartesian product, do not promote any staged IPA candidate to pronunciation_overrides.csv, and do not advance to Stage 3 until one of the following has happened explicitly in conversation:
pembrolizumab", etc.). Provide the afplay (macOS) or paplay/aplay (Linux) commands so the user can play them — then halt and wait for their reply after listening. Paper-only approval via an AskUserQuestion prompt — clicking "Promote all" or "Lock in" without auditioning — does not satisfy this gate. Magpie-validating an IPA proves it's in the phoneme inventory; it does not prove it matches the intended pronunciation. Only the user's ears do that.eval/cycle<N>/cycle_notes.md) so a future operator can see the audition was deferred.Magpie NVCF rate-limits aggressively on >100-row jobs, and a do-over costs both API credits and clock time — but the larger risk is shipping a manifest with mispronounced reference audio that quietly corrupts the Stage 3 KER signal. Time spent auditioning is cheaper than re-running the cycle.
After pronunciations are locked, generate the full Cartesian product |terms| × |voices| × |noise_levels| × |context_types|. Defaults: 2–4 Magpie en-US voices (Mia/Jason/Ray), [clean, snr_15db, snr_5db], [dictation, handoff, chart_note, history].
Self-contained synthesis — no /read-aloud required. The synthesize_row(row, all_overrides, out_dir, api_key) recipe — opens an NVCF gRPC stream, wraps overrides into SSML via render_sentence_with_overrides, writes 16-bit mono PCM to <out_dir>/audio/<slug>.wav — is in references/pronunciation-pipeline.md (§Synthesis call). Key invariant: all_overrides carries every entry from pronunciation_overrides.csv (including context-word overrides like intravenously) so the renderer wraps any override whose verbatim text appears in row['text']. Wrapping only row['term'] silently drops context-word overrides.
Noise-injection (clean → snr_15db → snr_5db) and the manifest schema (NeMo canonical fields + clinical extension, plus pre-flight schema and audio-existence checks) all live in references/manifest-schema.md.
Warn when product > 100 rows. Magpie NVCF rate-limits with ~5–10% RESOURCE_EXHAUSTED drops on big runs. Re-run the dropped rows.
Don't consider Stage 2 done until all five sub-steps ran. Agents commonly stop after 2a or 2b; the goal is a synthesized manifest plus a hand-off:
term_seed.csv, 4–10 terms, entity_category ∈ {drug, procedure, anatomy, condition, lab, role}context_type sentence variants per termipa_source ∈ {override, merriam-webster, magpie_g2p}manifest.jsonl + per-row audio for the Cartesian product/digital-health-clinical-asr-eval as the next skill and KER as its headline metricWrites go only into the user-chosen $EVAL_DIR/cycle<N>/. Don't write elsewhere, modify env, or install packages — those belong to /digital-health-clinical-asr-setup.
Scenario A — fresh oncology benchmark. User: "We're seeing chemo drug names mistranscribed. Where do I start?" → Step 2a: confirm specialty is oncology, ask about which drugs (immunotherapy biologics, platinum agents, taxanes). Propose ~10 candidates: cisplatin, paclitaxel, pembrolizumab, nivolumab, carboplatin, docetaxel, bevacizumab, trastuzumab, cetuximab, pemetrexed. Write term_seed.csv with all entity_category=drug. Step 2b: brief /data-designer for 4 context variants each = 40 sentences. Step 2c: MW lookup for each — biologics like pembrolizumab will likely fall to magpie_g2p; platinum agents likely hit MW. Step 2d: synthesize one QA wav per term, walk the user through the pembrolizumab etc. clips, propose IPA candidates with -mab suffix stress patterns. Step 2e: on approval, run 10 terms × 2 voices × 2 noise levels × 3 contexts = 120 rows.
Scenario B — appending to an existing cycle. User: "I have a cycle-1 manifest and I want to add 5 more procedures." → Re-run only Steps 2a (specialty interview just for the new terms), 2b (sentence gen for the additions), 2c (IPA pipeline for the additions), 2d (audition the new terms), and 2e (synthesize only the new term rows). Append to the existing manifest.jsonl. Do not regenerate audio for existing terms — cycle isolation is intentional so leaderboards diff cycle N vs cycle N+1 cleanly.
term_seed.csv — curated terms with entity_categorypronunciation_overrides.csv — verified IPA, appendable across cyclesmanifest.jsonl — NeMo format with clinical extension fields (one JSON object per line)audio/<slug>.wav — synthesized clips, one per manifest rowRESOURCE_EXHAUSTED) on >100-row generation → expected on Magpie NVCF. Confirm exponential backoff is active in /read-aloud; expect ~5–10% drops on big runs and re-run for the gaps.ipa_source rows tagged magpie_g2p → MW lookup is failing across the board, or candidate IPAs are failing phoneme validation. Re-verify whichever MW path you configured (DICTIONARY_API_KEY for A; HTTPS reachability + parser for B), then check candidates against Magpie's en-US phoneme inventory./read-aloud (/riva-tts) — route there for diagnosis. This skill provides the override mechanism but does not own the neural G2P or SSML parser./data-designer are bland / template-like → check the brief; the schema-only prompt sometimes produces stereotyped output. Add 1–2 in-context examples to the brief and re-run.manifest.jsonl is short → manifest writer skipped rows whose synthesis returned a NVCF error. Re-run the build with only the missing rows.For anything not in this list, identify which upstream skill is implicated and route there. The digital-health-clinical-asr-build skill owns the methodology, not the TTS or DataDesigner internals.
entity_category is a deliberate methodology change, not a one-off tweak — KER breakdowns, leaderboard sections, and downstream finetune scripts all key off the vocabulary.ipa_source leaderboard split won't have enough rows in each bucket to be statistically meaningful. Build a meaningful cycle even if it costs a session./digital-health-clinical-asr-eval — transcribe the manifest, score WER/CER/KER/SER, produce the five-section leaderboard./digital-health-clinical-asr-setup./read-aloud or /riva-tts.references/manifest-schema.md — NeMo canonical fields + clinical extension; pre-flight schema and audio-existence checks; cross-cycle stability rules© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (references) in skills/digital-health-clinical-asr-build of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Digital Health Clinical Asr Build next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Digital Health Clinical Asr Build this skillNVIDIA/skills | 3.6k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| Parakeet Sttsundial-org/awesome-openclaw-skills | 663 | — | ~771 | Automated safety check: Pass | None | |
| 9Router Speech-to-Textdecolua/9router | 31k | — | ~914 | Automated safety check: Pass | MIT | |
| TriageTalAter/annyang | 6.8k | — | ~810 | Automated safety check: Notes | MIT | |
| Yichen Asrmcncarl/yichen-skills | 4.4k | — | ~780 | Automated safety check: Pass | Custom licence | |
| Dingtalk MinutesDingTalk-Real-AI/dingtalk-workspace-cli | 3.2k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 |
sundial-org/awesome-openclaw-skills
Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU).
decolua/9router
Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.
TalAter/annyang
Triage and close GitHub issues on TalAter/annyang. An agent skill from TalAter/annyang.
mcncarl/yichen-skills
逸尘自用的统一音视频转写入口,在 StepFun Step ASR 与火山引擎豆包 ASR 之间按输出需求、安全边界和可用状态路由。用于本地音频或视频的纯文本转写、时间戳、SRT 字幕、口播粗剪,以及转写前体检;用户明确指定服务商时不得静默切换。Use when a local audio or video file needs transcription and the correct…
DingTalk-Real-AI/dingtalk-workspace-cli
钉钉 AI 听记。Use when 查询或修改听记摘要、完整逐字稿、关键词、标签、行动项、录音、上传、思维导图、发言人洞察、ASR 热词/识别词配置或分享权限。写文档走 dingtalk-doc;建待办走 dingtalk-todo;日程走 dingtalk-calendar。命令前缀:dws minutes。
JimmySadek/youtube-fetcher-to-markdown
Retrieve transcripts from YouTube, Instagram, TikTok, X, Vimeo and other video sites, summarize or analyze what was said (and shown on screen), or save an Obsidian-ready Markdown knowledge-base note…
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Stage 2 of the Clinical ASR Flywheel. An agent skill from NVIDIA/skills. Digital Health Clinical Asr Build is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Stage 2 of the Clinical ASR Flywheel.
Digital Health Clinical Asr Build fits situations like: curating clinical terms; synthesizing a NeMo manifest.
Run `npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a claude-code`. Or copy the skill folder (skills/digital-health-clinical-asr-build in NVIDIA/skills) into .claude/skills/digital-health-clinical-asr-build in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a codex`. Or copy the skill folder (skills/digital-health-clinical-asr-build in NVIDIA/skills) into .agents/skills/digital-health-clinical-asr-build in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/digital-health-clinical-asr-build, .gemini/skills/digital-health-clinical-asr-build, .github/skills/digital-health-clinical-asr-build and .opencode/skills/digital-health-clinical-asr-build in your project.
Going by SKILL.md and its folder, Digital Health Clinical Asr Build needs credentials named DICTIONARY_API_KEY and NVIDIA_API_KEY. Our summary lists: Python 3; Docker; A credential in NVIDIA_API_KEY; A credential in DICTIONARY_API_KEY. Compatibility (from SKILL.md): NVIDIA_API_KEY (required) for hosted Magpie TTS via NVCF. DICTIONARY_API_KEY (optional) for Merriam-Webster Medical Dictionary lookup. Stage 1 (/digital-health-clinical-asr-setup) must have been completed first. All TTS, IPA, and synthesis recipes are inlined — no sibling agent skill required..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Digital Health Clinical Asr Build is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Digital Health Clinical Asr Build: Parakeet Stt (sundial-org/awesome-openclaw-skills, 663 stars), 9Router Speech-to-Text (decolua/9router, 31k stars), Triage (TalAter/annyang, 6.8k stars) and Yichen Asr (mcncarl/yichen-skills, 4.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.