Official agent skill

Digital Health Clinical Asr Build

by NVIDIA in NVIDIA/skills

Stage 2 of the Clinical ASR Flywheel. An agent skill from NVIDIA/skills.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Digital Health Clinical Asr Build

skills CLI
$ npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills digital-health-clinical-asr-build --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/digital-health-clinical-asr-build .claude/skills/digital-health-clinical-asr-build && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
digital-health-clinical-asr-build
GitHub stars
3.6k
Token cost
~4.9k tokens
SKILL.md length
2,356 words
Files
7 (incl. references)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Stage 2 of the Clinical ASR Flywheel. An agent skill from NVIDIA/skills.

  • Works in 3 steps: What specialty / workflow is this for?… → What ASR failure modes have you seen? —… → Which terms come up daily vs which are…
  • Curating clinical terms
  • SKILL.md covers Data leaves your environment —…, Purpose, When to use this skill and Prerequisites, plus 7 more sections
  • Needs DICTIONARY_API_KEY and NVIDIA_API_KEY

What it does

Digital Health Clinical Asr Build is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Stage 2 of the Clinical ASR Flywheel. Use when curating clinical terms, tagging IPA, and synthesizing a NeMo manifest. NOT for scoring (use /digital-health-clinical-asr-eval).

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/manifest-schema.md`). Compatibility notes: NVIDIAAPIKEY (required) for hosted Magpie TTS via NVCF. DICTIONARYAPIKEY (optional) for Merriam-Webster Medical Dictionary lookup. Stage 1…

It sits in AI & LLM Engineering, covering Speech recognition and synthesis. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Curating clinical terms
  • Synthesizing a NeMo manifest

Example prompts

  • “/digital-health-clinical-asr-build”

Requirements

  • Python 3
  • Docker
  • A credential in NVIDIA_API_KEY
  • A credential in DICTIONARY_API_KEY
  • Compatibility (from SKILL.md): NVIDIA_API_KEY (required) for hosted Magpie TTS via NVCF. DICTIONARY_API_KEY (optional) for Merriam-Webster Medical Dictionary lookup. Stage 1 (/digital-health-clinical-asr-setup) must have been completed first. All TTS, IPA, and synthesis recipes are inlined — no sibling agent skill required.

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. What specialty / workflow is this for? (oncology dictation, ICU handoff, psych intake, ortho post-op, …)
  2. What ASR failure modes have you seen? — drug names, multi-word procedures, abbreviations, compound conditions.
  3. Which terms come up daily vs which are the hard ones? — daily-common terms become the sanity baseline; daily-hard terms become the signal.

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are csv).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DICTIONARY_API_KEY
    • NVIDIA_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    NVIDIA_API_KEY (required) for hosted Magpie TTS via NVCF. DICTIONARY_API_KEY (optional) for Merriam-Webster Medical Dictionary lookup. Stage 1 (/digital-health-clinical-asr-setup) must have been completed first. All TTS, IPA, and synthesis recipes are inlined — no sibling agent skill required.

    From compatibility in the SKILL.md frontmatter.

Context cost

Digital Health Clinical Asr Build loads about 4.9k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 2,356 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 2,356 words, ~4,906 tokens.

Download SKILL.mdSave it as .claude/skills/digital-health-clinical-asr-build/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
digital-health-clinical-asr-build
description
Stage 2 of the Clinical ASR Flywheel. Use when curating clinical terms, tagging IPA, and synthesizing a NeMo manifest. NOT for scoring (use /digital-health-clinical-asr-eval).
compatibility
NVIDIA_API_KEY (required) for hosted Magpie TTS via NVCF. DICTIONARY_API_KEY (optional) for Merriam-Webster Medical Dictionary lookup. Stage 1 (/digital-health-clinical-asr-setup) must have been completed first. All TTS, IPA, and synthesis recipes are inlined — no sibling agent skill required.
version
1.1.0
author
Ben Randoing <brandoing@nvidia.com>
tags
clinical-asr, dataset, ipa, magpie, nemo-manifest, flywheel
tools
Read, Write, Bash, Skill
license
Apache-2.0
metadata.author
Ben Randoing <brandoing@nvidia.com>
metadata.tags
clinical-asr, flywheel, dataset, ipa, magpie
metadata.team
healthcare-tme
metadata.domain
ai-ml
metadata.stage
2
<!--
SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: Apache-2.0
-->

Clinical ASR Flywheel — Stage 2 (Build the benchmark)

⚠ Agent: read this entire SKILL.md before answering. This stage is conversational and gated. Specifically: ask the user 1–2 specialty-aware clarifying questions before proposing terms (Step 2a), walk them through the two-tier IPA pipeline (override → merriam-webster → magpie_g2p) in Step 2c, hit the explicit QA-mode audition gate in Step 2d before full Cartesian synthesis, and name KER as the headline metric they'll see in Stage 3. Skipping any of these defeats the methodology.

You are the curate-and-synthesize stage. The user arrives from /digital-health-clinical-asr-setup and leaves with a NeMo-format manifest.jsonl plus the audio it references — both ready for scoring at /digital-health-clinical-asr-eval.

Be conversational. This is the warmest, most domain-aware step in the flywheel: you're asking a clinician (or someone who works with them) which terms hurt today and shaping a benchmark around their reality. Ask short, focused questions. Show the user what's being added. Don't lecture.

Data leaves your environment — disclose this to the user before any term is sent

This stage transmits user-curated content to two external services. Surface this to the user before invoking either call:

ServiceWhat gets sentWhen
Merriam-Webster (dictionaryapi.com API or merriam-webster.com public site)One HTTP request per term in the seed list — term goes in URL pathStep 2c — see MW path bullets below
NVIDIA NVCF Magpie TTS (grpc.nvcf.nvidia.com)Each generated clinical sentence (text, plus any SSML IPA wrappers)Steps 2d and 2e, every synthesis call

Both endpoints expect non-PHI synthetic content — the term list you curate, the sentences /data-designer (or your fallback templates) generates from it. Do not pass real patient records, real ASR transcripts, or any PHI through this skill. If the term list itself is sensitive (proprietary drug codenames, unreleased product names, customer-confidential indications), confirm with the user that external-API transmission is acceptable under their organization's data-governance policy before proceeding.

If no MW transmission is acceptable: take Path C below (skip MW; pipeline falls through to Magpie G2P with reduced coverage on long-tail terms).

Purpose

Curate a clinical-specialty term list, generate eval audio for it through Magpie TTS with a two-tier IPA pipeline, and write a NeMo-format manifest tagged with the clinical-extension fields (term, entity_category, ipa_source, voice_id, noise_level, context_type). The output is the input to Stage 3.

By the end the user has:

$EVAL_DIR/cycle<N>/
├── audio/<slug>.wav        synthesized clips
├── manifest.jsonl          NeMo format + clinical extension
├── term_seed.csv           the curated input
└── pronunciation_overrides.csv   appendable across cycles

($EVAL_DIR is the user's own choice — this skill does not impose a layout. The structure above is a recommendation, not a requirement.)

When to use this skill

Activate on user phrases like:

  • "Build a clinical ASR benchmark"
  • "Curate drug names / procedure names for ASR eval"
  • "Generate eval audio for medical terms"
  • "Create a NeMo manifest from clinical terms"
  • "Add oncology / cardiology / ortho terms to my benchmark"
  • "Audition the TTS pronunciation for these drug names"
  • "Make me a cycle-N manifest"

Do not activate when (also: if the message mentions auth, API key, gRPC, streaming, riva-build, NIM deploy, NGC, or Docker, route per the bullets below and stop):

  • The user already has a manifest and wants to score it → /digital-health-clinical-asr-eval
  • The user wants to fine-tune on an existing manifest → /digital-health-clinical-asr-finetune
  • The user is asking generic TTS / SSML / voice-cloning / voice-catalog questions → /read-aloud (or /riva-tts)
  • TTS/ASR auth / API keys / gRPC / streaming → /riva-tts or /riva-asr
  • NIM deploy or riva-build / riva-deploy flags → /riva-asr-custom or /riva-tts-custom
  • NGC / Docker / NVIDIA Container Toolkit → /riva-nim-setup
  • The user is asking generic synthetic-data questions → /data-designer

Prerequisites

  • /digital-health-clinical-asr-setup completed — NVIDIA_API_KEY exported, Python deps installed, the six upstream skills confirmed.
  • /read-aloud (or /riva-tts) reachable. Hosted Magpie via NVCF is the default. Self-hosted Magpie NIM works but adds /riva-nim-setup to the prerequisite chain.
  • /data-designer reachable. Template fallback is acceptable for a first cycle if /data-designer is unavailable, but tag those rows so future cycles can re-generate.
  • A working directory the user owns. The skill recommends $EVAL_DIR/cycle<N>/ but does not enforce it.

Instructions

2a. Specialty interview → term_seed.csv

Ask one question at a time. The goal is to surface 4–10 candidate terms with the right entity_category, not to write a textbook.

Questions, in order:

  1. What specialty / workflow is this for? (oncology dictation, ICU handoff, psych intake, ortho post-op, …)
  2. What ASR failure modes have you seen? — drug names, multi-word procedures, abbreviations, compound conditions.
  3. Which terms come up daily vs which are the hard ones? — daily-common terms become the sanity baseline; daily-hard terms become the signal.

Propose 4–10 candidate terms with entity_category. Confirm with the user before writing. Then write term_seed.csv:

csv
term,entity_category
cefazolin,drug
acetabular reamer,procedure
tibial plateau,anatomy
femoroacetabular impingement,condition
hemoglobin a1c,lab
respiratory therapist,role

The category vocabulary is fixed. KER keys off it. Allowed values:

drug | procedure | anatomy | condition | lab | role

If the user proposes a new category, push back: either it maps to one of the six, or the methodology needs a deliberate extension (which is a future cycle's job, not a one-off ad-hoc add).

2b. Sentence generation via /data-designer

Brief /data-designer with:

For each row in term_seed.csv, generate one or more natural English sentences embedding term in a way that fits the row's entity_category. Output schema: {term, entity_category, sentence, context_type}. Generate 3–5 context_type variants per term. Initial context_type vocabulary: dictation, handoff, chart_note, history. Sentence length 10–30 words.

The output of this step is a per-term sentence variants file. Any filename is fine — pick one and use it consistently across the cycle directory.

Template fallback. If /data-designer is unavailable, use a 4-template fallback (one per context_type) and substitute term mechanically. Tag those rows in the manifest (context_type is set, the sentence is just less natural) so a future cycle can regenerate.

2c. Two-tier IPA tagging (the load-bearing quality lever)

Every term passes through a 3-tier pipeline, in order:

  1. Override — pronunciation_overrides.csv carries verified IPA the team has audited. If term matches a row here, the override wins.
  2. Merriam-Webster — for un-overridden terms, fetch the MW respelling, convert to IPA, validate against Magpie's en-US phoneme set. If both succeed, the term is tagged merriam-webster.
  3. Magpie G2P (fall-through) — if neither override nor MW produces a valid IPA, the plain text is passed to Magpie's neural G2P at synthesis time. The row is tagged magpie_g2p.

Every manifest row carries the ipa_source tag (override | merriam-webster | magpie_g2p). The delta between merriam-webster and magpie_g2p rows in the Stage 3 leaderboard is the proof the pronunciation strategy is working — call it out explicitly when you produce the leaderboard.

Three MW lookup choices — all tag merriam-webster. A: dictionaryapi.com JSON API + DICTIONARY_API_KEY (free at dictionaryapi.com) — recommended for standalone use. B: HTML scrape of merriam-webster.com — no key, brittle to site HTML changes; recipe inlined in references/pronunciation-pipeline.md. C: skip MW, fall through to Magpie G2P with weaker long-tail coverage. Both recipes + the full respelling→IPA table live in references/pronunciation-pipeline.md. The Path A function takes api_key as an arg (never reads os.environ); pass None to skip MW.

pronunciation_overrides.csv schema:

csv
term,ipa,verified_by,verified_at,notes
cefazolin,sɛfəˈzoʊlɪn,brandoing,2026-05-13,confirmed against MW respelling + ear test

Append-only across cycles. Re-running the build later picks up new entries automatically.

2d. QA-mode synthesis (do not skip this gate)

Before running the full Cartesian product, synthesize one wav per term with: first voice, clean noise, default context. Audition each clip with the user.

For every term tagged magpie_g2p, propose an IPA candidate using clinical suffix patterns and validate against Magpie's en-US phoneme set before suggesting:

SuffixStress pattern (example)
-mycin…ˈmaɪsɪn (vancomycin, gentamicin)
-prazole…ˈpreɪzoʊl (esomeprazole, omeprazole)
-statin…ˈstætɪn (atorvastatin, rosuvastatin)
-sartan…ˈsɑːrtən (losartan, valsartan)
-azole…ˈeɪzoʊl (fluconazole, ketoconazole)
-cillin…ˈsɪlɪn (amoxicillin, piperacillin)
-parin…ˈpɛərɪn (enoxaparin, heparin)

Phoneme-validation pattern — live-probe Magpie's en-US neural G2P with a candidate IPA. If Magpie accepts the SSML, the IPA is in its inventory. Use the suffix patterns above as a pre-filter (cheap heuristic) and the live probe to confirm before committing to an override. The magpie_validates_ipa(ipa, api_key, voice_id) recipe — a minimal NVCF gRPC synthesis call that returns True/False fail-closed — is in references/pronunciation-pipeline.md.

Call it once per candidate IPA before showing it to the user. On user approval, append the verified IPA to pronunciation_overrides.csv. The row's ipa_source flips from magpie_g2p to override on the next manifest generation.

HITL audition gate before Step 2e — fail-closed. Do not synthesize the full Cartesian product, do not promote any staged IPA candidate to pronunciation_overrides.csv, and do not advance to Stage 3 until one of the following has happened explicitly in conversation:

  1. The user confirms they have auditioned the QA clips and reports their verdict per clip (or per bucket: "the MW set sounds fine", "fix pembrolizumab", etc.). Provide the afplay (macOS) or paplay/aplay (Linux) commands so the user can play them — then halt and wait for their reply after listening. Paper-only approval via an AskUserQuestion prompt — clicking "Promote all" or "Lock in" without auditioning — does not satisfy this gate. Magpie-validating an IPA proves it's in the phoneme inventory; it does not prove it matches the intended pronunciation. Only the user's ears do that.
  2. The user explicitly opts to skip audition for this cycle, in deliberate language (e.g. "skip audition, accept the risk that mispronunciations may dilute the Stage 3 KER signal — log it as a cycle-N caveat"), not as a side-effect of a single click-through. Record the skip in a cycle-level note (e.g. eval/cycle<N>/cycle_notes.md) so a future operator can see the audition was deferred.

Magpie NVCF rate-limits aggressively on >100-row jobs, and a do-over costs both API credits and clock time — but the larger risk is shipping a manifest with mispronounced reference audio that quietly corrupts the Stage 3 KER signal. Time spent auditioning is cheaper than re-running the cycle.

Show full SKILL.md (831 more words)Show less
2e. Full benchmark generation

After pronunciations are locked, generate the full Cartesian product |terms| × |voices| × |noise_levels| × |context_types|. Defaults: 2–4 Magpie en-US voices (Mia/Jason/Ray), [clean, snr_15db, snr_5db], [dictation, handoff, chart_note, history].

Self-contained synthesis — no /read-aloud required. The synthesize_row(row, all_overrides, out_dir, api_key) recipe — opens an NVCF gRPC stream, wraps overrides into SSML via render_sentence_with_overrides, writes 16-bit mono PCM to <out_dir>/audio/<slug>.wav — is in references/pronunciation-pipeline.md (§Synthesis call). Key invariant: all_overrides carries every entry from pronunciation_overrides.csv (including context-word overrides like intravenously) so the renderer wraps any override whose verbatim text appears in row['text']. Wrapping only row['term'] silently drops context-word overrides.

Noise-injection (clean → snr_15db → snr_5db) and the manifest schema (NeMo canonical fields + clinical extension, plus pre-flight schema and audio-existence checks) all live in references/manifest-schema.md.

Warn when product > 100 rows. Magpie NVCF rate-limits with ~5–10% RESOURCE_EXHAUSTED drops on big runs. Re-run the dropped rows.

Stage 2 completion checklist

Don't consider Stage 2 done until all five sub-steps ran. Agents commonly stop after 2a or 2b; the goal is a synthesized manifest plus a hand-off:

  • 2a — term_seed.csv, 4–10 terms, entity_category ∈ {drug, procedure, anatomy, condition, lab, role}
  • 2b — 3–5 context_type sentence variants per term
  • 2c — every term tagged ipa_source ∈ {override, merriam-webster, magpie_g2p}
  • 2d — QA wavs auditioned, IPA overrides locked with explicit user approval
  • 2e — manifest.jsonl + per-row audio for the Cartesian product
  • Hand-off — name /digital-health-clinical-asr-eval as the next skill and KER as its headline metric

Writes go only into the user-chosen $EVAL_DIR/cycle<N>/. Don't write elsewhere, modify env, or install packages — those belong to /digital-health-clinical-asr-setup.

Examples

Scenario A — fresh oncology benchmark. User: "We're seeing chemo drug names mistranscribed. Where do I start?" → Step 2a: confirm specialty is oncology, ask about which drugs (immunotherapy biologics, platinum agents, taxanes). Propose ~10 candidates: cisplatin, paclitaxel, pembrolizumab, nivolumab, carboplatin, docetaxel, bevacizumab, trastuzumab, cetuximab, pemetrexed. Write term_seed.csv with all entity_category=drug. Step 2b: brief /data-designer for 4 context variants each = 40 sentences. Step 2c: MW lookup for each — biologics like pembrolizumab will likely fall to magpie_g2p; platinum agents likely hit MW. Step 2d: synthesize one QA wav per term, walk the user through the pembrolizumab etc. clips, propose IPA candidates with -mab suffix stress patterns. Step 2e: on approval, run 10 terms × 2 voices × 2 noise levels × 3 contexts = 120 rows.

Scenario B — appending to an existing cycle. User: "I have a cycle-1 manifest and I want to add 5 more procedures." → Re-run only Steps 2a (specialty interview just for the new terms), 2b (sentence gen for the additions), 2c (IPA pipeline for the additions), 2d (audition the new terms), and 2e (synthesize only the new term rows). Append to the existing manifest.jsonl. Do not regenerate audio for existing terms — cycle isolation is intentional so leaderboards diff cycle N vs cycle N+1 cleanly.

Artifacts produced

  • term_seed.csv — curated terms with entity_category
  • pronunciation_overrides.csv — verified IPA, appendable across cycles
  • manifest.jsonl — NeMo format with clinical extension fields (one JSON object per line)
  • audio/<slug>.wav — synthesized clips, one per manifest row

Troubleshooting

  • TTS rate-limit drops (RESOURCE_EXHAUSTED) on >100-row generation → expected on Magpie NVCF. Confirm exponential backoff is active in /read-aloud; expect ~5–10% drops on big runs and re-run for the gaps.
  • All ipa_source rows tagged magpie_g2p → MW lookup is failing across the board, or candidate IPAs are failing phoneme validation. Re-verify whichever MW path you configured (DICTIONARY_API_KEY for A; HTTPS reachability + parser for B), then check candidates against Magpie's en-US phoneme inventory.
  • Magpie mispronounces a term even with the IPA override → first verify the IPA is in the Magpie en-US phoneme inventory and the SSML wrapping is syntactically valid. If both check out, the underlying TTS bug is owned by /read-aloud (/riva-tts) — route there for diagnosis. This skill provides the override mechanism but does not own the neural G2P or SSML parser.
  • Sentence variants from /data-designer are bland / template-like → check the brief; the schema-only prompt sometimes produces stereotyped output. Add 1–2 in-context examples to the brief and re-run.
  • Audio files exist but manifest.jsonl is short → manifest writer skipped rows whose synthesis returned a NVCF error. Re-run the build with only the missing rows.

For anything not in this list, identify which upstream skill is implicated and route there. The digital-health-clinical-asr-build skill owns the methodology, not the TTS or DataDesigner internals.

Limitations

  • English-only by default. Magpie's en-US phoneme inventory is what the two-tier IPA pipeline validates against. Other locales need a different upstream phoneme set + override CSV format.
  • Six fixed entity categories. Extending entity_category is a deliberate methodology change, not a one-off tweak — KER breakdowns, leaderboard sections, and downstream finetune scripts all key off the vocabulary.
  • Tiny first cycles. Below ~20 terms, the by-ipa_source leaderboard split won't have enough rows in each bucket to be statistically meaningful. Build a meaningful cycle even if it costs a session.
  • Magpie NVCF rate-limits. ~5–10% drops on large jobs; budget a re-run pass.

Next steps

  • Forward: /digital-health-clinical-asr-eval — transcribe the manifest, score WER/CER/KER/SER, produce the five-section leaderboard.
  • Back to setup (if anything in the env is broken): /digital-health-clinical-asr-setup.
  • Lateral for TTS-specific debugging: /read-aloud or /riva-tts.

References

  • references/manifest-schema.md — NeMo canonical fields + clinical extension; pre-flight schema and audio-existence checks; cross-cycle stability rules

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in skills/digital-health-clinical-asr-build of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • references/manifest-schema.md
  • references/pronunciation-pipeline.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Digital Health Clinical Asr Build next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Digital Health Clinical Asr Build compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Digital Health Clinical Asr Build this skillNVIDIA/skills3.6k—~4.9kAutomated safety check: PassApache-2.0
Parakeet Sttsundial-org/awesome-openclaw-skills663—~771Automated safety check: PassNone
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT
TriageTalAter/annyang6.8k—~810Automated safety check: NotesMIT
Yichen Asrmcncarl/yichen-skills4.4k—~780Automated safety check: PassCustom licence
Dingtalk MinutesDingTalk-Real-AI/dingtalk-workspace-cli3.2k—~2.3kAutomated safety check: PassApache-2.0

Similar skills

  • Parakeet Stt

    sundial-org/awesome-openclaw-skills

    Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU).

    663 GitHub stars~771 tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Triage

    TalAter/annyang

    Triage and close GitHub issues on TalAter/annyang. An agent skill from TalAter/annyang.

    6.8k GitHub stars~810 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes
  • Yichen Asr

    mcncarl/yichen-skills

    逸尘自用的统一音视频转写入口,在 StepFun Step ASR 与火山引擎豆包 ASR 之间按输出需求、安全边界和可用状态路由。用于本地音频或视频的纯文本转写、时间戳、SRT 字幕、口播粗剪,以及转写前体检;用户明确指定服务商时不得静默切换。Use when a local audio or video file needs transcription and the correct…

    4.4k GitHub stars~780 tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Dingtalk Minutes

    DingTalk-Real-AI/dingtalk-workspace-cli

    钉钉 AI 听记。Use when 查询或修改听记摘要、完整逐字稿、关键词、标签、行动项、录音、上传、思维导图、发言人洞察、ASR 热词/识别词配置或分享权限。写文档走 dingtalk-doc;建待办走 dingtalk-todo;日程走 dingtalk-calendar。命令前缀:dws minutes。

    3.2k GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Youtube Fetcher

    JimmySadek/youtube-fetcher-to-markdown

    Retrieve transcripts from YouTube, Instagram, TikTok, X, Vimeo and other video sites, summarize or analyze what was said (and shown on screen), or save an Obsidian-ready Markdown knowledge-base note…

    485 GitHub stars~3.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Digital Health Clinical Asr Build

What does Digital Health Clinical Asr Build do?

Stage 2 of the Clinical ASR Flywheel. An agent skill from NVIDIA/skills. Digital Health Clinical Asr Build is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Stage 2 of the Clinical ASR Flywheel.

When should I use Digital Health Clinical Asr Build?

Digital Health Clinical Asr Build fits situations like: curating clinical terms; synthesizing a NeMo manifest.

How do I install Digital Health Clinical Asr Build in Claude Code?

Run `npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a claude-code`. Or copy the skill folder (skills/digital-health-clinical-asr-build in NVIDIA/skills) into .claude/skills/digital-health-clinical-asr-build in your project. Claude Code loads it when a task matches its description.

How do I install Digital Health Clinical Asr Build in Codex?

Run `npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a codex`. Or copy the skill folder (skills/digital-health-clinical-asr-build in NVIDIA/skills) into .agents/skills/digital-health-clinical-asr-build in your project. Codex loads it when a task matches its description.

Can I use Digital Health Clinical Asr Build in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill digital-health-clinical-asr-build -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/digital-health-clinical-asr-build, .gemini/skills/digital-health-clinical-asr-build, .github/skills/digital-health-clinical-asr-build and .opencode/skills/digital-health-clinical-asr-build in your project.

What does Digital Health Clinical Asr Build need to run?

Going by SKILL.md and its folder, Digital Health Clinical Asr Build needs credentials named DICTIONARY_API_KEY and NVIDIA_API_KEY. Our summary lists: Python 3; Docker; A credential in NVIDIA_API_KEY; A credential in DICTIONARY_API_KEY. Compatibility (from SKILL.md): NVIDIA_API_KEY (required) for hosted Magpie TTS via NVCF. DICTIONARY_API_KEY (optional) for Merriam-Webster Medical Dictionary lookup. Stage 1 (/digital-health-clinical-asr-setup) must have been completed first. All TTS, IPA, and synthesis recipes are inlined — no sibling agent skill required..

Does Digital Health Clinical Asr Build access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Digital Health Clinical Asr Build safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Digital Health Clinical Asr Build use?

Digital Health Clinical Asr Build is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Digital Health Clinical Asr Build use?

About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.

What are the alternatives to Digital Health Clinical Asr Build?

Skills that share tags, products or a category with Digital Health Clinical Asr Build: Parakeet Stt (sundial-org/awesome-openclaw-skills, 663 stars), 9Router Speech-to-Text (decolua/9router, 31k stars), Triage (TalAter/annyang, 6.8k stars) and Yichen Asr (mcncarl/yichen-skills, 4.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Digital Health Clinical Asr Build?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.