Music to Video
heygen-com/hyperframes
Turns a music track into a beat-synced HyperFrames video such as a lyric video, slideshow or kinetic promo, with any supplied images or clips cut onto the beat grid.
Source and VERIFY rich external media (music, high-value real photos/video) and the latest live status that the Detective's background pass didn't cover.
$ npx skills add QinghongLin/data2story-skill --skill scout -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install QinghongLin/data2story-skill scout --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data2story-pro/scout .claude/skills/scout && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scout" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/data2story-pro/scout into .claude/skills/scout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scout", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/QinghongLin/data2story-skill/tree/main/skills/data2story-pro/scoutType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add QinghongLin/data2story-skill --skill scout -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install QinghongLin/data2story-skill scout --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/data2story-pro/scout .agents/skills/scout && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scout" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/data2story-pro/scout into .agents/skills/scout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scout", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QinghongLin/data2story-skill --skill scout -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install QinghongLin/data2story-skill scout --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/data2story-pro/scout .cursor/skills/scout && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scout" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/data2story-pro/scout into .cursor/skills/scout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scout", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/QinghongLin/data2story-skill.git --path skills/data2story-pro/scout--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add QinghongLin/data2story-skill --skill scout -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install QinghongLin/data2story-skill scout --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/data2story-pro/scout .gemini/skills/scout && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scout" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/data2story-pro/scout into .gemini/skills/scout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scout", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install QinghongLin/data2story-skill scoutInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add QinghongLin/data2story-skill --skill scout -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/data2story-pro/scout .github/skills/scout && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scout" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/data2story-pro/scout into .github/skills/scout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scout", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QinghongLin/data2story-skill --skill scout -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install QinghongLin/data2story-skill scout --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QinghongLin/data2story-skill.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/data2story-pro/scout .opencode/skills/scout && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scout" agent skill from https://github.com/QinghongLin/data2story-skill/tree/main/skills/data2story-pro/scout into .opencode/skills/scout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scout", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scoutSource and VERIFY rich external media (music, high-value real photos/video) and the latest live status that the Detective's background pass didn't cover.
Scout is an agent skill from QinghongLin/data2story-skill. Source and VERIFY rich external media (music, high-value real photos/video) and the latest live status that the Detective's background pass didn't cover. Every asset carries a checked license + identity block; nothing unlicensed or misidentified passes downstream. Outputs scout.json (sctxx) after the Detective, before analysis.
Its SKILL.md is about 7.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/license_allowlist.json`, `references/manifest_schema.json` and `references/media_verification.json`).
The repository describes itself as: Data Journalist Agent: Transforming Data into Verifiable Multimodal Story. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 63a55c1. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
Bash(*)ReadWriteGlobGrepWebSearchWebFetchFrom allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
commons.wikimedia.orgvimeo.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
UNSPLASH_ACCESS_KEYPEXELS_API_KEYOPENROUTER_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scout loads about 7.1k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 3,384 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
` and/or `PEXELS_API_KEY` (same env / `~/.env` pattern as `OPENROUTER_API_KEY`); if no key is set it exits with a clearallowed-tools: Bash(*), Read, Write, Glob, Grep, WebSearch, WebFetchAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from QinghongLin/data2story-skill at commit 63a55c1, republished under its MIT licence (© QinghongLin). 3,384 words, ~7,060 tokens.
.claude/skills/scout/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Premium-profile stage. The orchestrator runs the Scout only in the
premiumprofile; thefastprofile skips it. The "always runs / mandatory BGM, no exemption" rules below apply within premium.
Your job is rich media + freshness, with proof. The Detective already gathered background context and basic reference photos; you go further — you find the emotionally strong media (the star player, the packed stadium), the music that sets the mood, and the latest real-world status — and you are the pipeline's media verifier: every asset you pass on has a checked license that permits republication and a checked identity (it really is what the caption says).
You do not generate media — that is the Designer's job. You find real, license-clean media and you prove it.
DATA_DIR = first argumentPROJECT_DIR = second argumentSKILL_DIR = the directory containing this SKILL.md (.../skills/data2story-pro/scout)PROJECT_DIR/detective.json — its items give you the subjects/topics; its reference_media + instances tell you what's already covered, so you don't duplicate.PROJECT_DIR/assets/ (wikimedia_manifest.json, flags_manifest.json, logos_manifest.json) for the same reason.python3 SKILL_DIR/../detective/scripts/fetch_images.py (and fetch_flags.py, fetch_logos.py, fetch_openverse.py).PROJECT_DIR/scout.json (write incrementally). Assets → PROJECT_DIR/assets/scout_* (prefix scout_ to distinguish from the Detective's ref_*).The Cinematographer scroll background and the front BGM are MANDATORY pipeline stages on every blog — there is no "off" / opt-out, and BGM has no exemption (not even privacy) — so the Scout always runs and always sources a real-image set + a fitting real track, on every topic. Key the flavour off the shared topic_profile (the S3 classifier the Detective resolved into detective.json; if detective.json carries no resolved topic_profile, the Scout MUST write one into scout.json itself — explicit is_visual + is_computational booleans — because an absent profile is now a hard contract error (topic_profile_unresolved), so it cannot be left unresolved): when is_visual is true (any of visual_subject / event / sport / culture / place / emotional) you source the obvious strong subject photos; when the classifier marked the topic non-visual (abstract, text-only, statistical — economics, elections, public-health stats, finance), you still source a relevant real-image set — historical / archival / atmospheric real photos of the era and subject (for an industrial-revolution / economics story: real factory, loom, worker, steam-engine, trading-floor photos from Wikimedia Commons / public domain). Every topic gets a real-image set for the cinematic backing and a fitting real BGM. The only IMAGE exception is privacy_sensitive: there you do not source real-person imagery even if other visual tags are set (lean on non-person archival / atmospheric photos for the backing). A privacy-sensitive topic still gets a BGM — pick a quiet, non-intrusive, mood-appropriate real track (a restrained classical recording fits well); BGM is mandatory on every blog with no audio.used=false escape.
BGM is MANDATORY on EVERY blog with NO exemption: every blog opens with a fitting real-sourced track — there is no audio.used=false and no "skip audio for a sober / abstract / privacy topic" branch. Your job is not to decide whether there is a soundtrack but to source the track whose mood fits this story's tone. Match the mood word to the tone: a sober / computational story (economics, elections, public-health stats, finance) wants a pensive / ambient / minimal / orchestral / nocturne track, not a generic upbeat loop; a celebratory / sport / event story wants epic / anthem / fanfare; a somber story wants elegy / adagio / requiem with no especially strong emotion (quiet, non-triumphant). A fitting restrained track on a sober topic is the right BGM — it is NOT "tonally-wrong filler" to set a quiet ambient bed under a numbers story. You ALWAYS return a license-clean BGM track — if no topic-fitting real track exists, you fall to the classical-recording fallback (rung C below), which always yields a license-clean recording. You never leave a blog without a BGM.
The BGM is a real audio track presented in a self-hosted cover-art card at the TOP of the article, directly below the title, that starts on the reader's first gesture — the album/track art is a spinning vinyl disc (a circular cover that rotates only while playing). The track you source here IS the BGM that plays; the Designer never AI-composes a BGM. (text2music is SFX only — atmospheric sound-design beds for an un-findable sound — and is handled by the Designer, never as the front BGM.) FIT FIRST: source the most recognizable best-FIT real track; license-tier is only a tiebreak among comparably-fitting tracks. If the story HAS a signature anthem/track — an official anthem, the artist the post profiles, the song the story is about — source THAT (rung 2) rather than a generic unrelated mood loop, even though the signature track is the demo-gated rung; a recognizable signature track beats a clean-but-unrelated CC0 loop. Only when no signature track fits the story do you reach for a clean-but-generic mood track (rung 1) or, failing that, the classical floor (rung C). Walk this BGM ladder by fit (not blindly top-down), and never AI-compose the BGM:
License-clean real track (publishable) — find a freely-licensed instrumental that fits the story's mood / place / era (CC0 / CC-BY / public-domain / explicitly royalty-free), self-hosted so it can ship publicly. Use list → pick → download, not blind first-result:
python3 SKILL_DIR/scripts/fetch_music.py --list --query <word> --limit 8 prints a JSON array (each with id, title, license, spdx, duration_s, source_url) to stdout. Commons audio is sparse — search with a SINGLE broad mood word matched to the story's actual tone (epic, anthem, fanfare for a triumphant story; pensive, ambient, minimal, orchestral, nocturne for a sober/analytical one; elegy, adagio for a somber one); multi-word queries usually return nothing (the script auto-falls-back to single words, but a broad word is more reliable). Pick the mood word from the STORY's emotion, not a default-celebratory one — a flat economics/elections/health-stats story wants a restrained pensive/ambient track (still a real BGM, never "no BGM").spdx is on references/license_allowlist.json and whose duration_s suits a loopable BGM. Do not just take #1.id: python3 SKILL_DIR/scripts/fetch_music.py --query <word> --outdir PROJECT_DIR/assets --download scout_bgm --id "<the File: id you picked>". The fetcher also downloads the Commons file's cover-art thumbnail alongside the audio (a *_cover.<ext> next to the track, recorded as cover_path / cover_source_url in music_manifest.json) so the now-playing card has a license-clean square cover. If a Commons file has no usable thumbnail, fetch a representative license-clean image for the card via fetch_stock.py / Commons (Step 3), or leave a designed CSS cover to the Designer (NOT an AI image).Record the full license + attribution (track and cover). This is the track that actually plays, registered as a Scout sct_ audio item (license-clean → it passes the validate.py license-allowlist gate).
Copyrighted best-fit real track — self-host for the DEMO, publish-gated — when the song the story is about is itself the right BGM (e.g. an official anthem, the artist the post profiles) and no license-clean track fits as well, self-host that real track + its cover for the demo rather than AI-composing one. Fetch the track + a representative cover from its source (a small --track-url/--cover-url helper on fetch_music.py, or grab them by hand — the World Cup anthem + cover were grabbed manually), then record it with an explicit publish-gate so it is never silently treated as clean:
license.spdx = "All Rights Reserved — demo-only", license.permits_republication = false, and a real source_url (where the track came from).des_ audio asset with publish_blocker: true — NOT as a clean sct_ item — so it does not pass the validate.py license-allowlist gate as clean. A publish_note is MANDATORY (the swap target — the clean track or embed to switch to before publishing): validate.py Section 8 hard-errors a gated asset with no swap target, so hand the Designer the publish_note along with the track + cover + the gate fields. Note in your scout.json (e.g. a live_status/note item or the relevant sct_ notes) that the BGM is the copyrighted demo track to be registered as a des_ publish-blocker. The Auditor raises an advisory publish-blocker and the Programmer renders a "demo-only — must license or swap before publishing" credit line; the demo build is flagged, never blocked.Embed the official player — if you can neither find a license-clean track nor self-host the copyrighted one, surface the real song as an oEmbed-verified embed (the official Spotify/YouTube player carries its own rights). For an embed: put the /embed/ player URL in embed_url and the watch/track URL you oEmbed-verified in source_url; set identity.method="oembed", identity.verified=true, and license.permits_republication=false (you are not re-hosting — the platform player carries the license; license.spdx may be "All Rights Reserved") per ../detective/references/instance_verification.json. The validate.py license gate skips embeds. An embed does NOT replace a self-hosted now-playing card if rung 1 or 2 was available.
Classical-recording fallback ladder — the GUARANTEED license-clean floor (rung C). When no topic-fitting real track (rung 1) and no signature track (rung 2/3) lands, you do not stop with no BGM — you source a license-clean classical RECORDING. The key correctness point: a public-domain composition (Beethoven / Bach / Chopin / Tchaikovsky / Mozart / Haydn / Brahms / Debussy / Satie…) is NOT automatically a public-domain recording — the score may be PD while a modern performance is fully copyrighted. So you must source a license-clean RECORDING of the piece and verify the recording's own license, from a PD/CC recording library:
fetch_music.py --list … --download … flow; record the recording's spdx (must be on references/license_allowlist.json), permits_republication, and attribution_text. Verify the recording (not the composition) is what passes the gate.sct_ audio item (license-clean → it passes the validate.py license-allowlist gate), with its cover (the album/portrait art the library or Commons provides, else a representative license-clean image for the disc, else a designed CSS disc — never AI). This rung always succeeds, so every blog ends with a license-clean BGM.You may also record the real songs the story references (an anthem, a viral hit) as oEmbed embed instances for a "listen ↗" link in context even when the BGM is a rung-1 / rung-C track — that is separate from the BGM itself.
License gate: never pass a copyrighted commercial track off as a license-clean sct_ BGM. A copyrighted self-hosted BGM is only the rung-2 des_ publish-blocker path above (flagged, demo-only); a clean sct_ BGM is rung 1 or the rung-C classical recording. A PD composition with a copyrighted recording is NOT clean — verify the recording's license, and if the only available recording is copyrighted, treat it like any copyrighted track (rung 2 demo-gate or rung 3 embed), then keep climbing toward a clean classical recording so the blog ends license-clean.
Weight note: Commons audio is often a multi-MB WAV/FLAC. Pass it on as-is (don't degrade the source), but the Designer will transcode it to a web-weight streaming copy (~128 kbps mp3/opus, < 3 MB) before referencing it — so the heavy original never ships. If the downloaded track is very large, note its size in the sct_xx item so the Designer knows to optimize it.
If the dataset is about an ongoing / recent event, fetch the current real-world status (latest results, standings, counts) with python3 SKILL_DIR/scripts/fetch_live_status.py. Write a timestamped data file for the Analyst (mirrors the Detective's fetch_venue_weather.py → *_source.json), and add a live_status[] entry with a dated source.
Leakage guard: live status is display context only — always dated "as of <date>". It is NEVER fed to a forecasting / training model. Keep it on the Analyst's data path with its as_of, not as free-floating page text.
Keep it compact (presentation restraint). Record live status as a short "since the snapshot" summary, not a long log: a count of what changed plus the single latest result (with its as_of date) is enough. Do not dump many specific forward-dated results — a long forward-dated list reads like new data and confuses the dataset snapshot the story is built on. Give the Designer a tight, dated badge to render ("as of <date>: N updates, latest = …"), nothing more. (Shared with the Editor/Designer work-streams; topic-agnostic.)
For the subjects that carry the story emotionally (named people, specific stadiums / places, key objects), fetch a strong, specific real photo / video the Detective missed or got only weakly. You have three complementary image sources — use whichever lands the better, more specific shot, and you may try more than one:
python3 SKILL_DIR/../detective/scripts/fetch_images.py --qids <Wikidata-QID> --props P18 --outdir PROJECT_DIR/assets --prefix scout_ --append (find the subject's Wikidata QID; P18 is the entity's photo). Writes assets/scout_* directly.python3 SKILL_DIR/../detective/scripts/fetch_openverse.py --list --q "<keyword>" --limit 8 returns JSON candidates (each with id, spdx, permits_republication, attribution_text, license_url, foreign_landing_url, source_url); pick one whose spdx is on the allowlist (permits_republication: true), then ... --download --id <openverse-id> --q "<keyword>" --outdir PROJECT_DIR/assets --prefix scout_.Unsplash-License / Pexels-License (both on the allowlist, genuinely re-hostable); best for atmospheric / generic / cinematic-background shots (a floodlit stadium, a city skyline, an empty arena) where Commons/Openverse are thin — this is the channel the gold blog's cinematic backdrops drew on. Same list → pick → download: python3 SKILL_DIR/scripts/fetch_stock.py --list --q "<keyword>" --limit 8 --source both returns S2-shaped candidates; pick one, then ... --download --id <candidate_id> --q "<keyword>" --outdir PROJECT_DIR/assets --prefix scout_. Needs a free key — UNSPLASH_ACCESS_KEY and/or PEXELS_API_KEY (same env / ~/.env pattern as OPENROUTER_API_KEY); if no key is set it exits with a clear message and you fall back to the two no-key sources above. The fetcher emits permits_republication:true / requires_attribution:false but leaves identity.verified:false — it can't confirm the subject, so the Step 4 identity check below is mandatory before any specific-real-subject stock photo ships.Either way, run every candidate through the same Step 4 license + identity gate below, and make sure each downloaded asset's provenance record matches the shared references/manifest_schema.json S2 block ({id, file, source_url, site, license{…}, identity{…}}) — the one shape validate.py and the Designer's registration rule both read. Always cover different subjects (no duplicates) and prefer specific, verified shots over generic fills.
Image-count target — MANDATORY on every topic (a real-image set to back the cinematic). Because the cinematic scroll background is a mandatory stage, every topic gets a relevant real-image set — aim for around 5–6 verified, license-clean, mostly landscape / cover-able real images across distinct subjects (a soft target on the count, but sourcing the set itself is not optional). This is what feeds the Cinematographer (it needs ≥5 registered verified landscape backgrounds for the scroll background; under-supply sends it back to you to source ≥5 cover-able backgrounds before it re-runs, not accepted as final). Lean on fetch_stock.py for the atmospheric, cover-able shots (an empty arena, a skyline, a moody landscape) that round the set out even when Commons/Openverse are thin on a subject — these full-bleed-friendly stills are exactly what the cinematic background layer stages.
Abstract / historical / economic topics still get a real-image set — they are not exempt. When the classifier marked the topic non-visual (economics, elections, public-health stats, finance, history), do not "skip lightly" — source relevant real historical / archival / atmospheric photos of the era and subject: for an industrial-revolution / economics story, real factory, loom, mill-worker, steam-engine, and trading-floor photos from Wikimedia Commons / public domain; for an elections story, real polling-station / ballot-box / campaign-rally archival photos; for a public-health story, real hospital / clinic / lab archival photos. These are abundant in the public domain and on Commons, are relevant (not decorative), and back the mandatory cinematic. The line is relevance, not subject-type: a relevant historical/archival real photo is exactly right; what stays banned is purely-decorative stock that says nothing about the story (a random smiling-businessperson stock photo on an inflation piece). Source the relevant real set on every topic; only the genuinely privacy-sensitive topic stays light on imagery (no real-person photos). (Quantity never buys past Step 4, and never overrides "is this image relevant to the topic". Only if relevant real images genuinely cannot be found does the Cinematographer fall back to generative/data-driven backgrounds — but try hard here first.)
Video channel. A clip may enter the page two ways: (a) a verified oEmbed embed — YouTube or Vimeo — where the platform player carries the license and you re-host nothing (verify per ../detective/references/instance_verification.json, now covering Vimeo's https://vimeo.com/api/oembed.json?url=... endpoint; set kind: "embed", identity.method="oembed", license.permits_republication=false); or (b) best-effort, a CC-licensed clip from Wikimedia Commons / Openverse, re-hosted only if its license is on the allowlist and it passes the identity gate. Prefer an embed for rights-encumbered footage.
For every media item you add, fill a license block and an identity block — the exact S2 manifest shape in references/manifest_schema.json ({id, file, source_url, site, license{spdx, permits_republication, requires_attribution, attribution_text}, identity{method, verified, subject}}), shared verbatim with the Designer + every fetch script + validate.py:
License: spdx, permits_republication, requires_attribution, attribution_text (non-empty, footer-ready). Only licenses on references/license_allowlist.json may be re-hosted; anything else → drop it, or downgrade to an embed.
Identity: prove the asset is what the caption claims, by the cheapest sufficient method (see references/media_verification.json):
oembed — for embeds (HTTP 200 + title match); reuse the Detective's workflow.trusted_source — a Wikidata-QID / Commons file whose page names the subject.vlm_view — open the image with the Read tool and confirm it depicts <subject> (and is a real photo, not an AI render of a real thing); record one line of what you saw.Set identity.verified = true only when one method passed. A real-subject asset with identity.verified = false is hard-rejected by validate.py — so don't pass it on.
Designer data_source grammar for scouted sources. When a Designer item draws on a scouted source, the only resolvable forms of its data_source are data_source: "scout.<sct_id>" (the suffix — or sct_ + the suffix — MUST name a registered sct_ item in this scout.json) or data_source: "scout.live_status" (valid only when this scout.json has a non-empty live_status list). A free-text scout.<anything-else> now hard-errors at the contract gate (des_data_source_scout_dangling) — so register the sct_ item (or supply a live_status entry) before the Designer points at it.
scout.jsonWrite incrementally (read-add-write), same as the Detective. Shape (validator-enforced): items is a dict keyed by sct_xx id (NOT a list) — validate.py iterates scout.items as {id: {...}}; live_status stays a list. Full schema in references/schema.json:
{
"meta": { "role": "scout", "version": "1.0" },
"items": {
"sct_01": {
"kind": "image",
"label": "Messi lifting the trophy",
"filename": "scout_messi.jpg",
"caption": "Lionel Messi after the 2022 final.",
"caption_claims": ["this is Lionel Messi"],
"source_url": "https://commons.wikimedia.org/wiki/File:...",
"retrieved_at": "2026-06-21T14:03:00Z",
"license": { "spdx": "CC-BY-SA-4.0", "permits_republication": true, "requires_attribution": true,
"attribution_text": "Photo: <author> / Wikimedia Commons (CC BY-SA 4.0)" },
"identity": { "method": "vlm_view", "verified": true, "verified_title": "a man in an Argentina shirt holding the trophy", "subject": "Lionel Messi" },
"relates_to": ["det_03"], "purpose": "INFORM"
},
"sct_02": {
"kind": "audio",
"label": "Pensive ambient BGM for the inflation story",
"filename": "scout_bgm_web.mp3",
"caption": "License-clean ambient instrumental sourced for the top-of-article spinning-vinyl BGM card.",
"source_url": "https://commons.wikimedia.org/wiki/File:...",
"retrieved_at": "2026-06-21T14:04:00Z",
"license": { "spdx": "CC0-1.0", "permits_republication": true, "requires_attribution": false,
"attribution_text": "Music: <author> / Wikimedia Commons (CC0)" },
"identity": { "method": "trusted_source", "verified": true, "subject": "ambient instrumental track" },
"cover_path": "scout_bgm_cover.jpg",
"note": "Lead finding is the CPI inflation series; tone is sober/computational, so the mood word was `pensive`/`ambient`, not a celebratory loop. A fitting restrained track is the right BGM — BGM is mandatory, not opt-out.",
"relates_to": ["det_01"], "purpose": "IMMERSE"
}
},
"live_status": [
{ "subject": "Group C standings, matchday 3", "as_of": "2026-06-21",
"status": "...", "source": { "url": "https://...", "title": "...", "fetched_at": "2026-06-21T14:05:00Z" },
"relates_to": ["det_05"] }
]
}For embeds, replace filename with embed_url and set identity.method = "oembed". For the BGM track, kind: "audio", purpose: "IMMERSE" — record the cover image too (the *_cover file fetch_music.py saved, or a representative license-clean image) so the Designer's top-of-article BGM card has a square cover for the spinning vinyl disc. If the BGM is the rung-2 copyrighted demo track, do not record it as a clean sct_ item — flag it for registration as a Designer des_ publish-blocker (see Step 1). If no topic-fitting real track exists, the classical-recording fallback (rung C) always lands a license-clean recording, so a clean sct_ BGM is always present.
references/schema.json — full scout.json structure + field notes.references/license_allowlist.json — SPDX (+ stock) licenses that may be re-hosted.references/manifest_schema.json — the S2 per-asset provenance block (id/file/source_url/site/license/identity) shared by Scout, Designer, and every fetch script.references/media_verification.json — the identity-check methods.../references/topic_profile.json — the shared S3 topic classifier the "When to run" gate keys off (is_visual / is_computational).scripts/fetch_stock.py — keyword stock-photo search across Unsplash + Pexels (free key; emits the S2 block).../detective/references/instance_verification.json (oEmbed: Spotify / YouTube / Vimeo), ../detective/scripts/fetch_images.py (Commons fetch by QID), and ../detective/scripts/fetch_openverse.py (keyword image search across the Openverse aggregator).Done when the Designer has strong, verified, license-clean media to work with (music + photos + any live-status), every item carries a license + identity block, and no real-subject asset is left unverified.
© QinghongLin, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts, references) in skills/data2story-pro/scout of QinghongLin/data2story-skill.
Open the folder on GitHubat commit 63a55c1
Scout next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scout this skillQinghongLin/data2story-skill | 155 | — | ~7.1k | Automated safety check: Notes | MIT | |
| Music to Videoheygen-com/hyperframes | 60k | 3 repos | ~4.7k | Automated safety check: Notes | Apache-2.0 | |
| MiniMax Music Generationbytedance/deer-flow | 84k | — | ~717 | Automated safety check: Pass | MIT | |
| Venice Audio Musicnexu-io/open-design | 100k | — | ~297 | Automated safety check: Pass | Apache-2.0 | |
| AI Music Albumnexu-io/open-design | 100k | — | ~328 | Automated safety check: Pass | Apache-2.0 | |
| Music Generation Studiosickn33/agentic-awesome-skills | 47k | 1 repos | ~3.1k | Automated safety check: Pass | MIT-0 |
heygen-com/hyperframes
Turns a music track into a beat-synced HyperFrames video such as a lyric video, slideshow or kinetic promo, with any supplied images or clips cut onto the beat grid.
bytedance/deer-flow
Generates songs, jingles or instrumental tracks as MP3 files from a style prompt and optional lyrics through the MiniMax music API.
nexu-io/open-design
Music generation queueing, retrieval, and completion endpoints via Venice.ai.
nexu-io/open-design
Full-lifecycle AI music album production — concept, lyric drafting, track sequencing, and export.
sickn33/agentic-awesome-skills
Install and use the official AI Music Generator package, pinned by digest, for paid hosted work on the Beatra service.
ruvnet/ruflo
List the account's saved music productions with metadata and audiourl
QinghongLin/data2story-skill
Run sentence-level traceability verification on a Data2Story blog (verify.py - verifier.json), then emit the in-page Inspector panel (the reader-facing runnable verifier) + the verify/ artifacts…
QinghongLin/data2story-skill
Audit a generated Data2Story blog for build correctness across ALL modalities by ACTUALLY RENDERING it in a real headless browser (when available) — catching blank/0-width charts, broken/oversized…
QinghongLin/data2story-skill
Review a finished Data2Story blog against the 5 quality rubric dimensions (visualdesign, narrativepacing, datamethodtransparency, claimdataalignment, insightvalue), score each 1-7 with on-page…
QinghongLin/data2story-skill
Research external context for a dataset — domain background, history, related studies, and why this data matters.
QinghongLin/data2story-skill
Run sentence-level traceability verification on a blog, then generate viewer.html with interactive evidence panel.
QinghongLin/data2story-skill
A skill your agent uses to turn a dataset into a verifiable multimedia blog (a data story / data-driven article / interactive dashboard from a dataset).
Source and VERIFY rich external media (music, high-value real photos/video) and the latest live status that the Detective's background pass didn't cover. Scout is an agent skill from QinghongLin/data2story-skill. Source and VERIFY rich external media (music, high-value real photos/video) and the latest live status that the Detective's background pass didn't cover.
Run `npx skills add QinghongLin/data2story-skill --skill scout -a claude-code`. Or copy the skill folder (skills/data2story-pro/scout in QinghongLin/data2story-skill) into .claude/skills/scout in your project. Claude Code loads it when a task matches its description.
Run `npx skills add QinghongLin/data2story-skill --skill scout -a codex`. Or copy the skill folder (skills/data2story-pro/scout in QinghongLin/data2story-skill) into .agents/skills/scout in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QinghongLin/data2story-skill --skill scout -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scout, .gemini/skills/scout, .github/skills/scout and .opencode/skills/scout in your project.
Going by SKILL.md and its folder, Scout needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named UNSPLASH_ACCESS_KEY, PEXELS_API_KEY and OPENROUTER_API_KEY. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(*), Read, Write, Glob, Grep, WebSearch, WebFetch.
SKILL.md names 2 domains. In commands or code: commons.wikimedia.org and vimeo.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Scout is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.1k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Scout: Music to Video (heygen-com/hyperframes, 60k stars), MiniMax Music Generation (bytedance/deer-flow, 84k stars), Venice Audio Music (nexu-io/open-design, 100k stars) and AI Music Album (nexu-io/open-design, 100k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
QinghongLin (a GitHub user) maintains it in QinghongLin/data2story-skill, which has 155 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on July 5, 2026.
Source: QinghongLin/data2story-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.