Agent skill

AI Media Generator

by Hao0321 in Hao0321/ai-media-generator

為使用者產生高品質的 AI 生圖、生影片、生音樂提示詞,並在需要時透過瀏覽器自動化實際送到目標平台。涵蓋 OiiOii、Kling 3.0/O-series、Seedance 2.0/2.5、Suno v5.5、Seedream 5.0/4.0、Vidu Q3、Midjourney V8.1、Flux 1.1 Pro / Kontext、Runway Gen-4.5 /…

MITAuto-check passedMedia & Creative

Install AI Media Generator

skills CLI
$ npx skills add Hao0321/ai-media-generator --skill ai-media-generator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Hao0321/ai-media-generator ai-media-generator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-media-generator
GitHub stars
258
Token cost
~5.6k tokens
SKILL.md length
1,805 words
Files
71 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

為使用者產生高品質的 AI 生圖、生影片、生音樂提示詞,並在需要時透過瀏覽器自動化實際送到目標平台。涵蓋 OiiOii、Kling 3.0/O-series、Seedance 2.0/2.5、Suno v5.5、Seedream 5.0/4.0、Vidu Q3、Midjourney V8.1、Flux 1.1 Pro / Kontext、Runway Gen-4.5 /…

  • Works in 5 steps: 選平台 → 讀對應的 reference 檔 → 根據任務類型參考 template → …
  • Tasks that involve AI video generation
  • SKILL.md covers 🤖 Auto-Pilot Mode (超傻瓜一句話到成品), 🎯 快查表:3 秒找對檔案(更直覺入口), 🔴 硬規則 (Mandatory) and 核心原則, plus 5 more sections
  • Calls bash

What it does

AI Media Generator is an agent skill from Hao0321/ai-media-generator. 為使用者產生高品質的 AI 生圖、生影片、生音樂提示詞,並在需要時透過瀏覽器自動化實際送到目標平台。涵蓋 OiiOii、Kling 3.0/O-series、Seedance 2.0/2.5、Suno v5.5、Seedream 5.0/4.0、Vidu Q3、Midjourney V8.1、Flux 1.1 Pro / Kontext、Runway Gen-4.5 / Aleph、Google Veo 3.1、Ideogram 3、Nano Banana Pro、Stable Diffusion 3.5(⚠️ OpenAI Sora 2 已於 2026-04-26 停運,API 撐到 2026-09-24,預設改推 Runway/Veo/Kling)。只要使用者提到「AI 生圖」「AI 影片」「AI 音樂」「做 MV」「做 storyboard」「寫 prompt 給 XXX」「我想用 Kling/Suno/Midjourney/Runway/Veo...」「幫我操作 OiiOii / 即夢 / 可靈」「txt2img / img2video / 文生圖 / 文生影片 /…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 73 other files, including reference files (for example `CHANGELOG.md`, `CONTRIBUTING.md` and `README.md`).

It sits in Media & Creative, covering AI video generation. It works with Midjourney, Suno, Seedance and Google Gemini. The repository describes itself as: Zero-skill cinema. Senior-director prompts on autopilot. A Claude Code Skill for high-quality AI image / video / music prompt crafting and browser-based execution across 14+… The licence is MIT.

When your agent uses it

  • Tasks that involve AI video generation

Example prompts

  • “/ai-media-generator”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. 選平台
  2. 讀對應的 reference 檔
  3. 根據任務類型參考 template
  4. 組出 prompt 並輸出
  5. (選配) — 瀏覽器自動化執行

What it can do on your machine

Read from SKILL.md and the folder at commit edc8aa5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Media Generator loads about 5.6k tokens when it runs, and up to ~160k if it reads all its reference files. Until then it costs about 164 tokens; SKILL.md has 1,805 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~164
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~160k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Hao0321/ai-media-generator at commit edc8aa5, republished under its MIT licence (© Hao0321). 1,805 words, ~5,641 tokens.

Download SKILL.mdSave it as .claude/skills/ai-media-generator/SKILL.md (or your agent's skills folder). This skill also uses 70 other files; get the full folder from GitHub.
name
ai-media-generator
description
為使用者產生高品質的 AI 生圖、生影片、生音樂提示詞,並在需要時透過瀏覽器自動化實際送到目標平台。涵蓋 OiiOii、Kling 3.0/O-series、Seedance 2.0/2.5、Suno v5.5、Seedream 5.0/4.0、Vidu Q3、Midjourney V8.1、Flux 1.1 Pro / Kontext、Runway Gen-4.5 / Aleph、Google Veo 3.1、Ideogram 3、Nano Banana Pro、Stable Diffusion 3.5(⚠️ OpenAI Sora 2 已於 2026-04-26 停運,API 撐到 2026-09-24,預設改推 Runway/Veo/Kling)。只要使用者提到「AI 生圖」「AI 影片」「AI 音樂」「做 MV」「做 storyboard」「寫 prompt 給 XXX」「我想用 Kling/Suno/Midjourney/Runway/Veo...」「幫我操作 OiiOii / 即夢 / 可靈」「txt2img / img2video / 文生圖 / 文生影片 / 圖生影片」「角色一致性」「多鏡頭分鏡」「運鏡」「結果有瑕疵 / 不夠精緻 / 怎麼修」,或任何跟上述平台或影像/影片/音樂生成工作流相關的任務,都要用這個 skill。即使他們沒講明平台,只要任務是要餵給某個生成模型的 prompt,就用這個 skill 幫他們選對的平台、寫對的格式。

AI Media Generator

Codex compatibility: References that still say “Claude in Chrome” or “Chrome MCP” describe legacy browser behavior. On Codex, load chrome:control-chrome and use the currently available Chrome tools; never invent or call legacy mcp__Claude_in_Chrome__*, javascript_tool, or browser_batch names.

幫使用者把想法變成高品質的 AI 生成內容 (圖片、影片、音樂),核心工作是 寫對每個平台的 prompt 以及 必要時自動操作網站。

AI 短劇邊界: 使用者要做 AI 短劇/漫劇、重生/系統/逆襲/隱藏大佬、分集連載或整季架構時,先由 ai-short-drama 建題材、Bible、pilot、分集狀態與 production pack;本 Skill 只接下游角色圖、場景圖、逐鏡影片/聲音 prompt 與模型執行,不自行重寫季弧。

🤖 Auto-Pilot Mode (超傻瓜一句話到成品)

觸發: 使用者說「幫我做 X」「做一個 X」「生個 X」「我要 X」這類命令式、且 X 含媒體類型 (圖/影片/音樂/MV/短片/海報/動畫) 或風格關鍵字 → 直接進 auto-pilot 不問。

流程 (見 templates/auto-pilot.md 完整版):

  1. Intent Parser — 一句話拆 9 slot (媒體/長度/畫面比/主題/風格/角色/場景/音訊/語言)
  2. Fill Defaults — 沒講的用預設 (video 預設 10s 16:9;動漫預設 Shinkai+Ghibli;電影預設 Deakins DP…)
  3. Platform Decider — 按 references/selector.md + quick pick 矩陣自動選
  4. Script + Storyboard Auto-gen — 自動寫 logline + 角色卡 + 分鏡
  5. Prompt Crafting — 強制套語彙庫 (見下方硬規則)
  6. Preview + Go — 30-秒 override 視窗,使用者回「go/確認/OK」才執行;無回 + 需花錢 → 停;無回 + 免費 → 30 秒後自動繼續
  7. Execution + Report — 按 automation/click-protocol.md + site-profiles/ 協議操作

Auto-Pilot 必停的 checkpoint (不代做):

  • ⛔ 付費 paywall / 升級提示
  • ⛔ 送出前最終一次確認
  • ⛔ 下載 / 分享 / 公開發布
  • ⛔ 敏感內容 (真人照片疑似裸露/暴力/政治人物)
  • ⛔ click-protocol.md 定義的「irreversible action」

Auto-Pilot 禁止行為:

  • ❌ 幫 Claude 已知的事問使用者 (「你要幾秒的影片?」— 10s 預設就好)
  • ❌ 一次拋多選項 — 選 1 個最佳 commit,讓使用者在 Preview 階段才決定改不改
  • ❌ 寫純 generic 填詞 prompt (beautiful / masterpiece / detailed) — 違反下方硬規則(注意:cinematic / 4K / 8K 在新一代模型如 Seedance 2.0 是有效 token,看平台分流)
  • ❌ 代付款

使用者自然語言 flags 可 override 預設: 使用者一句話可帶「用 Kling / 5 秒 / 豎屏 / 免費 / Ghibli 風 / 有對白」等自然語言 flag,auto-pilot 掃 templates/user-flags.md 對照表自動套用。使用者不懂術語也 OK — 「做個抖音」「可愛一點」「夢幻」都有對應翻譯。

⚡ Token-Efficient Mode (大專案必讀): 本 skill 55+ 檔 / 14000+ 行,全量讀會爆 context。Auto-Pilot 預設套 templates/token-efficient-mode.md 的 7 層策略 (lazy load / grep / 子代理 / preset 跳過 / cache polling)。一般任務目標 ~25-40k tokens,不是 100k+。豪華模式 (全量讀) 只在 benchmark / 學 skill / 陌生平台探索時啟用。

🎯 快查表:3 秒找對檔案(更直覺入口)

先看這張表 → 對到使用者說的話 → 直接讀那個檔。 找不到再往下走完整流程。

使用者說…讀這個檔為什麼
「AI短劇/AI漫劇/爽劇/重生/系統/逆襲/隱藏大佬/分集連載」ai-short-drama先鎖故事、角色、揭露梯與 state;媒體 prompt 再回本 Skill
「結果有瑕疵/爛掉/不夠精緻/閃爍/變形/塑膠感/文字怪」quality-control.md7+1 類瑕疵系統診斷 + 對症修法
「prompt 沒主題/畫面很空/不知道在拍什麼/效果差」concept-first-prompting.md先定 concept 再加技術詞
「怎麼把想法組成強 prompt/太慢/每次從零想/要打分」prompt-craft-engine.md6層骨架 + 填空模板 + 10分評分閘(更快更強)
「成品要更高級/更精美/更貴/像 editorial/反塑膠感」aesthetic-grade.md打光比/調色/材質/鏡頭 DP 級配方(更精美)
「不知道用哪個模型/該用哪個平台/不熟各家」model-picker.md決策樹 + 各家最強情境
「角色要一致/同一個人出現多鏡/風格鎖定」asset-library.md + multimodel-video-cheatsheet.md角色卡 + @element 身份鎖
「Seedance 2.5/30秒/社群案例/UGC/Vlog/多參考/硬切/提示詞案例」seedance-2-5-community-playbook.md + seedance.md證據分級 + capability gate + 社群案例蒸餾 + 六類模板
「Seedance 商業/企業/建商/建築/房地產/廣告/產品/汽車/美食/電影/動畫」seedance-domain-playbook.md + seedance.md八個垂直領域的 reference map、truth lock、物理帳本、原創模板與 QA
「Seedance 長片/參考圖角色/首鏡後鎖定/時間凍結或倒轉/複雜事件連續性」seedance-longform-continuity.md + seedance.md不變量帳本 + 時間碼 + 事件狀態機 + 15 秒切段
「圖生影片/i2v/img2video/參考圖鎖死/產品廣告鎖形狀」image-to-video-workflow.mdi2v 鎖形狀 SOP(產品廣告最穩)
「寫歌/純音樂」→suno.md|「對白/音效/配樂/原生音訊」→sound-design.md二選一純音樂 vs 影片內音訊分流
「做 MV/卡點/轉場/剪接節奏」music-video.md + editing-transitions.mdBPM 同步 + match cut/J-cut
「要現成 prompt/範本/某某風格(Wes Anderson/賽博龐克)」preset-packs.md30+ preset 換占位符即用
「想抄某導演/DP/品牌/製作公司風格」director-style-library.md簽名 token(⚠️ Flux/Nano 洗掉導演名)
「超短廣告/TikTok/10秒快剪」nano-banana.md<30 token 超精簡 prompt
「Wan 2.7/Kling Omni/HappyHorse/拉片復刻/多模型省成本」multimodel-video-cheatsheet.md2026 新模型招牌技巧
「幫我直接操作 OiiOii/Flow/Kling」click-protocol.md + site-profiles/自動化協議 + 各站 UI 地圖
「免費/不想花錢」selector.md Cost 欄 + 各平台檔 Free tier免費額度速查

🔴 OiiOii 專用: OiiOii 2026-06-08 大改版,最新 UI/流程/成本/i2v/選模型/新模型(Oii Image 2 / Gemini Omni / Oii X Imagine)一律以 automation/site-profiles/oiioii.md §0 為準(SOURCE OF TRUTH)。references/oiioii.md 只看 prompt 哲學,UI 別照它。

🔴 硬規則 (Mandatory)

Meta 優先序(2026-04-21 實戰鐵律)

Prompt 寫對一次 ≫ 操作快 10 次。

根因:寫錯 prompt → 重做 → 等 8-10 分鐘。操作 25 秒 vs 5 分鐘差距(~4.5 分鐘),遠遠小於「一次 prompt 失敗」成本(~10 分鐘等待 + token 浪費)。所以速度優化的真正優先序:

  1. 第一優先:查平台簽名 + 寫對 prompt → references/community-prompt-patterns.md(跨 X/Threads/Reddit/小紅書/Bilibili 社群驗證,單一 source of truth)
  2. 第二優先:單次提交極速化 → automation/site-profiles/
  3. 第三優先:等待不 polling → Bash run_in_background:true + sleep 400

順序反了 = 浪費 40+ 分鐘做 4-5 次嘗試才生出可用的。

Prompt 必備語彙

每次產 image / video / music prompt,都必須從 skill 進階語彙庫挑 token — 不是選配,是必做。

流程:

  1. 最優先查:references/community-prompt-patterns.md — 按目標模型查簽名 token + 長度甜蜜點 + 禁忌(⚠️ 平台吃不同 token,Seedance/Wan 吃導演名 = 災難)
  2. 決定任務類型 (攝影感/電影感/廣告/時尚/MV/VFX/社群短片)
  3. 挑對應 reference 檔讀:
  4. 快速路徑:先查 preset-packs.md 找最近的 preset,換占位符即可
  5. 每個 prompt 至少嵌入 5-8 個高訊號 token,從下列層挑(看平台分流):
    • 導演/DP 名 (Deakins、Lubezki、Hoytema、王家衛、新海誠…) — ✅ MJ/Sora 2/Veo;❌ Flux/Nano Banana Pro/Seedance/Wan
    • 鏡頭/焦段 (Panavision anamorphic / 85mm / C-Series 等) — 通吃
    • 底片/感光 (Kodak Vision3 500T / Cinestill 800T 等) — ✅ Flux/MJ;❌ Seedance
    • 光比/燈光 (Rembrandt / 4:1 contrast / volumetric god rays 等) — 通吃
    • 色彩分級 (teal-orange / bleach bypass / A24 indie 等) — 通吃
    • 構圖/景別 (rule of thirds / medium close-up 等) — 通吃
    • VFX/大氣 (halation / Tyndall effect / particles 等) — 通吃
    • (Veo/Sora) 音訊三層 (Dialogue / SFX / Soundtrack)
禁用模式(通用 + 平台特定)

通用原則: 寧可 5 個具體 token,不要 20 個泛詞。 但 generic 與否看平台。

真正全平台垃圾(任何時候都別用):

  • beautiful / masterpiece / detailed / high quality / professional(這幾個從沒在新一代模型有實證效果)
  • --no blur 等 flag-style 負面 prompt(多數模型不吃,用自然語言 no blur 反而 OK)

⚠️ 平台特定(注意「2026-04-21 vs 2026-05-18 推翻」— 模型升級會改變斷言):

  • ⚠️ cinematic / 4K / 8K / 35mm-50mm-85mm — 舊版 Seedance 1.0 弱,Seedance 2.0 大量吃(v1.1.0 修正)。Kling / Sora 2 / Veo 3.1 / MJ v7 / Flux 全吃。
  • ❌ Seedance 2.0:fast(改 extreme speed / kinetic / rapid)、多動詞同句、多主體獨立動作、chaotic wide 無時間區塊、個別 DP 名(藝術運動 / 品牌風格 OK)
  • ❌ Flux / Nano Banana Pro:artist names(訓練時被 scrub)、--ar 語法
  • ❌ Runway Gen-4:>60 字 prompt(最短的模型)、synonym drift
  • ❌ Kling:stacking 多個相機運動、>4-5 distinct nouns

先查 community-prompt-patterns.md — 該檔每個模型都有「禁忌」section,且註明 cross-platform 推翻歷史。

驗證自檢(送出前必過 10 分評分閘)

每個 prompt 送出前,用 prompt-craft-engine.md §3 的 10 分品質評分表 打分 —— ≥8 才送,<8 回去補。 每個跨平台垃圾詞(beautiful/masterpiece/detailed/high quality)額外 −1。

10 項速記:①具體主體 ②可視動詞 ③2-4場景元素 ④鏡頭精度 ⑤單一風格錨 ⑥長度甜蜜點 ⑦無自相矛盾 ⑧平台簽名對 ⑨負面對症 ⑩存在性測試。

快速 4 題版(趕時間): 查過 community-prompt-patterns.md 目標模型?Token 符合平台簽名?避開禁忌?長度甜蜜點?—— 4 過才送。

核心原則

  1. 先問清楚「要什麼」,再決定「用哪個」。同一個想法送到不同模型,prompt 寫法完全不同。先釐清:

    • 媒體類型:靜態圖 / 影片 / 音樂 / 複合 (MV、分鏡動畫)
    • 用途:社群貼文 / 廣告 / 電影感短片 / 角色一致性專案 / 文字海報
    • 手上資源:有沒有參考圖、首尾幀、角色圖、歌詞
    • 預算/可用性:免費網站 / 付費會員 / API 金鑰
  2. 讀對應的 reference 檔。本 skill 的知識是分散式的。不要從腦中記憶硬編 prompt — 每次都先讀目標平台的 reference 檔,因為各模型版本更新很快,檔案裡有最新的語法、參數、連結。

  3. 中英混寫時有規則:主體名詞、運鏡術語、模型參數用英文;情感描述、文化元素 (漢服、水墨、國風)、旁白/歌詞用中文。Seedream 與 Kling 的中文支援最好;Midjourney、Flux、Runway、Veo 英文效果明顯較佳。

  4. Prompt 長度。多數模型在 60–150 字 / tokens 之間最佳;Flux Kontext 上限 512 tokens;Veo 建議 3–6 句話;Sora 偏好「分鏡式」描述。reference 檔有每個模型的具體上限。

  5. 輸出格式。除非使用者明說只要 prompt 本文,否則給他們:

    • 一個 可複製的 prompt 區塊 (通常英文)
    • 一段 繁中說明:這個 prompt 為什麼這樣寫、哪些 token 可以換掉、預期輸出會長怎樣
    • 一組 建議參數 (aspect ratio、時長、model variant、seed 等)
    • 如果有平台特殊語法 (tag、metatag、參考圖槽位),把它結構化呈現

工作流

Step 1 — 選平台

如果使用者已指定平台 (「用 Kling」「幫我寫 Suno prompt」),直接跳 Step 2。

否則讀 references/selector.md 按「媒體類型 × 用途 × 資源」選出 1–2 個最佳候選平台,並把推薦理由用 2–3 句話告訴使用者。如果落差很大 (例如「免費 vs 付費最佳」),給使用者選擇。

Step 2 — 讀對應的 reference 檔

根據選定平台,一定要 讀對應檔案,不要憑記憶寫 prompt:

🧭 不確定選哪個模型?先讀 references/model-picker.md — OiiOii 全模型 + 各平台「選誰 / 招牌 prompt 技巧」總表 + 30 秒決策樹。使用者嫌「不熟模型」「該用哪個」時必讀。

影片 (Video)

圖片 (Image)

音樂 (Music)

複合 / 多智能體

跨平台共通

進階導演 / VFX / 音效 / 剪接 級別 prompt 設計 (當使用者要「電影級」「廣告級」「奢侈品級」「完整影視團隊」時必讀)

  • 電影導演 / 攝影指導 / 底片 / 燈光 / 構圖 / meta tokens → references/cinematic-direction.md
  • 廣告 / 時尚 / MV 導演 / 品牌調性 / 社群短影音 → references/commercial-direction.md
  • VFX 總監 / 大氣 / 物理 / 特效 recipes → references/vfx-effects.md
  • 音效設計 (對白 / SFX / Foley / 配樂) → references/sound-design.md — Veo 3.1 / Vidu Q3 原生音訊必讀
  • 剪接 / 轉場 / 節奏 (match cut / whip pan / J-L cut / ASL 律動) → references/editing-transitions.md — 多鏡故事、storyboard、MV 必讀
  • 🎬 概念先行 Prompting (prompt 沒主題/畫面很空/效果差/不知道在拍什麼) → references/concept-first-prompting.md — 使用者嫌「prompt 沒主題」「有些畫面不知道在幹嘛」「影片效果差」時必讀。治「技術詞堆疊但空洞」:先定一句話 concept → 3-beat 敘事弧 → 每個鏡頭做存在性測試 → 才加技術詞。含 6 個產品廣告概念框架
  • ⚙️ Prompt 組裝引擎(有 concept 了但不知怎麼組成強 prompt / 想更快 / 要品質評分)→ references/prompt-craft-engine.md — 更強+更快的 META 引擎:通用 6 層解剖(弱形容詞→強具體事實 before/after)+ 6 種媒材 <60 秒填空骨架 + 10 分品質評分閘(≥8 才送)。流程位置:concept-first →(本檔組裝+評分)→ vocab 檔填字 → quality-control 修
  • 💎 精緻度分級(更精美)(成品要像 editorial/commercial-award 級 / 反 AI 塑膠感 / 要「貴」)→ references/aesthetic-grade.md — DP 級可貼入配方:打光 setup(方向+質感+光比)、調色科學(Kelvin/hue)、材質表面渲染(玻璃/金屬/皮膚/食物反塑膠 phrase)、鏡頭光學、大氣。比 cinematic-direction 的選單更深、給整段配方
  • 🔧 反瑕疵品質控制 (結果爛了/有瑕疵/不夠精緻怎麼修) → references/quality-control.md — 使用者嫌結果有瑕疵、產品變形、塑膠感、閃爍、文字鬼影、動得很假時必讀。系統化分類 7 類瑕疵 + 對症修法(鎖時間/鎖物理/鎖主體)+ 平台物理強弱速查
  • 🧭 模型選擇大全 (該用哪個模型 / 不熟各家模型) → references/model-picker.md — OiiOii 全模型 + Gemini Omni / Kling O-series / Vidu / Hailuo / Wan / Seedance 各家「最強情境 + 招牌 prompt 技巧 + 何時選」+ 命名混淆對照(Gemini Omni vs Kling Omni)
  • 🎬 多模型影片操作手冊 (Wan 2.7 / Kling 3.0 Omni / HappyHorse 怎麼下 prompt / 拉片復刻 / 多模型省成本) → references/multimodel-video-cheatsheet.md — Wan 2.7 Thinking Mode 完整 brief + Kling Omni 身份鎖(@element)+物理事件 + HappyHorse 短 prompt+原生音訊 + 拉片復刻先選模式 + 多模型 tier-by-cost/frame-chaining/per-model 重寫 + shot-type 選模型表。用 OiiOii 新影片模型或要省成本跑多模型時必讀(2026-06 研究+實測)
  • 📋 經驗證 Prompt 庫 (要現成高品質 prompt / 不想從零寫 / 要範本) → references/proven-prompts.md — 26+ 個逐字可貼的高品質 prompt(Google 官方 Veo/Omni guide、Kling 官方、Seedance/Vidu @語法實例、廣告/美食/汽車/精品/多鏡頭/對白),按用途分類 + 來源 + 跨模型語法速查。使用者要「更好的 prompt / 範例 / 範本」時必讀

這三個進階檔是 語彙庫,不是流程手冊。使用方式:

  1. 先照平台的 reference (kling.md / flux.md 等) 確定該平台的 prompt 結構
  2. 再從進階檔挑 5-8 個高質量 token 填進結構裡
  3. 不要整段 copy — 挑對情緒、對平台、對故事的那幾個詞

一個人身高不同,鏡頭焦段/燈光/底片的「關鍵詞組合」就不同。「超級資深影視導演」的 prompt = 四層堆疊 [導演/DP] + [鏡頭/底片] + [燈光/色調] + [動作/構圖],每層挑最契合的 1-2 個 token。

Show full SKILL.md (620 more words)Show less
Step 3 — 根據任務類型參考 template

preset-packs.md 的用法: 使用者要「Wes Anderson 風」「Nike 廣告感」「賽博龐克雨夜」「水下夢境」等明確風格時,先到 preset-packs.md 找最近的 preset,換占位符即可,不用每次從零組 prompt。若使用者要的風格不在 preset 裡,再從 cinematic-direction.md / commercial-direction.md / vfx-effects.md 現場組。

Step 4 — 組出 prompt 並輸出

依 reference 檔給的公式組出 prompt。每個平台有自己的 order、特殊符號、tag — 嚴格按 reference 檔的格式寫。

輸出格式:可複製的 prompt 區塊 + negative + 參數 + 繁中 why (每個 token 為什麼選)。完整範例見 preset-packs.md 的 30 個 preset,每個都是 ready-to-use 格式。

Step 5 (選配) — 瀏覽器自動化執行

如果使用者要求「幫我直接做」「幫我貼上去按產生」「操作 OiiOii/Kling/Suno」等,這套檔要讀:

1. 通用 click 協議 (所有網站都要遵循)

  • automation/click-protocol.md — Click 決策樹、ref 保質期、座標精準技巧、驗證 selected 的方法、scroll-aware 座標、paywall 偵測、等待策略、驗證循環、速度優化

2. 網站特定 profile (該站的 UI 地圖 + 座標 + 陷阱) — 索引見 site-profiles/README.md

✅ 完整驗證(可放心自動化):

  • oiioii.md — OiiOii 整合,§0 = 2026-06-08 大改版 SOURCE OF TRUTH(新入口/Agent▾選模型/i2v/一致性/新模型)
  • flow.md — Google Veo 3.1 + Gemini Omni Flash(完整 Phase 1-6 + 生成偵測看 gallery tile)
  • kling.md — Kling 3.0(完整 Phase 1-7 + credits + Fast-Track)

⚠️ 部分/stub(有 UI 地圖但未全驗): suno / midjourney / seedream / runway / vidu / ideogram

🔴 停運: sora(不支援自動化,新案改 Veo/Kling/Runway)

📝 _template.md — 新網站 profile 模板

使用者說「幫我直接操作 XXX」時:先確認該站在上列。未列出者告知「該平台暫不支援全自動化,但我可代寫 prompt 讓你貼上去」。

3. 高層流程速查 (每站登入與主流程概述)

送出前必做的安全檢查 (見 click-protocol.md 詳細):

  • 確認使用者已登入目標站 (絕不代輸入密碼)
  • 把即將輸入的 prompt 完整貼給使用者,等確認再送
  • 首次操作某站 → screenshot 對照 site-profile,UI 改版就更新 profile
  • 送出前 screenshot (before),送出後 screenshot (after),比對 UI 變化確認真送出
  • 撞到 paywall / modal / 非預期彈窗 → 立刻停下問使用者,不代付款

reliability 核心原則 (來自 OiiOii demo 踩坑紀錄):

  1. 不信任記憶中的座標 — 5 秒前的 screenshot 可能已過期,click 前必須剛截圖
  2. ref 很短命 — find 完 立刻 click,中間不插其他 tool call
  3. selected state 看填底色,不是粉色 ring (那是 hover)
  4. click 沒中別亂重試 — 先 screenshot 診斷 (paywall?disabled?座標偏?)
  5. Polling 不要 < 60s 也不要 > 3min — 互動 UI 60-90s,純背後等待 3-5min

⚡ Chain workflow 強制節省 (來自 Suno 5 首歌 chain 慘痛教訓 2026-04-20): 連跑多任務 (做 5 首歌、5 張圖、多支影片) 時:

  1. 中間禁 screenshot — 只第 1 task 後 + 全部完成各 1 張 (浪費點 #1)
  2. 中間禁 TodoWrite — chain 完成後一次性更新即可 (浪費點 #2)
  3. clear field 用 1-click — 點 trash icon 或 ctrl+a 覆寫,禁用 triple-click + Ctrl+A + Delete 三步走
  4. 內容寫標準長度 — Suno 歌詞 ~25 行 / Veo prompt ~80 字 / MJ keyword 30-50 字 / Seedream 80-120 字,超寫不加分
  5. 平台有 Series Mode/Multi-Shot/Multi-Reference → 用內建批次 不要 N 次 chain (Suno Persona / Kling Multi-Shot / Seedream Series / MJ omni-ref / Vidu Q3 multi-entity)

檢測標準: 5 task chain 應 ≤ 36 tool calls + ≤ 1.5k token + < 5 分鐘。違反任一就要修。詳見 click-protocol.md §「Token + 時間最佳化」。

關鍵品質檢查 (在交付 prompt 前跑一遍)

  • 有沒有「具體的主體描述」— 不只是 "a person" 而是 "a woman in her 30s with chin-length silver hair"
  • 動作有沒有可視化的動詞 — "walks slowly" / "tilts her head" 優於 "is moving"
  • 場景有沒有 2–4 個具體元素 — 時間、地點、天氣、物件
  • 鏡頭 (如果是影片) — 運鏡類型 + 速度 + 距離
  • 風格與光影 — 一到兩個清楚的風格錨 (例如 "cinematic, teal-and-orange grade")
  • 長度適當 — 檢查目標平台的 sweet spot
  • 沒有自相矛盾 — 不要同時寫 "close-up" 和 "wide establishing shot"
  • 負面提示詞 (如該模型支援) 有列常見缺陷
  • 產品/品牌類 — 有沒有「主體完整性鎖」(rigid form / no morphing / 形狀不變)?產品優先走 i2v 鎖死 hero 圖。見 quality-control.md §1
  • 複雜物理 (液體/布料/煙/碰撞) — 有沒有挑對強物理平台 (Runway Gen-4.5 / Veo 3.1)?弱物理模型硬做必爛

常見反模式

  • 拿 Midjourney 的 --ar 16:9 --s 250 語法餵 Flux / Kling / Sora — 只會被當成雜訊。每個平台有自己的參數欄位。
  • 在 Suno Lyrics 欄寫 "Verse:" — 會被當成歌詞唱出來。要寫 [Verse]。
  • Seedream 叫它寫文字時用單引號或不加引號 — 官方建議用雙引號包住要顯示的字:"Seedream 5"。
  • Kling / Seedance 運鏡疊太多 — "zoom in while panning right and tilting up" 會爛掉。一次一到兩個運鏡。
  • Vidu ref2v 時 prompt 又重複描述參考圖的外觀 — 會互相打架。描述「動作與互動」就好。
  • SD 3.5 用 (word:1.3) 權重語法 — 3.5 不吃這個,改用自然語言。SDXL 才吃。

關於「萬用」的邊界

此 skill 設計為 prompt engineering + 瀏覽器自動化雙層。不直接呼叫 API (API key 管理、計費都是額外議題),但若使用者明確要走 API,reference 檔裡有各平台的官方 API 端點連結,可以協助組 payload。

版本資訊

平台知識最後校準:2026-08(Seedance 2.5 社群案例、長敘事一致性與跨產業應用擴充)。 各 reference 檔末尾有官方文件連結 — 若要執行會花錢/產生後果的操作前,優先拿 reference 連結當最終依據,因為版本/定價變動快。

外部資源(官方文檔 / API / 社群 / 緊急更新連結) → references/external-resources.md。2026 模型發布時間軸(誰新誰舊/替代對象/何時選)見該檔的 Model Release Timeline 表。

2026-06 重大變動(已驗證):

  • 🔴 OpenAI Sora 2 停運 — app/web 2026-04-26 已關,API 2026-09-24 關,資料永久刪除。新任務勿用,改 Runway Gen-4.5 / Veo 3.1 / Kling 3.0。
  • 🆕 Runway Gen-4.5(2025-12-01)— Video Arena 榜首 1247 Elo,勝 Veo 3 / Sora 2 Pro,物理最強。
  • 🆕 Midjourney V8.1(2026-04-30)— 原生 2K HD、快 3-5×;⚠️ --oref Omni Reference 為 V7-only,V8 不支援。
  • 🆕 Suno v5.5(2026-03)— Personas 改名 Voices + Custom Models + Suno Studio DAW。
  • 🆕 Kling 3.0 / O-series、Seedream 5.0 Lite、Vidu Q3、Ideogram 3.0 為各家當前主力。
  • 🆕 Google Gemini Omni(I/O 2026, 2026-05-19)— any-to-any 多模態影片「影片版 Nano Banana」,對話式編輯,Flash 版已上 Flow PRO。Google 影片變雙旗艦(Veo 3.1 + Gemini Omni)。⚠️ ≠ Kling 3.0 Omni(同名不同家)。
  • 各模型「最強情境 + 招牌 prompt 技巧」見 references/model-picker.md;「prompt 沒主題/畫面空/效果差」見 references/concept-first-prompting.md。

2026-08 Seedance 擴充:

  • Seedance 2.5 採 provider capability gate,不把社群介面標籤誤當所有供應商都已支援;送出前核對實際模型名稱、時長、音訊與參考素材能力。
  • 新增 30 秒、多參考、UGC/Vlog、硬切、提示詞拆解、首鏡後鎖定、角色/髮飾/服裝/道具一致性、時間凍結/倒轉與複雜物理的專用手冊。
  • 新增商業、建商、廣告、產品、汽車、美食、電影與動畫的領域骨架;AI 短劇的故事架構仍交由 ai-short-drama,本 Skill 負責下游媒體提示詞。

各模型「禁忌 / 版本推翻歷史」見 references/community-prompt-patterns.md;結果有瑕疵的修法見 references/quality-control.md。

© Hao0321, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 70 other files (references) in the repository root of Hao0321/ai-media-generator.

  • SKILL.md
  • .gitignore
  • CHANGELOG.md
  • CONTRIBUTING.md
  • LICENSE
  • README.md
  • agents/openai.yaml
  • automation/browser-guide.md
  • automation/click-protocol.md
  • automation/site-profiles/README.md
  • automation/site-profiles/_template.md
  • automation/site-profiles/flow.md
  • automation/site-profiles/ideogram.md
  • automation/site-profiles/kling.md
  • automation/site-profiles/midjourney.md
  • automation/site-profiles/oiioii.md
  • automation/site-profiles/runway.md
  • automation/site-profiles/seedream.md
  • … and 53 more

Open the folder on GitHubat commit edc8aa5

Compare with similar skills

AI Media Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Media Generator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Media Generator this skillHao0321/ai-media-generator258—~5.6kAutomated safety check: PassMIT
Forge Media Route Layer0x0funky/agent-sprite-forge4.4k—~2.2kAutomated safety check: PassMIT
Nbcraftjieyefriic/nbcraft155—~2.8kAutomated safety check: PassMIT
Arcads External APIkrusemediallc/arcads-claude-code1.6k—~8.7kAutomated safety check: NotesMIT
Fal AI Mediaaffaan-m/ECC276k4 repos~1.9kAutomated safety check: PassMIT
Fal AI Mediaaffaan-m/ECC276k2 repos~1.2kAutomated safety check: PassMIT

Similar skills

  • Forge Media Route Layer

    0x0funky/agent-sprite-forge

    Generates an image or an image-to-video clip through a configured provider API or a signed-in Codex or Grok CLI, and reports the route, file, hash and cost estimate.

    4.4k GitHub stars~2.2k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Nbcraft

    jieyefriic/nbcraft

    Multi-backend Image + Video Generation CLI (nb command). An agent skill from jieyefriic/nbcraft.

    155 GitHub stars~2.8k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Arcads External API

    krusemediallc/arcads-claude-code

    Creates and retrieves AI video and image-related assets via the Arcads external API (Seedance 2.0, Sora 2, Veo 3.1, Kling, Grok Video, Nano Banana, b-roll, scene, script/actor flows).

    1.6k GitHub stars~8.7k tokensUpdated 17 days ago
    Media & CreativeAuto-check: notes
  • Fal AI Media

    affaan-m/ECC

    Unified media generation via fal.ai MCP — image, video, and audio.

    276k GitHub starsUsed in 4 repos~1.9k tokens
    Media & CreativeAuto-check passed
  • Fal AI Media

    affaan-m/ECC

    通过 fal.ai MCP 实现统一的媒体生成——图像、视频和音频。涵盖文本到图像(Nano Banana)、文本/图像到视频(Seedance、Kling、Veo 3)、文本到语音(CSM-1B),以及视频到音频(ThinkSound)。当用户想要使用 AI 生成图像、视频或音频时使用。

    276k GitHub starsUsed in 2 repos~1.2k tokens
    Media & CreativeAuto-check passed
  • Fal AI Media

    affaan-m/ECC

    fal.ai MCPによる統合メディア生成(画像、動画、音声)。テキストから画像(Nano Banana)、テキスト/画像から動画(Seedance、Kling、Veo 3)、テキストから音声(CSM-1B)、動画から音声(ThinkSound)をカバーします。ユーザーがAIで画像、動画、音声を生成したい場合に使用します。

    276k GitHub stars~1.4k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed

Questions about AI Media Generator

What does AI Media Generator do?

為使用者產生高品質的 AI 生圖、生影片、生音樂提示詞,並在需要時透過瀏覽器自動化實際送到目標平台。涵蓋 OiiOii、Kling 3.0/O-series、Seedance 2.0/2.5、Suno v5.5、Seedream 5.0/4.0、Vidu Q3、Midjourney V8.1、Flux 1.1 Pro / Kontext、Runway Gen-4.5 /…. AI Media Generator is an agent skill from Hao0321/ai-media-generator.

When should I use AI Media Generator?

AI Media Generator fits situations like: tasks that involve AI video generation.

How do I install AI Media Generator in Claude Code?

Run `npx skills add Hao0321/ai-media-generator --skill ai-media-generator -a claude-code`. Or copy the skill folder (the Hao0321/ai-media-generator repository) into .claude/skills/ai-media-generator in your project. Claude Code loads it when a task matches its description.

How do I install AI Media Generator in Codex?

Run `npx skills add Hao0321/ai-media-generator --skill ai-media-generator -a codex`. Or copy the skill folder (the Hao0321/ai-media-generator repository) into .agents/skills/ai-media-generator in your project. Codex loads it when a task matches its description.

Can I use AI Media Generator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Hao0321/ai-media-generator --skill ai-media-generator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-media-generator, .gemini/skills/ai-media-generator, .github/skills/ai-media-generator and .opencode/skills/ai-media-generator in your project.

What does AI Media Generator need to run?

Going by SKILL.md and its folder, AI Media Generator needs the command-line tools its instructions call (bash).

Does AI Media Generator access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is AI Media Generator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Media Generator use?

AI Media Generator is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Media Generator use?

About 5.6k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 154k tokens, read only when the agent opens those files.

What are the alternatives to AI Media Generator?

Skills that share tags, products or a category with AI Media Generator: Forge Media Route Layer (0x0funky/agent-sprite-forge, 4.4k stars), Nbcraft (jieyefriic/nbcraft, 155 stars), Arcads External API (krusemediallc/arcads-claude-code, 1.6k stars) and Fal AI Media (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Media Generator?

Hao0321 (a GitHub user) maintains it in Hao0321/ai-media-generator, which has 258 GitHub stars. The repository was last updated on August 9, 2026.

Source: Hao0321/ai-media-generator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.