Agent skill

Multimodal Gen

by aAAaqwq in aAAaqwq/AGI-Super-Team

多模态内容生成(图片、视频)。当用户需要生成图片、生成图像、生成视频、AI绘画、AI作图、画一张图、做个视频、文生图、文生视频时使用此技能。自动调用 multimodal-agent 进行生成。

MITAuto-check: notesMedia & Creative

Install Multimodal Gen

skills CLI
$ npx skills add aAAaqwq/AGI-Super-Team --skill multimodal-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aAAaqwq/AGI-Super-Team multimodal-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aAAaqwq/AGI-Super-Team.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/multimodal-gen .claude/skills/multimodal-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multimodal-gen
GitHub stars
105
Used in
1 other repo
Token cost
~580 tokens
SKILL.md length
107 words
Files
5
Skills in repo
167
Repo updated
First seen
Licence
MIT

At a glance

多模态内容生成(图片、视频)。当用户需要生成图片、生成图像、生成视频、AI绘画、AI作图、画一张图、做个视频、文生图、文生视频时使用此技能。自动调用 multimodal-agent 进行生成。

  • Works in 4 steps: 提示词优化: multimodal-agent 会自动优化用户的描述 → 模型选择: 如果用户没指定,默认使用 flux (图片) 或 veopro (视频) → 异步执行: 视频生成可能需要较长时间,会在后台执行 → …
  • Tasks that involve AI video generation
  • SKILL.md covers 触发场景, 使用方式, 执行流程 and 示例, plus 2 more sections
  • Runs Python scripts from its folder

What it does

Multimodal Gen is an agent skill from aAAaqwq/AGI-Super-Team. 多模态内容生成(图片、视频)。当用户需要生成图片、生成图像、生成视频、AI绘画、AI作图、画一张图、做个视频、文生图、文生视频时使用此技能。自动调用 multimodal-agent 进行生成。

Its SKILL.md is about 580 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `generate.py`, `generate_image.py` and `generate_video.py`).

It sits in Media & Creative, covering AI video generation. The repository describes itself as: An installable, cross-framework AI organization: C-suite agents, expert subagents, curated skills, independent review, and one-command setup across 18 AI client/runtime adapters. The licence is MIT.

When your agent uses it

  • Tasks that involve AI video generation

Example prompts

  • “/multimodal-gen”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): sessions_spawn, Bash, Read, Write, Edit

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. 提示词优化: multimodal-agent 会自动优化用户的描述
  2. 模型选择: 如果用户没指定,默认使用 flux (图片) 或 veopro (视频)
  3. 异步执行: 视频生成可能需要较长时间,会在后台执行
  4. 结果返回: 生成完成后会自动发送结果给用户

What it can do on your machine

Read from SKILL.md and the folder at commit 331ecd3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • sessions_spawn
    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multimodal Gen loads about 580 tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 107 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~580

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: sessions_spawn, Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aAAaqwq/AGI-Super-Team at commit 331ecd3, republished under its MIT licence (© aAAaqwq). 107 words, ~580 tokens.

Download SKILL.mdSave it as .claude/skills/multimodal-gen/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
multimodal-gen
description
多模态内容生成(图片、视频)。当用户需要生成图片、生成图像、生成视频、AI绘画、AI作图、画一张图、做个视频、文生图、文生视频时使用此技能。自动调用 multimodal-agent 进行生成。
allowed-tools
sessions_spawn, Bash, Read, Write, Edit

多模态内容生成

当用户需要生成图片或视频时,自动调用 multimodal-agent 进行处理。

触发场景

图片生成
  • "生成一张图片"
  • "画一张..."
  • "AI 作图"
  • "文生图"
  • "帮我生成图像"
  • "用 flux/imagen/dalle 生成"
视频生成
  • "生成一个视频"
  • "做个视频"
  • "文生视频"
  • "用 veo/sora/kling 生成视频"

使用方式

自动调用 multimodal-agent
python
sessions_spawn(
    agentId="multimodal-agent",
    task="生成图片: {用户描述}, 使用 {模型} 模型"
)
可用模型
图片生成
别名模型特点
fluxflux-pro-max高质量,推荐
imagengoogle/imagen-4-ultraGoogle 最强
dallegpt-image-1DALL-E 3
doubaodoubao-seedream-4-5中式美学
klingimgkling-image可灵生图
视频生成
别名模型特点
veoproveo3.1-proGoogle 专业版
veo4kveo3.1-pro-4k4K 高清
sorasora-2-pro-allOpenAI Sora
klingkling-video可灵视频

执行流程

用户请求 "生成一张猫咪图片"
    │
    ▼
┌─────────────────────────────────────┐
│ 1. 识别为图片生成请求              │
│ 2. 提取描述: "猫咪"                │
│ 3. 选择默认模型: flux-pro-max      │
└─────────────────────────────────────┘
    │
    ▼
┌─────────────────────────────────────┐
│ sessions_spawn(                     │
│   agentId="multimodal-agent",       │
│   task="生成图片: 一只可爱的猫咪,   │
│         使用 flux 模型"             │
│ )                                   │
└─────────────────────────────────────┘
    │
    ▼
multimodal-agent 执行生成并返回结果

示例

生成图片
用户: 帮我生成一张日落海滩的图片

执行:
sessions_spawn(
    agentId="multimodal-agent",
    task="生成图片: 日落时分的海滩,金色阳光洒在海面上,使用 flux 模型"
)
生成视频
用户: 用 sora 生成一个猫咪玩耍的视频

执行:
sessions_spawn(
    agentId="multimodal-agent",
    task="生成视频: 一只可爱的猫咪在草地上玩耍,使用 sora 模型"
)
指定模型
用户: 用 doubao 画一张中国风山水画

执行:
sessions_spawn(
    agentId="multimodal-agent",
    task="生成图片: 中国风山水画,云雾缭绕的山峰,使用 doubao 模型"
)

模型选择建议

场景推荐模型
通用高质量flux
中式风格doubao
写实照片imagen
创意艺术dalle
高清视频veo4k
电影级视频sora
快速视频kling

注意事项

  1. 提示词优化: multimodal-agent 会自动优化用户的描述
  2. 模型选择: 如果用户没指定,默认使用 flux (图片) 或 veopro (视频)
  3. 异步执行: 视频生成可能需要较长时间,会在后台执行
  4. 结果返回: 生成完成后会自动发送结果给用户

© aAAaqwq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/multimodal-gen of aAAaqwq/AGI-Super-Team.

  • SKILL.md
  • generate.py
  • generate_image.py
  • generate_video.py
  • prompt_optimizer.py

Open the folder on GitHubat commit 331ecd3

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in aAAaqwq/AGI-Super-Team, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Multimodal Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multimodal Gen compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multimodal Gen this skillaAAaqwq/AGI-Super-Team1051 repos~580Automated safety check: NotesMIT
Video Generationbytedance/deer-flow84k3 repos~1.4kAutomated safety check: PassMIT
Video Cover Imageitwanger/toBeBetterJavaer18k—~3.3kAutomated safety check: PassNone
Seedancesongguoxs/seedance-prompt-skill2.9k1 repos~2.5kAutomated safety check: PassNone
HyperFrames Video Entry Pointheygen-com/hyperframes59k3 repos~5.2kAutomated safety check: PassApache-2.0
Lanshu Create AI Presenter Videocclank/lanshu-create-ai-presenter-video2.6k—~3.6kAutomated safety check: PassMIT

Similar skills

  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    84k GitHub starsUsed in 3 repos~1.4k tokens
    Media & CreativeAuto-check passed
  • Video Cover Image

    itwanger/toBeBetterJavaer

    Generate matched 3:4, 16:9, and 4:3 short-video cover images from toBeBetterJavaer video scripts or AI/Java technical topics.

    18k GitHub stars~3.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Seedance

    songguoxs/seedance-prompt-skill

    This skill should be used when the user asks to "generate video prompts", "create Seedance prompts", "write video descriptions", mentions "Seedance", "seedance", "即梦", "即梦平台", "视频提示词", "视频生成"…

    2.9k GitHub starsUsed in 1 repo~2.5k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    59k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Lanshu Create AI Presenter Video

    cclank/lanshu-create-ai-presenter-video

    Turn a topic or finished script into a complete, publish-ready explainer video — led by an AI presenter from an authorized adult presenter image, or performed in one of nine visual explainer styles…

    2.6k GitHub stars~3.6k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Video Shots

    eternityspring/reelbench-skills

    拉片:把一条成片拆成逐镜头的分析表——每个镜头的时长、景别、类别、运镜、画面. An agent skill from eternityspring/reelbench-skills.

    872 GitHub starsUsed in 1 repo~1.8k tokens
    Media & CreativeAuto-check: notes

More from aAAaqwq/AGI-Super-Team

All 167 skills in this repo
  • Content Creator

    aAAaqwq/AGI-Super-Team

    Create SEO-optimized marketing content with consistent brand voice.

    105 GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • Financial Calculator

    aAAaqwq/AGI-Super-Team

    Advanced financial calculator with future value tables, present value, discount calculations, markup pricing, and compound interest.

    105 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Bankr Signals

    aAAaqwq/AGI-Super-Team

    Transaction-verified trading signals on Base blockchain. An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub starsUsed in 2 repos~3.3k tokens
    Auto-check passed
  • Erc 8004

    aAAaqwq/AGI-Super-Team

    Register AI agents on Ethereum mainnet using ERC-8004 (Trustless Agents).

    105 GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Frontend Design Ultimate

    aAAaqwq/AGI-Super-Team

    Create distinctive, production-grade static sites with React, Tailwind CSS, and shadcn/ui — no mockups needed.

    105 GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Zsxq Smart Publish

    aAAaqwq/AGI-Super-Team

    Publish and manage content on 知识星球 (zsxq.com). An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Questions about Multimodal Gen

What does Multimodal Gen do?

多模态内容生成(图片、视频)。当用户需要生成图片、生成图像、生成视频、AI绘画、AI作图、画一张图、做个视频、文生图、文生视频时使用此技能。自动调用 multimodal-agent 进行生成。. Multimodal Gen is an agent skill from aAAaqwq/AGI-Super-Team.

When should I use Multimodal Gen?

Multimodal Gen fits situations like: tasks that involve AI video generation.

How do I install Multimodal Gen in Claude Code?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill multimodal-gen -a claude-code`. Or copy the skill folder (skills/multimodal-gen in aAAaqwq/AGI-Super-Team) into .claude/skills/multimodal-gen in your project. Claude Code loads it when a task matches its description.

How do I install Multimodal Gen in Codex?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill multimodal-gen -a codex`. Or copy the skill folder (skills/multimodal-gen in aAAaqwq/AGI-Super-Team) into .agents/skills/multimodal-gen in your project. Codex loads it when a task matches its description.

Can I use Multimodal Gen in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aAAaqwq/AGI-Super-Team --skill multimodal-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multimodal-gen, .gemini/skills/multimodal-gen, .github/skills/multimodal-gen and .opencode/skills/multimodal-gen in your project.

What does Multimodal Gen need to run?

Going by SKILL.md and its folder, Multimodal Gen needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: sessions_spawn, Bash, Read, Write, Edit.

Does Multimodal Gen access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Multimodal Gen safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Multimodal Gen use?

Multimodal Gen is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multimodal Gen use?

About 580 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Multimodal Gen?

Skills that share tags, products or a category with Multimodal Gen: Video Generation (bytedance/deer-flow, 84k stars), Video Cover Image (itwanger/toBeBetterJavaer, 18k stars), Seedance (songguoxs/seedance-prompt-skill, 2.9k stars) and HyperFrames Video Entry Point (heygen-com/hyperframes, 59k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multimodal Gen?

aAAaqwq (a GitHub user) maintains it in aAAaqwq/AGI-Super-Team, which has 105 GitHub stars. The repository holds 167 skills in this directory. The repository was last updated on October 8, 2026.

Source: aAAaqwq/AGI-Super-Team on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.