Agent skill

Huashu Slides

by alchaincyf in alchaincyf/huashu-skills

从内容到成品PPTX的端到端演示文稿制作,含AI插画生成和18种设计风格。当用户提到"做PPT"、"做幻灯片"、"演示文稿"、"Keynote"、"slides"时使用。

MITAuto-check: notesDocuments & Office

Install Huashu Slides

skills CLI
$ npx skills add alchaincyf/huashu-skills --skill huashu-slides -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alchaincyf/huashu-skills huashu-slides --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alchaincyf/huashu-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/huashu-slides .claude/skills/huashu-slides && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
huashu-slides
GitHub stars
1.7k
Token cost
~7.2k tokens
SKILL.md length
2,759 words
Files
27 (incl. scripts, references, assets)
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

从内容到成品PPTX的端到端演示文稿制作,含AI插画生成和18种设计风格。当用户提到"做PPT"、"做幻灯片"、"演示文稿"、"Keynote"、"slides"时使用。

  • Works in 6 steps: Choose Workflow Settings → Content Structuring → Design System → …
  • Tasks that involve Slides and decks
  • SKILL.md covers Step 0: Choose Workflow Settings, Step 1: Content Structuring, Step 2: Design System and Step 3: Build Slides, plus 5 more sections
  • Calls uv and npx; needs GEMINI_API_KEY

What it does

Huashu Slides is an agent skill from alchaincyf/huashu-skills. 从内容到成品PPTX的端到端演示文稿制作,含AI插画生成和18种设计风格。当用户提到"做PPT"、"做幻灯片"、"演示文稿"、"Keynote"、"slides"时使用。

Its SKILL.md is about 7.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 28 other files, including scripts, reference files and assets.

It sits in Documents & Office, covering Slides and decks and PowerPoint presentations. It works with Microsoft PowerPoint. The repository describes itself as: 花叔全部开源 Agent Skills 总目录:16 旗舰 + 14 人物视角 + 22 内置共 52 个 skill,分层分类 + AI Agent 安装协议 + 机器可读 skills.json + 更新检查机制. The licence is MIT.

When your agent uses it

  • Tasks that involve Slides and decks
  • Tasks that involve PowerPoint presentations

Example prompts

  • “Keynote”
  • “slides”
  • “/huashu-slides”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Choose Workflow Settings
  2. Content Structuring
  3. Design System
  4. Build Slides
  5. PPTX Assembly
  6. Preview & Polish

What it can do on your machine

Read from SKILL.md and the folder at commit 2efea35. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Huashu Slides loads about 7.2k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 25 tokens; SKILL.md has 2,759 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~7.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:336
    export $(grep GEMINI_API_KEY ~/.claude/.env) && \
  • NoteMentions a .env fileSKILL.md:503
    export $(grep GEMINI_API_KEY ~/.claude/.env) && \

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alchaincyf/huashu-skills at commit 2efea35, republished under its MIT licence (© alchaincyf). 2,759 words, ~7,199 tokens.

Download SKILL.mdSave it as .claude/skills/huashu-slides/SKILL.md (or your agent's skills folder). This skill also uses 26 other files; get the full folder from GitHub.
name
huashu-slides
description
从内容到成品PPTX的端到端演示文稿制作,含AI插画生成和18种设计风格。当用户提到"做PPT"、"做幻灯片"、"演示文稿"、"Keynote"、"slides"时使用。

AI Presentation Workflow

Create professional presentations: Content → Design → Build → Assembly → Polish.

Step 0: Choose Workflow Settings

At the start of every presentation task, ask the user TWO choices:

0-A. Collaboration Mode
ModeDescriptionCheckpoints
Full AutoMinimal interaction. Confirm topic only, deliver final PPTX.1 checkpoint
Guided (recommended)Confirm outline, pick design, preview before assembly.3 checkpoints
CollaborativeReview every slide, approve every illustration, full control.Per-slide

If the user doesn't specify, default to Guided mode.

0-B. Assembly Method
MethodHow it worksBest for
Editable HTML (Path A)HTML slides + selective AI illustrations → html2pptx → editable PPTXNeed to edit text later, precise layout, corporate decks
Full AI Visual (Path B)Every slide as a complete AI-generated image → create_slides.py → image PPTXMaximum visual impact, artistic presentations, quick drafts

Trade-offs:

Path A: Editable HTMLPath B: Full AI Visual
TextEditable in PPTBaked into image (not editable)
Visual qualityGood with illustrationsExcellent — cohesive design
Layout controlPixel-preciseAI-interpreted
File sizeSmaller (~5-25MB)Larger (~30-80MB)
Chinese textPerfect (font rendering)Usually good (AI may occasionally misrender)
SpeedFaster (HTML creation)Slower (image generation per slide)

If the user doesn't specify, default to Path A (Editable HTML).


Step 1: Content Structuring

Turn raw material into a slide-by-slide outline.

Per slide, define:

  • Title — a complete assertion sentence (not a topic word)
  • Key points — 3-4 maximum
  • Visual type — illustration / chart / diagram / icon / quote
  • Path A: Illustration needed? — Yes/No. If yes, one-line description.
  • Path B: Visual scene description — one paragraph describing the complete slide visual (layout + imagery + mood).

Assertion-Evidence rule:

Bad titleGood title
Q3 SalesQ3销售增长23%,新用户是主要驱动力
Methodology我们通过双盲实验验证了这个结论

语言规则:slide内容一律用中文,仅保留必要的英文术语(人名、品牌名、技术专有名词)。 Section label(如 INSIGHT、TAKEAWAY)可用英文作为设计元素。

✅ Checkpoint 1 (Guided + Collaborative)

Present the outline as a table:

Path A:

| # | Title (assertion) | Key Points | Visual Type | Illustration? |
|---|-------------------|------------|-------------|---------------|
| 1 | Cover: ... | — | Decorative | Yes: ... |
| 2 | ... | 1. ... 2. ... | Chart | No |
| 3 | ... | 1. ... 2. ... | Illustration | Yes: ... |

Path B:

| # | Title (assertion) | Key Points | Visual Scene Description |
|---|-------------------|------------|--------------------------|
| 1 | Cover: ... | — | Dark gradient bg, large title centered, abstract network nodes |
| 2 | ... | 1. ... 2. ... | Split layout: text left, bar chart right, clean white bg |
| 3 | ... | 1. ... 2. ... | Full illustration: person at crossroads with floating clocks |

Ask the user:

  • Approve / adjust slide count
  • Path A: Approve / adjust which slides get illustrations
  • Path B: Approve / adjust visual scene descriptions
  • Any content to add or remove

Step 2: Design System

Present 3 design system options for the user to choose from. Each is a complete visual language, not just a color palette.

CRITICAL: A design system is NOT just colors. It defines visual philosophy, typography ratios, composition rules, and emotional intent. This is the difference between "boring PPT" and "magazine-quality deck."

🗣️ Style Discussion (Optional, if user wants to explore)

If the user says things like:

  • "我想要XX风格"(田中一光、瑞士国际主义、包豪斯、蒙德里安...)
  • "我不确定想要什么风格"
  • "能给我看看不同风格的例子吗"

Then consult the design movements reference: references/design-movements.md — 设计运动与风格参考库

This file maps classic design movements (Neo-Brutalism, Swiss Style, Bauhaus, etc.) to our AI-ready style presets. Use it to:

  1. Translate user's aesthetic language into actionable prompts
  2. Build shared vocabulary ("这个方向偏田中一光" vs "那个偏构成主义")
  3. Reference when designing new custom styles from scratch

After discussing movements, proceed to recommend 3 concrete presets below.


Design System Presets

⚠️ CRITICAL INSIGHT: 插画/漫画类风格的AI生成效果远好于「专业极简」类风格。 漫画/插画风格有明确的视觉语言(线条、角色、色块),AI可以充分发挥;极简风格(暗色底+发光文字+大量留白)缺乏视觉元素,生成出来「空」且「平」。

Pick 3 that match the topic/mood. Use the topic recommendation table below, then present each with its full description.

按主题自动推荐(优先从此表选):

主题类型第一推荐第二推荐第三推荐
品牌/产品介绍Snoopy温暖漫画Neo-Pop新波普浮世绘/敦煌(东方品牌)
教育/培训Neo-Brutalism学習漫画Snoopy温暖漫画
技术分享xkcd白板Neo-BrutalismLigne Claire
数据报告Pentagram编辑Fathom数据Ligne Claire
年轻受众Neo-Pop像素画孔版印刷
创意/艺术达达拼贴孔版印刷The Oatmeal
国风/东方敦煌壁画浮世绘Takram思辨
正式商务Pentagram编辑Müller-Brockmann网格Build极简
产品发布/keynote苏联构成主义Neo-PopPentagram编辑
内部分享Neo-BrutalismThe Oatmealxkcd白板
行业分析/咨询Fathom数据Pentagram编辑Müller-Brockmann网格
培训课件/教材Takram思辨温暖叙事学習漫画
投资/融资路演Build极简Pentagram编辑苏联构成主义

完整18种风格详细参考: references/proven-styles-gallery.md 风格样例图片: assets/style-samples/ 目录


第一梯队(强烈推荐,效果极好):

1. Warm Comic Strip — Snoopy温暖漫画风

  • Philosophy: Peanuts漫画的温暖与哲理感——简单的角色说着深刻的话,日常场景中蕴含人生智慧
  • Visual world: 圆头小孩、小狗、小鸟组成一个温暖的小世界。背景极简(草地、天空、狗屋、树)。色调像泛黄的报纸漫画
  • Reference: "Like a Peanuts comic strip — warm, philosophical, charming"
  • Style guide: references/proven-styles-snoopy.md
  • ⚠️ 关键经验: 不要在prompt中过度约束视觉细节(颜色比例、构图位置、角色姿势),否则会严重降低多样性。只描述情绪和内容,让AI自由发挥

2. Manga Educational — 学習漫画风

  • Philosophy: Japanese educational manga (学習漫画) — a character GUIDES you through the concept with reactions and drama
  • Colors: Bright and warm palette, white bg with selective color panels, screen-tone gray for emphasis areas
  • Ratio: 60% illustration / 30% text (in bubbles) / 10% effects
  • Typography: Bold manga-style titles with impact, body text in speech/thought bubbles, onomatopoeia as decorative elements. Size contrast 3:1
  • Composition: Dynamic manga panel layouts (3-5 panels per slide), character reactions drive emphasis, speed lines for energy, dramatic angles
  • Visual language: Expressive anime-style characters, reaction faces (surprise, confusion, eureka!), manga effects (sweat drops, sparkles, speed lines), panel borders with varied thickness
  • Reference: "Like a 'Manga Guide to Statistics' page — a character walks you through the concept, reacting with surprise and delight"

3. Ligne Claire Comics — 清线漫画风

  • Philosophy: Hergé's Tintin tradition — maximum information clarity through visual restraint
  • Colors: White/cream (#FFFDF7) bg, black (#000000) outlines, flat saturated fills (3-5 solid colors, no gradients)
  • Ratio: 70% clean bg / 20% illustration / 10% text
  • Typography: Hand-lettered feel for titles, clean sans-serif for body. Speech bubbles for key quotes. Title:body = 2.5:1
  • Composition: Panel-based layouts (2-4 panels per slide), sequential left-to-right reading flow, clear gutters between panels
  • Visual language: Uniform-weight outlines, flat colors without shading or hatching, no gradients, precise details but zero visual noise
  • Reference: "Like a Tintin page explaining a concept — every panel advances understanding, nothing is decorative"

4. Neo-Pop Magazine — 新波普杂志风

  • Philosophy: Youth media / streetwear brand aesthetic, bold and playful
  • Colors: Cream (#FFF8E7) bg, black (#000000) text, color-blocking with hot pink (#FF1493) + cyan (#00CED1) + yellow (#FFD700)
  • Ratio: 50% bg / 25% color blocks / 25% content
  • Typography: Headlines 40-50% of slide area (typography AS the visual), thick black borders around text blocks, 10:1 size ratio vs body
  • Composition: Modular color blocks with "controlled chaos", stacked asymmetric layouts, thick borders
  • Visual language: Pixel-art 8-bit icons, cutout photography, speech bubbles, bold graphic surfaces
  • Reference: "Like a Supreme lookbook meets a HYPEBEAST article — treats typography as graphic art"

第二梯队(推荐,特定场景效果好):

5. Whiteboard Sketch — xkcd白板手绘风

  • Philosophy: xkcd meets a professor's whiteboard — extreme minimalism forces focus on the idea itself
  • Colors: White (#FFFFFF) bg, black (#000000) ink, ONE accent color for emphasis (red #FF4444 or blue #4488FF)
  • Ratio: 85% white space / 10% sketch / 5% accent highlight
  • Typography: Hand-drawn/handwritten feel for everything, rough uneven baselines, arrows and annotations everywhere. Key numbers can be large (60pt+)
  • Composition: Freeform whiteboard layout, hand-drawn arrows connecting concepts, diagrams and stick figures, informal and alive
  • Visual language: Stick figures, hand-drawn charts and graphs, wobbly lines, annotation arrows, circled keywords, equation-style layouts
  • Reference: "Like an xkcd 'What If?' explanation — simple drawings that make complex ideas instantly click"

6. Soviet Constructivism — 苏联构成主义

  • Philosophy: Revolutionary propaganda poster — power through geometry and limited color
  • Colors: Revolutionary red (#CC0000) 40% + black (#1A1A1A) 25% + cream white (#F5E6D3) 30%
  • Typography: All text rotated 15-30 degrees, NO horizontal lines, bold condensed
  • Composition: Diagonal wedge from bottom-left to top-right, geometric shapes growing small to large (visual crescendo)
  • Visual language: NO gradients, pure flat fills + sharp edges, three-color limit, propaganda poster energy
  • Reference: "Like a 1920s Rodchenko poster — power, urgency, and geometric precision"

7. Warm Narrative — 温暖叙事风

  • Philosophy: Friendly storytelling, like a TED talk visual or Airbnb pitch deck
  • Colors: Warm cream (#FDF6EC) bg, dark charcoal (#3D3D3D) text, coral (#E17055) accent
  • Ratio: 60% warm bg / 25% content / 15% illustration
  • Typography: Headlines bold and warm, 3:1 ratio to body. Short sentences, not bullets
  • Composition: Illustration occupies 40-50% of slide, text wraps around visuals, rounded shapes
  • Visual language: Flat vector illustrations with warm palette, people-centric imagery, storytelling flow
  • Reference: "Like a Mailchimp or Notion brand presentation — approachable and human"

更多风格(第二/三梯队) 详见 references/proven-styles-gallery.md,包括:The Oatmeal信息图漫画、敦煌壁画、浮世绘、孔版印刷Risograph、等轴测Isometric、Bauhaus包豪斯、工程蓝图Blueprint、复古广告Vintage Ad、达达拼贴Collage、像素画Pixel Art


第四类:Professional / Editorial 设计系统(Path A 专用)

⚠️ 以下风格 强烈建议使用 Path A(HTML→PPTX)。它们依赖精确排版、数据可视化和网格系统,AI图片生成无法达到所需精度。口腔行业分析案例已验证 Path A + Pentagram编辑风格的出色效果。

8. Pentagram Editorial — 编辑杂志风(信息建筑派)

  • Philosophy: Pentagram/Michael Bierut — 字体即语言,网格即思想。用极度克制的设计让数据和内容自己说话
  • Colors: 奶油白(#FFFDF7) bg, 近黑(#1A1A1A) text, ONE accent color (如橙红#D4480B或品牌色)
  • Ratio: 60% whitespace / 30% content / 10% accent
  • Typography: 粗黑标题(28pt+) + 轻正文(10-13pt), 英文section label作为设计元素 (INSIGHT / PART 03)
  • Composition: 瑞士网格系统, 2px黑色边框卡片, 精确的水平分隔线, 数据可视化内嵌
  • Visual language: 极简图标, 条形图/饼图/趋势线, callout框, tag标签
  • Reference: "Like a McKinsey insight report meets Monocle magazine — data-rich but editorially elegant"
  • 执行路径: Path A only(HTML→PPTX)
  • 实战验证: 口腔行业分析15页deck(_temp/口腔行业分析/slides/)

9. Fathom Data Narrative — 数据叙事风(科学期刊派)

  • Philosophy: Fathom Information Design — 每一个像素都必须承载信息。科学严谨+设计优雅
  • Colors: 白(#FFFFFF) bg, 深灰(#333) text, 海军蓝(#1A365D) primary + 一个highlight color
  • Ratio: 50% charts/data / 30% text / 20% whitespace
  • Typography: GT America/Graphik风格的sans-serif, 大数字(60pt+)作为视觉锚点, 精确的脚注/来源标注
  • Composition: 高信息密度但不拥挤, 注释系统嵌入布局, small multiples图表阵列, 精确的时间线
  • Visual language: 散点图, 热力图, timeline, 带注释的图表, 数据标签精确到小数
  • Reference: "Like a Nature paper's data supplement meets a Bloomberg data feature"
  • 执行路径: Path A only(HTML→PPTX)

10. Müller-Brockmann Grid — 瑞士网格风(纯粹主义派)

  • Philosophy: Josef Müller-Brockmann — 客观性即美。数学精确的网格系统让任何混乱的信息变得有序
  • Colors: 白(#FFFFFF) bg, 黑(#000) text, 最多一个强调色
  • Ratio: 70% structured grid / 20% text / 10% accent
  • Typography: Akzidenz-Grotesk/Helvetica, 严格的8pt基线网格, 绝对左对齐, 字重对比(300 vs 700)
  • Composition: 8列数学网格, 所有元素对齐到网格线, 绝对不允许装饰元素, 功能主义至上
  • Visual language: 纯几何图形, 黑色线条表格, 精确对齐的列表, 无图标无插画
  • Reference: "Like the original Swiss Style poster — timeless, rational, zero decoration"
  • 执行路径: Path A only(HTML→PPTX)

11. Build Luxury Minimal — 奢侈极简风(当代品牌派)

  • Philosophy: Build Studio — 精致的简单比复杂更难。用大量留白和微妙字重变化传达高端感
  • Colors: 纯白(#FFFFFF) bg, 深灰(#2D2D2D) text, 单一accent(品牌色)极少量使用
  • Ratio: 75% whitespace / 15% text / 10% accent
  • Typography: 字重变化极微妙(200-600), 标题巨大(48pt+)但轻, 正文小而精(12pt), 字间距宽松
  • Composition: 黄金比例构图, 元素极少, 每页只说一件事, 呼吸感优先
  • Visual language: 高端产品图(如果有), 极简图标线条, 大面积纯色块, 圆角卡片
  • Reference: "Like an Apple keynote meets a Celine lookbook — confident restraint"
  • 执行路径: Path A(HTML→PPTX)

12. Takram Speculative — 日式思辨风(东方哲学派)

  • Philosophy: Takram — 技术是思考的媒介。用柔和的科技感和概念原型图传达深度思考
  • Colors: 暖灰(#F5F3EF) bg, 深灰(#3D3D3D) text, 鼠尾草绿(#8B9D77) accent
  • Ratio: 55% warm bg / 25% diagrams / 20% text
  • Typography: 圆润的sans-serif, 标题不用粗体而用大尺寸(36pt+), 正文温暖(14pt), 行高宽松(1.8)
  • Composition: 柔和阴影(blur 20px+), 圆角(16px+), 概念图/流程图作为核心视觉, 卡片式布局
  • Visual language: 概念原型图, 柔和渐变, 流程图即艺术, 手绘感图标, 自然色调
  • Reference: "Like a Takram project page — where technology feels thoughtful, not aggressive"
  • 执行路径: Path A(HTML→PPTX,配图可AI辅助生成)

更深入的风格细节:参考 design-philosophy skill 的 references/design-styles.md,包含20种设计哲学的完整提示词DNA

🎨 Custom Character Style (User-Defined)

Users may want to reference specific cartoon/anime aesthetics. When a user says "do it in Doraemon style" or "like Studio Ghibli", treat this as a style reference, not a request to draw copyrighted characters. Build a custom Design System by extracting the visual DNA of that style.

How to convert a character reference into a Design System:

User saysExtract these visual traits
"Doraemon style"Round shapes, bright primary blue + white + red, simple backgrounds, cute proportions, magical gadget reveals
"Studio Ghibli"Watercolor textures, natural greens and sky blues, detailed backgrounds with simple characters, warmth and wonder
"Calvin and Hobbes"Dynamic ink brushwork, expressive motion lines, philosophical contrast between fantasy and reality, lush outdoor scenes
"One Piece manga"Bold dynamic lines, exaggerated proportions, dramatic action poses, high energy, thick outlines
"Crayon Shin-chan"Crude crayon-like lines, flat bright colors, comedic proportions, everyday scenarios made absurd
"Adventure Time"Geometric simple shapes, pastel candy colors, thin outlines, whimsical surreal backgrounds

Template for custom style:

[User Style]: "[reference name]"
→ Shape language: [round/angular/geometric/organic]
→ Line quality: [thin uniform / thick varied / sketchy / brushwork]
→ Color palette: [specific colors extracted from that aesthetic]
→ Character style: [proportions, expressiveness level]
→ Background treatment: [detailed/minimal/abstract]
→ Emotional tone: [warm/energetic/philosophical/surreal]
Typography Rules (All Presets)
  • Max 2 font families (1 heading + 1 body)
  • Heading: bold, personality — ≥36pt (trend: even larger, as graphic surface)
  • Body: clean, readable — ≥18pt
  • Chinese: system default (PingFang SC / Microsoft YaHei)
  • Key principle: Typography is a DESIGN ELEMENT, not just an information container
Show full SKILL.md (1,090 more words)Show less
✅ Checkpoint 2 (Guided + Collaborative)

Ask the user to pick one of the 3 proposed design systems, or describe their own preference. Show the full description including philosophy, visual language, and reference.


Step 3: Build Slides


Step 3-A: HTML + Selective Illustrations (Path A)

Generate AI illustrations for key slides, then create HTML slide files.

Which slides need illustrations? Prioritize:

  1. Cover slide — always. Sets the visual tone.
  2. Key insight slides — the "aha moment" slides benefit most.
  3. Closing slide — optional but impactful.
  4. Data-heavy slides — charts/diagrams instead of AI art.

Illustration Generation — use nano-banana-pro skill:

⚠️ nano-banana-pro 是一个独立的外部 skill,不随本仓库分发。如果你的 ~/.claude/skills/ 下没有它,直接用本仓库自带的等价脚本(同样调用 Gemini 3 Pro Image API,默认 16:9,正好适配 slides): uv run <本仓库路径>/huashu-wechat-image/scripts/generate_image.py --prompt "..." --filename "..." --resolution 2K

bash
export $(grep GEMINI_API_KEY ~/.claude/.env) && \
uv run ~/.claude/skills/nano-banana-pro/scripts/generate_image.py \
  --prompt "[description]" \
  --filename "[timestamp]-slide-[N]-[name].png" \
  --resolution 2K

Base Style Prompt — define ONE style suffix, append to every illustration:

[Base Style]: flat vector illustration, [palette background color] background,
[accent color] highlight elements, clean minimalist aesthetic,
professional presentation style, no text in image

Per-slide prompt = [specific content] + [Base Style]

Key rules:

  • Always include "no text in image" — text will be added as editable elements
  • Use descriptive paragraphs, not keyword lists
  • Specify hex colors explicitly
  • Use "flat vector" / "flat illustration" for consistency

Embedding in HTML slides:

html
<!-- Side illustration (recommended) -->
<div class="left"><!-- text content --></div>
<div class="right"><img src="illustration.png" style="width: 280pt; height: 280pt;"></div>

<!-- Background illustration -->
<body style="background-image: url('illustration.png'); background-size: cover;">

✅ Checkpoint 3-A (Guided: preview 2-3 key illustrations; Collaborative: every one)

Show generated illustrations. Ask: Approve / regenerate / style consistent?


Step 3-B: Full AI Slide Generation (Path B)

Generate EVERY slide as a complete AI image — layout, text, visuals, all in one.

⚠️ THE #1 MISTAKE: Over-constraining the prompt with layout details and visual restrictions. More constraints = LESS creativity and diversity. The AI generates best when given mood + reference + content, NOT specific positions, color ratios, or character restrictions.

The Golden Rule of AI Image Prompts

SHORT prompts > LONG prompts. A 3-sentence prompt describing mood and content produces better results than a 30-line prompt specifying every visual detail. Specifically:

DON'T (kills diversity)DO (enables creativity)
Specify color ratios (60%/25%/15%)Describe the mood ("warm like a Sunday comic page")
Dictate layout positions ("title centered, image on right")Reference a specific aesthetic ("Peanuts comic strip")
Restrict characters ("NOT Snoopy — an original character")Let AI interpret the style naturally
List every visual element to includeDescribe what the viewer should FEEL
Repeat the base style in every per-slide promptDefine base style once, keep per-slide prompts short
Base Style Prompt — Keep it SHORT

Define a base style once, append to every slide. Keep it under 5 lines. The base style sets the mood; per-slide prompts add the content.

[Base Style]:
VISUAL REFERENCE: [Specific art/design aesthetic in one sentence]
CANVAS: 16:9 aspect ratio, 2048x1152 pixels, high quality rendering.
COLOR SYSTEM: [Describe the mood/feel of colors, not exact ratios]

Example (good — concise):

VISUAL REFERENCE: Charles Schulz Peanuts comic strip — warm, philosophical, charming.
Characters include round-headed kids, a lovable beagle dog, and a small yellow bird.
CANVAS: 16:9 aspect ratio, 2048x1152 pixels, high quality rendering.
COLOR SYSTEM: Warm cream/newspaper tone background, soft muted pastels, warm ink lines.

Anti-pattern (bad — over-specified): Do NOT include typography sizes, color ratios, composition percentages, margin specifications, or visual weight distributions in the base style. These constraints reduce diversity without improving quality.

Per-Slide Prompt Structure

Keep per-slide prompts short and focused. Do NOT repeat base style details or over-specify visual layout.

Create a [style] slide about [topic].

[Base Style]

DESIGN INTENT: [1 sentence — what the viewer should FEEL]

TEXT TO RENDER:
- Title: "[exact text]"
- Body: "[exact text]"

[Optional: 1-2 sentences describing mood or scene. Let AI decide composition.]
Example — GOOD vs BAD

BAD (traditional PPT — boring):

Design a professional presentation slide.
Professional presentation slide, 16:9 aspect ratio, 2048x1152 pixels.
Dark navy background, light gray text, gold accent.
Slide type: content. Layout: Title at top-left, two columns below.
Title: "看涨期权收益结构"
Body: "行权价: 100元, 权利金: 10元"
Visual: a line chart showing call option payoff

→ Result: Generic PPT that could come from any template

GOOD (magazine-level — stunning):

Create a slide that feels like a Bloomberg terminal data visualization
brought to life as editorial art.

VISUAL REFERENCE: Bloomberg Businessweek data feature meets cinematic lighting.
CANVAS: 16:9, 2048x1152, sharp rendering.
COLOR SYSTEM: Deep black (#0A0A0A) background 75%, white text 15%,
gold (#BF9A4A) accent 10%. The gold represents profit — it should GLOW.
TYPOGRAPHY: The number "110" rendered at 100pt as the dominant visual anchor
(the break-even point IS the story). Supporting text at 14pt, muted gray.

DESIGN INTENT: The viewer should instantly FEEL the asymmetry of options —
limited downside, unlimited upside. The visual must make this visceral,
not just informational.

TEXT TO RENDER:
- Hero metric: "110" (giant, gold, the break-even price)
- Title: "盈亏平衡点" (medium, white, above the number)
- Left data: "行权价 100" "权利金 10" (small, gray, understated)
- Insight: "亏损有底 盈利无限" (accent color, bottom)

VISUAL NARRATIVE: A single golden curve emerges from the left side of the slide,
flat and muted in gray at -10 (the maximum loss), then suddenly bending upward
at the strike price, transitioning from gray to brilliant gold as it rises
into the profit zone. The curve should feel like a ray of light breaking
through darkness. The profitable area above zero glows with warm gold
atmospheric lighting, like sunrise. The chart has NO grid lines, NO axes labels
cluttering the visual — just the pure, dramatic curve and the giant "110"
floating at the inflection point.

→ Result: An editorial data visualization that tells a story

Key Rules for Path B Prompts

Prompt Quality Checklist (verify before every generation):

  1. Visual Reference — Does the prompt name a specific art style or publication? (NOT just "professional" or "modern")
  2. Mood, not Layout — Does the prompt describe what the viewer should FEEL, not where elements should be PLACED?
  3. Text Content — Are all texts to render listed clearly and accurately?
  4. Short Enough — Is the prompt concise? Long prompts with detailed specs REDUCE diversity. Remove anything the AI can decide on its own.
  5. NO Micro-Management — No hex color ratios, no typography sizes, no composition percentages, no character pose instructions.

Technical Rules:

  • Always specify resolution: 2048x1152 (2K, 16:9) for crisp text
  • Include ALL text verbatim — AI must render exact words
  • 中文优先: slide上的文字一律用中文,仅保留必要英文术语
  • Chinese text tip: Keep titles short (≤8 characters) for best rendering
  • Use descriptive paragraphs, not keyword lists
  • Generate in parallel: Run 3-5 slide generations concurrently for speed
  • Consistency: The Base Style is applied to EVERY slide. It's a system, not a suggestion

Generation command (same tool, but full-slide prompts):

bash
export $(grep GEMINI_API_KEY ~/.claude/.env) && \
uv run ~/.claude/skills/nano-banana-pro/scripts/generate_image.py \
  --prompt "[full slide prompt]" \
  --filename "slide-[NN]-[name].png" \
  --resolution 2K

没有 nano-banana-pro 时同样可换用本仓库的 huashu-wechat-image/scripts/generate_image.py(见 Step 3-A 的说明)。

Quality check after generation:

  1. Text accuracy — verify all Chinese/English text rendered correctly
  2. Layout — elements positioned as described
  3. Style consistency — colors and design language match across slides
  4. If a slide has text errors → regenerate with adjusted prompt (simplify text or shorten)

✅ Checkpoint 3-B (Guided: preview all slides as a set; Collaborative: approve each)

Show ALL generated slide images to the user. Ask:

  • Text readable and accurate?
  • Visual style consistent across slides?
  • Any slides to regenerate?

Step 4: PPTX Assembly

4-A: html2pptx Workflow (Path A)

Create HTML files per slide, convert with html2pptx.js:

javascript
const pptxgen = require('pptxgenjs');
const html2pptx = require(process.env.HOME + '/.agents/skills/pptx/scripts/html2pptx.js');

const pptx = new pptxgen();
pptx.layout = 'LAYOUT_16x9';
await html2pptx('slide1.html', pptx);
await html2pptx('slide2.html', pptx);
await pptx.writeFile({ fileName: 'output.pptx' });

HTML rules (from pptx skill):

  • Body dimensions: width: 720pt; height: 405pt (16:9)
  • ALL text must be in <p>, <h1>-<h6>, <ul>, <ol> tags
  • Backgrounds/borders only on <div> elements
  • No CSS gradients — pre-render as PNG with Sharp
  • Use web-safe fonts only (Arial, Helvetica, Georgia, etc.)
  • Images: <img src="illustration.png" style="width: Xpt; height: Ypt;">

Known issue: Chinese characters in file paths can break image loading. Use symlinks to ASCII paths if needed:

bash
ln -sf "/path/with/中文/" /tmp/ascii-path
4-B: Image Assembly (Path B)

Assemble generated slide images into PPTX using create_slides.py:

bash
uv run ~/.claude/skills/image-to-slides/scripts/create_slides.py \
  slide-01-cover.png slide-02-intro.png slide-03-definition.png ... \
  --layout fullscreen \
  --bg-color 000000 \
  -o output.pptx

Recommended layout for Path B: fullscreen — images fill the entire slide since they already contain all layout, text, and visuals.

LayoutUse case
fullscreenAI-generated full-page slides (Path B default)
title_aboveImage + editable title (hybrid approach)
title_leftSplit: text + visual
centerCentered image with padding
gridMultiple images per slide

Step 5: Preview & Polish

Preview

Path A: Screenshot 3-4 key HTML slides with Playwright:

bash
npx playwright screenshot "file:///path/to/slide.html" preview.png \
  --viewport-size=960,540 --wait-for-timeout=1000

Path B: Show the generated slide images directly (they ARE the slides). Use Read tool to display 3-4 key PNGs.

✅ Checkpoint 4 (All modes)

Show preview to the user. The PPTX file is ready — ask:

  • Any slides to adjust?
  • Ready to open in Keynote/PowerPoint?
Final Polish (in Keynote/PowerPoint)
  • Transitions and animations
  • Speaker notes
  • Brand logo placement
  • Path A: Final text adjustments (editable)
  • Path B: Text NOT editable — if text errors found, regenerate the slide image

Design Quick Reference

5/5/5 rule: ≤5 words/line, ≤5 bullets/slide, ≤5 text-heavy slides in a row

Cognitive load: One idea per slide. ~1 min per slide. Slides complement speech, never duplicate it.

Visual hierarchy: F/Z-pattern reading flow. Title:body size ≈ 3:1. Every slide should have a visual element.

Detailed references:

  • references/proven-styles-gallery.md — 17 tested visual styles with tiered recommendations
  • references/proven-styles-snoopy.md — Snoopy/Peanuts style detailed per-slide templates
  • references/prompt-templates.md — Content generation and image prompts
  • references/design-principles.md — Full design framework, color palettes, typography
SkillRole
pptxAdvanced PPTX creation/editing (html2pptx, templates)
nano-banana-proAI illustration generation (Gemini 3 Pro Image)。外部 skill 不随本仓库分发,未安装时用本仓库 huashu-wechat-image/scripts/generate_image.py 代替
multi-modelExternal AI for content drafting
design-philosophy20种设计哲学深度参考(风格DNA + 场景模板 + 评审标准)。Professional/Editorial风格的详细提示词和评审指南在此

Output

  • .pptx files compatible with PowerPoint, Keynote, Google Slides
  • Web-safe fonts for cross-platform compatibility
  • AI illustrations as separate PNG files (reusable)

花叔出品 | AI Native Coder · 独立开发者 公众号「花叔」| 30万+粉丝 | AI工具与效率提升 代表作:小猫补光灯(AppStore付费榜Top1)·《一本书玩转DeepSeek》

© alchaincyf, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 26 other files (scripts, references, assets) in huashu-slides of alchaincyf/huashu-skills.

  • SKILL.md
  • assets/.DS_Store
  • assets/style-samples/slide04-01-苏联构成主义-constructivism.png
  • assets/style-samples/slide04-02-浮世绘-ukiyo-e.png
  • assets/style-samples/slide04-03-包豪斯-bauhaus.png
  • assets/style-samples/slide04-04-孔版印刷-risograph.png
  • assets/style-samples/slide04-05-等轴测-isometric.png
  • assets/style-samples/slide04-06-工程蓝图-blueprint.png
  • assets/style-samples/slide04-07-复古广告-vintage-ad.png
  • assets/style-samples/slide04-08-达达拼贴-collage.png
  • assets/style-samples/slide04-09-像素画-pixel-art.png
  • assets/style-samples/slide04-10-敦煌壁画-dunhuang.png
  • assets/style-samples/slide04-13-温暖叙事-warm-narrative.png
  • assets/style-samples/slide04-ligne-claire.png
  • assets/style-samples/slide04-manga.png
  • assets/style-samples/slide04-neo-pop.png
  • assets/style-samples/slide04-oatmeal.png
  • assets/style-samples/slide04-snoopy.png
  • assets/style-samples/slide04-xkcd.png
  • … and 8 more

Open the folder on GitHubat commit 2efea35

Compare with similar skills

Huashu Slides next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Huashu Slides compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Huashu Slides this skillalchaincyf/huashu-skills1.7k—~7.2kAutomated safety check: NotesMIT
Image To Editable Pptningzimu/image-to-editable-ppt-skill2.9k—~4.3kAutomated safety check: PassMIT
Slidesfcakyon/claude-codex-settings1.2k1 repos~1.1kAutomated safety check: PassMIT
Ppt Image FirstNyxTides/ppt-image-first1.2k—~1.6kAutomated safety check: PassApache-2.0
Gpt Image2 PptJuneYaooo/gpt-image2-ppt-skills1.3k—~8.9kAutomated safety check: NotesApache-2.0
Paper Deckzsyggg/paper-craft-skills1.3k—~1.4kAutomated safety check: PassNone

Similar skills

  • Image To Editable Ppt

    ningzimu/image-to-editable-ppt-skill

    Rebuild slide images, scanned or image-based PPT/PPTX files, and PDF decks into object-level editable PowerPoint (.pptx), preserving speaker notes when supplied.

    2.9k GitHub stars~4.3k tokensUpdated 25 days ago
    Documents & OfficeAuto-check passed
  • Slides

    fcakyon/claude-codex-settings

    Create and edit presentation slide decks (.pptx) with PptxGenJS, bundled layout helpers, and render/validation utilities.

    1.2k GitHub starsUsed in 1 repo~1.1k tokens
    Documents & OfficeAuto-check passed
  • Ppt Image First

    NyxTides/ppt-image-first

    Build presentation plans for PPT / slides / decks through a conversation-first workflow, then propose multiple visual directions with preview images before writing deck specs.

    1.2k GitHub stars~1.6k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Gpt Image2 Ppt

    JuneYaooo/gpt-image2-ppt-skills

    Generate visually striking PPT slides via OpenAI's gpt-image-2 -- use any style in styles/<collection/STYLEID.md or mimic a user-supplied .pptx template; outputs high-res slide PNGs and a 16:9 .pptx.

    1.3k GitHub stars~8.9k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check: notes
  • Paper Deck

    zsyggg/paper-craft-skills

    将论文、技术文章或知识内容制作成高真实感的 AIGC 幻灯片。先做叙事结构和逐页视觉导演,再调用生图模型生成每一页 16:9 slide image,最后合成为 PPTX/PDF。适合论文汇报、组会、公开课、技术分享、商业化研究展示;当用户提到“论文PPT”“AI生成PPT”“不像AI的PPT”“高质感幻灯片”“逐页生图PPT”时使用。

    1.3k GitHub stars~1.4k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • HTML Slide To PPTX

    mucsbr/ppt-agent-workflow-san

    Convert structured single-slide or small deck HTML files into editable PPTX slides with native text boxes, shapes, chips, arrows, and panels.

    645 GitHub stars~1.1k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from alchaincyf/huashu-skills

All 21 skills in this repo
  • Huashu Agent Swarm

    alchaincyf/huashu-skills

    多Agent蜂群并行协作,纯git自组织,适合大型项目开发。当用户提到"蜂群模式"、"多agent"、"并行开发"、"agent swarm"时使用。

    1.7k GitHub stars~576 tokensUpdated 18 days ago
    Auto-check passed
  • Huashu Data Pro

    alchaincyf/huashu-skills

    数据分析与办公提效全能助手。覆盖数据处理、分析洞察、报告撰写、PPT制作、数据可视化的端到端工作流. An agent skill from alchaincyf/huashu-skills.

    1.7k GitHub stars~928 tokensUpdated 18 days ago
    Auto-check passed
  • Huashu Douyin Script

    alchaincyf/huashu-skills

    抖音爆款脚本创作工作流。从竞品视频拆解到脚本生成的完整流程:下载抖音视频→Gemini视频分析→爆款公式提炼→脚本+分镜生成→AI味审校。

    1.7k GitHub stars~1.7k tokensUpdated 18 days ago
    Auto-check passed
  • Huashu Wechat Image

    alchaincyf/huashu-skills

    为微信公众号文章生成高质量配图。支持封面图(2.35:1)、正文插图(16:9/4:3)、信息图。提供两条路径:AI生成(视觉创意型)和HTML渲染(文字精确型)。当用户提到"公众号配图"、"公众号封面"、"文章配图"、"正文插图"、"公众号图片"时使用此技能。

    1.7k GitHub stars~2.4k tokensUpdated 18 days ago
    Auto-check: notes
  • Huashu Xhs Image

    alchaincyf/huashu-skills

    为小红书笔记生成高质量配图。默认AI生成(Gemini),仅精确数据表格用HTML兜底。当用户提到"小红书配图"、"小红书封面"、"小红书图片"、"做张小红书图"、"笔记配图"时使用此技能。

    1.7k GitHub stars~1.5k tokensUpdated 18 days ago
    Auto-check: notes
  • Huashu Md To PDF

    alchaincyf/huashu-skills

    将 Markdown 文档转换为专业的 PDF 白皮书,采用苹果设计风格. An agent skill from alchaincyf/huashu-skills.

    1.7k GitHub stars~824 tokensUpdated 18 days ago
    Auto-check passed

Questions about Huashu Slides

What does Huashu Slides do?

从内容到成品PPTX的端到端演示文稿制作,含AI插画生成和18种设计风格。当用户提到"做PPT"、"做幻灯片"、"演示文稿"、"Keynote"、"slides"时使用。. Huashu Slides is an agent skill from alchaincyf/huashu-skills.

When should I use Huashu Slides?

Huashu Slides fits situations like: tasks that involve Slides and decks; tasks that involve PowerPoint presentations.

How do I install Huashu Slides in Claude Code?

Run `npx skills add alchaincyf/huashu-skills --skill huashu-slides -a claude-code`. Or copy the skill folder (huashu-slides in alchaincyf/huashu-skills) into .claude/skills/huashu-slides in your project. Claude Code loads it when a task matches its description.

How do I install Huashu Slides in Codex?

Run `npx skills add alchaincyf/huashu-skills --skill huashu-slides -a codex`. Or copy the skill folder (huashu-slides in alchaincyf/huashu-skills) into .agents/skills/huashu-slides in your project. Codex loads it when a task matches its description.

Can I use Huashu Slides in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alchaincyf/huashu-skills --skill huashu-slides -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/huashu-slides, .gemini/skills/huashu-slides, .github/skills/huashu-slides and .opencode/skills/huashu-slides in your project.

What does Huashu Slides need to run?

Going by SKILL.md and its folder, Huashu Slides needs the command-line tools its instructions call (uv and npx) and credentials named GEMINI_API_KEY.

Does Huashu Slides access the network?

SKILL.md contains no URLs. Its commands use uv and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Huashu Slides safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Huashu Slides use?

Huashu Slides is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Huashu Slides use?

About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to Huashu Slides?

Skills that share tags, products or a category with Huashu Slides: Image To Editable Ppt (ningzimu/image-to-editable-ppt-skill, 2.9k stars), Slides (fcakyon/claude-codex-settings, 1.2k stars), Ppt Image First (NyxTides/ppt-image-first, 1.2k stars) and Gpt Image2 Ppt (JuneYaooo/gpt-image2-ppt-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Huashu Slides?

alchaincyf (a GitHub user) maintains it in alchaincyf/huashu-skills, which has 1,695 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on September 22, 2026.

Source: alchaincyf/huashu-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.