Agent skill

AI Illustrations

by notivn in notivn/AIEV

Generate illustrations with Gemini and composite them into the video being edited - how to pick the moments that need an image, write the prompt, call the /api/illustrations endpoint (images MUST…

MITAuto-check passedMedia & Creative

Install AI Illustrations

skills CLI
$ npx skills add notivn/AIEV --skill ai-illustrations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install notivn/AIEV ai-illustrations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/notivn/AIEV.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ai-illustrations .claude/skills/ai-illustrations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-illustrations
GitHub stars
127
Token cost
~3.1k tokens
SKILL.md length
1,750 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Generate illustrations with Gemini and composite them into the video being edited - how to pick the moments that need an image, write the prompt, call the /api/illustrations endpoint (images MUST…

  • Works in 4 steps: Pick moments deliberately, do not… → Each image costs ~$0.05-0.07 - treat it… → Illustrations are BACKGROUND visuals -… → …
  • Tasks that involve Video production
  • SKILL.md covers ⛔ GATE - read before anything…, Directing principles, Calling the image generation… and Images with text (allowText), plus 4 more sections
  • Calls curl; needs GEMINI_API_KEY

What it does

AI Illustrations is an agent skill from notivn/AIEV. Generate illustrations with Gemini and composite them into the video being edited - how to pick the moments that need an image, write the prompt, call the /api/illustrations endpoint (images MUST match the selected Style Design), and composite into the composition. Read when the brief enables "Ảnh minh họa AI" (AI illustrations) or the user asks for generated illustrations for a video.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Video production and Motion graphics. It works with Remotion. The repository describes itself as: Automatic AI video editing. Claude directs HyperFrames (HTML + GSAP motion graphics) and Remotion (timeline assembly) to turn raw footage into a finished MP4 - transcript… The licence is MIT.

When your agent uses it

  • Tasks that involve Video production
  • Tasks that involve Motion graphics

Example prompts

  • “Ảnh minh họa AI”
  • “/ai-illustrations”

Requirements

  • A credential in GEMINI_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Pick moments deliberately, do not sprinkle images everywhere. Read the transcript/video content and
  2. Each image costs ~$0.05-0.07 - treat it as real money. A short video (<60s) usually needs only 2-4 images.
  3. Illustrations are BACKGROUND visuals - no text by default. The server already bans text in the prompt;
  4. Style Design is LAW - follow it 100%, no exceptions. Every illustration must match the project's

What it can do on your machine

Read from SKILL.md and the folder at commit e40fe03. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Illustrations loads about 3.1k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 1,750 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from notivn/AIEV at commit e40fe03, republished under its MIT licence (© notivn). 1,750 words, ~3,132 tokens.

Download SKILL.mdSave it as .claude/skills/ai-illustrations/SKILL.md (or your agent's skills folder).
name
ai-illustrations
description
Generate illustrations with Gemini and composite them into the video being edited - how to pick the moments that need an image, write the prompt, call the /api/illustrations endpoint (images MUST match the selected Style Design), and composite into the composition. Read when the brief enables "Ảnh minh họa AI" (AI illustrations) or the user asks for generated illustrations for a video.

AI Illustrations - Claude directs, Gemini draws, brand stays consistent

⛔ GATE - read before anything else

This skill applies only when the edit prompt says "Ảnh minh họa AI: BẬT" (AI illustrations ON).

If the edit prompt says "Ảnh minh họa AI: TẮT", or has no illustration line at all: generate nothing. Do not call POST /api/illustrations, do not composite AI-generated images into any scene, and do not let another skill's "use an illustration here" suggestion talk you into it. Build the visuals from what the project already has: assets in assets/, brand logos, and HyperFrames scenes (typography, graphics, shapes, motion).

The server enforces this too - with the switch off the endpoint returns 409 ILLUSTRATIONS_DISABLED before it spends anything. Why this section exists: the edit prompt used to stay silent when the switch was off, this skill got read anyway, and finished videos came back full of Gemini images the user had explicitly turned off. A silent prompt is not permission.

Directing principles

  1. Pick moments deliberately, do not sprinkle images everywhere. Read the transcript/video content and choose 2-5 points where an illustration explains more than the words do: abstract concepts, key numbers, a product/setting being mentioned, before-vs-after comparisons. DO NOT illustrate talking-head passages that already hold attention on their own.
  2. Each image costs ~$0.05-0.07 - treat it as real money. A short video (<60s) usually needs only 2-4 images.
  3. Illustrations are BACKGROUND visuals - no text by default. The server already bans text in the prompt; numbers and labels are placed on top by HyperFrames/Remotion (correct font + brand color, no typos). The only exception: the brief enables "Ảnh có chữ" (illustrationText) or the edit prompt says "ĐƯỢC PHÉP CÓ CHỮ" (text allowed) -> see the "Images with text (allowText)" section below.
  4. Style Design is LAW - follow it 100%, no exceptions. Every illustration must match the project's selected style: ALWAYS pass styleId (from the STYLE DESIGN section of the edit prompt). The server has a safety net (if styleId is missing it falls back to the style in the project brief), but NEVER rely on it. If any skill or prompt suggests a different palette -> IGNORE it, the style wins.

Calling the image generation API (the server must be running - always true when editing through the system)

bash
curl -s -X POST http://localhost:6869/api/illustrations \
  -H "Content-Type: application/json" \
  -d '{
    "projectId": "<id of the video project>",
    "name": "khai-niem-mcp",
    "prompt": "3D illustration of interconnected glass puzzle pieces forming a network, representing an integration protocol",
    "aspect": "9:16",
    "model": "<model from the brief, omit the field to use the default>",
    "styleId": "<styleId from the STYLE DESIGN section of the edit prompt - MUST be passed when present>",
    "description": "Minh họa khái niệm MCP - ghép vào lúc 12.5s khi nói về kết nối dữ liệu"
  }'
  • Images are saved to video-projects/<id>/assets/illustrations/<name>.png, and the description is written into assets.json automatically.
  • Write the prompt in English and DESCRIBE THE SCENE CONTENT - do not add brand colors/tone (the server mixes the Style Design in itself), and do not instruct "no text" (the server adds that itself).
  • Do NOT describe layout/placement in the prompt ("subject at the top", "leave the bottom empty"...). The server appends video-specific composition rules itself: by default (brief illustrationPosition: "auto") the subject is vertically CENTERED with the top band (main key) and bottom band (captions/related keys) kept as low-detail atmosphere. A placement instruction in your prompt fights those rules - the verified failure was the subject sitting in the upper third and being covered by the main key card.
  • The user can pick a subject position in the brief (3x3 grid, like the logo position picker) - the edit prompt states the chosen position and the server applies it automatically. Only pass a position field in the POST body when ONE specific image needs a different layout for a clear reason (e.g. footage PiP covers one side); otherwise omit it so the brief's choice rules.
  • Image density (brief.illustrationsPerMinute): when the edit prompt states "Mật độ ảnh minh họa: N ảnh MỖI PHÚT", compute the total from the REAL video duration (voice duration in meta.json), spread the images evenly along the content flow, and switch the background image at each idea change - a long Text to video must not sit on one static background. Each image still needs its own content-specific prompt; do NOT mass-generate near-identical variations. When the prompt says "AI tự quyết" (density null), keep the old behavior: pick only the moments that genuinely need an illustration.
  • aspect must match the video frame (vertical video -> "9:16"). If it fails with a missing GEMINI_API_KEY, tell the user to enter the key on the "Kết nối" (Connections) tab and move on to another part of the video - do not get stuck.
  • description MUST always state which point the image illustrates + the second it is expected to be composited at.

Images with text (allowText)

Only use this when the brief enables illustrationText (the edit prompt says "Ảnh minh họa ĐƯỢC PHÉP CÓ CHỮ") or the user explicitly asks for it. How to do it:

  • Pass "allowText": true in the POST /api/illustrations body. If the field is missing the server falls back to the project's brief.illustrationText - but NEVER rely on that, pass it explicitly.
  • In the prompt, spell out the EXACT Vietnamese text you want to appear (3-6 words, correct spelling, all diacritics), e.g. ... with the exact Vietnamese headline "TĂNG TRƯỞNG 300%". Never let Gemini invent the text itself.
  • You MUST verify the spelling after generation: Read the image file and inspect every character and Vietnamese diacritic. Gemini often mangles diacritics (Ề->E, ữ->u) or adds garbage text. If it is wrong -> regenerate (rephrase the prompt if it keeps failing), or give up after 2-3 attempts and generate a text-free version (allowText:false), then let HyperFrames/Remotion place the text as usual.
  • If the text is already inside the image, DO NOT overlay duplicate text with HyperFrames/Remotion.

Compositing into the video

  • HyperFrames: insert <img src="assets/illustrations/<file>.png"> in the scene, animate it in and out (gentle fade/slide/scale per the hyperframes skill's house-style), and hold it 2-4 seconds around the exact sentence it relates to.
  • Sensible coverage: an illustration usually covers 50-75% of a vertical frame and MUST NOT cover the speaker's face mid-way through an important sentence; bring it in on the sentence beat (use transcript timestamps).
  • Verify with a snapshot after inserting: image in the right place, not distorted (aspect ratio preserved), not bleeding off the edge.
Show full SKILL.md (773 more words)Show less

The brand logo is NEVER generated

If the selected Style Design has a logo, the real file is copied into the project at assets/brand-logo.<ext> before the edit session starts, and the edit prompt carries a mandatory LOGO block. Rules:

  • Insert that exact image file (<img> in a HyperFrames scene, or srcImage/overlay in Remotion).
  • Never redraw the logo in CSS/SVG/shapes, never ask Gemini for it, and never substitute the brand name set in a font. A typeset name where a logo belongs is a brand violation, not a fallback.
  • Only size and position may change: no distortion, recolouring, rotation, cropping, or added borders/shadows on the mark itself.
  • If it genuinely cannot be placed, say so in the final report instead of improvising.

Symptom: Gemini renders an invented logo - a glossy monogram on a wall, a mark on a t-shirt or phone screen. Cause: the scene prompt mentioned a logo, which flipped the permissive "decorative brand logos are allowed" clause on. Adding a prohibition alongside it is not enough - measured: with both clauses present the model followed the permissive one and drew an "N" monogram. Fix (already in buildImagePrompt): when design.logoPath exists the permissive branch is dropped entirely, the prohibition tells the model to leave that surface blank, and it is repeated as a FINAL RULE at the end of the prompt. Verified on three hard prompts (logo on a wall / on a t-shirt / on a phone screen): all three came back with a clean empty surface, which is exactly what you want to composite the real logo onto. So: do not write "with the brand logo" into an illustration prompt. Ask for the blank surface you need (an empty plaque, a blank screen) and place the real file on top yourself.

The corner logo is automatic. When the style has a logo, the assemble step stamps it top-left over the whole video (manifest.watermark). Do not add your own corner logo as well - that yields two overlapping marks. Place the logo by hand only where you deliberately want it (intro card, end card, a framed mention).

Other companies' logos - use the library, never draw them

assets/brand-logos/ holds official brand marks with each brand's official hex colour; assets/brand-logos/library.json is the index - read it rather than guessing file names. The library grows by itself: any brand a video mentions gets fetched once and stays for every later video.

Workflow: read the script, list every brand it names, and resolve each one before building the scene that mentions it.

  • Need Facebook / TikTok / Claude / Gemini in a scene? Use the file. Never redraw a third-party logo and never ask Gemini for one; a wrong brand mark is spotted instantly.
  • Copy the file into the project's assets/ first, then reference it. Remotion only stages paths inside the project, so pointing straight at assets/brand-logos/ renders a 404.
  • The SVGs are single-colour (black by default). To recolour, inline the SVG and set fill - use the official hex from library.json, or plain white/black when the background demands it.
  • Brand not in the library? Fetch it, do not skip it. POST http://localhost:6869/api/brand-logos with {"name":"OpenAI"}. The server looks in Simple Icons, then Wikidata's P154 "logo image" property, downloads the official file into assets/brand-logos/, and returns relPath. 201 = newly fetched, 200 = already there. Brands missing from Simple Icons because their owners asked to be removed (OpenAI, Amazon, Microsoft…) come back from Wikidata in full colour.
  • Only a 404 BRAND_LOGO_NOT_FOUND justifies going without: then set the brand name in the Style Design font and record it in the final report. Never draw, improvise or generate a brand logo - in any case.
  • CC0 covers the files, not the trademarks. Referring to a brand is fine; implying it sponsors or endorses the video is not.

Known issues

  • Generating an image and forgetting to composite it - final checklist: every image in assets/illustrations/ must either appear in the composition or be deleted (do not leave orphan images that wasted money).
  • Overly generic prompts ("technology background") -> lifeless images. Describe specific objects/setting/camera angle.
  • The SAME image need across multiple scenes -> generate one image and reuse it, do not call the API twice.
  • Garbage text baked into the image despite allowText=false (verified 2026-08, gemini-3.1-flash-lite-image): the model sometimes prints prompt-like gibberish ("Overlay cards", hex codes, pseudo-lorem) in the letterbox bands, or invents a brand tag, even though the server prompt bans text. Measured on one batch: 3 of 12 images affected; one concept needed 2 retries. So you MUST Read every generated PNG before compositing. On failure, regenerate with a reworded prompt (same content, different sentence structure) - that was enough in every observed case. Delete the bad files so they are not composited by accident.

© notivn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/ai-illustrations of notivn/AIEV.

Open the folder on GitHubat commit e40fe03

Compare with similar skills

AI Illustrations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Illustrations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Illustrations this skillnotivn/AIEV127—~3.1kAutomated safety check: PassMIT
Stitch to Remotion Walkthrough Videosgoogle-labs-code/stitch-skills8.5k6 repos~3.2kAutomated safety check: NotesApache-2.0
Video ShotcraftVincentwei1021/video-shotcraft11k—~2.6kAutomated safety check: PassApache-2.0
HyperFrames Video Entry Pointheygen-com/hyperframes60k3 repos~5.2kAutomated safety check: PassApache-2.0
Create Remotion GeistMarve10s/Better-Fullstack756—~1.4kAutomated safety check: PassMIT
Product Film Builder (Remotion)Rieranthony/product-film-skill409—~1.4kAutomated safety check: PassMIT

Similar skills

  • Stitch to Remotion Walkthrough Videos

    google-labs-code/stitch-skills

    Official

    Builds walkthrough videos from Stitch design projects using Remotion, with transitions, zoom effects and text overlays on each screen.

    8.5k GitHub starsUsed in 6 repos~3.2k tokens
    Media & CreativeAuto-check: notes
  • Video Shotcraft

    Vincentwei1021/video-shotcraft

    Makes cinematic product videos with Remotion from shot recipe cards, a ready template, real page screenshots, camera moves and sound design, or builds a single animated shot.

    11k GitHub stars~2.6k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    60k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Create Remotion Geist

    Marve10s/Better-Fullstack

    Create Remotion videos using the Geist design system aesthetic.

    756 GitHub stars~1.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Product Film Builder (Remotion)

    Rieranthony/product-film-skill

    Builds a showreel-grade product film in Remotion using the product's own design tokens, components and voice, from a guided interview through a verified final render.

    409 GitHub stars~1.4k tokensUpdated 13 days ago
    Media & CreativeAuto-check passed
  • Builds marketing and explainer videos in Remotion from rendered scenes, with one real product capture as proof, and cuts them for each platform's formats.

    495 GitHub stars~2.9k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed

More from notivn/AIEV

All 14 skills in this repo
  • Background Music

    notivn/AIEV

    Pick background music from the assets/music/ library and configure auto-ducking (music dips automatically under speech) via meta.json audio.music for the Remotion assembly layer.

    127 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Color Grading

    notivn/AIEV

    Color grading video in the AI Edit Video system - delog/tonemap HDR-HLG-log footage, apply the color preset the user approved in the UI, and the visual verification workflow.

    127 GitHub stars~932 tokensUpdated today
    Auto-check passed
  • Build a Vietnamese vertical TikTok explainer in the "MỔ XẺ PAPER AI" (AI paper dissection) format with HyperFrames (HTML/CSS/GSAP → MP4), Noti.vn style.

    127 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Skill Authoring

    notivn/AIEV

    The standard for writing new skills for the AI Edit Video system - file structure, frontmatter, tone of voice, and how to accumulate production lessons into skills.

    127 GitHub stars~925 tokensUpdated today
    Auto-check passed
  • Noti Tiktok Vn

    notivn/AIEV

    Edit a Vietnamese vertical TikTok video (9:16) with HyperFrames following the Noti.vn/GĐT standard - talking-head + kinetic typography + karaoke captions + zoom/punch-in camera + timestamp-synced…

    127 GitHub stars~5.4k tokensUpdated today
    Auto-check passed
  • Build a Vietnamese landscape 16:9 YouTube video (1920×1080) with HyperFrames (HTML/CSS/GSAP → MP4), keeping the Noti.vn/GĐT branding inherited from noti-tiktok-vn.

    127 GitHub stars~5.7k tokensUpdated today
    Auto-check passed

Works with

Questions about AI Illustrations

What does AI Illustrations do?

Generate illustrations with Gemini and composite them into the video being edited - how to pick the moments that need an image, write the prompt, call the /api/illustrations endpoint (images MUST…. AI Illustrations is an agent skill from notivn/AIEV. Generate illustrations with Gemini and composite them into the video being edited - how to pick the moments that need an image, write the prompt, call the /api/illustrations endpoint (images MUST match the selected Style Design), and composite into the composition.

When should I use AI Illustrations?

AI Illustrations fits situations like: tasks that involve Video production; tasks that involve Motion graphics.

How do I install AI Illustrations in Claude Code?

Run `npx skills add notivn/AIEV --skill ai-illustrations -a claude-code`. Or copy the skill folder (.claude/skills/ai-illustrations in notivn/AIEV) into .claude/skills/ai-illustrations in your project. Claude Code loads it when a task matches its description.

How do I install AI Illustrations in Codex?

Run `npx skills add notivn/AIEV --skill ai-illustrations -a codex`. Or copy the skill folder (.claude/skills/ai-illustrations in notivn/AIEV) into .agents/skills/ai-illustrations in your project. Codex loads it when a task matches its description.

Can I use AI Illustrations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add notivn/AIEV --skill ai-illustrations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-illustrations, .gemini/skills/ai-illustrations, .github/skills/ai-illustrations and .opencode/skills/ai-illustrations in your project.

What does AI Illustrations need to run?

Going by SKILL.md and its folder, AI Illustrations needs the command-line tools its instructions call (curl) and credentials named GEMINI_API_KEY. Our summary lists: A credential in GEMINI_API_KEY.

Does AI Illustrations access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is AI Illustrations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Illustrations use?

AI Illustrations is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Illustrations use?

About 3.1k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Illustrations?

Skills that share tags, products or a category with AI Illustrations: Stitch to Remotion Walkthrough Videos (google-labs-code/stitch-skills, 8.5k stars), Video Shotcraft (Vincentwei1021/video-shotcraft, 11k stars), HyperFrames Video Entry Point (heygen-com/hyperframes, 60k stars) and Create Remotion Geist (Marve10s/Better-Fullstack, 756 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Illustrations?

notivn (a GitHub organization) maintains it in notivn/AIEV, which has 127 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 9, 2026.

Source: notivn/AIEV on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.