Agent skill

Picture It

by geongeorge in geongeorge/picture-it

Generate and edit images from the CLI using picture-it. An agent skill from geongeorge/picture-it.

MITAuto-check passedMedia & Creative

Install Picture It

skills CLI
$ npx skills add geongeorge/picture-it --skill picture-it -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install geongeorge/picture-it picture-it --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/geongeorge/picture-it.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skill/picture-it .claude/skills/picture-it && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
picture-it
GitHub stars
132
Token cost
~3.7k tokens
SKILL.md length
1,413 words
Files
3 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Generate and edit images from the CLI using picture-it. An agent skill from geongeorge/picture-it.

  • Works in 3 steps: Understand the purpose → Plan the composition → Plan the pipeline
  • The user asks to create
  • SKILL.md covers Prerequisites, Core Concept, Before You Generate Anything —… and Commands Quick Reference, plus 8 more sections
  • Calls npm, pnpm and bun; needs FAL_KEY

What it does

Picture It is an agent skill from geongeorge/picture-it. Generate and edit images from the CLI using picture-it. Use this skill whenever the user asks to create, edit, or manipulate images — blog headers, social cards, hero images, product comparisons, YouTube thumbnails, movie posters, magazine covers, Instagram edits, background removal, or any visual content. Also trigger when the user mentions picture-it by name, wants to composite images, apply color grading, add text to images, remove or replace backgrounds, crop/resize photos, or needs any kind of image…

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/composition-guide.md` and `references/prompt-library.md`). Compatibility notes: Requires Node.js 18+ and picture-it CLI (npm package). FALKEY environment variable needed for AI operations. Network access to fal.ai for image…

It sits in Media & Creative, covering Image editing, Image generation and Social media graphics. It works with Instagram and YouTube. The licence is MIT.

When your agent uses it

  • The user asks to create
  • Manipulate images — blog headers
  • Product comparisons
  • YouTube thumbnails

Example prompts

  • “/picture-it”

Requirements

  • Node.js
  • A credential in FAL_KEY
  • Compatibility (from SKILL.md): Requires Node.js 18+ and picture-it CLI (npm package). FAL_KEY environment variable needed for AI operations. Network access to fal.ai for image generation/editing.

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Understand the purpose
  2. Plan the composition
  3. Plan the pipeline

What it can do on your machine

Read from SKILL.md and the folder at commit 5a3247d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • pnpm
    • bun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • npmjs.com
    • fal.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FAL_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Node.js 18+ and picture-it CLI (npm package). FAL_KEY environment variable needed for AI operations. Network access to fal.ai for image generation/editing.

    From compatibility in the SKILL.md frontmatter.

Context cost

Picture It loads about 3.7k tokens when it runs, and up to ~8.8k if it reads all its reference files. Until then it costs about 167 tokens; SKILL.md has 1,413 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~167
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from geongeorge/picture-it at commit 5a3247d, republished under its MIT licence (© geongeorge). 1,413 words, ~3,713 tokens.

Download SKILL.mdSave it as .claude/skills/picture-it/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
picture-it
description
Generate and edit images from the CLI using picture-it. Use this skill whenever the user asks to create, edit, or manipulate images — blog headers, social cards, hero images, product comparisons, YouTube thumbnails, movie posters, magazine covers, Instagram edits, background removal, or any visual content. Also trigger when the user mentions picture-it by name, wants to composite images, apply color grading, add text to images, remove or replace backgrounds, crop/resize photos, or needs any kind of image generation or photo editing from the terminal. This skill covers multi-pass AI image editing workflows that chain composable operations together.
compatibility
Requires Node.js 18+ and picture-it CLI (npm package). FAL_KEY environment variable needed for AI operations. Network access to fal.ai for image generation/editing.
license
MIT
metadata.author
geongeorge
metadata.version
0.2.1
metadata.homepage
https://github.com/geongeorge/picture-it
metadata.source
https://github.com/geongeorge/picture-it
metadata.package
https://www.npmjs.com/package/picture-it

picture-it

Photoshop for AI agents. Composable image operations from the CLI.

Source: https://github.com/geongeorge/picture-it | npm: https://www.npmjs.com/package/picture-it

Prerequisites

picture-it must be installed and configured. Requires Node.js 18+.

bash
# Install (pick one)
npm install -g picture-it
pnpm add -g picture-it
bun install -g picture-it

# Setup
picture-it download-fonts
Credentials

The FAL API key is required for AI operations (generate, edit, remove-bg, upscale). Set it via environment variable or the CLI:

bash
# Option 1: Environment variable (preferred — use platform-managed secrets)
export FAL_KEY=your-key-here

# Option 2: CLI config (stored in ~/.picture-it/config.json with 0600 permissions)
picture-it auth --fal <fal-api-key>

NEVER paste API keys into chat. Always use environment variables or the CLI auth command. Get a FAL key from https://fal.ai.

Note: User images are uploaded to fal.ai for AI processing when using generate, edit, remove-bg, or upscale commands. Local-only commands (crop, grade, grain, vignette, text, compose, template, info) do not transmit data.

Core Concept

Every command takes an image in and outputs an image. Chain them to build anything. The agent calling picture-it IS the planner — there is no AI planner inside the tool.

Before You Generate Anything — Think First

Image generation costs real money ($0.03–$0.15 per FAL call). A 4-pass workflow is $0.10+. Don't burn budget on a vague idea — spend time planning before running any commands.

Step 1: Understand the purpose

Before touching picture-it, get full clarity on what the user wants. Ask yourself:

  • What is this image for? (blog header, Instagram ad, YouTube thumbnail, product comparison, poster)
  • Who is the audience? (developers, consumers, enterprise buyers)
  • What should someone FEEL when they see it? (excitement, trust, urgency, curiosity)
  • What's the one message? Every good image communicates exactly one thing.
  • Where will it be displayed? This determines size, text sizing, and composition rules.

If any of these are unclear, ask the user before proceeding. A 30-second question saves $0.15 in wasted generation.

Step 2: Plan the composition

Think through at least 3 different approaches before picking one. Consider:

  • Can this be done without FAL? Templates and Satori compose are free. A solid gradient + good typography is often enough.
  • What's the minimum number of FAL calls? Each call costs money. Plan the fewest passes that achieve the goal.
  • Which technique fits? Text-behind-subject for thumbnails, remove-bg + compose for product photos, multi-pass for cinematic scenes.

Present your top 2-3 ideas to the user briefly — one sentence each — and let them pick before generating. Example:

"Here are a few directions:

  1. Dramatic product shot — generate a dark stage, edit to place your logo as a glowing 3D object ($0.07)
  2. Clean comparison — remove-bg from both products, compose on gradient with text ($0.01)
  3. Text-behind-subject — generate an action scene, edit to weave the title behind the subject ($0.07)

Which direction, or a mix?"

Step 3: Plan the pipeline

Before running the first command, write out the full pipeline:

1. generate (flux-dev $0.03) — dark stage scene
2. edit (seedream $0.04) — place logo into scene
3. compose (free) — add text overlay
4. grade + vignette (free) — post-process
Total: ~$0.07

This avoids discovering mid-way that you need a different approach and wasting the earlier calls.

Commands Quick Reference

CommandWhat it doesNeeds FAL?
generateCreate image from text promptYes
editEdit image(s) with AIYes
remove-bgRemove backgroundYes
replace-bgRemove bg + generate new oneYes
cropResize/crop to exact dimensionsNo
gradeApply color gradingNo
grainAdd film grainNo
vignetteAdd edge darkeningNo
textRender text onto image (Satori)No
composeOverlay images/text/shapes from JSONNo
templateBuilt-in templates (no AI)No
infoAnalyze image dimensions/colorsNo

Model Selection

Choose the right model for the job — don't overspend.

Generation-only models:

ModelCostBest for
flux-schnell$0.003Default. Fast drafts, backgrounds, base scenes
imagineart$0.03High-fidelity realism, accurate text rendering
flux-dev$0.03Detailed scenes, portraits, cinematic quality
recraft-v3$0.04Text in images, vector art, brand-style graphics
fibo$0.04Enterprise, structured/controlled generation
recraft-v4$0.25Premium. Best composition, lighting, materials. Use sparingly

Edit-only models:

ModelCostBest for
reve-fast$0.02Cheapest. Quick iterations, speed over refinement
kontext-lora$0.035Edits with LoRA styles, brand-consistent modifications
kontext$0.04Default. Targeted local edits, scene transforms, text placement
reve$0.04Style transforms, product variations, context-aware edits
fibo-edit$0.04Precise control with JSON + masks, object add/remove, restyling

Both generate AND edit (use --model with either generate or edit):

ModelCostBest for
seedream-v4$0.03Budget option. Good multi-image compositing
seedream$0.04Multi-image compositing (up to 10 inputs), placing objects in scenes
banana2$0.08Better image preservation, >10 inputs, extreme aspect ratios, web search
banana-pro$0.15Premium. Best realism, typography, character consistency for up to 5 people

How to pick the right model:

TaskBest modelCostWhy
Quick background/draftflux-schnell$0.003Fastest, cheapest
Quality hero imageflux-dev or imagineart$0.03Good balance of quality/cost
Text-heavy generationrecraft-v3$0.04Best text rendering in generated images
Quick edit iterationreve-fast$0.02Half the price, good enough for drafts
Single image edit (bg swap, add text)kontext$0.04Best targeted edits
Compose multiple images into one sceneseedream$0.04Handles up to 10 inputs
Style transfer / product variationsreve$0.04Context-aware transforms
Precise masked edit (add/remove object)fibo-edit$0.04JSON + mask control
Subject must stay very faithfulbanana2$0.08Best preservation
Premium quality, complex scenebanana-pro$0.15Best overall but expensive

Background removal:

  • bria (default) — Best edge quality, clean cutouts
  • birefnet — Good general purpose
  • pixelcut — Alternative
  • rembg — Cheapest
Show full SKILL.md (598 more words)Show less

How to Write Good Prompts

This is the difference between mediocre and professional output. Read references/prompt-library.md for a full library of tested prompts you can copy and adapt. Key rules:

For generation: Be specific about lighting ("dramatic side lighting from upper right"), camera ("shot on Canon R5 70-200mm f2.8"), and atmosphere ("dust particles visible in the light beam"). Vague prompts produce generic results.

For text-behind-subject: The key phrase is: "Add '[TEXT]' in large bold [color] letters BEHIND the [subject] — the [subject's] body overlaps and partially covers the letters." Without "BEHIND" and the occlusion instruction, the text floats on top.

For edits: Always end with "Keep everything else exactly the same" and list what to preserve. Without this, the AI changes things you didn't want changed.

For background replacement: Use realistic, specific locations ("modern upscale mall entrance during daytime, natural warm daylight"). Over-dramatic backgrounds ("city at night with neon reflections") look obviously fake.

Typography

For big titles and hero text: Use the FAL model via edit — it handles large text well and integrates it into the scene naturally. No font size math needed, just say "very large bold" in the prompt.

For precise small text (credits, URLs, badges, coverlines): Use compose or text with Satori. This is where font sizing matters — images display much smaller on phones. Quick rule: on a 1080px Instagram image, nothing under 36px is readable. Run picture-it download-fonts first if fonts aren't installed.

Hierarchy: Max 3 text sizes per image. Brand name should be larger than tagline.

Font pairing: Serif + sans-serif works best. For FAL model text, just describe the style in the prompt. For Satori, 3 fonts are bundled — drop more .ttf files into ~/.picture-it/fonts/. Run picture-it download-fonts if fonts aren't installed. See references/composition-guide.md for pairing suggestions.

Composition Techniques

Read references/composition-guide.md for detailed multi-pass workflows, product photography, magazine covers, and overlay composition.

Common Workflows

Simple: Generate an image
bash
picture-it generate --prompt "dark cosmic background with nebula" --size 1200x630 -o bg.png
Simple: Add text to an image
bash
picture-it text -i bg.png --title "Hello World" --font "Space Grotesk" --color white --font-size 64 -o hero.png
Medium: Blog header with AI background + text
bash
picture-it generate --prompt "abstract dark tech background" --size 1200x630 -o bg.png
picture-it text -i bg.png --title "My Blog Post" --font "DM Serif Display" --font-size 72 -o header.png
picture-it grade -i header.png --name cinematic -o header-graded.png
Medium: Edit a photo background
bash
picture-it edit -i photo.jpg --prompt "replace background with modern hotel entrance, keep subject identical" --model banana-pro -o edited.jpg
Advanced: Text behind subject (YouTube thumbnail style)
bash
# 1. Generate a scene
picture-it generate --prompt "runner on mountain trail at golden hour" --model flux-dev --size 1280x720 -o runner.png

# 2. Use FAL edit to add text BEHIND the subject
picture-it edit -i runner.png --prompt "Add 'RUN FASTER' in large bold black letters BEHIND the runner — the runner's body overlaps the text" --model seedream -o thumbnail.png
Advanced: Product comparison with real photos
bash
# 1. Remove backgrounds from product photos
picture-it remove-bg -i product-a.png --model bria -o a-cutout.png
picture-it remove-bg -i product-b.png --model bria -o b-cutout.png

# 2. Generate a background
picture-it generate --prompt "split gradient, blue left to orange right" --size 1200x630 -o bg.png

# 3. Compose cutouts onto background with text
picture-it compose -i bg.png --overlays overlays.json -o comparison.png
Advanced: Multi-pass cinematic composition
bash
# 1. Generate base scene
picture-it generate --prompt "dark stage with green spotlight" --model flux-dev --size 2048x1080 -o stage.png

# 2. Edit scene to place objects
picture-it edit -i stage.png -i logo.png --prompt "Place Figure 2 as glowing 3D cube in the spotlight" --model seedream -o composed.png

# 3. Post-process
picture-it crop -i composed.png --size 1200x630 --position attention -o cropped.png
picture-it grade -i cropped.png --name cinematic -o graded.png
picture-it vignette -i graded.png --opacity 0.3 -o final.png

Platform Presets

Use --platform <name> with generate or crop:

PresetSize
blog-featured1200x630
og-image1200x630
youtube-thumbnail1280x720
instagram-square1080x1080
instagram-story1080x1920
twitter-header1500x500

Output Behavior

  • stdout: only the output file path
  • stderr: progress logs
  • Exit 0 on success, Exit 1 on failure

Read stdout to get the file path. This is how you chain commands.

Gotchas

  • Always use --model bria for remove-bg — the default birefnet leaves rectangular artifacts that cause ugly glow/shadow halos when compositing.
  • The glow effect in compose mode blurs the entire rectangular buffer, not the shape. Avoid using glow on cutout images — use the background color/lighting to create the glow effect instead.
  • The shadow effect has the same rectangular artifact issue. For cutout images on clean backgrounds, skip shadows entirely.
  • When editing with FAL, the model may alter product details (logos, text, design elements). For product images where accuracy matters, use remove-bg + compose instead of edit to preserve the original exactly.
  • SeedDream takes ~60 seconds per generation. Don't assume it failed if it's slow.
  • For edit with banana-pro, don't pass resolution or limit_generations params — it auto-detects.
  • Always crop to exact dimensions after FAL generation — FAL models output approximate sizes.
  • Use flux-dev ($0.03) not flux-schnell ($0.003) when image quality matters (hero images, portraits). The quality difference is significant.
  • Satori does NOT support: display:grid, transforms, animations, box-shadow, filters. Use flexbox only.
  • When adding text behind a subject with edit, be very explicit in the prompt: "the text is BEHIND the subject — the subject's body overlaps and partially covers the letters."

© geongeorge, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skill/picture-it of geongeorge/picture-it.

  • SKILL.md
  • references/composition-guide.md
  • references/prompt-library.md

Open the folder on GitHubat commit 5a3247d

Compare with similar skills

Picture It next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Picture It compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Picture It this skillgeongeorge/picture-it132—~3.7kAutomated safety check: PassMIT
Yuv Reel Covershoodini/ai-agents-skills282—~1.1kAutomated safety check: PassNone
Native Transparent ImagegenZSeven-W/craft-skills2251 repos~1.1kAutomated safety check: PassApache-2.0
Youtube Thumbnailhassancs91/claude-youtube-editor328—~2.1kAutomated safety check: NotesMIT
Generating Imagesspencerpauly/awesome-cursor-skills844—~3.7kAutomated safety check: NotesCC0-1.0
Youtube Thumbnailcharlie947/social-media-skills3.8k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Yuv Reel Covers

    hoodini/ai-agents-skills

    Generate unified, on-brand Instagram Reel covers for Yuval (YUV.AI Neon Phoenix system) — the signature look is a giant Hebrew headline BEHIND the subject cutout + a punch line IN FRONT (depth…

    282 GitHub stars~1.1k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Native Transparent Imagegen

    ZSeven-W/craft-skills

    Generate new raster assets that must contain native pixel transparency, then verify the untouched PNG or WebP before delivery.

    225 GitHub starsUsed in 1 repo~1.1k tokens
    Media & CreativeAuto-check passed
  • Youtube Thumbnail

    hassancs91/claude-youtube-editor

    Dedicated YouTube thumbnail generator — interviews you for exactly the style elements you want (environment, text budget, extras, accent color), then renders high-contrast, vibrant, face-consistent…

    328 GitHub stars~2.1k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Generating Images

    spencerpauly/awesome-cursor-skills

    Generate or edit images using the OpenAI Image API (gpt-image-2).

    844 GitHub stars~3.7k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Youtube Thumbnail

    charlie947/social-media-skills

    Generate a branded YouTube thumbnail from a video title. An agent skill from charlie947/social-media-skills.

    3.8k GitHub stars~1.9k tokensUpdated 25 days ago
    Media & CreativeAuto-check passed
  • Canva Designer

    asgeirtj/system_prompts_leaks

    Creates and edits Canva designs, generates images and works with brand kits through the canva CLI and Canva's official MCP server, from decks to social posts.

    69k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check passed

Questions about Picture It

What does Picture It do?

Generate and edit images from the CLI using picture-it. An agent skill from geongeorge/picture-it. Picture It is an agent skill from geongeorge/picture-it. Generate and edit images from the CLI using picture-it.

When should I use Picture It?

Picture It fits situations like: the user asks to create; manipulate images — blog headers; product comparisons; youTube thumbnails.

How do I install Picture It in Claude Code?

Run `npx skills add geongeorge/picture-it --skill picture-it -a claude-code`. Or copy the skill folder (skill/picture-it in geongeorge/picture-it) into .claude/skills/picture-it in your project. Claude Code loads it when a task matches its description.

How do I install Picture It in Codex?

Run `npx skills add geongeorge/picture-it --skill picture-it -a codex`. Or copy the skill folder (skill/picture-it in geongeorge/picture-it) into .agents/skills/picture-it in your project. Codex loads it when a task matches its description.

Can I use Picture It in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add geongeorge/picture-it --skill picture-it -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/picture-it, .gemini/skills/picture-it, .github/skills/picture-it and .opencode/skills/picture-it in your project.

What does Picture It need to run?

Going by SKILL.md and its folder, Picture It needs the command-line tools its instructions call (npm, pnpm and bun) and credentials named FAL_KEY. Our summary lists: Node.js; A credential in FAL_KEY. Compatibility (from SKILL.md): Requires Node.js 18+ and picture-it CLI (npm package). FAL_KEY environment variable needed for AI operations. Network access to fal.ai for image generation/editing..

Does Picture It access the network?

SKILL.md names 2 domains. As links in the text: npmjs.com and fal.ai. This is read from the text; nothing was executed.

Is Picture It safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Picture It use?

Picture It is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Picture It use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.1k tokens, read only when the agent opens those files.

What are the alternatives to Picture It?

Skills that share tags, products or a category with Picture It: Yuv Reel Covers (hoodini/ai-agents-skills, 282 stars), Native Transparent Imagegen (ZSeven-W/craft-skills, 225 stars), Youtube Thumbnail (hassancs91/claude-youtube-editor, 328 stars) and Generating Images (spencerpauly/awesome-cursor-skills, 844 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Picture It?

geongeorge (a GitHub user) maintains it in geongeorge/picture-it, which has 132 GitHub stars. The repository was last updated on April 5, 2026.

Source: geongeorge/picture-it on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.