Agent skill

Visual Identity

by letta-ai in letta-ai/skills

Build and maintain a persistent visual identity for your agent using Flux Kontext Pro.

MITAuto-check passedMedia & Creative

Install Visual Identity

skills CLI
$ npx skills add letta-ai/skills --skill visual-identity -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install letta-ai/skills visual-identity --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/letta-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tools/visual-identity .claude/skills/visual-identity && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
visual-identity
GitHub stars
149
Token cost
~2.6k tokens
SKILL.md length
1,055 words
Files
6 (incl. scripts, references, assets)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Build and maintain a persistent visual identity for your agent using Flux Kontext Pro.

  • Works in 4 steps: Establish the reference → Generate scenes with the reference → Anchor the prompt → …
  • The user asks the agent to generate selfies
  • SKILL.md covers When to use, Environment, Dependencies and Workflow 1: Visual Identity…, plus 7 more sections
  • Runs Python scripts from its folder; calls python3, uv and pip3; reaches api.openai.com; needs OPENAI_API_KEY and BFL_API_KEY

What it does

Visual Identity is an agent skill from letta-ai/skills. Build and maintain a persistent visual identity for your agent using Flux Kontext Pro. Use when the user asks the agent to generate selfies, avatars, character art, or any image that should look like the same person across generations.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts, reference files and assets (for example `references/api.md` and `scripts/generate_image.py`).

It sits in Media & Creative, covering Logo and visual identity and Image generation. It works with OpenAI. The repository describes itself as: A shared repository for skills. Intended to be used with Letta Code, Claude Code, Codex CLI, and other agents that support skills. The licence is MIT.

When your agent uses it

  • The user asks the agent to generate selfies
  • Any image that should look like the same person across generations

Example prompts

  • “/visual-identity”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY
  • A credential in BFL_API_KEY

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Establish the reference
  2. Generate scenes with the reference
  3. Anchor the prompt
  4. Iterate with the user

What it can do on your machine

Read from SKILL.md and the folder at commit 6785511. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • uv
    • pip3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.openai.com

    Also links to:

    • api.bfl.ai
    • platform.openai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • BFL_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Visual Identity loads about 2.6k tokens when it runs, and up to ~3.5k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 1,055 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from letta-ai/skills at commit 6785511, republished under its MIT licence (© letta-ai). 1,055 words, ~2,619 tokens.

Download SKILL.mdSave it as .claude/skills/visual-identity/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
visual-identity
description
Build and maintain a persistent visual identity for your agent using Flux Kontext Pro. Use when the user asks the agent to generate selfies, avatars, character art, or any image that should look like the same person across generations.

Visual Identity

Build a persistent visual identity that stays consistent across sessions. Supports OpenAI (gpt-image-1) and Flux Kontext Pro.

Two workflows:

  1. Visual identity (primary): Establish a reference appearance, then generate new scenes that preserve the same face and features. Identity persists in the agent's memory across sessions.
  2. Text-to-image (secondary): One-off image generation from a text prompt, no identity persistence.

When to use

  • The user asks "show me what you look like" or wants agent selfies
  • The user wants an avatar, profile picture, or character art
  • The user wants to create a visual identity (a consistent character across scenes)
  • The user provides a reference photo and wants variations or new scenes
  • The user asks you to generate or create any image

Environment

The script auto-detects which provider to use based on environment variables:

PriorityEnv varProviderNotes
1stOPENAI_API_KEYOpenAI gpt-image-1Recommended. Most users already have this.
2ndBFL_API_KEYFlux Kontext ProBetter face consistency. Requires BFL account.

You can override with --provider openai or --provider flux.

If neither key is set, guide the user:

Never ask the user to paste the full key in chat.

Dependencies

Install if missing (prefer uv):

bash
uv pip install requests Pillow

If uv is unavailable:

bash
pip3 install requests Pillow

Workflow 1: Visual Identity (Character Consistency)

This is the primary workflow. The goal is to establish a reference appearance and then generate new scenes that preserve the same face, bone structure, and features.

Step 1: Establish the reference

Either the user provides a photo, or you generate a base character:

Option A -- User provides a reference photo: The user pastes or specifies an image file. Save it to the persistent identity directory (see "Persisting Visual Identity" below).

Option B -- Generate a base character from text: Use text-to-image to create the initial character. Be very specific about physical features. Example prompt:

A portrait of a young woman with shoulder-length auburn hair, green eyes, light freckles, wearing a black leather jacket. Clean background, studio lighting, 3:4 portrait.

Save the result as the reference image.

Step 2: Generate scenes with the reference

Pass the reference image as base64-encoded input_image:

bash
python3 <path-to-skill>/scripts/generate_image.py edit \
  --reference /path/to/canonical.jpg \
  --prompt "The same person is sitting at a desk coding late at night, lit by monitor glow" \
  --out /tmp/identity_coding.jpg
Step 3: Anchor the prompt

Always include an identity-anchoring phrase in every prompt that uses a reference. This tells the model to preserve facial features:

Keep his/her exact face, bone structure, eye color, and hair.

Or more naturally woven into the prompt:

The same man is relaxing on a tropical beach at sunset, wearing a linen shirt. Golden hour lighting. Keep his exact face, bone structure, eye color, and hair.

Step 4: Iterate with the user
  • Show each result and ask for feedback
  • Adjust scene, clothing, lighting, or setting based on feedback
  • Always reuse the same reference image for consistency
  • If the user wants to change the base appearance, go back to Step 1
Example session flow
  1. User: "Create a visual identity for me -- here's my photo"
  2. Agent: Saves reference, generates 2-3 scenes (beach, office, hiking)
  3. User: "I like the beach one but make me wearing a hat"
  4. Agent: Regenerates beach scene with hat, same reference
  5. User: "Now make one of me cooking"
  6. Agent: New scene with same reference

Persisting Visual Identity

Two things persist across sessions: the reference image (binary) and the identity metadata (markdown). They live in different places.

Reference image: agent data directory

Save the canonical reference image to ~/.letta/agents/$AGENT_ID/reference/visual-identity/canonical.jpg. This is outside memfs because binary images would bloat the git-backed memory repo. The reference/ directory persists across sessions.

bash
mkdir -p ~/.letta/agents/$AGENT_ID/reference/visual-identity
cp /tmp/generated_portrait.jpg ~/.letta/agents/$AGENT_ID/reference/visual-identity/canonical.jpg
Identity metadata: memfs

After establishing a visual identity, create a memory file at reference/visual-identity.md in the agent's memory filesystem. This syncs via git like all other memory files.

Use the Memory tool to create it:

memory(command="create", reason="Store visual identity metadata",
  file_path="reference/visual-identity.md",
  description="Agent's persistent visual identity -- reference image path and appearance description.",
  file_text="## Reference Image\n~/.letta/agents/$AGENT_ID/reference/visual-identity/canonical.jpg\n\n## Appearance\n- Hair: shoulder-length auburn, slight wave\n- Eyes: green\n- Skin: light with freckles\n- Build: athletic\n- Distinguishing: small scar above left eyebrow\n\n## Anchoring Phrase\nKeep the exact same face, bone structure, eye color, and hair from the reference image.\n\n## History\n- Established: 2026-04-15\n- User feedback: \"make the hair a bit darker\" -> regenerated, approved")
Show full SKILL.md (424 more words)Show less
Auto-detect on load

When this skill is loaded, check the agent's memory tree for reference/visual-identity.md. If it exists:

  • The agent already has an established identity
  • Use the stored reference image path for all image generation requests
  • Prepend the stored anchoring phrase to every prompt
  • Do not ask the user to re-establish their identity

If it does not exist, the agent has no visual identity yet. Offer to create one if the user asks for images.

Updating the identity

If the user wants to change their visual identity:

  1. Generate or receive the new reference image
  2. Overwrite canonical.jpg in the reference directory
  3. Update the memory file with new appearance details
  4. Note the change in the History section

Workflow 2: Text-to-Image

For one-off image generation that does not need identity persistence.

bash
python3 <path-to-skill>/scripts/generate_image.py generate \
  --prompt "A corgi wearing a tiny space helmet on the moon" \
  --out /tmp/corgi_moon.jpg

Or inline with requests (OpenAI):

python
import requests, base64, os

resp = requests.post(
    "https://api.openai.com/v1/images/generations",
    headers={
        "Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}",
        "Content-Type": "application/json",
    },
    json={
        "model": "gpt-image-1",
        "prompt": "A corgi wearing a tiny space helmet on the moon",
        "n": 1,
        "size": "1024x1024",
        "quality": "medium",
    },
).json()

img = base64.b64decode(resp["data"][0]["b64_json"])
with open("/tmp/corgi_moon.png", "wb") as f:
    f.write(img)

Parameters

ParameterValuesDefaultProviderNotes
--promptstringrequiredBothScene description
--referencefile pathnoneBothReference photo for identity mode (edit only)
--provideropenai, fluxautoBothOverride provider auto-detection
--aspect-ratio1:1, 3:4, 4:3, 16:9, 9:163:4BothUse 3:4 for portraits
--output-formatpng, jpeg, webppngBoth
--qualitylow, medium, highmediumOpenAIImage quality
--seedintegerrandomFluxFix for reproducible results
--safety-tolerance0-62FluxHigher = more permissive
--guidance1.5-100variesFluxPrompt adherence strength

Prompting best practices

  • Be specific about physical setting, lighting, clothing, and pose
  • For portraits, specify aspect ratio 3:4 or 4:3
  • For landscapes/scenes, use 16:9
  • Include lighting direction: "golden hour", "studio lighting", "neon-lit"
  • Describe clothing and accessories explicitly
  • For identity mode, always include the anchoring phrase about preserving facial features
  • Avoid contradicting the reference photo (e.g., don't say "blonde hair" if the reference has dark hair)

Rate limits and costs

  • Maximum 6 concurrent requests per API key
  • Queue times: typically 4-10 seconds, can spike to 2+ minutes under load
  • If a request stays in Pending for over 120 seconds, retry once
  • Polling interval: 2 seconds is sufficient
  • Download URLs in the Ready response are signed and expire; save images immediately

CLI reference

Full CLI documentation: references/api.md

Common commands:

bash
# Text-to-image
python3 <path-to-skill>/scripts/generate_image.py generate \
  --prompt "..." --out output.jpg

# Reference-based editing (visual identity)
python3 <path-to-skill>/scripts/generate_image.py edit \
  --reference photo.jpg --prompt "..." --out output.jpg

# Dry run (show request without sending)
python3 <path-to-skill>/scripts/generate_image.py generate \
  --prompt "..." --dry-run

# Custom aspect ratio and seed
python3 <path-to-skill>/scripts/generate_image.py generate \
  --prompt "..." --aspect-ratio 16:9 --seed 42 --out wide.jpg

Error handling

OpenAI:

  • HTTP 400/422: Usually a malformed request or content policy violation
  • HTTP 429: Rate limited -- wait and retry
  • Missing OPENAI_API_KEY: Guide the user to https://platform.openai.com/api-keys

Flux:

  • Insufficient credits: Direct user to https://api.bfl.ai/credits
  • HTTP 422: Usually a malformed request -- check prompt and parameters
  • Pending timeout: Retry the request; the queue may be congested
  • Missing BFL_API_KEY: Guide the user to https://api.bfl.ai

Both:

  • Image too large for base64: Resize to under 10MB before encoding
  • No API key found: See Environment section

© letta-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references, assets) in tools/visual-identity of letta-ai/skills.

  • SKILL.md
  • LICENSE
  • assets/flux.png
  • assets/visual-identity-small.svg
  • references/api.md
  • scripts/generate_image.py

Open the folder on GitHubat commit 6785511

Compare with similar skills

Visual Identity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Visual Identity compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Visual Identity this skillletta-ai/skills149—~2.6kAutomated safety check: PassMIT
Image Generationonyx-dot-app/onyx32k1 repos~1.7kAutomated safety check: PassCustom licence
Design Masterminhnv0807/ai-business-skills609—~4.6kAutomated safety check: PassMIT
CarouselsTheCraigHewitt/skills159—~2.3kAutomated safety check: PassMIT
Generate ImageK-Dense-AI/claude-scientific-writer2.4k1 repos~3.8kAutomated safety check: NotesMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Image Generation

    onyx-dot-app/onyx

    Generate or edit raster images (photos, illustrations, textures, sprites, mockups, logos, infographics) using the workspace's configured image-generation provider via onyx-cli image.

    32k GitHub starsUsed in 1 repo~1.7k tokens
    Media & CreativeAuto-check passed
  • Design Master

    minhnv0807/ai-business-skills

    Handles eight kinds of marketing visual requests, from logos and campaign key visuals to infographics and quote graphics, by generating images or writing paste-ready prompts.

    609 GitHub stars~4.6k tokensUpdated 27 days ago
    Media & CreativeAuto-check passed
  • Carousels

    TheCraigHewitt/skills

    Turns a piece of Craig's content (an email, a YouTube script, an essay, or pasted text) into a polished image carousel publishable to both LinkedIn and Instagram from one set of 1080x1350 slides.

    159 GitHub stars~2.3k tokensUpdated 4 mo ago
    Media & CreativeAuto-check passed
  • Generate Image

    K-Dense-AI/claude-scientific-writer

    Generate or edit images with AI models through the OpenRouter Image API (Gemini, Seedream, Recraft, GPT-Image, Riverflow).

    2.4k GitHub starsUsed in 1 repo~3.8k tokens
    Media & CreativeAuto-check: notes
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • GPT Image Generation CLI

    wuyoscar/GPT-Image2-Skill

    Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.

    5.7k GitHub stars~2.5k tokensUpdated 9 days ago
    Media & CreativeAuto-check: notes

More from letta-ai/skills

All 22 skills in this repo
  • AI News

    letta-ai/skills

    Fetch and summarize recent AI news from curated RSS feeds (Hugging Face, VentureBeat, The Verge, OpenAI, Anthropic, DeepMind, etc.) and YouTube channels (Yannic Kilcher, Two Minute Papers, AI…

    149 GitHub stars~582 tokensUpdated 8 days ago
    Auto-check passed
  • Builds and debugs Letta Code channels, including first-party channel adapters and dynamic user channel plugins under ~/.letta/channels.

    149 GitHub stars~1.1k tokensUpdated 8 days ago
    Auto-check passed
  • Letta Configuration

    letta-ai/skills

    Configure LLM models and providers for Letta agents and servers.

    149 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check: notes
  • Migrates deprecated Letta Filesystem folders/files to MemFS using markdown document corpora, chunking, local lexical search, and QMD semantic search via the memfs-search skill.

    149 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Memfs Search

    letta-ai/skills

    Semantic search over agent memory files. An agent skill from letta-ai/skills.

    149 GitHub stars~901 tokensUpdated 8 days ago
    Auto-check passed
  • Navigates archived ChatGPT or Claude-style conversation exports and a MemFS reference archive on demand.

    149 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed

Works with

Questions about Visual Identity

What does Visual Identity do?

Build and maintain a persistent visual identity for your agent using Flux Kontext Pro. Visual Identity is an agent skill from letta-ai/skills. Build and maintain a persistent visual identity for your agent using Flux Kontext Pro.

When should I use Visual Identity?

Visual Identity fits situations like: the user asks the agent to generate selfies; any image that should look like the same person across generations.

How do I install Visual Identity in Claude Code?

Run `npx skills add letta-ai/skills --skill visual-identity -a claude-code`. Or copy the skill folder (tools/visual-identity in letta-ai/skills) into .claude/skills/visual-identity in your project. Claude Code loads it when a task matches its description.

How do I install Visual Identity in Codex?

Run `npx skills add letta-ai/skills --skill visual-identity -a codex`. Or copy the skill folder (tools/visual-identity in letta-ai/skills) into .agents/skills/visual-identity in your project. Codex loads it when a task matches its description.

Can I use Visual Identity in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add letta-ai/skills --skill visual-identity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/visual-identity, .gemini/skills/visual-identity, .github/skills/visual-identity and .opencode/skills/visual-identity in your project.

What does Visual Identity need to run?

Going by SKILL.md and its folder, Visual Identity needs Python for the scripts in its folder, the command-line tools its instructions call (python3, uv and pip3) and credentials named OPENAI_API_KEY and BFL_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in BFL_API_KEY.

Does Visual Identity access the network?

SKILL.md names 3 domains. In commands or code: api.openai.com; the agent is likely to contact it when it follows the instructions. As links in the text: api.bfl.ai and platform.openai.com. This is read from the text; nothing was executed.

Is Visual Identity safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Visual Identity use?

Visual Identity is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Visual Identity use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 837 tokens, read only when the agent opens those files.

What are the alternatives to Visual Identity?

Skills that share tags, products or a category with Visual Identity: Image Generation (onyx-dot-app/onyx, 32k stars), Design Master (minhnv0807/ai-business-skills, 609 stars), Carousels (TheCraigHewitt/skills, 159 stars) and Generate Image (K-Dense-AI/claude-scientific-writer, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Visual Identity?

letta-ai (a GitHub organization) maintains it in letta-ai/skills, which has 149 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 1, 2026.

Source: letta-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.