Agent skill

Editorial Vision Studio

by Yu-0312 in Yu-0312/editorial-vision-studio

Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters…

MITAuto-check passedMedia & Creative

Install Editorial Vision Studio

skills CLI
$ npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Yu-0312/editorial-vision-studio editorial-vision-studio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
editorial-vision-studio
GitHub stars
226
Token cost
~7.5k tokens
SKILL.md length
3,079 words
Files
94 (incl. assets)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters…

  • Works in 12 steps: Intent Engine → 5: Style Gate → Visual Analyzer → …
  • Visual concepting
  • SKILL.md covers When to Use, Architecture: Decision Engine…, Core Pipeline and Step 0: Intent Engine, plus 21 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Editorial Vision Studio is an agent skill from Yu-0312/editorial-vision-studio. Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters for GPT Image, Flux, Ideogram, and generic image backends. Use for visual concepting, image prompts, photo-to-design, posters, covers, zines, gallery prints, campaigns, brand key visuals, product/editorial imagery, social assets, website hero art, moodboards, presentation decks, Panter-style low-contrast recovery…

Its SKILL.md is about 7.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 97 other files, including assets (for example `README.en.md`, `README.md` and `adapters/_template.md`).

It sits in Media & Creative, covering Image generation. The licence is MIT.

When your agent uses it

  • Visual concepting
  • Photo-to-design
  • Brand key visuals
  • Product/editorial imagery

Example prompts

  • “/editorial-vision-studio”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Intent Engine
  2. 5: Style Gate
  3. Visual Analyzer
  4. Visual Language Engine
  5. Art Direction Engine
  6. Editorial Planner
  7. Recovery Engine
  8. Prompt Compiler + Model Adapter
  9. Prompt Reviewer
  10. 5: Prompt Optimizer
  11. Quality Evaluator
  12. Iteration Engine

What it can do on your machine

Read from SKILL.md and the folder at commit efd81c2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Editorial Vision Studio loads about 7.5k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 3,079 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~165
When it runs · the whole SKILL.md, loaded when a task matches
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Yu-0312/editorial-vision-studio at commit efd81c2, republished under its MIT licence (© Yu-0312). 3,079 words, ~7,544 tokens.

Download SKILL.mdSave it as .claude/skills/editorial-vision-studio/SKILL.md (or your agent's skills folder). This skill also uses 93 other files; get the full folder from GitHub.
name
editorial-vision-studio
description
Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters for GPT Image, Flux, Ideogram, and generic image backends. Use for visual concepting, image prompts, photo-to-design, posters, covers, zines, gallery prints, campaigns, brand key visuals, product/editorial imagery, social assets, website hero art, moodboards, presentation decks, Panter-style low-contrast recovery, multi-image series with a consistent visual system, or switching image models while preserving the same creative direction.

Editorial Vision Studio

AI Creative Director for Visual Generation.

Philosophy: Do not decorate. Always interpret.

An extensible Editorial Design Engine: one decision pipeline, swappable model adapters. Every layer — intent, analysis, art direction, planning, recovery, compilation, review, iteration — is a separate module with a single job, and the contract between them is a model-agnostic spec.

Architecture: reference/architecture.md

When to Use

  • User asks for AI image direction, image prompts, art direction, visual concepting, or prompt adaptation across models
  • User uploads a photo and asks for photo-to-design, editorial poster, cover, zine, gallery print, campaign key visual, brand visual, product visual, or hero image
  • User gives a theme only and wants a poster, social asset, zine, campaign, moodboard, or conceptual image
  • User mentions low-contrast / gray photo recovery (Panter compensation)
  • User specifies a style: Swiss, Kinfolk, MUJI, Brutalist, Wallpaper*, Purple, Apartamento, POPEYE
  • User wants analysis → direction → prompt → image, not immediate generation
  • User specifies model: gpt-image, flux, ideogram — or asks to reuse direction with a different model
  • User wants a set, not an image: campaign at all sizes, a carousel, a multi-page deck
  • User wants a second image that matches the first — series consistency, brand system, "同一套視覺"
  • User names a preset, or wants a specific fixed look reproduced exactly

Architecture: Decision Engine + Model Adapters

DECISION ENGINE (fixed)          MODEL ADAPTER (swappable)
Intent → Style Gate                   VisionSpec / EditorialSpec
      → Analyzer             →      ↓
      → Visual Language      →   adapters/{model}.md
      → Art Direction        →      ↓
      → Planner              →   GenerationRequest → API
      → Recovery             →
      → Compiler → VisionSpec
                                 SHARED POST-LAYER
                                 Reviewer → Optimizer → Generate
                                          → Evaluator → Iteration
                                                             │
                                 one spec mutation,          │
                                 re-enter Compiler ◄─────────┘
                                 (max 3 passes)

Core Pipeline

User Request
    ↓
Intent Engine          → [prompts/intent.md](prompts/intent.md)
    ↓                    (series? → [prompts/series.md](prompts/series.md) · memory? → [prompts/visual-memory.md](prompts/visual-memory.md))
Style Gate             → [prompts/style-gate.md](prompts/style-gate.md)  ← ask the user once, up front
    ↓
Visual Analyzer        → [prompts/analyzer.md](prompts/analyzer.md)  (skip if theme-only / prompt-only)
    ↓
Visual Language Engine → [prompts/visual-language.md](prompts/visual-language.md)
    ↓
Art Direction Engine   → [prompts/art-direction.md](prompts/art-direction.md)  ← 2–3 candidates, commit to one
    ↓
Visual Planner         → [prompts/planner.md](prompts/planner.md)
    ↓
Recovery Engine        → [prompts/recovery.md](prompts/recovery.md) + [recovery/](recovery/)
    ↓
VisionSpec             → [spec/editorial-spec.schema.md](spec/editorial-spec.schema.md)
    ↓
Model Adapter          → [adapters/registry.md](adapters/registry.md)  ← swappable
    ↓
Prompt Reviewer        → [prompts/reviewer.md](prompts/reviewer.md)  ← is it legal?
    ↓
Prompt Optimizer       → [prompts/optimizer.md](prompts/optimizer.md)  ← is it well said?
    ↓
Image Generation
    ↓
Quality Evaluator      → [prompts/evaluator.md](prompts/evaluator.md)  → quality vector
    ↓
Iteration Engine       → [prompts/iteration.md](prompts/iteration.md)  ← fix one layer, loop (max 3)
    ↓
VisualManifest         → [spec/visual-manifest.schema.md](spec/visual-manifest.schema.md)

Each layer does one job. Never analyze in Compiler. Never generate in Analyzer. Never re-decide in Iteration — mutate the spec and re-run the layer that owns the failure.

Quick routing: reference/decision-tree.md

Step 0: Intent Engine

Before analyzing pixels, resolve user goal → output family:

User saysIntentAllowed outputs
art book coverArt Bookmagazine-cover, gallery-print, poster
TEDx key visualEvent Campaigncampaign-poster, brand-key-visual, social-asset
skincare brand launchBrandingbrand-key-visual, product-editorial, social-asset, campaign-poster, website-hero
app hero imageDigital Productwebsite-hero, interface-asset, social-asset
zine pageZinezine, poster, editorial-spread
gallery printGallerygallery-print, photo-abstract-diptych, poster
moodboardVisual Conceptmoodboard, poster, editorial-spread
10-page deckany family, purpose: presentationpresentation-deck

Also resolve the six intent dimensions — subject, purpose, audience, emotion, platform, aspect ratio. Infer them; ask only when a missing one is load-bearing. They are what every later layer is graded against.

Detect scope here, not later: series_id for a set (prompts/series.md), memory_id for a continuation (prompts/visual-memory.md).

Read prompts/intent.md. Reject mismatched formats (e.g. gallery print for TEDx campaign).

Step 0.5: Style Gate

The first question about how the image should look. Before analysing a single pixel, ask which look they want — a numbered menu of the compatible presets plus 「讓 AI 提案」. (Intent may have asked about purpose or platform first; that is a different question.)

Half a step because it computes nothing. It asks one question and turns the answer into a lock.

要什麼風格?

1. 米色明信片 —— 米白紙底,照片重畫成簡化色塊,大量留白。
2. 時代海報 —— 飽和油墨滿版,硬邊平面色塊,地名做成版面。
3. 紙雕明信片 —— 寫實照片,明信片上長出立體紙雕世界。1:1。
4. 自己描述 —— 你說想要的樣子,我拆解照片後照你的描述重畫。
5. 讓 AI 提案 —— 看過照片和用途後給你三個方向再選。

Nobody chooses a look twice unless they asked to see proposals. The gate emits a single field, style_gate.outcome, and Art Direction branches on that and nothing else:

Gate answeroutcomeArt Direction
A presetcommitBecomes the session's Visual Memory; auto-commits, asks nothing
自己描述, or just typing a descriptioncommitCaptured verbatim; the Analyzer deconstructs the photo, then prompts/style-brief.md rebuilds it to the description. A description is never answered with a menu
「讓 AI 提案」offerOffers 2–3 candidates. This is the one path with two prompts, and the user asked for it

It skips — outcome: commit, no menu — when a preset or style is already named, a memory or series is active, a direction was committed earlier this session, or the run is unattended. It skips to outcome: offer when fewer than two presets fit the intent, because a one-item menu is not a choice. Unattended runs pick the highest-fit preset with intent-only weights and say which in one line — never silently.

Presets declare intended_layouts; the gate offers only those that intersect intent.allowed_outputs.

Show looks, not internal vocabulary. Nobody outside this repo knows what Kinfolk means.

The gate settles ground and medium only. Layout, composition, abstraction, and copy stay with the Planner.

Read prompts/style-gate.md.

Step 1: Visual Analyzer

Produce structured Image Report with star ratings and Editorial Score (0–100).

Dimensions: subject, clarity, contrast, saturation, composition, negative space, geometry, texture, lighting, emotion.

Read prompts/analyzer.md.

Step 2: Visual Language Engine

Derive Visual Language first, then style/palette/layout — not the reverse.

Examples: Museum → Swiss or MUJI + fine serif; Quiet Human → Kinfolk + cream/sage; Indie Memory → Zine + riso anchor; Poster Graphic → Travel Poster + flat inks.

Read prompts/visual-language.md. User style: override skips auto-derivation but Reviewer still validates DNA fit.

Step 3: Art Direction Engine

Do not generate the first plausible reading of the brief. Draft competing directions, each with a one-sentence thesis and a named trade-off, score them for fit, commit to one, and keep the runner-up.

Branch on style_gate.outcome and nothing else. The gate already resolved every case — preset chosen, free text, 「你決定」, explicit style:, memory, series, unattended, too few presets — into one of two values:

outcomeBehaviour
commitBuild one direction, commit silently. Never offer candidates — that would be the user's second time choosing a look. Name the runner-up in one line when one exists
offerPresent 2–3 candidates and ask. Two for a narrow intent family, three for a theme-only brief or Editorial Score <50

The layer always runs, even on commit — every spec needs an art_direction block with a thesis.

Every candidate must set a different ground. Eight of the eleven style DNAs resolve to a light paper field, so candidates chosen on style alone come back as three shades of ivory. Beyond that, candidates must differ on at least two of: render mode, visual language, layout family, abstraction level, typography weight. Palette swaps are not directions.

The committed direction is a hard constraint on the Planner. Switching to the runner-up later re-runs Planner onward only — never the Analyzer.

Read prompts/art-direction.md.

Step 4: Editorial Planner

Decide layout, typography direction, abstraction level inside the committed direction — not the final prompt.

Key rules (full matrix in prompts/planner.md):

  • Portrait + negative space >50% → Magazine Cover
  • Architecture + strong geometry → Swiss Poster
  • Landscape + quiet mood → Gallery Print
  • Street + human story → Documentary Zine
  • Food/object + minimal → Product Editorial

The Planner also fills direction.spatial_plan — projection, shared ground plane, light direction, arrangement. On a photo run it is a verbatim copy of image_report.spatial; on a theme-only run the Planner authors it from intent.subject. Spatial fields locked by an active preset or memory are applied over the result field by field — a locked projection: flat-elevation is an authored decision that overrides what the photograph shows. It is required either way: prompts/reviewer.md rejects a prompt with no named projection whether or not a photograph existed, and without the field those clauses have no provenance for prompts/optimizer.md to protect.

If user specifies style: kinfolk, load styles/kinfolk.md DNA.

Step 5: Recovery Engine

Apply only when Image Report flags weakness. Each recovery is one atomic fix — see recovery/.

ProblemRecovery
Low contrast / gray (saturation <30%)Panter Mode: warm/cool conflict hues, high-sat anchor, wider tonal separation
Weak subjectIncrease silhouette / scale
Flat lightingDirectional light
Busy backgroundSimplify geometry
Too many colorsLimit palette to 4
No focal pointEditorial color anchor
No rhythmAbstract panel

Panter Mode: discard dull grays; boost warm to 75% / cool to 70% saturation; add 8% high-chroma anchor block; widen tonal separation and mark scale. Panter is a colour compensation and never adds texture on its own. See recovery/contrast.md.

Texture Permission — single source of truth: assets/texture.md. Three tiers: PRINT (riso/halftone/scan defects) is zine only; SURFACE (substrate character such as cotton paper) is allowed on CLEAN layouts whose style DNA rates Texture ★★★+; FLAT (zero texture words) covers the photo-abstract-diptych panel ground, interface-asset, the website-hero copy-safe area, the product-editorial background, and every page of a presentation-deck that contains a data page. Recoveries never raise a layout's tier.

Never redesign the entire image unless Editorial Score <50 (Concept Reconstruction).

Step 6: Prompt Compiler + Model Adapter

Phase 1: Assemble VisionSpec / EditorialSpec — read prompts/compiler.md

Phase 2: Route to adapter by target.model:

ModelWhenAdapter
gpt-image (default for photo upload)Diptych, photo fidelityadapters/gpt-image.md
fluxZine texture, atmosphereadapters/flux.md
ideogramCover/campaign typographyadapters/ideogram.md
genericUnknown backendadapters/generic.md

User: model: flux or "用 Flux 生成" → set adapter, do not re-analyze.

Same direction, different model: reuse VisionSpec / EditorialSpec, swap adapter only.

Step 7: Prompt Reviewer

Before generation, run conflict detection. Read prompts/reviewer.md.

Examples:

  • Swiss grid + Kinfolk organic → reject or resolve
  • MUJI + heavy typography → reject
  • Brutalist + soft pastoral palette → warn

Auto-correct incompatible pairings.

Step 7.5: Prompt Optimizer

The Reviewer answers is this prompt legal. It cannot fix a prompt that is legal but weakly worded — the ground clause buried in sentence four, a spec number softened to "a small accent", the same instruction stated twice.

Half a step because it decides nothing. Every clause it emits must trace to a field the spec already holds; a clause with no spec field behind it is a rejection back to the Compiler, not an invention. It never sets a ground, style, subject, layout, palette, or type scale.

Seven ordered ops: provenance_strip · ground_first · concretize · bind_numbers · dedupe · route_negatives · compress. Most gain comes from concretize and bind_numbers — a model renders nouns and numbers, never an adjective about taste.

This does not break the "mutate the spec, never the prompt string" rule in prompts/iteration.md. That rule forbids hand patches after generation. The Optimizer is deterministic, pre-generation, and recorded: same spec + same adapter + same ruleset_version produces the same string, so the manifest still replays.

No op firing is the normal outcome. The same op firing twice means the adapter emits that defect every time — fix adapters/{model}.md, not the string.

Variant mode emits 2–3 prompts from one spec, differing only on clause_density, specificity, and emphasis_order. A variant that differs on any spec field is a second direction, not a variant. Offer it only when the user asks to compare wordings, when the model is generic, or when a prompt-attributed dimension failed last run.

Read prompts/optimizer.md.

Step 8: Quality Evaluator

After generation, score a quality vector — not one number. Ten weighted dimensions: subject, composition, focal_point, palette, typography, texture, style_coherence, photo_fidelity, intent_fit, platform_fit. Each 0.00–1.00; inapplicable dimensions are null, never 0.

overall is the weighted mean and maps to grade A–D. Any single dimension below 0.60 fails the run regardless of overall.

The evaluator must name lowest_failing and responsible_layer — without them the iteration loop has nothing to act on.

Read prompts/evaluator.md.

Step 9: Iteration Engine

A low score points at one layer, not at the whole image. Fix that layer, recompile, regenerate.

lowest failing dimension → responsible layer → one minimal spec mutation → recompile → re-score
  • One mutation per iteration. Two at once makes the next score unattributable.
  • Mutate the spec, never the prompt string. A hand-patched prompt cannot be replayed from the manifest.
  • Escalate when the same dimension fails twice: Compiler → Recovery → Planner → Art Direction → Intent.
  • Re-enter at the Compiler after every mutation. "Responsible layer" names whose decision changed, not an entry point.
  • Stop at overall ≥ 0.85 with no dimension below 0.60, at 3 iterations, on score regression, or when escalation reaches Intent — then report, don't guess.

Taste disagreement ("I don't like the blue") is a direction change, not a QC failure — route it to prompts/art-direction.md.

Read prompts/iteration.md.

Show full SKILL.md (1,228 more words)Show less

Ground & Render Mode

Two axes, required, with no default:

AxisValuesQuestion it answers
design_tokens.groundpaper-light · neutral-gray · dark · saturated · full-bleed-photo · duotoneWhat is the canvas field itself?
direction.render_modephotographic · photo-plus-graphic · graphic · painterly · mixedWhat medium is the image made of?

Enum values are contract tokens, never prompt words — every adapter resolves them to prose through assets/ground.md.

An unset value is a rejection, not a fallback. This matters more than it looks: before these fields existed, an undecided ground fell through to warm ivory paper on every run, because that was the most-repeated value in the repo. A field with no default cannot be skipped.

render_mode is not abstraction_level. Medium and subject literalness are independent — a photographic image can still be full-abstract.

Presets

A preset is a shipped VisualMemory with source: preset — a locked partial spec authored in the repo instead of established from a run. It reuses the entire memory mechanism; nothing new enforces it.

PresetGroundRender mode
ivory-postcardpaper-lightpainterly
vintage-travel-postersaturatedgraphic
papercraft-diorama-postcardfull-bleed-photophotographic

The Style Gate offers these as a numbered menu at the start of a run; preset: ivory-postcard is the shortcut that skips the menu, for automation and series work. Three different grounds is deliberate — it is what gives the candidate-diversity rule something to draw on. Registry and authoring guide: presets/registry.md.

Unlike a series memory, a preset may lock composition, aspect_ratio, layout, and — field by field, under the spec's own key names — spatial_plan: a preset is avowedly a template, which is what it is for. Locked spatial fields override the photographed values; unlocked fields still come from the photo or the Planner.

Visual Memory & Series

The engine keeps a visual system across images, not just within one.

AskRead
"now do the next one," "同一套視覺," brand assets suppliedprompts/visual-memory.md
"一套 10 頁簡報," "campaign 全尺寸," carousel, all platform sizesprompts/series.md

Visual Memory locks the identity fields — style, visual language, palette, typography, texture tier — and leaves layout, composition, abstraction level, and aspect ratio free per image. Locking composition produces a template, not a system. On a memory run, Art Direction still runs but auto-commits to the locked DNA rather than offering candidates. A lock that no longer fits gets forked, never silently mutated. Contract: spec/visual-memory.schema.md.

Series runs Intent → Analyzer → Visual Language → Art Direction once, then fans out per output. Generate the hero first and pass QC before derivatives — a weak hero multiplies into N weak frames. Cross-image QC scores palette drift, typographic identity, texture tier, anchor legibility, and compositional variety; a single outlier is fixed alone, never by re-running the set.

Run Manifest

Every run emits spec/visual-manifest.schema.md: direction taken and runner-up, visual system, quality vector, iteration history, prompt hash, provenance.

This is what makes the cheap paths cheap:

User saysRe-run
"同一張,改用 Flux"Adapter → Reviewer → Generate
"改用 B 那個方向"Planner → downstream
"一模一樣再生一次"Generate only
"同一套視覺,換主題"Analyzer → Art Direction (auto-commit) → Planner → downstream

Never re-run the Analyzer when a valid image_report for the same source image already exists. Keep the manifest internal unless the user asks for it or the run is part of a series.

Editorial Score & Modes

This is the Editorial Score — an input measure of the source photo, produced by the Analyzer. It is not the Quality Score from Step 8, which measures the generated image on different bands.

Editorial ScoreMode
90+Premium Editorial — refined extraction, minimal recovery
(no photo)Theme-only: no score exists. editorial_mode: standard, recoveries: [] — never invent a score
70–89Standard Editorial
50–69Compensation Mode — apply Recovery stack
<50Concept Reconstruction — abstract reinterpretation

Output Contract

Match the requested depth. Default to a concise direction summary plus GenerationRequest.

  • Include an Image Report only when a source image is analyzed.
  • Include full VisionSpec / EditorialSpec when the user asks for a reusable direction, comparison, or model switch.
  • Include a generated image only when an image-generation tool is available and the user asks for generation; otherwise return the model-ready prompt.
  • Include Quality Grade and evaluator notes after generating, or when the user requests review.
  • Include the full quality vector only when the user asks why, or when a dimension failed.
  • Include the optimizer op log only when the user asks what changed, or when it reverted. In variant mode show the prompts, never the op log — the user is choosing wordings, not reviewing edits.
  • Include the VisualManifest only on request, or when the run belongs to a series.
  • Present Art Direction candidates as a 3-line table — name, thesis, trade-off — never as raw YAML.
Model switch without re-analysis

User: "同一份方向,改用 Ideogram" → reuse VisionSpec / EditorialSpec, run adapters/ideogram.md only.

Direction switch without re-analysis

User: "改用 B 那個方向" → reuse Intent + Image Report, set the runner-up as selected, re-run Planner onward.

Series and continuation

User: "同一套視覺,做東京街景" → load the Visual Memory, run Analyzer on the new photo, let Art Direction auto-commit to the locked DNA, and Planner sets free fields only.

Bilingual output
  • Image prompt: English (model-optimized)
  • Analysis/direction summary: match user's language (中文/English)

Guardrails

Never:

  • Redraw, filter, or stylize the original photo region when photo_policy.fidelity: required
  • Blindly copy fixed 60/30/10 layout — adapt proportions to subject
  • Mix style languages without Reviewer pass
  • Overload typography or decorative elements
  • Let the Optimizer add a clause no spec field backs, or lengthen a prompt to strengthen it

Always:

  • Preserve the source's arrangement, overlaps, and relative scale — abstraction removes detail, not relationships
  • Name one projection, one ground plane, one light direction in direction.spatial_plan; give every grounded object a contact shadow
  • Preserve visual identity of source photo when one is provided
  • Make every abstract mark traceable to a photo fact, theme fact, brand cue, or stated goal
  • Keep prompts imageable and concrete
  • Apply Recovery only when Image Report warrants it
  • Give each direction a thesis before giving it a palette
  • Fix one layer per iteration, and stop at three
  • Fork a Visual Memory rather than mutating a lock to rescue one image

Style & Layout Reference

Asset moduleFile
Ground & Render Modeassets/ground.md
Scene construction (space)assets/scene-construction.md
Texture permissionassets/texture.md
Paletteassets/palette.md
Typographyassets/typography.md

Extending the Engine

ExtendActionTouch Decision Engine?
New style (Aesop, NYT Mag)Add styles/foo.mdNo
New presetAdd presets/foo.md + registerNo
New layoutAdd layouts/foo.mdNo
New recoveryAdd recovery/foo.mdNo
New image modelAdd adapters/foo.md + registerNo
New intent familyEdit prompts/intent.mdYes (minimal)
New QC dimensionEdit prompts/evaluator.md + map a layer in prompts/iteration.mdYes (minimal)
New prompt-wording opAdd a row to prompts/optimizer.md + bump ruleset_versionNo
New lockable DNA fieldEdit spec/visual-memory.schema.md + enforce in prompts/reviewer.mdYes (minimal)

See adapters/_template.md for new models.

Extending Styles

Add new magazines/brands by creating styles/your-style.md with Style DNA table + compiler clauses. No need to rewrite SKILL.md.

Engine Reference

LayerFile
Intent Engineprompts/intent.md
Style Gateprompts/style-gate.md
Style Briefprompts/style-brief.md
Visual Analyzerprompts/analyzer.md
Visual Language Engineprompts/visual-language.md
Art Direction Engineprompts/art-direction.md
Visual Plannerprompts/planner.md
Recovery Engineprompts/recovery.md
Prompt Compilerprompts/compiler.md
Prompt Reviewerprompts/reviewer.md
Prompt Optimizerprompts/optimizer.md
Quality Evaluatorprompts/evaluator.md
Iteration Engineprompts/iteration.md
Visual Memoryprompts/visual-memory.md
Series Plannerprompts/series.md
Preset Registrypresets/registry.md
ContractFile
VisionSpec / EditorialSpecspec/editorial-spec.schema.md
VisualManifestspec/visual-manifest.schema.md
VisualMemoryspec/visual-memory.schema.md

Agent Config

Model parameters: agents/openai.yaml

© Yu-0312, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 93 other files (assets) in the repository root of Yu-0312/editorial-vision-studio.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.en.md
  • README.md
  • adapters/_template.md
  • adapters/flux.md
  • adapters/generic.md
  • adapters/gpt-image.md
  • adapters/ideogram.md
  • adapters/registry.md
  • agents/openai.yaml
  • assets/examples/autumn-walk-postcard.webp
  • assets/examples/harbor-postcard.webp
  • assets/examples/mountain-dawn-postcard.webp
  • assets/examples/mountain-valley-postcard.webp
  • assets/examples/osaka-castle-postcard.webp
  • … and 77 more

Open the folder on GitHubat commit efd81c2

Compare with similar skills

Editorial Vision Studio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Editorial Vision Studio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Editorial Vision Studio this skillYu-0312/editorial-vision-studio226—~7.5kAutomated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
Structured Image Generationbytedance/deer-flow84k4 repos~2.9kAutomated safety check: PassMIT
Canghe Comicfreestylefly/canghe-skills4618 repos~3.2kAutomated safety check: PassNone
Generate Imageynulihao/AgentSkillOS61810 repos~1.7kAutomated safety check: NotesNone
GPT Image Generation CLIwuyoscar/GPT-Image2-Skill5.7k—~2.5kAutomated safety check: NotesMIT

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Structured Image Generation

    bytedance/deer-flow

    Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.

    84k GitHub starsUsed in 4 repos~2.9k tokens
    Media & CreativeAuto-check passed
  • Canghe Comic

    freestylefly/canghe-skills

    Knowledge comic creator supporting multiple art styles and tones.

    461 GitHub starsUsed in 8 repos~3.2k tokens
    Media & CreativeAuto-check passed
  • Generate Image

    ynulihao/AgentSkillOS

    Generate or edit images using AI models (FLUX, Gemini). An agent skill from ynulihao/AgentSkillOS.

    618 GitHub starsUsed in 10 repos~1.7k tokens
    Media & CreativeAuto-check: notes
  • GPT Image Generation CLI

    wuyoscar/GPT-Image2-Skill

    Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.

    5.7k GitHub stars~2.5k tokensUpdated 10 days ago
    Media & CreativeAuto-check: notes
  • Minimal Zine Poster Generator

    LiamGvchi/gc-minimal-zine-poster

    Creates or analyzes quiet, paper-texture zine posters with big negative space, one color accent and experimental type, returning an image prompt and the generated poster.

    7.3k GitHub stars~2.9k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

Questions about Editorial Vision Studio

What does Editorial Vision Studio do?

Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters…. Editorial Vision Studio is an agent skill from Yu-0312/editorial-vision-studio. Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters for GPT Image, Flux, Ideogram, and generic image backends.

When should I use Editorial Vision Studio?

Editorial Vision Studio fits situations like: visual concepting; photo-to-design; brand key visuals; product/editorial imagery.

How do I install Editorial Vision Studio in Claude Code?

Run `npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a claude-code`. Or copy the skill folder (the Yu-0312/editorial-vision-studio repository) into .claude/skills/editorial-vision-studio in your project. Claude Code loads it when a task matches its description.

How do I install Editorial Vision Studio in Codex?

Run `npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a codex`. Or copy the skill folder (the Yu-0312/editorial-vision-studio repository) into .agents/skills/editorial-vision-studio in your project. Codex loads it when a task matches its description.

Can I use Editorial Vision Studio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/editorial-vision-studio, .gemini/skills/editorial-vision-studio, .github/skills/editorial-vision-studio and .opencode/skills/editorial-vision-studio in your project.

What does Editorial Vision Studio need to run?

SKILL.md names no scripts, command-line tools or credentials: Editorial Vision Studio is instructions for the agent only.

Does Editorial Vision Studio access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Editorial Vision Studio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Editorial Vision Studio use?

Editorial Vision Studio is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Editorial Vision Studio use?

About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Editorial Vision Studio?

Skills that share tags, products or a category with Editorial Vision Studio: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), Structured Image Generation (bytedance/deer-flow, 84k stars), Canghe Comic (freestylefly/canghe-skills, 461 stars) and Generate Image (ynulihao/AgentSkillOS, 618 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Editorial Vision Studio?

Yu-0312 (a GitHub user) maintains it in Yu-0312/editorial-vision-studio, which has 226 GitHub stars. The repository was last updated on September 4, 2026.

Source: Yu-0312/editorial-vision-studio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.