AI Image Generation and Editing
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters…
$ npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Yu-0312/editorial-vision-studio editorial-vision-studio --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "editorial-vision-studio" agent skill from https://github.com/Yu-0312/editorial-vision-studio/tree/main into .claude/skills/editorial-vision-studio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "editorial-vision-studio", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Yu-0312/editorial-vision-studio editorial-vision-studio --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "editorial-vision-studio" agent skill from https://github.com/Yu-0312/editorial-vision-studio/tree/main into .agents/skills/editorial-vision-studio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "editorial-vision-studio", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Yu-0312/editorial-vision-studio editorial-vision-studio --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "editorial-vision-studio" agent skill from https://github.com/Yu-0312/editorial-vision-studio/tree/main into .cursor/skills/editorial-vision-studio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "editorial-vision-studio", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Yu-0312/editorial-vision-studio editorial-vision-studio --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "editorial-vision-studio" agent skill from https://github.com/Yu-0312/editorial-vision-studio/tree/main into .gemini/skills/editorial-vision-studio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "editorial-vision-studio", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Yu-0312/editorial-vision-studio editorial-vision-studioInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "editorial-vision-studio" agent skill from https://github.com/Yu-0312/editorial-vision-studio/tree/main into .github/skills/editorial-vision-studio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "editorial-vision-studio", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Yu-0312/editorial-vision-studio editorial-vision-studio --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "editorial-vision-studio" agent skill from https://github.com/Yu-0312/editorial-vision-studio/tree/main into .opencode/skills/editorial-vision-studio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "editorial-vision-studio", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
editorial-vision-studioUniversal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters…
Editorial Vision Studio is an agent skill from Yu-0312/editorial-vision-studio. Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters for GPT Image, Flux, Ideogram, and generic image backends. Use for visual concepting, image prompts, photo-to-design, posters, covers, zines, gallery prints, campaigns, brand key visuals, product/editorial imagery, social assets, website hero art, moodboards, presentation decks, Panter-style low-contrast recovery…
Its SKILL.md is about 7.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 97 other files, including assets (for example `README.en.md`, `README.md` and `adapters/_template.md`).
It sits in Media & Creative, covering Image generation. The licence is MIT.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit efd81c2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Editorial Vision Studio loads about 7.5k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 3,079 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Yu-0312/editorial-vision-studio at commit efd81c2, republished under its MIT licence (© Yu-0312). 3,079 words, ~7,544 tokens.
.claude/skills/editorial-vision-studio/SKILL.md (or your agent's skills folder). This skill also uses 93 other files; get the full folder from GitHub.AI Creative Director for Visual Generation.
Philosophy: Do not decorate. Always interpret.
An extensible Editorial Design Engine: one decision pipeline, swappable model adapters. Every layer — intent, analysis, art direction, planning, recovery, compilation, review, iteration — is a separate module with a single job, and the contract between them is a model-agnostic spec.
Architecture: reference/architecture.md
gpt-image, flux, ideogram — or asks to reuse direction with a different modelDECISION ENGINE (fixed) MODEL ADAPTER (swappable)
Intent → Style Gate VisionSpec / EditorialSpec
→ Analyzer → ↓
→ Visual Language → adapters/{model}.md
→ Art Direction → ↓
→ Planner → GenerationRequest → API
→ Recovery →
→ Compiler → VisionSpec
SHARED POST-LAYER
Reviewer → Optimizer → Generate
→ Evaluator → Iteration
│
one spec mutation, │
re-enter Compiler ◄─────────┘
(max 3 passes)User Request
↓
Intent Engine → [prompts/intent.md](prompts/intent.md)
↓ (series? → [prompts/series.md](prompts/series.md) · memory? → [prompts/visual-memory.md](prompts/visual-memory.md))
Style Gate → [prompts/style-gate.md](prompts/style-gate.md) ← ask the user once, up front
↓
Visual Analyzer → [prompts/analyzer.md](prompts/analyzer.md) (skip if theme-only / prompt-only)
↓
Visual Language Engine → [prompts/visual-language.md](prompts/visual-language.md)
↓
Art Direction Engine → [prompts/art-direction.md](prompts/art-direction.md) ← 2–3 candidates, commit to one
↓
Visual Planner → [prompts/planner.md](prompts/planner.md)
↓
Recovery Engine → [prompts/recovery.md](prompts/recovery.md) + [recovery/](recovery/)
↓
VisionSpec → [spec/editorial-spec.schema.md](spec/editorial-spec.schema.md)
↓
Model Adapter → [adapters/registry.md](adapters/registry.md) ← swappable
↓
Prompt Reviewer → [prompts/reviewer.md](prompts/reviewer.md) ← is it legal?
↓
Prompt Optimizer → [prompts/optimizer.md](prompts/optimizer.md) ← is it well said?
↓
Image Generation
↓
Quality Evaluator → [prompts/evaluator.md](prompts/evaluator.md) → quality vector
↓
Iteration Engine → [prompts/iteration.md](prompts/iteration.md) ← fix one layer, loop (max 3)
↓
VisualManifest → [spec/visual-manifest.schema.md](spec/visual-manifest.schema.md)Each layer does one job. Never analyze in Compiler. Never generate in Analyzer. Never re-decide in Iteration — mutate the spec and re-run the layer that owns the failure.
Quick routing: reference/decision-tree.md
Before analyzing pixels, resolve user goal → output family:
| User says | Intent | Allowed outputs |
|---|---|---|
| art book cover | Art Book | magazine-cover, gallery-print, poster |
| TEDx key visual | Event Campaign | campaign-poster, brand-key-visual, social-asset |
| skincare brand launch | Branding | brand-key-visual, product-editorial, social-asset, campaign-poster, website-hero |
| app hero image | Digital Product | website-hero, interface-asset, social-asset |
| zine page | Zine | zine, poster, editorial-spread |
| gallery print | Gallery | gallery-print, photo-abstract-diptych, poster |
| moodboard | Visual Concept | moodboard, poster, editorial-spread |
| 10-page deck | any family, purpose: presentation | presentation-deck |
Also resolve the six intent dimensions — subject, purpose, audience, emotion, platform, aspect ratio. Infer them; ask only when a missing one is load-bearing. They are what every later layer is graded against.
Detect scope here, not later: series_id for a set (prompts/series.md), memory_id for a continuation (prompts/visual-memory.md).
Read prompts/intent.md. Reject mismatched formats (e.g. gallery print for TEDx campaign).
The first question about how the image should look. Before analysing a single pixel, ask which look they want — a numbered menu of the compatible presets plus 「讓 AI 提案」. (Intent may have asked about purpose or platform first; that is a different question.)
Half a step because it computes nothing. It asks one question and turns the answer into a lock.
要什麼風格?
1. 米色明信片 —— 米白紙底,照片重畫成簡化色塊,大量留白。
2. 時代海報 —— 飽和油墨滿版,硬邊平面色塊,地名做成版面。
3. 紙雕明信片 —— 寫實照片,明信片上長出立體紙雕世界。1:1。
4. 自己描述 —— 你說想要的樣子,我拆解照片後照你的描述重畫。
5. 讓 AI 提案 —— 看過照片和用途後給你三個方向再選。Nobody chooses a look twice unless they asked to see proposals. The gate emits a single field, style_gate.outcome, and Art Direction branches on that and nothing else:
| Gate answer | outcome | Art Direction |
|---|---|---|
| A preset | commit | Becomes the session's Visual Memory; auto-commits, asks nothing |
| 自己描述, or just typing a description | commit | Captured verbatim; the Analyzer deconstructs the photo, then prompts/style-brief.md rebuilds it to the description. A description is never answered with a menu |
| 「讓 AI 提案」 | offer | Offers 2–3 candidates. This is the one path with two prompts, and the user asked for it |
It skips — outcome: commit, no menu — when a preset or style is already named, a memory or series is active, a direction was committed earlier this session, or the run is unattended. It skips to outcome: offer when fewer than two presets fit the intent, because a one-item menu is not a choice. Unattended runs pick the highest-fit preset with intent-only weights and say which in one line — never silently.
Presets declare intended_layouts; the gate offers only those that intersect intent.allowed_outputs.
Show looks, not internal vocabulary. Nobody outside this repo knows what Kinfolk means.
The gate settles ground and medium only. Layout, composition, abstraction, and copy stay with the Planner.
Read prompts/style-gate.md.
Produce structured Image Report with star ratings and Editorial Score (0–100).
Dimensions: subject, clarity, contrast, saturation, composition, negative space, geometry, texture, lighting, emotion.
Read prompts/analyzer.md.
Derive Visual Language first, then style/palette/layout — not the reverse.
Examples: Museum → Swiss or MUJI + fine serif; Quiet Human → Kinfolk + cream/sage; Indie Memory → Zine + riso anchor; Poster Graphic → Travel Poster + flat inks.
Read prompts/visual-language.md. User style: override skips auto-derivation but Reviewer still validates DNA fit.
Do not generate the first plausible reading of the brief. Draft competing directions, each with a one-sentence thesis and a named trade-off, score them for fit, commit to one, and keep the runner-up.
Branch on style_gate.outcome and nothing else. The gate already resolved every case — preset chosen, free text, 「你決定」, explicit style:, memory, series, unattended, too few presets — into one of two values:
outcome | Behaviour |
|---|---|
commit | Build one direction, commit silently. Never offer candidates — that would be the user's second time choosing a look. Name the runner-up in one line when one exists |
offer | Present 2–3 candidates and ask. Two for a narrow intent family, three for a theme-only brief or Editorial Score <50 |
The layer always runs, even on commit — every spec needs an art_direction block with a thesis.
Every candidate must set a different ground. Eight of the eleven style DNAs resolve to a light paper field, so candidates chosen on style alone come back as three shades of ivory. Beyond that, candidates must differ on at least two of: render mode, visual language, layout family, abstraction level, typography weight. Palette swaps are not directions.
The committed direction is a hard constraint on the Planner. Switching to the runner-up later re-runs Planner onward only — never the Analyzer.
Read prompts/art-direction.md.
Decide layout, typography direction, abstraction level inside the committed direction — not the final prompt.
Key rules (full matrix in prompts/planner.md):
The Planner also fills direction.spatial_plan — projection, shared ground plane, light direction, arrangement. On a photo run it is a verbatim copy of image_report.spatial; on a theme-only run the Planner authors it from intent.subject. Spatial fields locked by an active preset or memory are applied over the result field by field — a locked projection: flat-elevation is an authored decision that overrides what the photograph shows. It is required either way: prompts/reviewer.md rejects a prompt with no named projection whether or not a photograph existed, and without the field those clauses have no provenance for prompts/optimizer.md to protect.
If user specifies style: kinfolk, load styles/kinfolk.md DNA.
Apply only when Image Report flags weakness. Each recovery is one atomic fix — see recovery/.
| Problem | Recovery |
|---|---|
| Low contrast / gray (saturation <30%) | Panter Mode: warm/cool conflict hues, high-sat anchor, wider tonal separation |
| Weak subject | Increase silhouette / scale |
| Flat lighting | Directional light |
| Busy background | Simplify geometry |
| Too many colors | Limit palette to 4 |
| No focal point | Editorial color anchor |
| No rhythm | Abstract panel |
Panter Mode: discard dull grays; boost warm to 75% / cool to 70% saturation; add 8% high-chroma anchor block; widen tonal separation and mark scale. Panter is a colour compensation and never adds texture on its own. See recovery/contrast.md.
Texture Permission — single source of truth: assets/texture.md. Three tiers: PRINT (riso/halftone/scan defects) is zine only; SURFACE (substrate character such as cotton paper) is allowed on CLEAN layouts whose style DNA rates Texture ★★★+; FLAT (zero texture words) covers the photo-abstract-diptych panel ground, interface-asset, the website-hero copy-safe area, the product-editorial background, and every page of a presentation-deck that contains a data page. Recoveries never raise a layout's tier.
Never redesign the entire image unless Editorial Score <50 (Concept Reconstruction).
Phase 1: Assemble VisionSpec / EditorialSpec — read prompts/compiler.md
Phase 2: Route to adapter by target.model:
| Model | When | Adapter |
|---|---|---|
gpt-image (default for photo upload) | Diptych, photo fidelity | adapters/gpt-image.md |
flux | Zine texture, atmosphere | adapters/flux.md |
ideogram | Cover/campaign typography | adapters/ideogram.md |
generic | Unknown backend | adapters/generic.md |
User: model: flux or "用 Flux 生成" → set adapter, do not re-analyze.
Same direction, different model: reuse VisionSpec / EditorialSpec, swap adapter only.
Before generation, run conflict detection. Read prompts/reviewer.md.
Examples:
Auto-correct incompatible pairings.
The Reviewer answers is this prompt legal. It cannot fix a prompt that is legal but weakly worded — the ground clause buried in sentence four, a spec number softened to "a small accent", the same instruction stated twice.
Half a step because it decides nothing. Every clause it emits must trace to a field the spec already holds; a clause with no spec field behind it is a rejection back to the Compiler, not an invention. It never sets a ground, style, subject, layout, palette, or type scale.
Seven ordered ops: provenance_strip · ground_first · concretize · bind_numbers · dedupe · route_negatives · compress. Most gain comes from concretize and bind_numbers — a model renders nouns and numbers, never an adjective about taste.
This does not break the "mutate the spec, never the prompt string" rule in prompts/iteration.md. That rule forbids hand patches after generation. The Optimizer is deterministic, pre-generation, and recorded: same spec + same adapter + same ruleset_version produces the same string, so the manifest still replays.
No op firing is the normal outcome. The same op firing twice means the adapter emits that defect every time — fix adapters/{model}.md, not the string.
Variant mode emits 2–3 prompts from one spec, differing only on clause_density, specificity, and emphasis_order. A variant that differs on any spec field is a second direction, not a variant. Offer it only when the user asks to compare wordings, when the model is generic, or when a prompt-attributed dimension failed last run.
Read prompts/optimizer.md.
After generation, score a quality vector — not one number. Ten weighted dimensions: subject, composition, focal_point, palette, typography, texture, style_coherence, photo_fidelity, intent_fit, platform_fit. Each 0.00–1.00; inapplicable dimensions are null, never 0.
overall is the weighted mean and maps to grade A–D. Any single dimension below 0.60 fails the run regardless of overall.
The evaluator must name lowest_failing and responsible_layer — without them the iteration loop has nothing to act on.
Read prompts/evaluator.md.
A low score points at one layer, not at the whole image. Fix that layer, recompile, regenerate.
lowest failing dimension → responsible layer → one minimal spec mutation → recompile → re-scoreoverall ≥ 0.85 with no dimension below 0.60, at 3 iterations, on score regression, or when escalation reaches Intent — then report, don't guess.Taste disagreement ("I don't like the blue") is a direction change, not a QC failure — route it to prompts/art-direction.md.
Read prompts/iteration.md.
Two axes, required, with no default:
| Axis | Values | Question it answers |
|---|---|---|
design_tokens.ground | paper-light · neutral-gray · dark · saturated · full-bleed-photo · duotone | What is the canvas field itself? |
direction.render_mode | photographic · photo-plus-graphic · graphic · painterly · mixed | What medium is the image made of? |
Enum values are contract tokens, never prompt words — every adapter resolves them to prose through assets/ground.md.
An unset value is a rejection, not a fallback. This matters more than it looks: before these fields existed, an undecided ground fell through to warm ivory paper on every run, because that was the most-repeated value in the repo. A field with no default cannot be skipped.
render_mode is not abstraction_level. Medium and subject literalness are independent — a photographic image can still be full-abstract.
A preset is a shipped VisualMemory with source: preset — a locked partial spec authored in the repo instead of established from a run. It reuses the entire memory mechanism; nothing new enforces it.
| Preset | Ground | Render mode |
|---|---|---|
| ivory-postcard | paper-light | painterly |
| vintage-travel-poster | saturated | graphic |
| papercraft-diorama-postcard | full-bleed-photo | photographic |
The Style Gate offers these as a numbered menu at the start of a run; preset: ivory-postcard is the shortcut that skips the menu, for automation and series work. Three different grounds is deliberate — it is what gives the candidate-diversity rule something to draw on. Registry and authoring guide: presets/registry.md.
Unlike a series memory, a preset may lock composition, aspect_ratio, layout, and — field by field, under the spec's own key names — spatial_plan: a preset is avowedly a template, which is what it is for. Locked spatial fields override the photographed values; unlocked fields still come from the photo or the Planner.
The engine keeps a visual system across images, not just within one.
| Ask | Read |
|---|---|
| "now do the next one," "同一套視覺," brand assets supplied | prompts/visual-memory.md |
| "一套 10 頁簡報," "campaign 全尺寸," carousel, all platform sizes | prompts/series.md |
Visual Memory locks the identity fields — style, visual language, palette, typography, texture tier — and leaves layout, composition, abstraction level, and aspect ratio free per image. Locking composition produces a template, not a system. On a memory run, Art Direction still runs but auto-commits to the locked DNA rather than offering candidates. A lock that no longer fits gets forked, never silently mutated. Contract: spec/visual-memory.schema.md.
Series runs Intent → Analyzer → Visual Language → Art Direction once, then fans out per output. Generate the hero first and pass QC before derivatives — a weak hero multiplies into N weak frames. Cross-image QC scores palette drift, typographic identity, texture tier, anchor legibility, and compositional variety; a single outlier is fixed alone, never by re-running the set.
Every run emits spec/visual-manifest.schema.md: direction taken and runner-up, visual system, quality vector, iteration history, prompt hash, provenance.
This is what makes the cheap paths cheap:
| User says | Re-run |
|---|---|
| "同一張,改用 Flux" | Adapter → Reviewer → Generate |
| "改用 B 那個方向" | Planner → downstream |
| "一模一樣再生一次" | Generate only |
| "同一套視覺,換主題" | Analyzer → Art Direction (auto-commit) → Planner → downstream |
Never re-run the Analyzer when a valid image_report for the same source image already exists. Keep the manifest internal unless the user asks for it or the run is part of a series.
This is the Editorial Score — an input measure of the source photo, produced by the Analyzer. It is not the Quality Score from Step 8, which measures the generated image on different bands.
| Editorial Score | Mode |
|---|---|
| 90+ | Premium Editorial — refined extraction, minimal recovery |
| (no photo) | Theme-only: no score exists. editorial_mode: standard, recoveries: [] — never invent a score |
| 70–89 | Standard Editorial |
| 50–69 | Compensation Mode — apply Recovery stack |
| <50 | Concept Reconstruction — abstract reinterpretation |
Match the requested depth. Default to a concise direction summary plus GenerationRequest.
User: "同一份方向,改用 Ideogram" → reuse VisionSpec / EditorialSpec, run adapters/ideogram.md only.
User: "改用 B 那個方向" → reuse Intent + Image Report, set the runner-up as selected, re-run Planner onward.
User: "同一套視覺,做東京街景" → load the Visual Memory, run Analyzer on the new photo, let Art Direction auto-commit to the locked DNA, and Planner sets free fields only.
Never:
photo_policy.fidelity: requiredAlways:
direction.spatial_plan; give every grounded object a contact shadow| Asset module | File |
|---|---|
| Ground & Render Mode | assets/ground.md |
| Scene construction (space) | assets/scene-construction.md |
| Texture permission | assets/texture.md |
| Palette | assets/palette.md |
| Typography | assets/typography.md |
| Style | File |
|---|---|
| Swiss | styles/swiss.md |
| Kinfolk | styles/kinfolk.md |
| MUJI | styles/muji.md |
| Brutalist | styles/brutalist.md |
| Wallpaper* | styles/wallpaper.md |
| Apartamento | styles/apartamento.md |
| Purple Magazine | styles/purple.md |
| POPEYE | styles/popeye.md |
| Monocle | styles/monocle.md |
| COS | styles/cos.md |
| Period Travel Poster | styles/travel-poster.md |
| Layout | File |
|---|---|
| Editorial Poster | layouts/poster.md |
| Magazine Cover | layouts/magazine-cover.md |
| Gallery Print | layouts/gallery-print.md |
| Zine | layouts/zine.md |
| Editorial Spread | layouts/editorial-spread.md |
| Campaign Poster | layouts/campaign-poster.md |
| Brand Key Visual | layouts/brand-key-visual.md |
| Product Editorial | layouts/product-editorial.md |
| Website Hero | layouts/website-hero.md |
| Social Asset | layouts/social-asset.md |
| Moodboard | layouts/moodboard.md |
| Interface Asset | layouts/interface-asset.md |
| Photo-Abstract Diptych | layouts/photo-abstract-diptych.md |
| Presentation Deck | layouts/presentation-deck.md |
| Extend | Action | Touch Decision Engine? |
|---|---|---|
| New style (Aesop, NYT Mag) | Add styles/foo.md | No |
| New preset | Add presets/foo.md + register | No |
| New layout | Add layouts/foo.md | No |
| New recovery | Add recovery/foo.md | No |
| New image model | Add adapters/foo.md + register | No |
| New intent family | Edit prompts/intent.md | Yes (minimal) |
| New QC dimension | Edit prompts/evaluator.md + map a layer in prompts/iteration.md | Yes (minimal) |
| New prompt-wording op | Add a row to prompts/optimizer.md + bump ruleset_version | No |
| New lockable DNA field | Edit spec/visual-memory.schema.md + enforce in prompts/reviewer.md | Yes (minimal) |
See adapters/_template.md for new models.
Add new magazines/brands by creating styles/your-style.md with Style DNA table + compiler clauses. No need to rewrite SKILL.md.
| Layer | File |
|---|---|
| Intent Engine | prompts/intent.md |
| Style Gate | prompts/style-gate.md |
| Style Brief | prompts/style-brief.md |
| Visual Analyzer | prompts/analyzer.md |
| Visual Language Engine | prompts/visual-language.md |
| Art Direction Engine | prompts/art-direction.md |
| Visual Planner | prompts/planner.md |
| Recovery Engine | prompts/recovery.md |
| Prompt Compiler | prompts/compiler.md |
| Prompt Reviewer | prompts/reviewer.md |
| Prompt Optimizer | prompts/optimizer.md |
| Quality Evaluator | prompts/evaluator.md |
| Iteration Engine | prompts/iteration.md |
| Visual Memory | prompts/visual-memory.md |
| Series Planner | prompts/series.md |
| Preset Registry | presets/registry.md |
| Contract | File |
|---|---|
| VisionSpec / EditorialSpec | spec/editorial-spec.schema.md |
| VisualManifest | spec/visual-manifest.schema.md |
| VisualMemory | spec/visual-memory.schema.md |
Model parameters: agents/openai.yaml
© Yu-0312, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 93 other files (assets) in the repository root of Yu-0312/editorial-vision-studio.
Open the folder on GitHubat commit efd81c2
Editorial Vision Studio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Editorial Vision Studio this skillYu-0312/editorial-vision-studio | 226 | — | ~7.5k | Automated safety check: Pass | MIT | |
| AI Image Generation and Editingzhayujie/CowAgent | 47k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Structured Image Generationbytedance/deer-flow | 84k | 4 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Canghe Comicfreestylefly/canghe-skills | 461 | 8 repos | ~3.2k | Automated safety check: Pass | None | |
| Generate Imageynulihao/AgentSkillOS | 618 | 10 repos | ~1.7k | Automated safety check: Notes | None | |
| GPT Image Generation CLIwuyoscar/GPT-Image2-Skill | 5.7k | — | ~2.5k | Automated safety check: Notes | MIT |
zhayujie/CowAgent
Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.
bytedance/deer-flow
Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.
freestylefly/canghe-skills
Knowledge comic creator supporting multiple art styles and tones.
ynulihao/AgentSkillOS
Generate or edit images using AI models (FLUX, Gemini). An agent skill from ynulihao/AgentSkillOS.
wuyoscar/GPT-Image2-Skill
Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.
LiamGvchi/gc-minimal-zine-poster
Creates or analyzes quiet, paper-texture zine posters with big negative space, one color accent and experimental type, returning an image prompt and the generated poster.
Categories
Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters…. Editorial Vision Studio is an agent skill from Yu-0312/editorial-vision-studio. Universal visual direction engine for AI image, design, and layout work: model-agnostic decision pipeline (intent, analysis, visual language, planning, recovery/refinement) plus swappable adapters for GPT Image, Flux, Ideogram, and generic image backends.
Editorial Vision Studio fits situations like: visual concepting; photo-to-design; brand key visuals; product/editorial imagery.
Run `npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a claude-code`. Or copy the skill folder (the Yu-0312/editorial-vision-studio repository) into .claude/skills/editorial-vision-studio in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a codex`. Or copy the skill folder (the Yu-0312/editorial-vision-studio repository) into .agents/skills/editorial-vision-studio in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Yu-0312/editorial-vision-studio --skill editorial-vision-studio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/editorial-vision-studio, .gemini/skills/editorial-vision-studio, .github/skills/editorial-vision-studio and .opencode/skills/editorial-vision-studio in your project.
SKILL.md names no scripts, command-line tools or credentials: Editorial Vision Studio is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Editorial Vision Studio is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Editorial Vision Studio: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), Structured Image Generation (bytedance/deer-flow, 84k stars), Canghe Comic (freestylefly/canghe-skills, 461 stars) and Generate Image (ynulihao/AgentSkillOS, 618 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Yu-0312 (a GitHub user) maintains it in Yu-0312/editorial-vision-studio, which has 226 GitHub stars. The repository was last updated on September 4, 2026.
Source: Yu-0312/editorial-vision-studio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.