Agent skill

Generated Media QA

by calesthio in calesthio/generative-media-skills

Provider-independent quality assurance for AI-generated and AI-assisted media.

MITAuto-check passedFrontend & Design

Install Generated Media QA

skills CLI
$ npx skills add calesthio/generative-media-skills --skill generated-media-qa -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/generative-media-skills generated-media-qa --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/production/governance-delivery/generated-media-qa .claude/skills/generated-media-qa && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
generated-media-qa
GitHub stars
193
Token cost
~8k tokens
SKILL.md length
3,641 words
Files
4 (incl. scripts)
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

Provider-independent quality assurance for AI-generated and AI-assisted media.

  • Works in 6 steps: Verify consent and allowed languages for… → For each locale, compare translated… → Watch the dub once without captions for… → …
  • Reporting on images
  • SKILL.md covers Keep three evidence lanes…, Intake before review, First pass: acceptance matrix and Language-alignment QA lane, plus 14 more sections
  • Runs Python scripts from its folder; calls python

What it does

Generated Media QA is an agent skill from calesthio/generative-media-skills. Provider-independent quality assurance for AI-generated and AI-assisted media. Use when reviewing, accepting, revising, or reporting on images, video, audio, avatars, ads, product content, social clips, explainers, localization, mixed-source edits, captions, accessibility, provenance, model metadata, safety/policy, and delivery readiness.

Its SKILL.md is about 8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `EVAL.md`, `scripts/normalize_qa_report.py` and `tests/test_normalize_qa_report.py`).

It sits in Frontend & Design, covering Internationalization, QA and bug reports and Accessibility. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.

When your agent uses it

  • Reporting on images
  • Product content
  • Mixed-source edits
  • Delivery readiness

Example prompts

  • “/generated-media-qa”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Verify consent and allowed languages for the avatar/voice before watching outputs.
  2. For each locale, compare translated script to approved source for meaning, warnings, numbers, product names, and URLs.
  3. Watch the dub once without captions for naturalness and intelligibility; watch again with captions for timing and line breaks.
  4. Spot-check lip sync at dense consonant passages and at the final 20 seconds, where drift often accumulates.
  5. Confirm captions include speaker changes and meaningful non-speech audio if relevant.
  6. Record locale-specific issues separately; do not let a pass in one language imply pass in another.

What it can do on your machine

Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • partnerhelp.netflixstudios.com
    • w3.org
    • support.google.com
    • ftc.gov
    • tech.ebu.ch
    • itu.int
    • atsc.org
    • fcc.gov
    • spec.c2pa.org
    • cv.iptc.org
    • iptc.org
    • smpte.org
    • nist.gov
    • ai-challenges.nist.gov
    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Generated Media QA loads about 8k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 3,641 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 3,641 words, ~7,962 tokens.

Download SKILL.mdSave it as .claude/skills/generated-media-qa/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
generated-media-qa
description
Provider-independent quality assurance for AI-generated and AI-assisted media. Use when reviewing, accepting, revising, or reporting on images, video, audio, avatars, ads, product content, social clips, explainers, localization, mixed-source edits, captions, accessibility, provenance, model metadata, safety/policy, and delivery readiness.

Generated Media QA

Treat QA as a release decision, not a vibe check. Judge the deliverable against the approved brief, platform specifications, legal/safety constraints, and the audience context. Record enough evidence that another agent or producer can reproduce the decision.

Keep three evidence lanes separate

Documented facts are requirements from the brief, platform specs, legal/policy guidance, delivery standards, accessibility standards, or provider documentation. Cite them or name the source and verification date.

Empirical observations are what you directly measured or inspected in the asset: frame size, duration, loudness, sync offset, OCR output, transcript mismatch, visual artifact, metadata, or a timestamped defect.

Production heuristics are professional judgments used when no explicit spec exists: whether a hand artifact is audience-visible, whether a product packshot feels trustworthy, whether an accent is intelligible for the target market, whether a social caption is too fast for mobile. Label these as heuristics and avoid pretending they are universal standards.

Intake before review

Do not start with random artifact hunting. Build the acceptance frame first.

Collect:

  • Approved brief, prompt, storyboard, script, shot list, edit decision list, brand rules, product facts, target audience, target platform, duration, aspect ratio, language/locale, and known compromises.
  • Delivery spec: file format, codec, resolution, frame rate, color space/HDR, audio channels, loudness target, captions format, thumbnail and metadata requirements.
  • Source inventory: generated assets, human-shot assets, licensed stock, user-provided media, logos, fonts, music, SFX, voices, product claims, model releases, likeness/voice consent, and usage rights.
  • Generation metadata when available: provider, model, model version/date, prompt, negative prompt, reference images, seed, dimensions, duration, fps, voice ID, language, post tools, upscalers, edit tools, and safety filters.
  • Risk context: ads, health/finance/legal claims, political content, public figures, minors, regulated products, synthetic endorsements, realistic news-like scenes, localization, accessibility obligations, and platform disclosure requirements.

If the brief is missing, write a provisional QA basis such as: "Reviewing against supplied asset, target platform: Instagram Reels, inferred goal: 15 s product teaser. Product-claim and legal review cannot pass until facts and rights are supplied."

First pass: acceptance matrix

Create a small matrix before deep inspection:

AreaPass evidenceTypical reject evidence
Brief conformanceMessage, product, audience, tone, duration, format, and call to action match the approved briefWrong product, missing CTA, wrong locale, off-brand tone, materially changed promise
Technical deliveryFile opens, spec matches, no corruption, correct duration/fps/aspect, audio/captions present as requiredWrong aspect, corrupt frames, missing audio, bad encode, unsupported caption file
Generated-media realismNo audience-visible anatomy, physics, identity, text, logo, continuity, or temporal defectsWarped hands/faces, unreadable text, drifting identity, morphing product, flicker, impossible motion
Audio and speechIntelligible, synced, no clipping/noise/pops, correct language/voice, loudness meets targetDialogue buried, lip sync off, clipping, wrong voice, mistranslation
AccessibilityCaptions, transcript, alt text, audio description, safe flashing, and player affordances meet the delivery contextMissing captions where required, unreadable captions, unsafe flashes, visual-only information with no alternative
Rights and provenanceSource rights, consent, metadata, and AI disclosure are documentedUnlicensed source, unapproved likeness/voice, missing synthetic disclosure, unverifiable origin
Safety/policyNo disallowed deception, unsafe advice, discriminatory content, or platform-prohibited claimsMisleading synthetic person/event, unsupported health claim, unsafe instructions, regulated-product violation

Then inspect. Do not let a single attractive frame override a failed acceptance criterion.

Language-alignment QA lane

When the delivery includes generated shot descriptions, search metadata, prompt reconstructions, source captions, or dataset annotations, review language separately from pixel/audio quality.

  • Accuracy: every described entity, action, spatial relation, camera behavior, overlay, and state change matches the source; no hallucination or unsupported intent.
  • Completeness: coverage matches the declared use. Archive search, VFX handoff, prompt comparison, and training data require different levels of detail.
  • Terminology: approved project vocabulary and reference frames are used consistently; distinguish subject-left from frame-left and dolly/translation from zoom/rotation.
  • Temporal order: events, cuts, focus changes, entries/exits, and start/end states appear in source order.
  • Provenance: description author/model/version, prompt, source asset/hash, specification/glossary version, reviewer, and acceptance status are recorded.
  • Risk: false product, safety, legal, identity, medical, financial, or public-interest descriptions are Critical when they affect release or downstream decisions.

Accessibility captions, subtitles, and audio description have separate audience and standards requirements. Do not replace them with production descriptions. Use precise-video-description to define or repair the observation contract; use video-description-oversight for the complete pre-caption/critique/post-caption review workflow. This lane only ensures the final delivery's language assets align with the reviewed source and declared use.

Technical file QA

Use automated checks for measurable properties and human review for perceptual defects. Automated file QC catches many delivery failures, but it cannot decide whether a generated product label is semantically wrong or whether a synthetic spokesperson feels deceptive.

Check:

  • File integrity: opens in at least two players/viewers; no truncated tail, decode errors, missing streams, alpha surprises, or silent channels.
  • Container and streams: expected format, codec, profile, resolution, pixel aspect ratio, frame rate, duration, color primaries/transfer/matrix, bit depth, bitrate, audio sample rate, channel layout, and captions/subtitle streams.
  • Visual defects: black frames, freeze frames, dropped/duplicated frames, banding, macroblocking, moire, aliasing, unintended borders, watermark remnants, bad mattes, poor keying, crop errors, unsafe title area, unreadable small text, and compression damage.
  • Audio defects: clipping, inter-sample peaks, noise floor, hum, clicks, pops, gating artifacts, reverb wash, harsh sibilance, phase cancellation, channel imbalance, missing stems, and abrupt edits.
  • Versioning: filename, slate, burned-in timecode, metadata, export preset, and revision number match the delivery tracker.

For broadcast or premium delivery, use the client/platform spec as the authority. SMPTE describes IMF as a file-based media format for storing and delivering audiovisual masters across versions and territories. DPP/AS-11-style workflows use explicit technical and editorial metadata. These are not default requirements for a TikTok draft, but they are useful models for disciplined delivery tracking.

Visual and motion QA for generated images/video

Generated media needs targeted inspection beyond normal color and compression checks.

Inspect still frames at 100% and at the intended viewing size. Scrub video slowly, then watch once without pausing on the target device class. Many AI defects are only obvious during motion, and many "defects" disappear at normal mobile viewing size.

High-priority checks:

  • Anatomy and faces: hands, fingers, teeth, eyes, ears, joints, limb count, asymmetric pupils, skin texture, hairline shimmer, identity drift, age mismatch, and face-morphing between frames.
  • Text and symbols: product labels, UI text, subtitles burned into imagery, signs, legal copy, prices, dates, numbers, QR codes, logos, trademark shapes, and brand color placement. Use OCR when possible, but manually verify brand-critical text.
  • Product truth: package geometry, materials, scale, ports/buttons, SKU variant, colorway, dosage/amount, included accessories, ingredients, claims, safety warnings, and compatibility details.
  • Temporal continuity: subject identity, wardrobe, prop positions, lighting direction, shadows, reflections, weather, time of day, screen contents, labels, and geography across shots.
  • Motion plausibility: foot sliding, floating objects, rubber limbs, non-causal physics, warped vehicles, rolling-shutter hallucinations, facial/lip deformation, camera path discontinuity, interpolation smear, and background crawling.
  • Compositing and mixed-source edits: matte edges, grain/noise mismatch, perspective, lens blur, color temperature, contact shadows, reflection consistency, scale, eye lines, room tone, and source footage/generative insert seams.
  • Avatar and lip-sync: mouth closure on bilabials, jaw timing, head/neck motion, eye blinks, gaze, breathing, hand gestures, language/accent match, uncanny stillness, and whether the avatar is disclosed/approved.

Production heuristic: accept small artifacts only when they are not visible at intended size, do not affect the subject/product/message, and will not become a meme or trust-breaker. Reject or revise any defect on a face, hand, product, logo, legal copy, price, medical/financial statement, or safety instruction.

Audio, loudness, and intelligibility

Use the delivery spec first. If none exists, choose a target appropriate to the platform and document it as a production heuristic.

Documented facts:

  • ITU-R BS.1770 specifies algorithms for measuring programme loudness and true-peak audio level.
  • EBU R 128 recommends an average programme loudness of -23 LUFS and the use of Loudness Range and Maximum True Peak descriptors for audio signals.
  • ATSC A/85 provides methods for measuring and controlling loudness for digital television and is referenced by the FCC CALM Act context for TV commercials in the United States.
  • Netflix partner guidance for its own deliveries tells reviewers to use meters implementing ITU-R BS.1770 variants and describes how/when to flag loudness issues. Treat Netflix requirements as Netflix-specific, not universal.

Review:

  • Dialogue intelligibility: listen on headphones, laptop speakers, and a phone speaker when the target is social/mobile. Flag when music/SFX mask key words.
  • Sync: check lip sync at start, middle, and end; check that translated dubs preserve turn-taking and emotional timing. For avatars, inspect consonant closure, pauses, breaths, and facial expression alignment.
  • Loudness: measure integrated loudness, true peak, and loudness range. Record the tool and target. Do not "normalize by ear" for a final delivery.
  • Mix balance: narration, dialogue, SFX, music, ambience, and captions should support the message. Lower music under legal disclaimers, product claims, and instructions.
  • Localization: verify numbers, units, names, pronunciations, cultural references, idioms, and reading pace in the target language. Back-translation can catch gross errors but is not a substitute for native review on high-stakes content.

Captions, accessibility, and viewer safety

Accessibility is part of QA, not a last-minute export option.

Documented facts:

  • WCAG 2.2 requires captions for prerecorded audio content in synchronized media at Success Criterion 1.2.2 and audio description or a media alternative at 1.2.3; Level AA includes audio description at 1.2.5 for prerecorded synchronized media.
  • WCAG 2.2 Success Criterion 2.3.1 says content must not flash more than three times in any one-second period unless below the general and red flash thresholds.
  • Netflix timed-text guidance is a platform-specific professional reference for subtitle timing, duration, forced narratives, and style. Use the actual client/platform guide when delivering elsewhere.

Check:

  • Captions are present when required; every spoken word, meaningful sound cue, speaker ID, and off-screen speech is represented as appropriate.
  • Captions are timed to speech, not late; do not cover essential UI/product/legal information; remain readable over the background; stay within safe margins; and are not burned in if the platform/client requires sidecar captions.
  • Subtitle translation preserves meaning, tone, names, numbers, legal claims, humor, and safety instructions. Do not accept a localization solely because it "sounds fluent."
  • Audio description or text alternative exists when important visual-only information is necessary for comprehension.
  • Alt text or extended descriptions exist for still images when the delivery context requires accessible images.
  • Flashing/strobing sequences are tested or avoided. For public-facing video, treat fast flashes, saturated red flashes, and high-contrast repetitive patterns as a safety risk, not just a style choice.

Rights, provenance, and disclosure

Generated media QA must include origin and usage review. Do not present metadata as proof of truth; use it as one trust signal.

Documented facts:

  • C2PA defines technical specifications for content provenance; a C2PA manifest can contain assertions, claims, signatures, and content bindings for media provenance.
  • IPTC Digital Source Type includes trainedAlgorithmicMedia for media created using generative AI and compositeWithTrainedAlgorithmicMedia for composites containing generative-AI elements.
  • Google Merchant Center documentation says not to remove embedded metadata tags such as IPTC DigitalSourceType from AI-generated images; relevant NewsCodes include TrainedAlgorithmicMedia. Verified 2026-07-10; platform requirements are volatile.
  • YouTube Help says creators must disclose AI-generated or meaningfully AI-altered content that appears realistic through the AI-use setting, and disclosed content can be labeled for viewers. Verified 2026-07-10; platform requirements are volatile.
  • FTC advertising guidance does not create an "AI exemption" from truthful advertising rules; endorsements, testimonials, and product claims still need appropriate disclosure and substantiation.

Review:

  • Source rights: stock license, music license, font license, logo permission, brand assets, user-provided media permission, and commercial usage terms.
  • Likeness/voice: consent for synthetic or cloned voices, avatars, face swaps, public figures, employees, customers, and minors.
  • Claims: product performance, before/after imagery, health/financial/legal benefits, prices, availability, endorsements, awards, comparative claims, and environmental claims.
  • Provenance metadata: preserve C2PA/IPTC/XMP where required; record if metadata was stripped by editing or platform export; do not claim "verified real" just because a credential exists.
  • Disclosures: add human-visible AI/synthetic disclosure when the platform, law, client policy, or audience deception risk requires it. For realistic people/events, political content, news-like scenes, ads, and endorsements, escalate rather than guessing.
  • Model/provider metadata: capture provider, model, date, prompt lineage, reference sources, post tools, and revision history. If exact model version is unavailable, state that limitation.

Safety and policy review

Apply the relevant model, platform, legal, and client policies. Platform facts change; verify at delivery time for regulated or high-risk content.

Flag or reject:

  • Deceptive realistic synthetic people, events, endorsements, evidence, or news scenes.
  • Non-consensual intimate imagery, sexualized minors, harassment, hate, extremist persuasion, self-harm encouragement, or instructions for wrongdoing.
  • Medical, financial, legal, political, housing, employment, or credit claims without substantiation and required disclaimers.
  • Misrepresentation of a product's capabilities, certification, price, availability, environmental impact, or safety.
  • Privacy leaks: faces, license plates, addresses, screens, documents, account details, patient/customer data, location traces, or hidden metadata.
  • Localization harms: culturally offensive imagery, mistranslated warnings, wrong units, taboo gestures, or regionally illegal claims.

Revision triage

Use severity to decide whether to accept, revise, regenerate, or escalate.

Critical: must not ship. Examples: wrong product/claim, missing rights or consent, deceptive synthetic endorsement, disallowed safety/policy content, inaccessible required captions, unsafe flashing, corrupted final file, unintelligible core dialogue, wrong legal copy, or platform-required disclosure missing.

Major: revise before normal release unless stakeholder explicitly accepts risk. Examples: visible face/hand/logo defect, lip-sync drift, continuity break that harms comprehension, caption timing failures, audio clipping, noisy mix, bad localization, wrong brand color on key asset, or noticeable compositing seam.

Minor: fix if efficient; can ship with documented acceptance for low-stakes drafts. Examples: small background artifact, slight caption line break awkwardness, non-critical metadata typo, tiny compression artifact not visible at target size.

Observation: record without blocking. Examples: "AI-generated background texture visible on pause only", "music slightly more energetic than brief but message remains clear."

Do not use "minor" for defects on faces, bodies, product labels, logos, safety warnings, prices, legal copy, public figures, or accessibility requirements.

Show full SKILL.md (1,376 more words)Show less

QA report format

Write reports that make action obvious:

  • Asset ID, version, reviewer, date, target platform, intended use, and source package reviewed.
  • Verdict: Pass, Pass with notes, Revise, Reject, or Escalate.
  • Acceptance basis: brief/spec/policy documents used, with dates for volatile platform/provider facts.
  • Automated checks: tools used and measured results.
  • Manual checks: devices/viewers used and review coverage.
  • Findings table: severity, timestamp/frame/region, issue, evidence, likely cause, recommended fix, owner, and retest requirement.
  • Rights/provenance: source inventory, consent status, metadata status, and disclosure status.
  • Residual risks: what could not be verified and who must approve.

For revision loops, require a delta review plus spot checks around changed areas and one full playback/viewing pass. Fixing one generated defect often creates another nearby.

Deterministic report normalization helper

Use scripts/normalize_qa_report.py when a QA report needs a stable JSON handoff for automation, ticket creation, release gates, or producer review. The helper is Python 3.11+ and uses only the standard library. It validates the report shape, normalizes and sorts findings, summarizes severity counts, and computes a mechanical disposition from an explicit policy supplied in the report or via CLI.

Run it from a copied skill package or this repository:

bash
python scripts/normalize_qa_report.py report.json
python scripts/normalize_qa_report.py report.json --policy '{"dispositions":{"critical":"reject","major":"hold","minor":"ready","observation":"ready"},"waivable_severities":["major"],"non_waivable_areas":["rights","accessibility","safety","policy"]}'

Exit codes:

  • 0: valid report and mechanical disposition is ready.
  • 2: validly parsed report that results in hold or reject, or a report with schema validation findings.
  • 3: operational parse/read/write failure, such as invalid JSON or an unreadable file.

The input JSON must include:

  • asset_id: stable asset or delivery ID.
  • review: object with reviewer and ISO-compatible date; include target platform, intended use, and source package when available.
  • policy: explicit disposition rules. Provide dispositions mapping each severity (critical, major, minor, observation) to ready, hold, or reject. Optional waivable_severities lists severities that a supplied waiver may mechanically unblock. Optional non_waivable_areas is extended by the script to always include rights, accessibility, safety, and policy.
  • checks: array of automated or manual checks with id, status (pass, fail, not_applicable, or not_checked), required evidence_type, and evidence detail.
  • findings: array with severity, area, issue, evidence_type, evidence_detail, and a locator (timecode, frame, or region) whenever the evidence is an empirical observation or the issue is visual, audio, captions, accessibility, or safety related.

Evidence type must remain one of documented_fact, empirical_observation, or heuristic. Do not collapse these lanes during normalization: a documented platform requirement, a measured artifact, and a professional judgment have different authority.

Waivers are intentionally narrow. A blocking finding can be mechanically unblocked only when the policy allows that severity and the finding includes a waiver with non-empty owner and reference. The helper never waives rights, accessibility, safety, or policy findings, even if a report supplies a waiver.

The helper does not inspect media, judge subjective quality, verify claims, grant rights/accessibility/safety exceptions, or issue final approval. Its output includes mechanical_only: true and final_approval: false; a producer, reviewer, legal owner, accessibility owner, or release owner still has to make the actual acceptance decision.

Example minimal input:

json
{
	"asset_id": "skincare-ad-v004",
	"review": {
		"reviewer": "QA Agent",
		"date": "2026-07-11T12:20:00Z",
		"target_platform": "paid social",
		"intended_use": "public product ad",
		"source_package": "delivery-package-v004"
	},
	"policy": {
		"dispositions": {
			"critical": "reject",
			"major": "hold",
			"minor": "ready",
			"observation": "ready"
		},
		"waivable_severities": ["major"],
		"non_waivable_areas": ["rights", "accessibility", "safety", "policy"]
	},
	"checks": [
		{
			"id": "claim-substantiation",
			"status": "fail",
			"evidence_type": "documented_fact",
			"evidence_detail": "Product fact sheet does not approve disease-treatment claims."
		}
	],
	"findings": [
		{
			"severity": "critical",
			"area": "policy",
			"issue": "Overlay makes an unsupported eczema treatment claim.",
			"evidence_type": "documented_fact",
			"evidence_detail": "Approved product facts allow dry-skin support only; the video says 'repairs eczema overnight'.",
			"timecode": "00:06.000",
			"recommended_fix": "Replace with approved copy and send revised ad through legal and QA review.",
			"owner": "Legal review",
			"retest_required": true
		}
	]
}

Test matrix by deliverable

DeliverableMinimum QA focus
Still imageBrief fit, dimensions, crop/safe area, product/logo/text accuracy, anatomy, rights, metadata/provenance, platform disclosure, alt text if needed
Social clipHook timing, platform aspect/duration, mobile readability, captions, music/dialogue balance, generated motion artifacts, flashes, thumbnail/frame hold, disclosure
ExplainerFactual accuracy, script-to-visual alignment, narration intelligibility, captions/transcript, diagrams/text, pacing, source citations if factual, accessibility
Product adProduct truth, legal/claim substantiation, packshot fidelity, price/offer accuracy, brand rules, CTA, rights, synthetic endorsements, platform ad policy
Avatar/spokespersonConsent, identity/voice approval, lip sync, gaze/gesture naturalness, disclosure, language/accent, uncanny artifacts, claims
Localization/dubTranslation accuracy, cultural fit, units/names, subtitle timing, dub sync, voice approval, local legal claims, on-screen text replacement
Mixed-source editSource rights, continuity, color/grain match, compositing seams, provenance per ingredient, edit rhythm, captions/audio mix
Broadcast/premium masterClient spec, automated QC, loudness, color/HDR, captions/subtitles, PSE/flashing, metadata, IMF/AS-11 or client package rules where applicable

Example: QA report for a 15-second generated product ad

Context: 15 s vertical skincare ad for Instagram Reels. Approved brief says "show the blue 50 ml bottle, no medical claims, CTA: Shop the summer set." Source package includes generated video, brand logo, music license, product facts, and prompt metadata.

Acceptance basis: 1080x1920 vertical, 15 s max, brand color guide, product fact sheet, FTC truthful-advertising posture, platform synthetic-media disclosure checked on delivery date.

Findings:

SeverityTime/regionIssueEvidenceRecommended fix
Critical00:06 text overlayUnsupported claim "repairs eczema overnight"Product facts only allow "supports dry skin barrier"; health claim not approvedReplace with approved claim; legal review revised copy
Major00:08 bottle labelAI label reads "HYDRATlON" with wrong letterformManual inspection and OCR mismatchRegenerate packshot from locked product render or composite approved label
MajorFull mixMusic masks CTAPhone speaker playback; last three words hard to understandDuck music -6 dB under CTA and retest intelligibility
Minor00:03 backgroundSmall warped leaf visible on pauseNot visible at normal playbackAccept if no other render is needed

Verdict: Reject until claim and product label are fixed; retest final playback, captions, loudness, and disclosure after revision.

Why this example is structured this way: the critical issue is legal/claim risk, not visual polish. A visually beautiful ad with a false product claim fails QA.

Example: QA plan for an avatar localization batch

Context: English training video localized into Spanish, French, and Japanese with synthetic avatar dubbing. Target is internal LMS, not public ads. The client provided consent for the avatar performer and requires captions.

Workflow:

  1. Verify consent and allowed languages for the avatar/voice before watching outputs.
  2. For each locale, compare translated script to approved source for meaning, warnings, numbers, product names, and URLs.
  3. Watch the dub once without captions for naturalness and intelligibility; watch again with captions for timing and line breaks.
  4. Spot-check lip sync at dense consonant passages and at the final 20 seconds, where drift often accumulates.
  5. Confirm captions include speaker changes and meaningful non-speech audio if relevant.
  6. Record locale-specific issues separately; do not let a pass in one language imply pass in another.

Expected failure modes: avatar mouth over-opens on Japanese vowels, Spanish text overruns caption safe area, French dub starts 300 ms late after a pause, translated safety warning changes "must" to "may."

Verdict rule: a locale with mistranslated safety instructions is Critical even if the video file is technically perfect.

Example: QA checklist for a mixed-source explainer

Context: A 60 s explainer combines AI-generated diagrams, licensed stock footage, generated narration, and a human-edited timeline.

Review:

  • Brief conformance: each beat in the approved script appears in order; no unsupported statistic was added by a generated diagram.
  • Source ledger: stock clip license covers the distribution channel; AI diagrams are marked as generated; narration provider and voice are recorded.
  • Visual continuity: diagram labels match narration; colors mean the same thing across scenes; no arrows point to impossible flows.
  • Audio: narration is clean, music is under speech, no edits cut breaths unnaturally, integrated loudness target is documented.
  • Captions/accessibility: captions match final narration, diagram-only information is described in narration or transcript, no flashing transitions exceed safe thresholds.
  • Final package: export matches platform resolution/duration, filename/version is correct, and the QA report lists residual factual sources that were not independently verified.

Pass condition: all factual/statistical claims trace to approved source material, and no generated visual introduces a new unsupported claim.

Source notes

© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/production/governance-delivery/generated-media-qa of calesthio/generative-media-skills.

  • SKILL.md
  • EVAL.md
  • scripts/normalize_qa_report.py
  • tests/test_normalize_qa_report.py

Open the folder on GitHubat commit 8c85352

Compare with similar skills

Generated Media QA next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Generated Media QA compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Generated Media QA this skillcalesthio/generative-media-skills193—~8kAutomated safety check: PassMIT
Creative QA Checklistgtmagents/gtm-agents4131 repos~312Automated safety check: PassApache-2.0
Locale UI Patternsopenchamber/openchamber11k—~1.5kAutomated safety check: PassMIT
iOS Accessibilitydadederk/iOS-Accessibility-Agent-Skill174—~2.4kAutomated safety check: PassMIT
Instui Docsinstructure/instructure-ui482—~498Automated safety check: PassCustom licence
Accessibility I18ndevcodex-labs/devcodex439—~718Automated safety check: PassAGPL-3.0

Similar skills

  • Creative QA Checklist

    gtmagents/gtm-agents

    A skill your agent uses to verify creative assets meet brand, accessibility, and localization standards before launch.

    413 GitHub starsUsed in 1 repo~312 tokens
    Frontend & DesignAuto-check passed
  • Locale UI Patterns

    openchamber/openchamber

    A skill your agent uses when creating or modifying OpenChamber UI text, labels, buttons, placeholders, aria labels, empty states, toasts, dialogs, settings copy, navigation labels, or any…

    11k GitHub stars~1.5k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • iOS Accessibility

    dadederk/iOS-Accessibility-Agent-Skill

    Expert guidance on iOS accessibility best practices, patterns, and implementation.

    174 GitHub stars~2.4k tokensUpdated 7 mo ago
    Frontend & DesignAuto-check passed
  • Instui Docs

    instructure/instructure-ui

    Look up authoritative Instructure UI (InstUI, @instructure/ui-) documentation — component APIs, props, theme variables, usage examples, and guides — by querying instructure.design's plaintext docs.

    482 GitHub stars~498 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Accessibility I18n

    devcodex-labs/devcodex

    无障碍与国际化专家 Owner — 当任务涉及可访问性、键盘操作、焦点、屏幕阅读器、ARIA、语言地区、本地化、RTL、翻译资源、用户可见文案或多语言文档时使用;要求把包容性体验和本地化验证绑定到真实用户路径。

    439 GitHub stars~718 tokensUpdated 22 days ago
    Frontend & DesignAuto-check passed
  • Front-End Checklist CLI Review

    thedaviddias/Front-End-Checklist

    Runs the frontendchecklist CLI against local files or a deployed URL to check code against 380+ accessibility, performance, SEO and security rules, with fixable, cited findings.

    74k GitHub stars~692 tokensUpdated 3 days ago
    Frontend & DesignAuto-check passed

More from calesthio/generative-media-skills

All 26 skills in this repo
  • 3D Asset Production

    calesthio/generative-media-skills

    A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.

    193 GitHub stars~9.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Audio Mixing Mastering

    calesthio/generative-media-skills

    Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…

    193 GitHub stars~7k tokensUpdated 2 mo ago
    Auto-check passed
  • Captions Media Accessibility

    calesthio/generative-media-skills

    Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…

    193 GitHub stars~6.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Comfyui Media Workflows

    calesthio/generative-media-skills

    Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…

    193 GitHub stars~8.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Ffmpeg Media Finishing

    calesthio/generative-media-skills

    Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.

    193 GitHub stars~8.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Hyperframes Video Composition

    calesthio/generative-media-skills

    Provider-independent production workflow for AI agents assembling generated or source media into HyperFrames HTML/CSS/JS videos.

    193 GitHub stars~7.1k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Generated Media QA

What does Generated Media QA do?

Provider-independent quality assurance for AI-generated and AI-assisted media. Generated Media QA is an agent skill from calesthio/generative-media-skills. Provider-independent quality assurance for AI-generated and AI-assisted media.

When should I use Generated Media QA?

Generated Media QA fits situations like: reporting on images; product content; mixed-source edits; delivery readiness.

How do I install Generated Media QA in Claude Code?

Run `npx skills add calesthio/generative-media-skills --skill generated-media-qa -a claude-code`. Or copy the skill folder (skills/production/governance-delivery/generated-media-qa in calesthio/generative-media-skills) into .claude/skills/generated-media-qa in your project. Claude Code loads it when a task matches its description.

How do I install Generated Media QA in Codex?

Run `npx skills add calesthio/generative-media-skills --skill generated-media-qa -a codex`. Or copy the skill folder (skills/production/governance-delivery/generated-media-qa in calesthio/generative-media-skills) into .agents/skills/generated-media-qa in your project. Codex loads it when a task matches its description.

Can I use Generated Media QA in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill generated-media-qa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generated-media-qa, .gemini/skills/generated-media-qa, .github/skills/generated-media-qa and .opencode/skills/generated-media-qa in your project.

What does Generated Media QA need to run?

Going by SKILL.md and its folder, Generated Media QA needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Generated Media QA access the network?

SKILL.md names 15 domains. As links in the text: partnerhelp.netflixstudios.com, w3.org, support.google.com, ftc.gov, tech.ebu.ch, itu.int, atsc.org, fcc.gov, spec.c2pa.org, cv.iptc.org, iptc.org, smpte.org, nist.gov, ai-challenges.nist.gov and arxiv.org. This is read from the text; nothing was executed.

Is Generated Media QA safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Generated Media QA use?

Generated Media QA is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Generated Media QA use?

About 8k tokens (SKILL.md is roughly 32k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Generated Media QA?

Skills that share tags, products or a category with Generated Media QA: Creative QA Checklist (gtmagents/gtm-agents, 413 stars), Locale UI Patterns (openchamber/openchamber, 11k stars), iOS Accessibility (dadederk/iOS-Accessibility-Agent-Skill, 174 stars) and Instui Docs (instructure/instructure-ui, 482 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Generated Media QA?

calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 193 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.

Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.