skills/editorial-video-keyframes/SKILL.md
Workspace snapshot · 09/04 14:52
name: editorial-video-keyframes description: Create production-grade keyframes, style frames, image-to-video start/end frames, thumbnails, and campaign stills for short-form video. Use this whenever AI Video Lab needs to generate or edit a still image that will determine a video's visual quality, including editorial advertising, brand stories, product/service promotion, social video hooks, or converting first-party screenshots, photos, articles, meeting insights, and project artifacts into a coherent visual system. compatibility: Requires an image generation or image editing tool and a deterministic compositor such as Remotion for final Japanese typography.
Editorial Video Keyframes
Turn first-party evidence into images that can carry a film. Do not start from “make it stylish.” Start from what the audience should feel, what only this source can prove, and what must remain recognizable across shots.
Output contract
Produce the image, not a long prompt essay. Keep the accompanying note short:
- Intended role in the video
- Source assets used
- Generation or edit decision
- One strongest feature and one remaining risk
- Saved path and final generation instruction
For a sequence, deliver at least a hook frame, development frame, and payoff frame. Treat them as one visual family, not three unrelated posters.
Workflow
1. Ground the frame in evidence
Identify the primary source, the claim it supports, and its privacy class. Prefer real photos, work artifacts, screenshots, products, environments, and recognizable brand forms. Generated decoration must support the evidence rather than replace it.
Do not visualize inventory counts or abstract concepts when a specific event, decision, object, person, or before/after exists.
2. Make five decisions before prompting
Write these internally in one line each:
- Audience action: what should happen after viewing?
- Dominant emotion: choose one, such as curiosity, tension, aspiration, trust, surprise, or momentum.
- Visual protagonist: the single subject noticed first.
- Brand tension: the contrast that makes this story specific, such as craft versus deadline or philosophy versus ranking.
- Motion affordance: what can move, reveal, parallax, transform, or cut in the next stage?
If one of these is vague, the image will usually become generic.
3. Choose generation mode
- Edit a real source when identity, product truth, place, interface, or documentary credibility matters.
- Generate a new plate when the source provides an idea but no usable visual.
- Composite when several first-party artifacts must coexist.
- Keep a clean plate and add typography later when exact Japanese text matters.
4. Design the visual system
Choose one composition family from references/visual-grammar.md. Lock the following across a sequence:
- palette: one dominant, one support, one high-chroma accent, one neutral family
- light: one consistent light logic
- lens and perspective
- surface language and grain
- subject treatment and crop logic
- recurring graphic motif
Variation should come from scale, camera distance, crop, and action—not from changing the entire style every shot.
5. Separate image craft from typography
For video production, default to two layers:
- Image layer: subject, environment, light, material, color, depth, and deliberate negative space.
- Type layer: exact Japanese copy, brand name, labels, and CTA composed in Remotion or another deterministic layout tool.
Ask the image model for text only when baked-in lettering is itself the visual experiment. Otherwise prohibit fake characters, pseudo-logos, watermarks, and decorative English.
Japanese type should behave as an image: use scale contrast, cropping, outline/solid contrast, vertical/horizontal tension, repetition, or collision with the subject. Avoid ordinary centered subtitles and white bold text placed over a dark gradient.
6. Prompt with priorities
Use this compact order:
Use case: ads-marketing
Asset type: <hook keyframe / development keyframe / payoff keyframe / thumbnail>
Primary request: <specific event and desired audience response>
Input images: <role of each image>
Scene/backdrop: <concrete place or constructed set>
Subject: <single protagonist and action>
Style/medium: high-budget Japanese editorial campaign still, adapted to the brand
Composition/framing: <composition family, crop, depth planes, negative space>
Lighting/mood: <one light logic and one dominant emotion>
Color palette: <locked palette with one stop-color accent>
Materials/textures: <specific surfaces>
Motion affordance: <foreground/midground/background separation and intended movement>
Text: none; reserve <location> for deterministic Japanese typography
Constraints: preserve <identity/product/logo/source truth>; 9:16; mobile legibility
Avoid: generic AI futurism, stock-photo business scene, split hero layout, Canva/LP look, fake UI, fake text, excessive decoration, dark gray office, broken hands or faces
Prioritize in this order: source truth → protagonist → emotion → light → composition → color → styling → decoration.
7. Generate, inspect, and iterate narrowly
Inspect at full frame and phone-thumbnail size. Score with references/quality-gate.md.
Reject instead of polishing when the image is generic, the subject is not obvious in one second, identity or product truth drifted, the frame has no motion affordance, or it looks like a landing-page hero.
Iterate one issue at a time: light, crop, styling, background, accent color, or negative space. Do not rewrite the whole concept after every miss.
8. Run a dissent loop before video spend
Do not promote the first attractive image. For a new visual system:
- Build a minimum sequence of hook, development, turning point, and payoff.
- Score every frame yourself against the quality gate.
- Run an independent Claude
-preview against the brief, original assets, and quality gate. Instruct it to reject identity drift, invented artifacts, pseudo-text, generic AI-ad language, and a sequence that does not prove the single claim. - Verify the reviewer's factual claims against the filesystem; independent review can also hallucinate.
- Regenerate only the weakest frame and only the named defect.
- After two failures caused by generative corruption, change the production method. Preserve the real asset pixel-for-pixel and finish the shot with deterministic Remotion composition.
Only frames scoring at least 20/24 with no blocker may enter a paid image-to-video experiment.
Sequence rules for image-to-video
- Give the model real depth planes and partial occlusions so camera movement has something to reveal.
- Avoid tiny repeated details that shimmer during animation.
- Keep faces and hands large enough to evaluate before video generation.
- Use start and end frames when the intended transformation is specific.
- Do not bake motion blur into the keyframe unless the shot deliberately begins mid-action.
- Preserve a shot bible containing palette, lens, light direction, subject wardrobe, props, and seed/reference set.
Safety and brand integrity
- Do not expose confidential screenshots, client names, personal data, local paths, or unpublished meeting text.
- Do not imitate a named living artist, specific campaign, magazine cover, or competitor composition.
- Do not invent evidence, metrics, endorsements, product features, or brand claims.
- Preserve real identities when editing employee or customer photos; do not silently beautify, age-shift, or alter body shape.
- Keep external publication and paid generation behind the department approval boundary.