You type "beautiful sunset over the sea, cinematic" and get the same postcard the internet has seen a million times. The AI isn't the problem. A good prompt for an AI image or video is a short brief for a camera operator: who is in the frame, what they do, where, how the camera moves, what the light is like, and what we hear. The official guides from OpenAI, Google and Runway agree on the core rule: one scene, one action, one camera move. Below you'll find a 7-part formula, 12 copy-paste templates, and a bonus prompt that writes prompts for you.
Why the AI gives you the wrong picture
Because it's missing details, and it fills the gaps with the most typical option. A prompt (the text instruction you give an AI) works like ordering at a restaurant: ask for "something tasty" and you get the most popular dish on the menu, not the one you were craving.
Almost everyone who makes content now uses these tools. In September 2025, Adobe surveyed 16,000 creators across eight countries: the US, UK, France, Germany, Japan, South Korea, India and Australia. 86% use generative AI (models that create new images, video and text), and 52% use it specifically to generate images and video. Yet 34% say the output quality is unreliable.
Words like "beautiful" and "cinematic" are often the culprit. OpenAI's guide for its video model Sora says to replace a vague "cinematic look" with specifics: which lens, how shallow the depth of field (how blurred the background is), where the light comes from. The model can't know what looks good to you. It does understand "soft window light from the left, warm lamp behind."
So you need a fixed order for building the brief every time. Here it is.
The prompt formula: 7 parts that work everywhere
A universal prompt has seven parts: shot, subject, action, place, light and color, style, sound. Google's Veo 3.1 guide (October 2025) uses a similar formula: cinematography + subject + action + context + style. Runway puts the camera first, but the ingredients are the same.
- Shot. Framing and angle: wide shot, medium shot, close-up, low angle. It's the first thing the viewer sees, so decide it first.
- Subject. Who or what is at the center, with 2-3 details: "an old fisherman in a yellow raincoat," not "a man."
- Action. Verbs, broken into beats. OpenAI suggests writing not "actor walks across the room" but "takes four steps to the window, pauses, and pulls the curtain in the final second."
- Place and time. "Abandoned warehouse at dusk" beats "dark building."
- Light and color. The light source plus 3-5 color anchors (the main colors of the frame), like "amber, cream, walnut brown." Anchors keep a series of shots consistent.
- Style. Medium and era: 35mm film photo, 3D render, watercolor, 1980s print ad.
- Sound (video only). Dialogue, effects and music, each in its own sentence.
For photos, sound drops out. For video, camera movement comes in. It's one formula with different parts switched on:
- Image: shot - subject - place - light and color - style - aspect ratio.
- Video: shot and camera move - subject - one action - place - light - style - sound.
- Motion design in code: palette and fonts - scenes mapped to music beats - build rules - banned list.
Knowing the formula isn't enough, though. Five common mistakes can wreck even a well-built prompt.
5 mistakes that ruin a good prompt
The biggest one is describing what you don't want. Runway's guide explains that "no camera shake" makes the model focus on the word "shake." Write what you do want: "smooth, stable camera movement."
- Negatives. Google gives the same advice: instead of "no man-made structures," describe "a desolate landscape with no buildings or roads," painting the picture you need. The exception is Midjourney, which has a dedicated
--noparameter. - Too much in one clip. Runway asks you to skip "then" and "next": one generation, one action. OpenAI goes further and suggests stitching two 4-second clips instead of generating one 8-second clip.
- Empty words. "Masterpiece," "8k," "stunning" describe neither light nor framing. Swap them for details.
- Two camera moves at once. "Push in and orbit around the hero" rarely works. Runway: one primary camera move per prompt.
- Rewriting from scratch. OpenAI recommends changing one thing at a time: "same shot, switch to 85 mm." If a shot keeps failing, strip it back: lock the camera, simplify the action, clear the background.
Now the useful part: ready-made templates.
Image prompt templates: 5 to copy and paste
Swap anything in square brackets for your own details. The templates work with the image tools in ChatGPT and Gemini, with Midjourney, and with most other generators.
1. Profile photo or headshot
Photorealistic head-and-shoulders portrait of [who: age, appearance, clothing], [expression: calm half-smile], looking slightly off camera. Soft window light from the left, gentle fill, neutral [color] background. Shot on 85mm lens, shallow depth of field, natural skin texture. Vertical 4:5. No text, no logos.An 85mm lens is the classic portrait choice: no facial distortion, softly blurred background. OpenAI advises asking for "photorealistic" explicitly when you want a photo rather than an illustration.
2. Product shot
Product photo of [product, material, color] on [surface: wet black stone / linen / concrete], [angle: three-quarter view at eye level]. One hard key light from the right, soft reflection, [2-3 color anchors]. Clean premium commercial look, sharp focus on the product, square 1:1, empty space on the left for text. No extra objects, no text.3. Poster or banner with text
Poster for [what: workshop / sale / podcast]. Main image: [scene]. Headline in bold sans-serif at the top, render exactly once: "[YOUR TEXT]". Spelling: [Y-O-U-R T-E-X-T]. Colors: [3 colors]. Horizontal 16:9. No other text, no watermark.Garbled text is the most common complaint. OpenAI's three rules: put the exact copy in quotes, spell unusual words and brand names letter by letter, and ask for no other text.
4. Editing an existing photo
Edit this photo. Change only: [what: replace the grey sky with a warm sunset]. Keep identical: the person's face and identity, pose, clothing, framing, lighting direction. Do not add text, logos or extra people.The rule from OpenAI's guide: separate what changes from what must stay, and repeat the key constraints on every follow-up edit.
5. A series in one consistent style
Series style, keep for every image: [medium: flat editorial illustration], palette [#hex1, #hex2, #hex3], [light: soft top light], [mood], same line weight and texture. Image [N] of [total]: [scene].Paste the style block unchanged into every prompt and edit only the last line. The series will look like one artist made it.
Midjourney parameters worth knowing
In Midjourney, some settings go after the text, preceded by two hyphens. The most useful:
--ar 16:9- aspect ratio (default is square, 1:1);--s 0-1000- how much of its own aesthetic the model adds (default 100; lower follows your text more closely);--c 0-100- how different the variations are (default 0);--no text, logo- what to leave out;--seed 1234- locks the randomness, handy for comparing edits;--sref [link]- borrow the style of a reference image.
Example: [your prompt] --ar 16:9 --s 50 --c 10 --no text. Parameters change between model versions, so check Midjourney's docs before you start.
That covers images. Video adds something a photo doesn't have: time.
Video prompt templates: 5 to copy and paste
A video prompt describes one action and one camera move over 4-10 seconds. Google Veo 3.1 makes clips of 4, 6 or 8 seconds in 16:9 or 9:16, so longer stories are built from several short shots.
6. One cinematic shot
[Shot: medium shot, slow dolly forward]. [Subject with 2-3 details] [one action, in beats: takes three steps, stops, turns head to camera in the last second]. [Place and time: rain-soaked alley at night]. Light: [neon pink and cold blue, wet reflections]. Style: [35mm film, light grain]. Smooth, stable camera.7. Bring a photo to life
Animate this image. Camera: [slow push-in]. Motion: [hair and scarf move in light wind, she blinks and smiles slightly]. Background: [clouds drift slowly]. Keep the face, clothing and colors exactly as in the image. Calm, natural movement.The image already sets the look and style, so spend every word on motion: what moves, where the camera goes, and how fast.
8. A shot with dialogue and sound
Close-up of [character], [setting]. She looks at the camera and says: "[short line, up to 8 words]". SFX: [door creaks behind her]. Ambient noise: [quiet café chatter]. Warm, natural light.This is the syntax from Google's guide: dialogue in quotes, sound effects after SFX:, background after Ambient noise:. OpenAI adds that a 4-second shot fits one or two short lines, no more.
9. A mini-story with timestamps
[00:00-00:02] Wide shot: [setting], [character] enters frame.
[00:02-00:04] Close-up: [reaction, emotion].
[00:04-00:06] Tracking shot: [key action].
[00:06-00:08] Wide high-angle shot: [final image]. SFX: [music swells].Here the guides disagree. Google shows timestamps as a working technique for Veo, while Runway says to keep one action per prompt. The takeaway: use timestamps only if your model supports them. Otherwise, split the story into separate clips and join them in an editor.
10. Vertical clip for Reels, Shorts and TikTok
Vertical 9:16. First second: [strong visual hook: object falls into frame / extreme close-up of eyes]. [Subject] [one clear action], centered, large in frame, empty space at the top and bottom for captions. [Bright, high-contrast light]. Handheld but steady camera. Audio: [upbeat beat with a hit on the first second].Vertical video is watched on a phone, where small details disappear. Keep the subject big and centered, and leave room above and below for captions.
But if you need to show an interface, a logo or exact text, video models often scramble the letters. There's a better route for that.
AI motion design: when the animation is written in code
Motion design (animated graphics: text, icons, interfaces and logos in motion) works better with an AI coding assistant like Claude, ChatGPT or Gemini than with a video model. A video model paints pixels and can warp letters. A coding assistant builds the animation as a web page, and a program records it frame by frame. Text stays sharp, brand colors stay exact, and you make edits in plain words.
Tools for this already exist. Remotion, for example, is a library for making videos with React code, and its site says outright that you can make videos "agentically," through an AI assistant. A single HTML file is often enough, too.
Strong prompts from motion designers read like a six-part spec:
- Inputs: product, screenshots, logo, music, plus defaults if something is missing.
- Direction: palette in HEX codes (a digital color notation, like #0b1a24), fonts, and a banned list ("no crossfades, no particles, nothing that looks like a template").
- Structure in beats: 120 BPM (beats per minute) means 2 beats per second, so 15 seconds = 30 beats. Each scene gets its own beats.
- Build rules: every frame is computed from time alone, with no randomness, so the render comes out identical every run.
- Gotchas: what has broken before and how to avoid it.
- Start: a plan and 3-4 stills first, the full build only after you approve.
11. A 15-second product promo
Make a 15-second product promo for [product + URL] as one HTML file, 1920x1080, 60fps.
Inputs: [screenshots / logo / music]. If missing: study the website first, then ask me.
Direction: one continuous take, every scene grows out of the previous one (buttons become pages, text rises from a mask line). Palette: [#hex x4]. Fonts: [UI font] + [display font]. Banned: crossfades, particles, generic template look.
Structure at [120] BPM, [30] beats:
beats 0-6: [hook: the user's problem in one line of UI],
beats 6-14: [the product solves it: key feature in action],
beats 14-24: [result: numbers or outcome],
beats 24-30: [logo + call to action].
Build: every style is computed from time inside seek(t), no CSS transitions or timers; springs as closed-form functions; render frames with Playwright, add music and sound effects on the beats.
Start: show me the beat map and 4 stills before writing the full film.12. A 5-second animated logo
Animate my logo [file] as a 5-second intro, one HTML file, 1080x1080. The mark draws itself line by line, then fills with [#hex]; the wordmark "[NAME]" rises letter by letter from a mask; a final settle with a light spring. Background [#hex]. Every frame computed from time, so the render is identical each run. Show 3 stills first.The last line in both templates matters most. Stills before the full build save hours: fixing a plan is cheaper than fixing a finished video.
And if you'd rather not fill in the brackets yourself, you can hand that job to AI too.
A prompt that writes prompts for you
The fastest route is to ask the AI to interview you first, then assemble the prompt by the formula. It works in any chatbot: ChatGPT, Claude, Gemini, DeepSeek.
Bonus: an AI prompt assistant
You are a prompt director for AI image and video generation. Ask me up to 6 short questions, one at a time: what I want to create, where it will be published, format, mood, key details, what to avoid. Then write the final prompt in English using this order: shot, subject, one action, place and time, light and 3-5 color anchors, style, sound (video only). Rules: one camera move, one action, positive wording instead of "no ...", exact text in quotes. After the prompt, give 2 variations that change only one thing each.That final rule, variations that change just one detail, is OpenAI's own method: you compare two results and instantly see what made the difference. New to chatbots? Start with our beginner's guide to ChatGPT, Claude and Gemini.
One question remains: when do you have to label an AI-made video?
Do you have to label AI-generated video?
In the European Union, yes, if it resembles real people or events. Article 50 of the EU AI Act has applied since August 2, 2026. Companies that make generators must mark their output in a machine-readable way (a hidden tag software can detect). Anyone who publishes a deepfake (fake video, images or audio that look real) must disclose that it's artificial.
Creative work gets some leeway: if a piece is clearly artistic, satirical or fictional, the disclosure just has to avoid spoiling the viewing. Rules differ in other countries, so check the requirements of the platform where you post.
A simple rule that works anywhere: don't generate real people without their consent. We covered how easy it is to fake someone's voice in "How many seconds does it take to clone a voice?"
What to do right now
Open the last prompt that gave you the wrong result. Break it into the seven parts and see which one is missing. Most likely it's the light or the shot. Add that one part and generate again, changing nothing else.
That sunset over the sea becomes "wide shot, low angle, an old wooden pier running into the sea, the last ten minutes of sun, amber and plum, 35mm film." That's not a postcard anymore. That's your shot.
Sources
- OpenAI: Sora 2 Prompting Guide (2025)
- Google Cloud: Ultimate prompting guide for Veo 3.1 (2025)
- Runway: AI Video Prompting Guide
- OpenAI: Image prompting
- Adobe: Creators' Toolkit Report (2025)
- EU AI Act, Article 50: Transparency obligations
- Gradually: Midjourney Parameters (2026)
- Remotion: programmatic video with React






