Most AI videos fail for the same boring reasons. The subject morphs mid-clip. The camera does something nobody asked for. The product changes color between seconds two and four. The output is technically "a video" and practically unusable.
After looking at tens of thousands of real AI video prompts — the ones that produced clips people actually shipped, and the ones that produced garbage people deleted — a pattern emerges. Great prompts aren't longer or more poetic. They're more structured. They tell the model what changes, how the camera behaves, what must stay locked, and what they refuse to accept.
This is the craft companion to our data report on what 40,000 AI video prompts reveal about what people make. That post covers what creators generate. This one covers how the good ones write it. Five patterns, each with a weak version, a strong version, and why the difference matters.
Pattern 1: Lead With Subject, Action, and Change Over Time
Video is motion. The single biggest difference between prompts that produce living footage and prompts that produce a slow zoom on a photograph is whether you described something happening.
Weak prompts describe a scene. Strong prompts describe a scene that changes.
Weak: A coffee cup on a wooden table in a cafe.
Strong: A steaming coffee cup on a wooden cafe table; steam curls upward and drifts left as morning light slowly brightens across the surface over 5 seconds.
The weak version gives the model a still image and forces it to invent motion — usually a lazy push-in or some ambient jitter. The strong version names the subject (coffee cup), the action (steam curls and drifts), and the change over time (light brightening across the clip). The model now has a beginning and end state to interpolate between, which is exactly what a video model is built to do.
The fix is mechanical. For every prompt, ask: what is the one thing that is different at the end of this clip versus the start? If you can't answer, you're going to get a moving postcard. Bake that change into the sentence. Even a small one — a head turn, a door opening, fog rolling in — gives the model a job to do across the timeline.
Pattern 2: Direct the Camera Like a Cinematographer

If you don't specify the camera, the model picks one for you — and it picks badly, defaulting to a generic dolly-in or a drifting handheld wobble that screams "AI." The best prompts treat the camera as a deliberate creative choice, not an afterthought.
You need three things: shot size (wide, medium, close-up), lens or framing feel (35mm, wide-angle, shallow depth of field), and one motion (slow push-in, orbit, static lock-off). One motion. Not three.
Weak: A car driving down a coastal road, cinematic.
Strong: Wide tracking shot of a vintage convertible on a coastal highway, shot on a 35mm lens with shallow depth of field, camera tracks alongside the car at matching speed, golden hour.
"Cinematic" is a wish, not an instruction. The strong version tells the model the framing (wide tracking), the optical character (35mm, shallow depth of field), and a single coherent move (track alongside at matching speed). That coherence is what reads as professional. Conflicting camera instructions — "orbit while zooming and panning" — are where models fall apart and produce that swimmy, unstable look.
If you're new to thinking in camera terms, our guide on how to write AI video prompts breaks down the vocabulary. The shortcut: imagine you're handing a one-line instruction to a camera operator who will do exactly what you say and nothing more. Be that specific.
Pattern 3: Lock Your Continuity Tokens
This is the pattern that separates hobbyists from people producing usable footage. AI video models drift. Across a few seconds, a face subtly re-renders into a different person, a red logo shifts to orange, a product gains a button it didn't have. Continuity tokens are the specific, repeatable phrases you use to nail those elements down.
A continuity token is a short, distinctive description you commit to and reuse verbatim — for the subject's identity, the product, the color palette, and any branding.
Weak: A woman in a red jacket walks through a city, then we see her closer up.
Strong: A woman with shoulder-length curly black hair and a bright crimson leather jacket walks through a neon-lit city; same crimson jacket and same hairstyle held consistent throughout the clip.
"A woman in a red jacket" is an invitation for the model to reinvent her. "Shoulder-length curly black hair and a bright crimson leather jacket," repeated and explicitly flagged as consistent, gives the model an anchor to hold. When you generate multiple clips for one project, copy those exact tokens into every prompt — never paraphrase them. Paraphrasing is how the character in shot three stops looking like the character in shot one.
For brand work this is non-negotiable. Lock the exact hex-equivalent color name, the logo placement, and the product's defining feature in every single prompt. If your platform supports an image reference or text-to-video with a starting frame, use it — but back it up with locked text tokens, because the description is what carries identity through the motion, not just into the first frame.
Pattern 4: Match the Shot to Platform and Duration

A prompt that's great for a 12-second YouTube hero is wrong for a 4-second TikTok hook, and the difference isn't just aspect ratio. The best prompts are designed backward from where the video will live.
Three decisions get made before you write a word of description: aspect ratio (9:16 vertical for feeds, 16:9 for YouTube and landing pages), duration (and therefore how much can actually happen), and pacing (one calm beat for a short loop, a clear arc for a longer clip).
Weak: An energetic montage of a fitness product with lots of quick cuts and text, for social media.
Strong: 9:16 vertical, single continuous 5-second shot: a runner laces up bright orange sneakers and pushes off frame-left into a sprint, fast-paced, punchy, designed as a TikTok hook with the action landing in the first 2 seconds.
Asking for "lots of quick cuts" inside a single short generation is asking for a mess — most models produce one continuous shot per generation, so the request fights the tool. The strong version respects the format: vertical, one shot, an action engineered to hit in the first two seconds where the platform demands it. You'll often get a better result by generating several clean single-shot clips to this spec and cutting them together than by trying to cram an edit into one prompt.
Duration drives how much change you can ask for, too. In four seconds, one clear action lands. In twelve, you can stage a small arc. Asking for a three-act story in four seconds just smears everything together.
Pattern 5: Constrain With Negatives and a Clear Output Spec
The final pattern is the one almost nobody uses, which is exactly why it's an edge. Telling the model what you don't want is often more powerful than piling on more of what you do. Pair that with an explicit output spec and you stop leaving the unglamorous decisions to chance.
Two moves: negatives (the artifacts and clichés you refuse — warped hands, text gibberish, extra limbs, flickering, the unwanted slow zoom) and an output spec (frame rate feel, lighting, mood, and aspect ratio stated plainly at the end).
Weak: A chef plating a dish in a restaurant kitchen.
Strong: A chef precisely plating a dish in a warm restaurant kitchen; medium shot, soft key light from the left, calm and deliberate pacing, 16:9. Avoid: distorted hands, extra fingers, floating utensils, on-screen text, fast camera movement.
The negative list does real work. Hands are where video models embarrass themselves, so naming "distorted hands, extra fingers" tells the model to spend effort there. "Avoid on-screen text" kills the gibberish lettering models love to hallucinate. And closing with the output spec — shot size, lighting direction, pacing, aspect ratio — means you're not hoping the model guesses your intent; you've stated it.
Keep your negative list tight and relevant. Ten generic negatives dilute the signal. Three or four that target this prompt's likely failure points sharpen it. Different models have different weak spots, so it pays to know which one you're using — our AI model strengths map breaks down where each model excels and where it tends to break.
How to Combine All Five Into One Prompt

These patterns aren't a menu — the best prompts stack all five. Here's the order they naturally fall into:
- Subject + action + change ("a chef plates a dish; steam rises as she sets the final garnish")
- Camera ("medium shot, 50mm, slow push-in")
- Continuity tokens ("same chef in a white double-breasted jacket throughout")
- Platform + duration spec ("16:9, 8 seconds, calm pacing")
- Negatives + output ("warm key light from the left. Avoid: distorted hands, on-screen text")
Read top to bottom, that's a single coherent instruction a model can execute confidently. Each clause answers a question the model would otherwise answer for itself — and "for itself" is where bad AI video comes from.
You don't have to start from a blank page every time, either. A library of copyable prompt templates gives you proven skeletons for common shot types; you swap in your subject and tokens and you're already running all five patterns without thinking about it.
Your Next Step
Pick one prompt you've written that produced a disappointing clip. Run it through the five patterns: Does it name a change over time? Does it direct one clear camera move? Are your continuity tokens locked and repeated? Is it specced to a real platform and duration? Does it tell the model what to avoid?
Fix the two weakest answers and regenerate. That single edit pass is usually the difference between a clip you delete and a clip you ship.
When you're ready to put the patterns to work, open text-to-video in the app and write your first prompt the structured way — subject, camera, tokens, spec, negatives. And if you want the data behind what's actually working at scale, read the companion analysis of what 40,000 AI video prompts reveal. Craft plus evidence is how you stop guessing and start directing.
We Analyzed 40,000+ AI Video Prompts
Everyone has opinions about AI video. Pundits predict where it's going. Twitter debates whether it's "good enough yet." YouTube thumbnails scream about the latest model update.
But almost nobody talks about what people are actually making with these tools right now.
So we decided to find out.
We pulled data from over 120,000 AI-generated videos created on Vivideo, classified a sample of 40,000+ prompts using GPT-4o-mini, and crunched the numbers. What emerged is a surprisingly detailed portrait of how real people — not influencers, not researchers, but everyday creators and businesses — are using AI video in 2025.
Here's everything we found.
The Dataset: How We Got These Numbers
Let's get the methodology out of the way so you know exactly what you're looking at.
Our full dataset spans 120,000+ videos generated through Vivideo's platform. For the detailed prompt analysis, we took a stratified sample of 915 prompts and ran them through GPT-4o-mini for classification into use-case categories. The broader statistics — model usage, aspect ratios, durations, languages, and input types — come from the complete dataset.
We didn't cherry-pick. We didn't filter for "impressive" outputs. This is raw, unfiltered data from real users doing real work (and yes, some of it is people making birthday videos for their mom — and that's great).
A few caveats: prompt classification by AI isn't perfect. Some prompts are ambiguous. A "product video with a person talking" could be tagged as either a product demo or an avatar video. We optimized for the most likely intent, and spot-checked hundreds of classifications manually.
With that said, let's dive in.
The Big Picture: Text-to-Video vs. Image-to-Video
The first question we asked was simple: How are people starting their videos?
Are they typing a prompt from scratch? Or uploading an image and bringing it to life?
65.7% of all video orders are text-to-video. 32.6% are image-to-video. The remaining ~1.7% use other methods like avatar generation.
This was somewhat surprising. We expected image-to-video to be higher — after all, it's arguably "easier" since you're giving the AI a visual starting point. But the data tells a different story: two-thirds of users prefer to describe their vision in words and let the AI figure out the visuals.
Why? A few theories:
- Lower barrier to entry. You don't need to have or find the right image. You just type what you want. Text-to-video is the ultimate blank canvas.
- More creative control. Text prompts let you specify mood, camera movement, lighting, and style — things that are harder to communicate through a static image.
- The "imagination gap." Many users are creating scenes that don't exist yet — fantasy worlds, product concepts, narrative sequences. You can't upload a photo of something that hasn't been built.
That said, image-to-video has its own loyal audience. It's particularly popular for e-commerce product animations, real estate walkthroughs (start with a photo of the property), and bringing artwork to life.
What People Actually Create (The Use-Case Breakdown)
This is the section we were most excited about. When we classified all 915 sample prompts by use case, one category absolutely dominated.
| Use Case | Percentage |
|---|---|
| AI-generated video scenes | 88.2% |
| Avatar / talking head videos | 7.1% |
| Image animation | 4.7% |
Let that sink in. Nearly 9 out of 10 AI videos are fully generated scenes — not someone's face talking to camera, not a Ken Burns effect on a photo, but complete visual scenes conjured from text descriptions.
This is the real story of AI video in 2025: people are using it as a visual imagination engine.
What Those Scenes Actually Look Like
We dug deeper into the 88.2% to understand what kinds of scenes people are generating. While the categories overlap (a promotional video can also be a narrative), here are the primary patterns we observed:
- Promotional videos — Businesses creating ads, brand videos, and marketing content. Everything from local restaurant promos to SaaS product launches.
- Educational content — Explainer videos, tutorials, and "how it works" sequences. Teachers, course creators, and corporate trainers are early power users.
- Social media content — Short, punchy clips designed for TikTok, Instagram Reels, and YouTube Shorts. Often trend-driven and designed for maximum scroll-stopping impact.
- Storytelling and narrative — Short films, music video concepts, and narrative sequences. This is where the most creative prompts live — people building entire worlds in 4-12 seconds.
- Product demonstrations — E-commerce sellers showcasing products in lifestyle contexts. "Show my sneaker being worn by a runner on a mountain trail at sunset" — that kind of thing.
- Personal greetings and celebrations — Birthday messages, holiday cards, anniversary surprises. AI video as the new Hallmark card.
- Real estate tours — Virtual property walkthroughs, neighborhood showcases, and architectural visualizations.
- E-commerce product showcases — Product beauty shots, 360° style reveals, and lifestyle context videos that make products look premium.
The avatar/talking head category (7.1%) is smaller than you might expect given all the buzz around AI avatars. This is partly because avatar generation is a specialized use case — it requires a different workflow and appeals to a narrower audience (mostly corporate training and personalized sales outreach).
Image animation at 4.7% represents users who upload a still photo and add motion — a popular choice for bringing artwork, old photos, or product images to life.
The Language of AI Video: A 24-Language Phenomenon
Here's something that genuinely surprised us. If you assumed AI video creation is primarily an English-speaking activity, the data says otherwise.
English accounts for just 47.3% of all prompts. That means more than half of all AI video prompts on Vivideo are written in non-English languages.
This isn't just "a little international." This is a global phenomenon, with meaningful adoption across every continent.
| Language | % of Prompts |
|---|---|
| English | 47.3% |
| Vietnamese | 23.1% |
| Arabic | 11.4% |
| Russian | 3.2% |
| Turkish | 2.7% |
| German | 2.2% |
| Ukrainian | 1.9% |
| Indonesian | 1.7% |
| Spanish | 1.3% |
| Dutch | 0.9% |
| Hebrew | 0.7% |
| Polish | 0.7% |
| Chinese | 0.6% |
| Portuguese | 0.6% |
| Swedish | 0.5% |
| Greek | 0.4% |
A few things jump out:
Vietnamese at 23.1% is massive. Nearly a quarter of all prompts are in Vietnamese. This reflects Vietnam's booming digital creator economy and early adoption of AI tools for content creation. Vietnamese creators are using AI video for everything from e-commerce product videos to social media content at scale.
Arabic at 11.4% makes the MENA region one of the most active AI video markets. Given the rapid digital transformation happening across the Gulf states and the massive investment in AI infrastructure, this tracks.
The long tail is real. Beyond the top languages, there's meaningful activity in Russian, Turkish, German, Ukrainian, Indonesian, and many more. AI video isn't a Silicon Valley toy — it's a global creative tool.
This has huge implications for anyone building in this space: if your AI video tool only works well with English prompts, you're ignoring more than half your potential users.
Format Preferences: Aspect Ratios and Durations
How people format their videos tells you a lot about where those videos are going to end up.
Aspect Ratios
| Aspect Ratio | Percentage |
|---|---|
| 16:9 (Landscape) | 52.8% |
| 9:16 (Portrait/Vertical) | 43.7% |
| 1:1 (Square) | ~0% |
The landscape-vs-portrait split is remarkably close — 52.8% to 43.7% — which tells us something important: the battle between horizontal and vertical video is essentially a coin flip.
Landscape still leads, likely driven by YouTube, website embeds, presentations, and traditional marketing content. But vertical is right on its heels, fueled by TikTok, Instagram Reels, and YouTube Shorts.
The real shocker? Square video (1:1) is essentially dead. At roughly 0%, nobody is creating square videos anymore. Instagram's old square format, once the default for social media, has been completely abandoned in the AI video era.
Video Durations
| Duration | Percentage |
|---|---|
| 12 seconds | 30.1% |
| 4 seconds | 29.2% |
| 8 seconds | 23.3% |
| 6 seconds | 6.6% |
Duration preferences reveal a fascinating two-camp split:
Camp 1: The 12-second crew (30.1%). These users want the maximum available duration. They're creating narrative content, product demos, and promotional videos where every extra second counts. Twelve seconds is enough to tell a mini-story: setup, reveal, payoff.
Camp 2: The 4-second crew (29.2%). These users want quick, punchy clips — perfect for social media hooks, ad creatives, or stacking multiple clips into longer edits. Four seconds is basically one strong visual moment.
The 8-second middle ground (23.3%) captures users who want a bit more breathing room than 4 seconds but don't need the full 12. The relatively low popularity of 6-second videos (6.6%) is interesting — it seems people prefer to commit to either "short" or "long" rather than splitting the difference.
The Model Race: Veo 3.1 Runs Away With It
If there's a headline stat from this entire analysis, it might be this one:
Veo 3.1 powers 96.4% of all AI video generation on Vivideo.
That's not a typo. Google's Veo 3.1 model is the overwhelming choice for AI video creation.
| Model | % of Usage |
|---|---|
| Veo 3.1 | 96.4% |
| Sora 2 | 2.0% |
| HeyGen (Avatars) | 10.5% of all orders |
Note: HeyGen avatar generation is counted separately as it serves a different function (digital avatars vs. scene generation). Its 10.5% share overlaps with the avatar category in our use-case analysis.
Why does Veo 3.1 dominate so completely? Based on user feedback and our own testing:
- Visual quality. Veo 3.1 consistently produces the most photorealistic and visually coherent output.
- Prompt adherence. It follows complex prompts more faithfully — camera movements, lighting specifications, style directives.
- Speed. Generation times are competitive, and the quality-to-speed ratio is best-in-class.
- Consistency. Less "weird AI artifacts" — fewer melting hands, impossible physics, and uncanny valley moments.
Sora 2 at 2.0% still has its fans, particularly for more artistic and stylized content. But the market has spoken, at least for now: when people want reliable, high-quality AI video, they're choosing Veo 3.1.
Surprising Findings
Every good data analysis turns up things you didn't expect. Here are the patterns that made us do a double-take.
1. The 9% Content Moderation Rate
Approximately 9% of all prompts were flagged by content moderation systems as adult or inappropriate content. This is actually lower than what many in the industry expected — some estimates put the adult content attempt rate for AI image generators at 15-20%.
What does this mean? AI video creation skews more professional and purposeful than AI image generation. When you're paying for video generation (as opposed to playing with a free image tool), the intent is more serious and the use cases are more business-oriented.
2. The Birthday Card Effect
Personal greetings — birthdays, holidays, anniversaries — showed up far more than we expected. These aren't the flashy use cases that get featured in AI demo reels, but they represent a genuinely heartwarming application of the technology. People are creating personalized video messages that would have been impossible (or prohibitively expensive) just two years ago.
3. The Death of Square Video
We already mentioned this, but it bears repeating: 1:1 square video is at effectively 0%. The format that dominated Instagram from 2012-2019 has been completely abandoned. If your video tool still defaults to square, you're solving yesterday's problem.
4. The Vietnamese Creator Economy
At 23.1% of all prompts, Vietnamese isn't just represented — it's the second most popular language by a massive margin, more than doubling the third-place Arabic at 11.4%. Vietnam's creator economy is clearly at an inflection point, and AI video tools are a key accelerator.
5. No One Wants 6-Second Videos
With only 6.6% of orders, the 6-second format is the least popular duration. Users strongly prefer either short-and-punchy (4s) or longer-form (12s). The middle ground just doesn't resonate. This mirrors what we've seen in social media trends — content is either a quick hook or a mini-narrative, with little room for in-between.
What This Means for Creators
So you've seen the data. What should you actually do with it?
Whether you're a marketer, content creator, business owner, or just someone curious about AI video, here are the actionable takeaways:
1. Start with Text-to-Video
If you haven't tried AI video yet, text-to-video is where the action is. Two-thirds of users start here, and for good reason — you don't need any assets, just ideas. Describe what you want to see, and the AI builds it.
2. Think in 4s or 12s
When planning your AI videos, think in terms of 4-second punches or 12-second stories. The data shows these are the durations that resonate. For social media hooks and ad creatives, go with 4 seconds. For product demos, explainers, and narrative content, use the full 12.
3. Choose Your Orientation Deliberately
Don't default to landscape. If your content is heading to TikTok, Reels, or Shorts, go 9:16 vertical. If it's for YouTube, your website, or presentations, go 16:9. And forget about square — the market has moved on.
4. Don't Sleep on Non-English Markets
If you're building a business around AI video content, the data shows massive demand from Vietnamese, Arabic, Russian, and Turkish-speaking markets. These aren't niche audiences — they represent hundreds of millions of potential viewers.
5. Use Image-to-Video for Product Content
While text-to-video dominates overall, image-to-video is the secret weapon for e-commerce and product marketing. Upload your product photo and add motion, context, and life. It's faster than a photoshoot and infinitely more scalable.
6. Veo 3.1 Is the Safe Bet
If you're wondering which model to use, the data is clear: 96.4% of users choose Veo 3.1. It offers the best combination of quality, speed, and prompt adherence. Start there, and experiment with alternatives like Sora 2 for specific creative styles.
The bottom line: AI video isn't a novelty anymore. With 120,000+ videos generated, prompts in 24+ languages, and use cases spanning from birthday cards to real estate tours, it's a mainstream creative tool. The question isn't whether to use it — it's how to use it better than everyone else.
Ready to see what you can create? Try Vivideo free and add your prompts to the next dataset.
Explore More
Ready to Create Your Own AI Videos?
Try Vivideo free today - no credit card required. Create professional videos in minutes.
