How to Keep Characters Consistent in AI Video
Character drift is the first thing viewers notice in AI video. The narrator looks forty in scene three and twenty-five in scene nine. His tunic changes colour. The shopkeeper he met in the morning is a different man by the afternoon. Each shot on its own might look fine, but the story falls apart because the audience stops believing these are the same people.
This guide explains why that happens and what a working method looks like.
Why AI characters drift between shots
Image and video models do not remember anything between generations. Every new shot is a fresh guess based on whatever you gave it. If you gave it a paragraph of text, you gave it a range of possible people, and it picks one.
Three things make drift worse over a long video:
- More scenes. A 60-second short has maybe ten shots. A ten-minute episode can have eighty or more. Each one is another chance to drift.
- Different angles. A face described from the front says nothing reliable about the profile or the back of the head.
- Different contexts. The same character at night, in rain, or in a crowd pushes the model toward different defaults.
The fix is not a better adjective. It is a fixed reference that every shot is measured against.
Step 1: Define each character once, visually
Before you draw a single scene, build a character sheet for every person who appears more than once. A useful sheet includes:
- Front, side and back views of the character in the same neutral pose and lighting.
- Standard wardrobe, drawn on the character, not just described.
- Distinguishing features written down plainly: a scar, a missing tooth, a particular belt.
- Age, build and height relative to other characters, so group shots stay believable.
Do this once, approve it, and then stop changing it. If a character needs a second look (a work outfit and a festival outfit, for example), treat that as a named wardrobe variant on the same person, not a new character.
Step 2: Keep one entry per person
A common mistake is letting the same person exist twice: “the landlord” in one scene and “the lodging-house keeper” in another. The model has no idea they are the same man. Merge duplicates early so every mention of a person points to one reference.
The same rule applies to places and objects. One street should have one reference, drawn from different angles, rather than three slightly different streets.
Step 3: Write each scene against the narration
Consistency is not only about faces. If the narration says the character is tired, carrying a basket, and walking home at dusk, the shot description needs to say that too. When the description and the narration disagree, you get a correct-looking character doing the wrong thing, which reads as inconsistency to the viewer.
A good shot description names:
- which characters are in frame, by their reference name
- what each is doing and holding
- where they are, by location name
- the time of day and camera framing
Step 4: Check every still before you animate
Generate a few still-image candidates for each scene and check them against the references before anything moves. Check for:
- Identity: is this the same face, build and age as the sheet?
- Wardrobe: same garments, same colours, nothing added or missing?
- Props: is the character holding what the scene says, in the right hand?
- Location: does the background match the place reference?
- Framing: is the shot size and angle what the description asked for?
This is the cheapest place to catch drift. A still costs far less than a video clip, and a wrong face caught here never reaches the timeline.
Step 5: Animate from the approved still
If you generate video from text, the video model makes its own guess about the character and you lose everything you checked. Animating from the approved still keeps the face, costume and set anchored to a picture you already accepted.
Write the motion prompt from the still itself: what is in that picture, and what moves. Keep it small. Big camera moves and complicated action give the model room to reinvent things.
Step 6: Review the whole film, not just the shots
Drift that is invisible shot by shot often shows up when you watch scenes in order. A slightly different hair colour in two adjacent scenes is obvious at the cut. Review the full sequence with narration before you call it done.
A consistency checklist
- Every recurring character has front, side and back views
- Wardrobe is drawn, not only described
- Duplicate names for the same person are merged
- Each location has one reference with its fixed features listed
- Every shot description is checked against the narration
- Every still is checked for identity, wardrobe, props, location and framing
- Every clip is animated from its approved still
- The full film is reviewed in order with the narration
What is still hard
Be realistic about the limits. Hands, close physical interaction between two people, and crowded multi-person action are the hardest cases for current models. Shots with three or more named characters doing different things are where drift is most likely. A human still needs to watch the final film and make the call.
How VIA handles this
VIA is built around this method. In Setup you create the production bible once: character sheets with front, side and back views, wardrobe, locations with fixed features, recurring objects and creative rules. In Scenes, every shot description is checked against the narration by an independent model before an image is drawn, and stills are judged for identity, location, props and framing. Clips are animated from the approved still. In Film you review the whole video on a timeline with the narration.
In “24 Hours in Ancient Rome”, the narrator Marcus and four recurring characters (a lodging-house keeper, a bathhouse attendant, a dice gambler and a wealthy-house doorkeeper) appear across 83 scenes. You can see how that was set up in the case study, or read how it works.
If you are making narrated long-form video and character drift is what keeps breaking it, you can apply for the private beta.
Questions
Why does my AI character look different in every shot?
Each generation starts from a text prompt, and text describes a face loosely. Without a fixed visual reference, the model invents a slightly different person every time, and small changes add up across dozens of scenes.
Is a detailed text description enough to keep a character consistent?
No. Words like 'short dark hair, mid-thirties' match thousands of faces. A written description helps, but consistency comes from visual references plus a check that compares each new image against them.
How many reference images does a character need?
At minimum a front, side and back view with the character's standard wardrobe. Those three views cover most camera angles a scene will ask for, including over-the-shoulder and walking-away shots.
Can the video clip change a character even if the still was right?
Yes, if the clip is generated from text alone. Animating from the approved still keeps the face and costume in the clip anchored to the picture you already checked.