The most reliable camera instruction for an AI video take is a single committed move, described once, that runs for the whole generation. Most unstable, drifting or seasick output is not a model limitation. It is a prompt asking for a push in, a pan, an orbit and a focus change inside the same thirty seconds, which gives the model four objectives that cannot all be satisfied and no guidance about which to prioritise.
Choosing one camera idea and holding it is a discipline, not a compromise. It is also how most real coverage is shot.
One Idea, Held
A camera idea is a single spatial intention: the frame gets closer, the frame follows, the frame moves laterally past the subject, the frame circles, or the frame stays put. Pick one and let it run.
A slow push in that begins at a medium wide and ends at a medium close is one idea. A handheld follow behind a character walking a corridor is one idea. A locked frame where the subject moves through it is one idea, and it is the most underrated of the set, because it puts all the motion burden on the performance where the model handles it better.
Describe the Move by Start and End, Not by Verb
Push in tells the model a direction but not a distance. Begins at a medium wide from the doorway and ends at a medium close on her face tells it where it starts, where it finishes, and therefore how fast it has to travel. That last part is what produces smooth motion, because the pace is derived rather than guessed.
The same applies to lateral moves. State the start position, the end position, and what stays in frame throughout. A move with a stated anchor stays stable. A move with only a direction drifts.
Match the Move to the Beat
Camera should have a reason. A push in builds pressure and works for discovery and realisation beats. A follow creates urgency and suits pursuit or escape. A locked frame creates unease when something should be happening and is not. An orbit reveals space and works when the environment is part of the story, which in microdrama is rarely.
If a move cannot be justified by the beat, the correct choice is usually to lock the camera and let the actor carry it. Unmotivated movement reads as instability, not as style.
Vertical Framing Changes the Maths
A nine by sixteen frame has very little horizontal room, which means lateral moves and orbits lose their subject quickly and wide establishing moves show mostly floor and ceiling. Vertical microdrama favours push ins, pull outs, follows and locked frames, in roughly that order of usefulness.
It also favours tighter shot sizes generally. A medium wide in vertical carries about as much information as a wide in landscape, so a move that ends on a medium close is often ending in the right place rather than being too tight.
Worked Example: Rescuing a Drifting Take
A producer had a corridor confrontation that kept coming back unusable. The camera direction read: dynamic camera movement, sweeping cinematic shot, dramatic angles, push in on her face with a slow pan across the corridor and a rack focus to him in the background.
That is four moves. Every generation attempted all of them and completed none, which produced the wobble.
The rewrite: a single slow push in. Begins at a medium wide with both characters in frame, her in the foreground on the left, him twelve feet behind her. Ends at a medium on her shoulders and face. He stays visible over her shoulder throughout. Camera height stays at her eye level.
One idea, a stated start, a stated end, one anchor that must remain in frame, and a fixed height. The take came back stable on the second attempt, and the rack focus that had been requested explicitly happened naturally as a by product of the push, because depth behaves correctly when the move is coherent.
Common Camera Prompt Mistakes
Stacking moves. The primary cause of unstable output.
Using cinematic as a camera instruction. It specifies nothing and often triggers arbitrary movement.
Changing the camera at every timestamp. Timing controls action. Let the camera run across the beats.
Omitting camera height. Unstated height drifts, and drifting height is the most noticeable instability of all.
Requesting handheld without a reason. Handheld adds motion noise that competes with the performance unless the scene needs the energy.
Moving during an impact. A physical beat should be held on, not travelled through.
Shot Size Is a Camera Decision Too
Producers spend most of their camera direction on movement and almost none on size, which is inverted. Shot size carries more meaning per decision than movement does, and it is far more reliably generated.
State the size at the start of the take and at the end if it changes. Medium wide, medium, medium close, close. Avoid extreme close ups on faces in generation, since fine facial detail at that scale is where quality drops fastest, and avoid true wides in vertical, where they show mostly empty frame.
Camera Height and Angle Carry Power
Height is the most consistently omitted camera instruction and one of the most consequential. Below eye level makes a character dominant, above makes them vulnerable, and eye level makes them equal. In a confrontation those three choices are the whole subtext.
They also stabilise the take. An unstated height drifts across a long generation, and drifting height is far more noticeable to a viewer than a slightly imperfect move, because it changes the emotional reading of the shot midway through.
Lens Character Without Numbers
Asking for a specific focal length or aperture is unreliable across models. Describing the behaviour those numbers produce is much more dependable. A compressed background with the subject clearly separated. A wider field that keeps the room visible around her with everything reasonably sharp. Slight distortion at the frame edge because the camera is close.
These descriptions travel across models and across versions, which matters in 2026 when the model you are using may be replaced by a newer one before the season finishes shooting.
Building a Camera Grammar for a Show
Individual shot decisions made in isolation produce a show that looks different every episode. A camera grammar is a small written rulebook that fixes those decisions in advance: which move belongs to which beat type, what height each character is shot at, how tight the show goes, and what is never used.
A workable microdrama grammar can be four or five lines long. Discovery beats push in. Confrontations sit locked at eye level. Reactions are medium close and static. Exits pull back. No handheld, no orbits.
Restrictions of that kind look limiting written down and feel liberating in production, because they remove a decision from every single shot and they guarantee the season holds together visually even when different people are producing different episodes.
Let the Subject Move Instead of the Camera
The most reliable way to get motion into a shot without instability is to keep the camera still and move the subject through the frame. A character walking toward a locked camera produces the same increase in intimacy as a push in, with none of the drift risk, because all the motion is being generated by the performance rather than by the frame.
This is standard practice in low budget live action for the same reason, and it transfers directly. Entrances, exits, sitting, standing and turning all change the composition without asking the model to solve a spatial problem at the same time as everything else.
A useful discipline for a first pass on any scene is to lock every camera and let blocking do the work. Add movement afterwards only to the shots that clearly need it, which will usually be fewer than expected.
Frequently Asked Questions
Why does my camera drift when I did not ask for movement?
Because a locked frame has to be requested explicitly. State that the camera is static and that its height and position do not change.
Can I use two camera moves in a thirty second take?
Occasionally, if they are sequential and separated by a clear beat, such as a hold followed by a push. Simultaneous moves rarely work.
What is the safest camera choice for dialogue?
A locked medium or a very slow push. Dialogue scenes fail on performance, not on camera, so the camera should get out of the way.
How do I get shallow depth of field reliably?
Describe the lens character and the subject distance rather than naming an aperture. A tight lens on a subject well separated from the background produces it more consistently than a numeric request.
Do camera instructions work better at the start or end of a prompt?
Keep them in their own clearly delimited section. Position matters less than separation, because camera direction mixed into action description tends to be read as action.
How do I keep camera grammar consistent across episodes?
Write it down once as a fixed set of allowed moves per beat type and reuse it, rather than deciding shot by shot. Consistency of grammar is a large part of what makes a season feel like one show.
Start Directing Instead of Guessing
Every technique on this page gets easier when the take, the references and the continuity rules live in one place instead of scattered across tabs and text files. That is what Vertex is built for. Create a free account at MinionArts and start building your first microdrama scene today.




