Introducing Generative Video Editing

September 2, 2026
by Vishal Kumar

Today we're introducing Generative Video Editing on Caira. It starts with a question: what happens when a camera understands the scene it just captured?

For a hundred and twenty years, cameras have done one thing: record light, then stop. Everything after the shutter belonged somewhere else — the darkroom, the desktop, the post house. The camera saw the scene better than any device on earth, and then handed the hardest work to machines that had never seen it at all.

That handover has a cost, and every filmmaker knows it. An idea must survive a series of translations before anyone else can see it: from imagination to treatment, from treatment to budget, from budget to specialist software in a specialist's hands. Each step loses something. A written treatment is a lossy compression of the film in your head — and when a single VFX pass costs more than the commission pays, most ideas don't survive the journey. They don't fail as films. They fail as translations.

World models collapse that distance. Unlike earlier generative systems that edit one frame at a time, a world model builds an internal understanding of the scene itself — its geometry, lighting, materials, and motion. Google's Gemini Omni 1.1 Flash is physics-aware and remembers the scene across an entire clip, so footage can be relit or extended realistically while people, objects, and camera movement stay consistent from frame to frame. Paired with Runway's Aleph 2.0, this understanding now lives where the scene does: at the point of capture.

Generative Video Editing launches with Relight, Reframe, Camera Motion, and Set Extension, working through a curated library of templates designed with feedback from our community, refined further through conversational editing in plain English. Shoot a spec sequence in the morning; show a commissioner three treatments the same day. When the pitch lands, shoot the real production with a full crew. The idea finally travels from your head to their screen without dying in translation.

Left: the original clip captured with Caira in 4K. Right: a frame from the Omni-generated camera move of the same scene.
Camera Motion, via the Orbit template: the original clip captured with Caira (left) and a frame from the generated camera move of the same scene (right) — no drone, no second take.

Some things we designed deliberately, and they're worth stating plainly. There is no text-to-video on Caira: every edit begins from footage you captured, because a scene cannot — and on this camera, should not — be generated from nothing. Generated results are saved as separate, clearly named files, watermarked with Google's SynthID. Your originals are never modified, your footage is never used for model training, the on-camera pipeline remains non-generative, and the entire feature is optional.

The Caira app showing the Day to Night template with the instruction 'Turn the scene into golden hour'
How editing works on Caira: a curated template (here, Day to Night) steered in plain English — “Turn the scene into golden hour” — never an open prompt box.

We also know why we built this. During our Kickstarter we shipped generative photo editing, and learned a great deal from the conversation that followed. So before building the video version, we surveyed 111 of our backers: 77% told us generative video editing would be useful to their work. We build what the people using our cameras ask for — and we'll keep asking, monitoring how these tools are used and whether they genuinely help.

There are strong emotions around AI in filmmaking right now, and we respect the craft at the centre of them. Traditional filmmakers and VFX artists may prefer their existing workflows, and nothing about Caira asks them to change. This is for the solo creator and the small team — the people overwhelmed by professional tools but full of ideas that deserve to be seen before a budget decides their fate.

Generative interfaces, generative interfaces for images, and now scene-aware video: each generation of these models gets faster, more coherent, more controllable. We expect the same trajectory here. But the premise stays fixed: the camera still has to be pointed at something. What changes is how far what you point it at can now travel.

Caira camera held in a hand, shallow depth of field
Caira in hand. The camera still has to be pointed at something.

Generative Video Editing arrives as part of Caira Studio this autumn, in the US and UK. A private beta opens today for existing customers and 5 working filmmakers → cameraintelligence.com/generative-video-editing