Skip to main content
Runway and Google DeepMind logos with AI video generation interface, demonstrating real-time interactive capabilities.

Editorial illustration for Runway and Google DeepMind Pursue Real-Time, Interactive AI Video Generation

Runway and Google DeepMind Pursue Real-Time, Interactive...

3 min read

Runway published research this week outlining plans to make AI video generation happen in real time, more like a live stream than a request you send and wait on. The company's current models, like most in the field, work in discrete steps: type a prompt, wait several seconds to several minutes, get a clip back. Get it wrong, and you start the whole cycle over. Runway says its users keep flagging that generating and revising is where most of their time disappears.

The company laid groundwork for this in March with Runway Characters, and its December release of GWM-1, described as its first "General World Model," pushes further in that direction. GWM-1 builds on Gen-4.5 and generates video frame by frame, taking camera moves, robot commands, or audio as live controls. Runway's recent Solaris project applied a similar frame-by-frame approach to user interfaces that respond to clicks and voice.

Now the company is framing real-time generation as the next step, one aimed at cutting both the wait between idea and result and the GPU costs that come with slower, batch-style rendering.

Runway argues that real-time generation closes the gap between an idea and its execution. With instant feedback, users would spend most of their time actively steering the video rather than waiting.

Why this matters

Both companies are betting that the next fight in generative video isn't about fidelity, it's about latency. If Runway is right that revision cycles waste more time than generation itself, then a model you can steer live, frame by frame, changes the economics of production work rather than just the aesthetics. Genie 3's few-minutes-at-24fps benchmark is a real number to watch, not a marketing line: it tells us how far interactive world models still have to go before they're useful for anything beyond a demo reel.

For developers and founders building on top of these systems, the shift Runway describes, from compute spent training to compute spent at inference, means cost structures for video products are about to look a lot more like cloud gaming than like batch rendering. That's a different infrastructure bet, and a different pricing headache. We'd temper the excitement with a plain question: consistency over minutes is not consistency over hours, and nobody here has shown the latter yet.

Worth tracking who closes that gap first.

LIVE14:09Runway and Google DeepMind Pursue Real-Time, Interactive AI Video Generation