The Register · Thomas Claburn ·

Netflix debuts VOID, a vision language model that can erase objects from a scene and simulate how remaining objects would behave in the scene without them

Video-language model revises how objects interact when things get removed from a scene

In this story Netflix VOID
Netflix debuts VOID, a vision language model that can erase objects from a scene and simulate how remaining objects would behave in the scene without them

Lead Source

More

Tom's Guide: Tom's Guide
MobileSyrup: MobileSyrup

Discussion

TechSnif Coverage

Netflix Built an AI That Erases Objects and Rewrites Physics

Netflix's new VOID model can remove objects from video scenes and simulate how everything else would realistically behave without them.

Netflix just dropped VOID, a vision-language model that does something genuinely wild: it erases objects from a scene and then figures out how the remaining objects would physically behave without them.

This isn't just fancy Photoshop-style removal. VOID understands how objects interact. Pull a table out from under a vase, and the model simulates the vase falling. Remove a wall, and it recalculates lighting and shadows. The system revises the entire physical logic of a scene based on what's been taken away.

The implications for filmmaking and post-production are significant. Instead of expensive reshoots or painstaking manual VFX work, editors could potentially strip and reconstruct scene elements with a text prompt.

Netflix is positioning this as a tool that could fundamentally change how movies and shows get made. Whether studios actually adopt it at scale remains to be seen.