Google Deepmind argues video generators already contain the world models computer vision has been missing

Digital artwork: an urban street scene featuring buildings and a sidewalk, overlaid with colorful geometric AI overlays.

Google Deepmind’s GenCeption repurposes a video generator for classic vision tasks such as depth estimation and segmentation, matching state-of-the-art systems with far less training data. The model trained almost entirely on synthetic videos. Its results add to the debate over whether video generators already contain a kind of universal world model.

The article Google Deepmind argues video generators already contain the world models computer vision has been missing appeared first on The Decoder.