Skip to main content
AI-powered Vision Banana model outperforming SAM 3 and Depth Anything V3 in computer vision benchmark tests, showcasing advan

Editorial illustration for Google DeepMind's Vision Banana Outperforms SAM 3 and Depth Anything V3

DeepMind Vision Banana Beats SAM 3 in AI Perception

Google DeepMind's Vision Banana Outperforms SAM 3 and Depth Anything V3

Updated: 3 min read

Imagine a single model that looks at an image and does it all, segments objects, measures depth, reads surface angles, and does it better than the specialists built for each. That’s the leap Google DeepMind just pulled off with Vision Banana. By treating every vision task as an act of image generation, then instruction-tuning a lightweight backbone, the team has produced a model that beats SAM 3 on segmentation, Depth Anything V3 on metric depth estimation, and Lotus-2 on surface normals, all in zero-shot, all with the same set of weights.

No task-specific modules. No architectural gymnastics. Just a prompt switch.

This isn’t incremental progress; it’s a quiet declaration that visual intelligence emerges from learning how to generate, not just how to classify.

A team of Google DeepMind researchers introduced Vision Banana, a single unified model that surpasses or matches state-of-the-art specialist systems across a wide range of visual understanding tasks — including semantic segmentation, instance segmentation, monocular metric depth estimation, and surface normal estimation — while simultaneously retaining the original image generation capabilities of its base model.

Vision Banana doesn’t just beat benchmarks. It rewrites the rules of what a vision model can be. By treating every task, segmentation, depth, normals, as an act of image generation, Google DeepMind has collapsed a dozen specialized architectures into one.

The result is a system that outperforms SAM 3, Depth Anything V3, and Lotus-2 without a single task-specific module. That’s not incremental progress. That’s a blueprint.

The deeper insight is this: generative pretraining, long considered the domain of language, works just as powerfully for vision. Train a model to *create* images, and it learns to *understand* them, spontaneously, transferably, at scale. Vision Banana proves that perception doesn’t need dedicated heads or loss functions.

It needs a good decoder and the right prompt. What comes next? A future where every visual task is a single inference away.

Where the boundary between generation and understanding finally dissolves. Vision Banana is the first taste. The rest of the field will have to catch up.

Common Questions Answered

How did Vision Banana outperform specialized models like SAM 3 and Depth Anything V3?

Vision Banana achieved superior performance by leveraging image generation pretraining, which naturally develops powerful internal visual representations. Unlike purpose-built segmentation or depth estimation models, this instruction-tuned image generator demonstrated remarkable transfer learning capabilities across different visual perception tasks.

What does Vision Banana reveal about image generation pretraining?

The model shows that image generation pretraining can function as a generalist vision learner, similar to how large language models develop emergent linguistic abilities. By training on image generation, the model inherently develops sophisticated visual representations that can transfer effectively to tasks like segmentation, depth estimation, and surface normal estimation.

What makes Vision Banana's approach different from traditional specialized vision models?

Vision Banana was developed through lightweight instruction-tuning of a generative model, rather than being architecturally designed for specific perception tasks. This approach challenges traditional model development by demonstrating that generative pretraining can yield representations powerful enough to outperform specialist models without requiring task-specific architectural modifications.

LIVE20:23Cognition Buys Poke, an AI Agent for iMessage and SMS