Skip to main content
Google Gemini Carbon coding performance matches Anthropic Opus 5.5, showcasing advanced AI development.

Editorial illustration for Google's Gemini Carbon Reportedly Matches Anthropic Opus 5.5 in Coding

Google's Gemini Carbon Rivals Anthropic Opus 5.5

• 4 min read

Google hasn't shipped Gemini 4 "Argon" yet, and internal testing has already moved on to a model called Carbon. Business Insider reviewed documents, screenshots, and internal chats showing Google engineers running Carbon on Jetski, the company's internal coding platform, over the past several days. One employee working with the model drew a direct comparison to Anthropic's Opus 5.5, currently the strongest coding model on the market, though the comparison came with a caveat: testing isn't finished.

That's a notable jump from where Argon stood just weeks ago. Early versions of Argon reminded at least one Google employee more of Anthropic's older Opus 5 on certain coding tasks, even as overall internal feedback stayed positive. Carbon, by contrast, is reportedly beating Argon specifically on programming benchmarks, which raises questions about how Google plans to roll these variants into a public release. There's also a third name in the mix, Barium, adding to the confusion over what's a distinct model tier and what's just an internal checkpoint on the way to something else.

One employee compared Carbon to Anthropic's strongest coding model, Opus 5.5, though it still needs more testing. Early Argon versions, by contrast, reminded another employee of the older Opus 5 on some coding tasks, even though internal feedback was positive overall.

Why this matters

None of this is confirmed by Google, and that's the point worth sitting with. Business Insider's sourcing is internal chats and screenshots, not a product announcement, which means the Carbon-versus-Opus-5.5 comparison is really an internal benchmark leak, not a market signal yet. For developers picking a model today, the lesson is about cadence, not capability: Google is apparently iterating through Argon, Barium, and Carbon on a platform called Jetski before Argon, the one meant to ship, has even reached the public.

That's a lot of internal churn for a model still in testing, and it tracks with how Anthropic and OpenAI have been shipping incremental point releases rather than annual leaps. If you're building on Gemini, the practical move is to stop anchoring roadmaps to a specific named version. Treat "Gemini 4" as a moving target until Google confirms what actually ships.

The real story here isn't that Carbon beats Opus 5.5 on coding. It's that frontier labs are now testing three or more variants simultaneously before picking a winner, which should make everyone skeptical of leaked benchmarks until there's a product attached.

Common Questions Answered

What is Gemini Carbon and how does it compare to Anthropic's Opus 5.5?

Gemini Carbon is Google's internal coding model that reportedly matches Anthropic's Opus 5.5, which is currently the strongest coding model on the market. However, the comparison comes with a caveat that testing is not yet finished, and this information comes from internal documents and employee communications rather than official product announcements.

Why hasn't Google officially confirmed the Gemini Carbon coding performance claims?

The Carbon-versus-Opus-5.5 comparison is based on an internal benchmark leak from Business Insider's review of internal chats and screenshots, not a product announcement from Google. Since none of this information has been confirmed by Google officially, it represents internal testing data rather than a market signal or verified claim.

What is Jetski and how does it relate to Google's model development?

Jetski is Google's internal coding platform where engineers have been running and testing the Carbon model over the past several days. Google appears to be iterating through multiple model versions including Argon, Barium, and Carbon on this platform before official release.

How does Gemini Carbon's performance compare to earlier Argon versions?

While Carbon reportedly matches Opus 5.5's coding capabilities, earlier versions of Argon reminded some Google employees of the older Opus 5 model on certain coding tasks, despite receiving positive overall internal feedback. This suggests significant performance improvements between the Argon and Carbon iterations.

LIVE13:22Consumer AI Sees Broad Use But Little Paying, With Few Spending Big