Skip to main content
AI clinician from Google DeepMind competes in blind medical tests against GPT-5.4, outperforming AI but still trailing human

Editorial illustration for Google DeepMind AI co‑clinician beats GPT‑5.4 in blind tests, lags docs

Google DeepMind AI co‑clinician beats GPT‑5.4 in blind...

Updated: 4 min read

Medical AI has one job: get the answer right. A new system from Google DeepMind mostly does, which makes its single, critical failure that much more important.

In a blind test, physicians preferred its responses on 98 real-world primary care questions over a current clinical AI and a search-augmented GPT-5.4. The margins were solid. On medication specifics, its lead grew.

Then there was the error. One wrong answer out of ninety-eight. That's the line.

On a benchmark of 600 pharmacist-vetted drug questions, the AI scored 73.3%. Doctors with reference books got 61.3%. Without those books, they managed only 48.3%.

When questions were open-ended, mimicking a rushed clinic lookup, the AI's quality score hit 95.0%, nudging past an OpenAI model at 90.9%. The machine wins on volume and speed. It loses on the one thing that can't be missed.

In a blind comparison using 98 realistic primary care queries, doctors consistently picked the AI co-clinician's answers over leading evidence synthesis tools. It won 67 to 26 against an existing clinical AI system and 63 to 30 against GPT-5.4-thinking-with-search. In the objective analysis, the system logged a critical error in one of the 98 cases.

The lead was even bigger on medication questions. The RxQA benchmark covers 600 questions on active ingredients, interactions, and dosages, drawn from national drug directories in two countries and vetted by licensed pharmacists. These questions are tough for primary care doctors: with reference books, they got 61.3 percent right, and just 48.3 percent without.

The AI co-clinician scored 73.3 percent, just ahead of GPT-5.4-thinking-with-search at 72.7 percent. The gap widened when questions were asked open-ended rather than as multiple choice, the way doctors actually look things up on the job. Here the AI co-clinician hit a quality score of 95.0 percent, compared to 90.9 percent for OpenAI's model.

These results sketch a very specific, very useful role. This isn't an autonomous clinician. It's a formidable reference tool, one that can clearly outperform a human scrambling without resources.

The trust required for a doctor to use it, however, is eroded by that one critical mistake. Perfection is the only acceptable standard when the subject is a prescription. So the tool advances, but the hierarchy remains.

A doctor with a book still beats the AI. The goal is to get the doctor to trust the machine enough to use it as the book.

Common Questions Answered

How did Google DeepMind's AI co-clinician perform compared to GPT-5.4 in the blind test?

In blind tests evaluating 98 real-world primary care questions, physicians preferred Google DeepMind's AI co-clinician responses over both a current clinical AI system and search-augmented GPT-5.4 with solid margins. The system demonstrated superior performance across multiple clinical scenarios, particularly in medication-related decision support.

What is the critical limitation of Google DeepMind's medical AI system despite its strong performance?

Despite outperforming other AI systems in blind tests, the Google DeepMind AI co-clinician has one critical failure that undermines trust in medical applications. When the subject involves prescriptions and medical decisions, perfection is the only acceptable standard, and this single mistake erodes physician confidence in relying on the tool.

What role does Google DeepMind position its AI co-clinician to play in clinical practice?

Google DeepMind describes the AI co-clinician as a formidable reference tool rather than an autonomous clinician, designed to assist physicians who lack immediate access to resources. The system can clearly outperform a human clinician working without reference materials, but it remains subordinate to a doctor with proper resources and clinical judgment.

Why is physician trust essential for the adoption of Google DeepMind's medical AI?

Physician trust is critical for adoption because the AI's reliability directly impacts patient safety and clinical outcomes in medical decision-making. The presence of even a single critical error in prescription-related decisions creates hesitation among doctors, requiring the AI to demonstrate near-perfect accuracy before it can be fully integrated into clinical workflows.

LIVE21:39Google DeepMind Demos AI Orchestrating Boston Dynamics Spot Robot