The company reported that GPT-5.5 Instant, the free model available to everyday ChatGPT users, beat physician-written responses on accuracy, clarity, completeness, instruction-following and overall usefulness for health decisions. Researchers reached that conclusion after reviewing 3,500 ChatGPT health answers judged by a separate panel of doctors.
OpenAI released the results on June 18, and the timing matters. More than 230 million people already ask ChatGPT health questions every week, according to the company, covering everything from confusing lab reports to insurance paperwork to symptoms that show up at midnight.
The scale raises the stakes. A shaky answer about a headache barely registers. A shaky answer about chest pain could send someone down a dangerous path.
OpenAI insists GPT-5.5 Instant now catches urgent situations more reliably, asks follow-up questions when details are missing, and flags uncertainty instead of guessing with confidence. Still, the company stops short of calling the tool a diagnostic replacement for a physician, and the evaluation never touched physical exams, real treatment, or hands-on clinical decisions.
Inside the benchmark

OpenAI leans on two internal scoring systems, HealthBench and HealthBench Professional, to track progress. HealthBench alone runs through 5,000 realistic conversations built to mimic exchanges between patients, clinicians, and chatbots.
A network of 262 physicians spanning 60 countries helped write the rubric, which breaks every response into 48,562 individual grading criteria covering safety, accuracy, tone, context, and whether the model flags an emergency when one shows up.
Unlike a standard licensing exam, HealthBench skips multiple-choice questions entirely. Conversations run long, users volunteer incomplete details, and the topic sometimes shifts mid-thread, mirroring how people actually talk to a chatbot at 2 a.m.
One wrinkle deserves attention: OpenAI built the benchmark and uses another model to grade whether ChatGPT health answers meet physician-written standards. That setup produces a detailed, repeatable measurement, yet it falls short of independent clinical validation from an outside body.
HealthBench Professional narrows the focus further, testing the tasks doctors themselves bring to ChatGPT, including care consultations, medical research, and documentation. Researchers deliberately loaded the pool with harder cases, boosting difficult examples by roughly 3.5 times versus the original set, and physicians ran adversarial tests on about a third of the material.
Where the gap widened most

Completeness produced the widest split. GPT-5.5 Instant scored high marks in 81.5% of reviews, edging past the 77.3% mark physicians hit and well ahead of GPT-4o’s 55.7%.
Health decision helpfulness told a similar story. GPT-5.5 Instant landed high scores in 75.2% of reviews, compared with just 52.9% for physician-written answers and 67.2% for GPT-4o.
Those numbers suggest ChatGPT health answers now pack in more context and practical next steps than before. They don’t prove that the model reaches better diagnoses or picks safer treatments once a real patient sits in an exam room.
OpenAI also pulled data from live ChatGPT traffic rather than lab conditions. The company says flagged factual errors in ChatGPT health answers dropped 71% over two months, based on privacy-protected monitoring across billions of weekly messages. That figure tracks possible factual slip-ups only. It says nothing about recovery rates, correct diagnoses, or medication safety down the line.
A sciatica question shows the shift
OpenAI illustrated how ChatGPT health answers have changed with a real example: why would a doctor order an MRI before injecting steroids for sciatica?
An older model provided a brief answer, noting that the scan could locate the source of the pain and guide the needle. GPT-5.5 Instant went further, walking through nerve compression, injection placement, and safety concerns. and alternative treatments a patient might need instead. It also noted that not every sciatica case actually requires an MRI.
The newer response closed with a line a patient could carry straight into an appointment: “What are you looking for on the MRI, and how would the result change the injection plan?”
That question captures where chatbots fit best right now. They can sharpen what a patient asks. They shouldn’t make the final call.
Doctors still hold the line

Higher scores give people more reason to lean on ChatGPT for health answers, appointment prep, and cutting through medical jargon. Stronger ChatGPT health answers still don’t hand clinical authority to the tool.
Physicians examine patients directly, pull full histories, order tests, and track how a treatment plays out over weeks or months. A benchmark score, however impressive, can’t replicate that.
OpenAI has already moved on to its GPT-5.6 model family, announced July 9, but the company hasn’t published a matching health evaluation against physician answers for that release. For now, the medical claims rest entirely on GPT-5.5 Instant.
The dividing line for patients stays simple. Lean on ChatGPT health answers to decode a lab result, simplify medical terminology, or draft questions before an appointment. Don’t lean on them to decide whether a symptom is safe to ignore, whether to skip a medication, or whether a situation calls for urgent care.
Do better scores make you more comfortable bringing medical questions to an artificial intelligence bot, or would you rather consult a human doctor? Please post your comments below.

