Retirement Life
29 July 2026
Researchers warn AI chatbots often wrong on medical questions
If you’re tempted to treat AI like a stand in doctor, a new study suggests caution. Researchers testing five major chatbots have found that although their answers often sound convincing, about half were wrong.
A team of seven researchers tested five of the world’s most popular chatbots - ChatGPT, Gemini, Grok, Meta AI and DeepSeek - asking each 50 health and medical questions. Their results were recently published on BMJ Open.
In a recent article from the Conversation, Carsten Eickhoff, Professor of Medical Data Science at the University of Tübingen, says the questions covered everything from cancer, vaccines, stem cells, nutrition and athletic performance, with two experts then independently rating each answer.
“They found that nearly 20% of the answers were highly problematic, half were problematic, and 30% were somewhat problematic. None of the chatbots reliably produced fully accurate reference lists, and only two out of 250 questions were outright refused to be answered,” Professor Eickhoff says.
The five chatbots performed roughly the same, with Grok the worst performer, with 58% of its responses flagged as problematic, ahead of ChatGPT at 52% and Meta AI at 50%. However, Professor Eickhoff says performance varied by topic, with some topics handled better than others.
Planning what to do with your KiwiSaver in retirement?
“They stumbled most on nutrition and athletic performance, domains awash with conflicting advice online and where rigorous evidence is thinner on the ground,” he says.
Open-ended questions were particularly problematic, an important distinction as most real-world health queries are open-ended, Professor Eickhoff says.
“People do not ask chatbots neat true-or-false questions. They ask things like: ‘Which supplements are best for overall health?’
Why chatbots get things wrong
“There’s a simple reason why chatbots get medical answers wrong,” Professor Eickhoff says.
“Language models do not know things. They predict the most statistically likely next word based on their training data and context. They do not weigh evidence or make value judgments. Their training material includes peer-reviewed papers, but also Reddit threads, wellness blogs and social-media arguments”
He says there is a growing body of evidence that reinforces the dangers of taking medical advice from chatbots.
“Language models do not know things. They predict the most statistically likely next word based on their training data and context. They do not weigh evidence or make value judgments."
—Professor Eickhoff
A study published in February in Nature Medicine showed that chatbots could get the right medical answer almost 95% of the time. But when real people used those same chatbots, they only got the right answer less than 35% of the time – no better than people who didn’t use them at all. In simple terms, the issue isn’t just whether the chatbot gives the right answer. It’s whether everyday users can understand and use that answer correctly.
Another study found that chatbots readily repeated and even elaborated on made-up medical terms slipped into prompts.
“These chatbots are not going away, nor should they,” Professor Eickhoff says. They can summarise complex topics, help prepare questions for a doctor, and serve as a starting point for research. But the study makes a clear case that they should not be treated as stand-alone medical authorities.
“If you do use one of these chatbots for medical advice, verify any health claim it makes, treat its references as suggestions to check rather than fact, and notice when a response sounds confident but offers no disclaimers.”
Project how much you could get a fortnight with your retirement savings:
Invest with Lifetime for a retirement income managed for living.