AI Chatbots Show Cognitive Decline in BMJ MoCA Study
Why in the news
A study in the BMJ tested top AI chatbots with a human cognition test and found signs of decline in older model versions, casting doubt on using them for medical diagnosis.
Key facts
- Models tested: ChatGPT, Claude, Gemini (large language models).
- Tool: the MoCA test, normally used on people, adapted to probe attention, memory and executive function.
- Older versions scored lower, like human cognitive decline.
- Scores: Gemini 1.0 = 16 (lowest); ChatGPT 4 = 26; full marks are 30, and nobody got them.
- Weak spots: visuospatial skills and executive tasks; most scored under the 26 cut-off.
Significance
- Such deficits challenge the reliability of AI for medical diagnostics and the claim that AI will soon replace doctors.
- AI can help, but may not replace human expertise in areas such as neurology.
Exam angle
- Journal: BMJ; test: MoCA; pass mark: 26 out of 30.