Researchers at IIIT-Hyderabad have cautioned against blindly trusting medical AI models after auditing four vision-language models on chest X-rays. The study found that diagnostic information could influence image localisation, while radiologists preferred broader areas around abnormalities in some cases.
Updated On – 8 October 2026, 07:01 PM
Hyderabad: Medical AI tools are increasingly being used in medical imaging, where radiologists rely on them as a second pair of eyes to flag possible abnormalities in scans. However, a study by International Institute of Information Technology (IIIT) – Hyderabad researchers cautioned against blindly trusting output from such medical AI models.
The research team led by Prof. Parameswari Krishnamurthy, with Dr. Syed Faizan as principal investigator, at the Language Technologies Research Centre (LTRC), audited four vision-language models — MAIRA-2, MedGemma-4B, LLaVA-Med-1.5 and LLaVA-1.5 — on thousands of publicly available chest X-rays.
The team also conducted a reader study in which two radiologists assessed anonymised overlays, providing a human reference for the models’ localisation.
According to Dr. Faizan, an AI model may appear to highlight the correct part of an image, but that does not necessarily mean it has identified the disease in the same way a radiologist would. “The model may first arrive at a diagnosis and then use that diagnosis to determine where to place its heat map. In other words, it may be working backwards,” he said.
To test this possibility, the team removed the diagnostic information and examined what happened to the models’ localisation. The models’ performance dropped, suggesting that some of what looked like image-based reasoning could actually be influenced by an anatomical expectation learned from the diagnosis.
The researchers also had a 2-radiologist study, like a human-in-the-loop, to calibrate their findings and discovered that they rated MedGemma higher even though the audit had ranked Myra to be superior in overlapping heatmaps.
“This suggests that a model paying attention to a particular area – which may be the exact spot where the disorder lies – is not exactly helpful to a radiologist. A radiologist might want to look at not only areas with the disease, but also a broader surrounding area to know the extent of disease spread. MedGemma is doing that; it provides a broader area,” he explained.
The study was accepted at MICCAI 2026 (International Conference on Medical Image Computing and Computer Assisted Intervention) in Strasbourg at the iMIMIC Satellite Event.
The LTRC is also studying how reliably medical vision-language models respond when doctors phrase the same question differently. In collaboration with a radiologist from CMC Vellore, researchers examined the models’ responses to different paraphrases of medical questions, with the study accepted at EMNLP.

