TY - GEN
T1 - Evaluating the clinical plausibility of generative AI-Synthesized imaging
T2 - Medical Imaging 2026: Computer-Aided Diagnosis
AU - Bhattacharya, Moinak
AU - Pablo Garcia Camargo, Juan
AU - Chaudhry, Mohammad
AU - Prasanna, Prateek
AU - Singh, Gagandeep
N1 - Publisher Copyright:
© COPYRIGHT SPIE.
PY - 2026/4/2
Y1 - 2026/4/2
N2 - Generative artificial intelligence (AI) has advanced rapidly in medical imaging, enabling realistic synthesis of MRI, CT, and X-ray scans for applications such as data augmentation and privacy-preserving sharing. Yet, evaluating the clinical accuracy of generated images remains challenging, as conventional metrics like SSIM and FID measure pixel-level similarity but fail to capture diagnostic fidelity. We present a structured evaluation framework that incorporates radiologists' expertise across three domains: Anatomical fidelity, pathology plausibility, and overall image quality. Using two cohorts (multiple modalities and disease types)-50 patients from MIMIC-CXR and 20 from BraTS-we compared anatomically guided generative models (RadGazeGen, BrainMRDiff) against state-of-The-Art baselines (Stable Diffusion, ControlNet, MultiControlNet, DDPM). Quantitative analysis showed that RadGazeGen achieved the highest SSIM for chest X-rays (0.484 ± 0.060) and BrainMRDiff outperformed tumor-mask ControlNet for brain MRI (0.326 ± 0.074 vs. 0.215 ± 0.070). Radiologist scoring confirmed these results: our methods consistently achieved superior ratings in anatomy (3.5-3.7/4), pathology (3.3-3.7/4), and image quality (1.8-2.0/2). Baseline models, while visually plausible, frequently misrepresented anatomy or pathology, underscoring the limitations of computational metrics alone. Our findings establish that radiologist-informed, rubric-based evaluation provides a clinically grounded benchmark, ensuring generative models align with diagnostic priorities.
AB - Generative artificial intelligence (AI) has advanced rapidly in medical imaging, enabling realistic synthesis of MRI, CT, and X-ray scans for applications such as data augmentation and privacy-preserving sharing. Yet, evaluating the clinical accuracy of generated images remains challenging, as conventional metrics like SSIM and FID measure pixel-level similarity but fail to capture diagnostic fidelity. We present a structured evaluation framework that incorporates radiologists' expertise across three domains: Anatomical fidelity, pathology plausibility, and overall image quality. Using two cohorts (multiple modalities and disease types)-50 patients from MIMIC-CXR and 20 from BraTS-we compared anatomically guided generative models (RadGazeGen, BrainMRDiff) against state-of-The-Art baselines (Stable Diffusion, ControlNet, MultiControlNet, DDPM). Quantitative analysis showed that RadGazeGen achieved the highest SSIM for chest X-rays (0.484 ± 0.060) and BrainMRDiff outperformed tumor-mask ControlNet for brain MRI (0.326 ± 0.074 vs. 0.215 ± 0.070). Radiologist scoring confirmed these results: our methods consistently achieved superior ratings in anatomy (3.5-3.7/4), pathology (3.3-3.7/4), and image quality (1.8-2.0/2). Baseline models, while visually plausible, frequently misrepresented anatomy or pathology, underscoring the limitations of computational metrics alone. Our findings establish that radiologist-informed, rubric-based evaluation provides a clinically grounded benchmark, ensuring generative models align with diagnostic priorities.
KW - Anatomy
KW - Generative AI
KW - Radiologist evaluation
UR - https://www.scopus.com/pages/publications/105041095456
U2 - 10.1117/12.3088157
DO - 10.1117/12.3088157
M3 - Conference contribution
AN - SCOPUS:105041095456
T3 - Progress in Biomedical Optics and Imaging - Proceedings of SPIE
BT - Medical Imaging 2026
A2 - Wismuller, Axel
A2 - Deserno, Thomas Martin
PB - SPIE
Y2 - 15 February 2026 through 19 February 2026
ER -