Skip to main navigation Skip to search Skip to main content

REVEAL: Multimodal Vision–Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction

  • Seowung Leem
  • , Lin Gu
  • , Chenyu You
  • , Kuang Gong
  • , Ruogu Fang
  • University of Florida
  • Tohoku University

Research output: Contribution to journalConference articlepeer-review

Abstract

The retina provides a unique, noninvasive window into Alzheimer’s disease and dementia, capturing early structural changes through morphometric features, while systemic and lifestyle risk factors reflect well-established contributors to AD and dementia susceptibility long before clinical symptom onset. However, current retinal analysis frameworks typically model imaging and risk factors separately, preventing them from capturing the joint multimodal patterns that are critical for early risk prediction. Moreover, existing methods rarely incorporate mechanisms to organize or align patients with similar retinal and clinical characteristics, limiting their ability to learn coherent cross-modal associations. To address these limitations, we introduce REVEAL (REtinal-risk Vision-language Early Alzheimer’s Learning) that aligns color fundus photographs with individualized disease-specific risk profiles for incident AD and dementia prediction on average 8 years before diagnosis (range: 1–11 years). Because real-world risk factors are structured questionnaire data, we first translate them into clinically interpretable narratives compatible with pretrained vision-language models (VLMs). We further propose a group-aware contrastive learning (GACL) strategy that clusters patients with similar retinal morphometry and risk factors as positive pairs, strengthening multimodal alignment. This unified representation-learning framework substantially outperforms state-of-the-art retinal imaging models paired with clinical text encoders, as well as general VLMs, demonstrating the value of jointly modeling retinal biomarkers and clinical risk factors. By providing a generalizable, noninvasive approach for early AD and dementia risk stratification, REVEAL has the potential to enable earlier interventions and improve preventive care at the population level.

Original languageEnglish
Pages (from-to)1869-1889
Number of pages21
JournalProceedings of Machine Learning Research
Volume315
StatePublished - 2026
Event9th International Conference on Medical Imaging with Deep Learning, MIDL 2026 - Chientan, Taiwan, Province of China
Duration: Jul 8 2026Jul 10 2026

Keywords

  • Alzheimer’s disease and related dementia
  • Contrastive learning
  • Retinal morphometry
  • risk factors
  • Vision-language alignment

Fingerprint

Dive into the research topics of 'REVEAL: Multimodal Vision–Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction'. Together they form a unique fingerprint.

Cite this