Abstract
Medical Visual Question Answering (Med-Vqa) is a challenging task that requires a deep understanding of both medical images and textual questions. Although recent works leveraging Medical Vision-Language Pre-Training (Med-Vlp) have shown strong performance on the Med-Vqa task, there is still no unified solution for modality alignment, and the issue of hard negatives remains underexplored. Additionally, commonly used knowledge fusion techniques for Med-Vqa may introduce irrelevant information. In this work, we propose a framework to address these challenges through three key contributions: (1) a unified solution for heterogeneous modality alignments across multiple levels, modalities, views, and stages, leveraging methods like contrastive learning and optimal transport theory; (2) a hard negative mining method that employs soft labels for multi-modality alignments and enforces the hard negative pair discrimination; and (3) a Gated Cross-Attention Module for Med-Vqa that integrates the answer vocabulary as prior knowledge and selects relevant information from it. Our framework outperforms the previous state-of-the-art on widely used Med-Vqa datasets like RAD-VQA, SLAKE, PathVQA and VQA-2019. The code is available at https://github.com/AlexCo1d/AMiF
| Original language | English |
|---|---|
| Pages (from-to) | 29623-29633 |
| Number of pages | 11 |
| Journal | Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition |
| DOIs | |
| State | Published - 2025 |
| Event | 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, United States Duration: Jun 11 2025 → Jun 15 2025 |
Keywords
- medical imaging
- medical visual question answering
- representation learning
Fingerprint
Dive into the research topics of 'Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver