TY - GEN
T1 - Audio-visual Feature Fusion for Improved Thoracic Disease Classification
AU - Bhattacharya, Moinak
AU - Prasanna, Prateek
N1 - Publisher Copyright:
© 2023 SPIE.
PY - 2023
Y1 - 2023
N2 - In this work, we fuse imaging features from Chest X-Ray (CXR) scans and audio features from dictations of a radiologist to improve thoracic disease classification. Recent deep learning-based disease classification methods mostly use imaging modalities. Dictation audio from a radiologist contains rich auxiliary disease-related contextual information. The main hypothesis of this proposed work is that leveraging complementary imaging and audio representations improves disease classification. We use shifting window (Swin) transformer architectures as encoders for both visual and audio modalities and finally fuse the feature representations using cross-correlational feature multiplication fusion strategy. This fused feature representation is fed to a classification head for downstream disease classification. We experimentally show that the proposed fused model outperforms the individual modality models for multi-class thoracic disease classification that includes normal, pneumonia, and congestive heart failure cases. We report F1-score of 0.5415 and 0.5353 for shifting window transformer base and small architectures respectively, for fused modalities, while the corresponding baselines are reported at 0.5046 and 0.5076 for the audio modality and 0.4676 and 0.5261 for the imaging modality, respectively.
AB - In this work, we fuse imaging features from Chest X-Ray (CXR) scans and audio features from dictations of a radiologist to improve thoracic disease classification. Recent deep learning-based disease classification methods mostly use imaging modalities. Dictation audio from a radiologist contains rich auxiliary disease-related contextual information. The main hypothesis of this proposed work is that leveraging complementary imaging and audio representations improves disease classification. We use shifting window (Swin) transformer architectures as encoders for both visual and audio modalities and finally fuse the feature representations using cross-correlational feature multiplication fusion strategy. This fused feature representation is fed to a classification head for downstream disease classification. We experimentally show that the proposed fused model outperforms the individual modality models for multi-class thoracic disease classification that includes normal, pneumonia, and congestive heart failure cases. We report F1-score of 0.5415 and 0.5353 for shifting window transformer base and small architectures respectively, for fused modalities, while the corresponding baselines are reported at 0.5046 and 0.5076 for the audio modality and 0.4676 and 0.5261 for the imaging modality, respectively.
KW - audio-visual
KW - disease classification
KW - modality fusion
KW - transformers
UR - https://www.scopus.com/pages/publications/85160205676
U2 - 10.1117/12.2654571
DO - 10.1117/12.2654571
M3 - Conference contribution
AN - SCOPUS:85160205676
T3 - Progress in Biomedical Optics and Imaging - Proceedings of SPIE
BT - Medical Imaging 2023
A2 - Iftekharuddin, Khan M.
A2 - Chen, Weijie
PB - SPIE
T2 - Medical Imaging 2023: Computer-Aided Diagnosis
Y2 - 19 February 2023 through 23 February 2023
ER -