Abstract
Accurately classifying damage levels from structural inspection images is critical for automated infrastructure assessment. Although deep neural networks achieve impressive performance, their black-box nature limits explainability, and prior studies using Grad-CAM often yield coarse or inaccurate saliency maps. To overcome these limitations, this paper introduces XIDLE-Net, a multitask model that simultaneously performs damage classification and saliency map prediction to enhance explainability in structural damage assessment. Combining a Swin Transformer encoder with a convolutional neural network decoder, XIDLE-Net is trained with dual supervision using damage labels and inspector gaze-derived attention maps, enhancing both classification accuracy and model explainability. Experimental results show that XIDLE-Net outperforms state-of-the-art methods in both classification and saliency explainability, achieving 78.1% accuracy, 94.3% area under the curve (AUC), and a 39.7% improvement in saliency prediction over ResNet-50 with Grad-CAM. To our knowledge, this is one of the first investigations to employ large-scale inspector gaze data for supervision and to quantitatively evaluate Grad-CAM in structural image classification. The results highlight the promise of human gaze data for advancing explainable vision-based structural health monitoring.
| Original language | English |
|---|---|
| Pages (from-to) | 5824-5841 |
| Number of pages | 18 |
| Journal | Computer-Aided Civil and Infrastructure Engineering |
| Volume | 40 |
| Issue number | 30 |
| DOIs | |
| State | Published - Dec 19 2025 |
Fingerprint
Dive into the research topics of 'Inspector gaze-guided multitask learning for explainable structural damage assessment'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver