Abstract
Effectively combining different data types (RGB, depth) for 6D pose estimation in deep learning remains challenging. Effectively extracting complementary information from these modalities and achieving implicit alignment is crucial for accurate pose estimation. This work proposes a novel fusion module that utilizes Transformer-based architecture for cross-modal fusion. This design fosters feature combination and strengthens global information processing, reducing dependence on traditional convolutional methods. Additionally, a residual attentional structure tackles two key issues: (1) mitigating information loss commonly encountered in deep networks, and (2) enhancing modal alignment through learned attention weights. We evaluate our method on the LineMOD Hinterstoisser et al. (2011) and YCB-Video Xiang et al. (2018) datasets, achieving state-of-the-art performance on YCB-Video and outperforming most existing methods on LineMOD. These results demonstrate the effectiveness of our approach and its strong generalization capabilities.
| Original language | English |
|---|---|
| Article number | 111413 |
| Journal | Pattern Recognition |
| Volume | 162 |
| DOIs | |
| State | Published - Jun 2025 |
Keywords
- 6D pose estimation
- Feature supplementary
- Heterogeneous information fusion
- Modalities aligning
Fingerprint
Dive into the research topics of 'A novel 6DoF pose estimation method using transformer fusion'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver