Skip to main navigation Skip to search Skip to main content

A novel 6DoF pose estimation method using transformer fusion

  • Huafeng Wang
  • , Haodu Zhang
  • , Wanquan Liu
  • , Zhimin Hu
  • , Haoqi Gao
  • , Weifeng Lv
  • , Xianfeng Gu
  • North China University of Technology
  • Beihang University
  • Sun Yat-Sen University

Research output: Contribution to journalArticlepeer-review

3 Scopus citations

Abstract

Effectively combining different data types (RGB, depth) for 6D pose estimation in deep learning remains challenging. Effectively extracting complementary information from these modalities and achieving implicit alignment is crucial for accurate pose estimation. This work proposes a novel fusion module that utilizes Transformer-based architecture for cross-modal fusion. This design fosters feature combination and strengthens global information processing, reducing dependence on traditional convolutional methods. Additionally, a residual attentional structure tackles two key issues: (1) mitigating information loss commonly encountered in deep networks, and (2) enhancing modal alignment through learned attention weights. We evaluate our method on the LineMOD Hinterstoisser et al. (2011) and YCB-Video Xiang et al. (2018) datasets, achieving state-of-the-art performance on YCB-Video and outperforming most existing methods on LineMOD. These results demonstrate the effectiveness of our approach and its strong generalization capabilities.

Original languageEnglish
Article number111413
JournalPattern Recognition
Volume162
DOIs
StatePublished - Jun 2025

Keywords

  • 6D pose estimation
  • Feature supplementary
  • Heterogeneous information fusion
  • Modalities aligning

Fingerprint

Dive into the research topics of 'A novel 6DoF pose estimation method using transformer fusion'. Together they form a unique fingerprint.

Cite this