Skip to main navigation Skip to search Skip to main content

Not All Data are What You Need: A Data-Efficient Training Method Using Heterogeneous Hardware

  • Zulong Diao
  • , Mingyu Qiao
  • , Xin Wang
  • , Guangxing Zhang
  • , Wei Liang
  • , Jianguo Chen
  • , Changhua Pei
  • , Yanbiao Li
  • , Zhenyu Li
  • , Gaogang Xie
  • Hunan University of Science and Technology
  • CAS - Institute of Computing Technology
  • University of Chinese Academy of Sciences
  • Sun Yat-Sen University
  • CAS - Computer Network Information Center

Research output: Contribution to journalArticlepeer-review

3 Scopus citations

Abstract

Deep learning is applied to various tasks, such as image recognition and self-driving. Training acceleration is crucial for the further development of deep learning, as efficient training algorithms can help greatly reduce the time consumption and hardware usage while making real-time updates of large-scale deep learning models possible. The mainstream methods realize it through distributed training or network pruning. The former relies on abundant hardware resources and the latter may suffer from a non-negligible performance drop. In this paper, we propose CORESTR, a data-efficient training framework that asynchronously utilizes heterogeneous hardware resources. The framework consists of two major procedures. We first characterize the training status of each instance and propose a representative instance selection algorithm for reducing the total number of instances participating in each epoch of training. In the second procedure, we design a lightweight sample weighting mechanism based on meta-learning to closely approximate the convergence quality using a representative instance set selected from the full training dataset. We present the theoretical rationale for our approach and evaluate its training performance with several classical models and datasets. Experiment results demonstrate that our training method can achieve an average speedup of 4.8×4.8× and reach a higher final accuracy compared with state-of-the-art methods by only relying on a small part of the training data.

Original languageEnglish
Pages (from-to)2995-3008
Number of pages14
JournalIEEE Transactions on Knowledge and Data Engineering
Volume38
Issue number5
DOIs
StatePublished - May 1 2026

Keywords

  • Data-efficient training
  • heterogeneous hardware collaborative
  • high-performance computing

Fingerprint

Dive into the research topics of 'Not All Data are What You Need: A Data-Efficient Training Method Using Heterogeneous Hardware'. Together they form a unique fingerprint.

Cite this