Skip to main navigation Skip to search Skip to main content

Efficient Multi-Domain Text Recognition Deep Neural Network Parameterization With Residual Adapters

  • Stony Brook University

Research output: Contribution to journalArticlepeer-review

Abstract

Recent advancements in deep neural networks have markedly enhanced the performance of computer vision tasks, yet the specialized nature of these networks often necessitates extensive data and high computational power. Addressing these requirements, this study presents a novel neural network model adept at optical character recognition (OCR) across diverse domains, leveraging the strengths of multi-task learning to improve efficiency and generalization. The model is designed to achieve rapid adaptation to new domains, maintain a compact size conducive to reduced computational resource demand, ensure high accu-racy, retain knowledge from previous learning experiences, and allow for domain-specific performance improvements without the need to retrain entirely. Rigorous evaluation on open datasets has validated the model’s ability to significantly lower the number of train-able parameters without sacrificing performance, indicating its potential as a scalable and adaptable solution in the field of computer vision, particularly for applications in optical text recognition.

Original languageEnglish
Pages (from-to)1977-1990
Number of pages14
JournalAdvances in Artificial Intelligence and Machine Learning
Volume4
Issue number1
DOIs
StatePublished - 2024

Keywords

  • Continual learning
  • Deep neural network
  • Multi-domain adapter
  • Multi-task learning
  • Optical character recognition

Fingerprint

Dive into the research topics of 'Efficient Multi-Domain Text Recognition Deep Neural Network Parameterization With Residual Adapters'. Together they form a unique fingerprint.

Cite this