Skip to main navigation Skip to search Skip to main content

CoMA: Compositional Human Motion Generation with Multi-modal Agents

  • Shanlin Sun
  • , Jiaqi Xu
  • , Gabriel de Araujo
  • , Shenghan Zhou
  • , Hanwen Zhang
  • , Ziheng Huang
  • , Chenyu You
  • , Xiaohui Xie
  • University of California at Irvine
  • University of California at San Diego
  • Chongqing University
  • Huazhong University of Science and Technology
  • Columbia University

Research output: Contribution to journalConference articlepeer-review

Abstract

3D human motion generation has seen substantial advancement in recent years. While state-of-the-art approaches have improved performance significantly, they still struggle with complex and detailed motions unseen in training data, largely due to the scarcity of motion datasets and the prohibitive cost of generating new training examples. To address these challenges, we introduce CoMA, an agent-based solution for complex human motion generation, editing, and comprehension. CoMA leverages multiple collaborative agents powered by large language and vision models, alongside a mask transformer-based motion generator featuring body part-specific encoders and codebooks for fine-grained control. Our framework enables generation of both short and long motion sequences with detailed instructions, text-guided motion editing, and self-correction for improved quality. Evaluations on the HumanML3D dataset demonstrate competitive performance against state-of-the-art methods. Additionally, we create a set of context-rich, compositional, and long text prompts, where user studies show our method significantly outperforms existing approaches.

Original languageEnglish
Pages (from-to)9206-9214
Number of pages9
JournalProceedings of the AAAI Conference on Artificial Intelligence
Volume40
Issue number11
DOIs
StatePublished - 2026
Event40th AAAI Conference on Artificial Intelligence, AAAI 2026 - Singapore, Singapore
Duration: Jan 20 2026Jan 27 2026

Fingerprint

Dive into the research topics of 'CoMA: Compositional Human Motion Generation with Multi-modal Agents'. Together they form a unique fingerprint.

Cite this