TY - GEN
T1 - Beyond-Voice
T2 - 27th International Workshop on Mobile Computing Systems and Applications, ACM HotMobile 2026
AU - Srivastava, Tanmay
AU - Basu, Amartya
AU - Jain, Shubham
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/3/2
Y1 - 2026/3/2
N2 - Current Human-AI speech interfaces are limited to the acoustic signal, overlooking the rich information in articulatory motion. This paper explores how complementing the acoustic channel with signals from speech articulators can enhance Human-AI interaction. Captured via a range of sensors (e.g., IMU, EMG, or cameras), these non-acoustic signals offer a new layer of insight. These signals enable applications infeasible with audio alone, such as anticipating a user’s cognitive struggle before they speak, maintaining a silent channel for private human-AI interaction, providing real-time speech coaching, and even monitoring well-being by capturing subtle cues like jaw clenching. The integration of these “beyond-voice” interfaces with large language models presents an opportunity to move from voice-only interactions to a more complete model of human-AI communication. We outline the potential research directions and the associated challenges the research community must address to realize this vision, followed by key takeaways from our exploratory study.
AB - Current Human-AI speech interfaces are limited to the acoustic signal, overlooking the rich information in articulatory motion. This paper explores how complementing the acoustic channel with signals from speech articulators can enhance Human-AI interaction. Captured via a range of sensors (e.g., IMU, EMG, or cameras), these non-acoustic signals offer a new layer of insight. These signals enable applications infeasible with audio alone, such as anticipating a user’s cognitive struggle before they speak, maintaining a silent channel for private human-AI interaction, providing real-time speech coaching, and even monitoring well-being by capturing subtle cues like jaw clenching. The integration of these “beyond-voice” interfaces with large language models presents an opportunity to move from voice-only interactions to a more complete model of human-AI communication. We outline the potential research directions and the associated challenges the research community must address to realize this vision, followed by key takeaways from our exploratory study.
KW - AI Assistants
KW - Articulator Sensing
KW - Human-AI Interaction
KW - Mobile Interaction
KW - Silent Speech
UR - https://www.scopus.com/pages/publications/105033533504
U2 - 10.1145/3789514.3792054
DO - 10.1145/3789514.3792054
M3 - Conference contribution
AN - SCOPUS:105033533504
T3 - HotMobile 2026 - Proceedings of the 2026 ACM 27th International Workshop on Mobile Computing Systems and Applications
SP - 49
EP - 54
BT - HotMobile 2026 - Proceedings of the 2026 ACM 27th International Workshop on Mobile Computing Systems and Applications
PB - Association for Computing Machinery, Inc
Y2 - 25 February 2026 through 26 February 2026
ER -