TY - GEN
T1 - From Speech-to-Spatial
T2 - 33rd IEEE Conference on Virtual Reality and 3D User Interfaces, IEEE VR 2026
AU - Kim, Yoonsang
AU - Pradhan, Divyansh
AU - Jadeja, Devshree
AU - Kaufman, Arie E.
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gesture, gaze) or manual expert annotations, Speech-to-Spatial infers the intended target solely from spoken references (speech input). Motivated by our formative study of speech referencing patterns, we characterize recurring ways people specify targets (Direct Attribute, Relational, Remembrance, and Chained) and ground them to our object-centric relational graph. Given an utterance, referent cues are parsed and rendered as persistent in-situ AR visual guidance, reducing iterative micro-guidance (“a bit more to the right”, “now, stop.”) during remote guidance. We demonstrate the use cases of our system with remote guided assistance and intent disambiguation scenarios. Our evaluation shows that Speech-to-Spatial improves task efficiency, reduces cognitive load, and enhances usability compared to a conventional voice-only baseline, transforming disembodied verbal instruction into visually explainable, actionable guidance on a live shared view.
AB - We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gesture, gaze) or manual expert annotations, Speech-to-Spatial infers the intended target solely from spoken references (speech input). Motivated by our formative study of speech referencing patterns, we characterize recurring ways people specify targets (Direct Attribute, Relational, Remembrance, and Chained) and ground them to our object-centric relational graph. Given an utterance, referent cues are parsed and rendered as persistent in-situ AR visual guidance, reducing iterative micro-guidance (“a bit more to the right”, “now, stop.”) during remote guidance. We demonstrate the use cases of our system with remote guided assistance and intent disambiguation scenarios. Our evaluation shows that Speech-to-Spatial improves task efficiency, reduces cognitive load, and enhances usability compared to a conventional voice-only baseline, transforming disembodied verbal instruction into visually explainable, actionable guidance on a live shared view.
KW - Augmented Reality
KW - Large Language Models
KW - Remote Collaboration
KW - Spatial Interface
KW - Spatial Referencing
KW - Speech
UR - https://www.scopus.com/pages/publications/105036982868
U2 - 10.1109/VR67842.2026.00045
DO - 10.1109/VR67842.2026.00045
M3 - Conference contribution
AN - SCOPUS:105036982868
T3 - Proceedings - 2026 IEEE Conference on Virtual Reality and 3D User Interfaces, VR 2026
SP - 228
EP - 238
BT - Proceedings - 2026 IEEE Conference on Virtual Reality and 3D User Interfaces, VR 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 21 March 2026 through 25 March 2026
ER -