Skip to main navigation Skip to search Skip to main content

From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality

  • Stony Brook University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gesture, gaze) or manual expert annotations, Speech-to-Spatial infers the intended target solely from spoken references (speech input). Motivated by our formative study of speech referencing patterns, we characterize recurring ways people specify targets (Direct Attribute, Relational, Remembrance, and Chained) and ground them to our object-centric relational graph. Given an utterance, referent cues are parsed and rendered as persistent in-situ AR visual guidance, reducing iterative micro-guidance (“a bit more to the right”, “now, stop.”) during remote guidance. We demonstrate the use cases of our system with remote guided assistance and intent disambiguation scenarios. Our evaluation shows that Speech-to-Spatial improves task efficiency, reduces cognitive load, and enhances usability compared to a conventional voice-only baseline, transforming disembodied verbal instruction into visually explainable, actionable guidance on a live shared view.

Original languageEnglish
Title of host publicationProceedings - 2026 IEEE Conference on Virtual Reality and 3D User Interfaces, VR 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages228-238
Number of pages11
ISBN (Electronic)9798331559458
DOIs
StatePublished - 2026
Event33rd IEEE Conference on Virtual Reality and 3D User Interfaces, IEEE VR 2026 - Daegu, Korea, Republic of
Duration: Mar 21 2026Mar 25 2026

Publication series

NameProceedings - 2026 IEEE Conference on Virtual Reality and 3D User Interfaces, VR 2026

Conference

Conference33rd IEEE Conference on Virtual Reality and 3D User Interfaces, IEEE VR 2026
Country/TerritoryKorea, Republic of
CityDaegu
Period03/21/2603/25/26

Keywords

  • Augmented Reality
  • Large Language Models
  • Remote Collaboration
  • Spatial Interface
  • Spatial Referencing
  • Speech

Fingerprint

Dive into the research topics of 'From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality'. Together they form a unique fingerprint.

Cite this