The target is resolved from the spoken utterance through correct speech recognition.
The target is resolved by interpreting the speaker’s gesture in the visual scene.
The target is resolved by detecting and localizing the corresponding device’s sound source.
The target is resolved by reasoning about the speaker’s egocentric direction.
The target is resolved by identifying the speaker and interpreting the person-specific reference.
Smart-home assistants are expected to handle diverse, realistic requests that arise in daily life. In such interactions, users often rely on the surrounding multimodal context—pointing at objects or referring to what they see or hear, leaving their requests underspecified in language alone. Existing smart-home benchmarks, however, express user requests solely through language, leaving context-dependent real-world requests underexplored. To bridge this gap, we introduce OmniSmartHome, a multimodal smart-home benchmark where each spoken request is paired with the surrounding visual and spatial-audio context, providing complementary cues to disambiguate underspecified requests. OmniSmartHome comprises 1,360 synthetic and 272 real-world episodes. We evaluate 16 omnimodal large language models (Omni-LLMs) and reveal that, while they perform strongly when speech alone sufficiently conveys the user’s intent, performance drops substantially when resolving it requires reasoning over multimodal contextual cues. As a simple agent baseline, we provide PROME (PROcedural Memory for multimodal Evidence gathering), which equips agents with specialized audio-visual perception tools and procedural memory for orchestrating their use. PROME generally improves performance across six Omni-LLMs.
OmniSmartHome defines an episode profile along three dimensions: a grounding type, indicating which multimodal cue the user relies on to specify the target device; a query type, indicating what the user asks for; and a feasibility type, indicating whether the request can be fulfilled or not. Every episode provides a visual observation of the room where the speaker is, a first-order Ambisonics recording of the spoken request together with the sounds of active devices, and the initial home state: rooms, devices, and household members.
1,360
Synthetic episodes
272
Real-world episodes
5
Grounding types
4
Query types
2
Feasibility types
One explicit baseline (G0) and four implicit types (G1–G4) that rely on multimodal contextual cues.
| G0 | Speech | The user explicitly names the target device in the spoken utterance, so grounding reduces to speech recognition alone. |
| G1 | Gesture | The user refers to the target device deictically (“Turn that off”) while pointing toward it, requiring the agent to interpret the gesture and ground it in the visual scene. |
| G2 | Sound | The user refers to the target device by the sound it emits (“Silence that humming noise”), requiring the agent to detect and localize the sound source and associate it with the corresponding device. |
| G3 | Space | The user refers to the target device using an egocentric direction (“Turn off the light on my right”), requiring the agent to reason about that direction from the user’s perspective and identify the target located there. |
| G4 | User identity | The user refers to the target device through a possessive reference (“Turn off the AC in my room”), requiring the agent to identify the speaker by face or voice and resolve the reference against that person’s household profile. |
Please use your headphones 🎧 for the best spatial audio experience.
The released real-world videos are not anonymized; faces in the samples shown here are blurred for the anonymous submission.
Q1 State inquiry
“Hey, can you tell me what the remaining time is on the tv in the living room right now?”
Target device TV
Q2 Explicit device control
“Hey, turn on the dimmable light 1 in the living room and set its brightness to 100 percent.”
Target device Dimmable light
Q1 State inquiry
“Is the coffee machine in my room running?”
Target device Coffee machine · Grandpa's room
Q2 Explicit device control
“Turn the dishwasher in my room on and start it running.”
Target device Dishwasher · Dad's room
Q2 Explicit device control
“Hey, switch the channel of the TV in here to SBS.”
Target device TV
Q3 Workflow scheduling
“In 10 minutes, power this on.”
Target device Air purifier
Q1 State inquiry
“What’s the product name of the device making that beeping sound?”
Target device Microwave
Q1 State inquiry
“What brand makes the device in front of me?”
Target device TV
Q4 Implicit intent
“It is muggy in my room, my skin feels sticky.”
Target device Dehumidifier · Son's room
Before operating a device, an agent first resolves which device the user refers to from the audio-visual context. This involves actively gathering fine-grained multimodal cues and integrating them differently depending on how the request is grounded in the environment, which recent Omni-LLMs often struggle with. To support this process, we introduce PROME (PROcedural Memory for multimodal Evidence gathering) as a simple agent baseline. PROME equips the agent with a set of audio-visual perception tools — speech recognition, pointing estimation, sound localization, head-pose recognition, face recognition and speaker recognition — and augments its reasoning with a procedural memory that guides when such tools should be invoked and in what order.
Figure 4: Illustration of our agent framework. Components highlighted in orange are used only by PROME. The agent receives the visual and acoustic observations and iteratively gathers evidence to identify the target device using simulator-interaction and perception tools. PROME additionally provides procedural guidance. Once the target device is identified, the agent declares it and either queries its information (Q1) or executes the requested operation (Q2–Q4) via simulator-interaction tools. Grounding accuracy evaluates the declared target device, while goal accuracy is evaluated from the simulator state or by an LLM judge.
We evaluate 16 Omni-LLMs: seven open-source models with fewer than 10B parameters, five larger open-source models, and four proprietary models. For models that do not support spatial audio input, the first-order Ambisonics recordings are downmixed to mono. PROME is applied to six of the models, each with a memory built from the 680 synthetic training episodes.
| Model | All | G0 Speech | G1 Gesture | G2 Sound | G3 Space | G4 User identity | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gr. | Goal | Gr. | Goal | Gr. | Goal | Gr. | Goal | Gr. | Goal | Gr. | Goal | |
| Human Level | 94.8 | 82.3 | 97.8 | 82.1 | 97.3 | 89.4 | 93.1 | 91.7 | 93.4 | 84.1 | 91.7 | 80.4 |
| Open-source Omni-LLMs (<10B) | ||||||||||||
| Video-LLaMA2 7B | 5.8 | 4.9 | 6.8 | 3.4 | 6.2 | 6.9 | 6.2 | 6.2 | 7.6 | 6.2 | 2.9 | 2.9 |
| Qwen2.5-Omni 7B | 33.0 | 17.6 | 55.5 | 25.5 | 30.9 | 19.4 | 28.1 | 14.2 | 27.2 | 17.4 | 20.3 | 10.9 |
| Spatial-Omni 7B | 20.0 | 11.5 | 39.1 | 14.1 | 14.9 | 11.5 | 16.3 | 10.4 | 16.0 | 11.8 | 10.4 | 9.4 |
| video-SALMONN2+ 7B | 27.0 | 13.7 | 48.2 | 20.6 | 20.1 | 11.8 | 23.6 | 13.5 | 24.7 | 14.9 | 15.4 | 7.6 |
| video-SALMONN-o1 7B | 24.5 | 10.2 | 44.5 | 13.5 | 24.7 | 10.4 | 17.4 | 9.4 | 25.3 | 11.8 | 9.1 | 6.5 |
| OmniVinci 9B | 17.8 | 8.8 | 31.2 | 9.9 | 9.0 | 6.9 | 14.9 | 9.7 | 17.0 | 9.7 | 13.5 | 7.6 |
| MiniCPM-o-4.5 9B | 60.6 | 30.9 | 85.2 | 51.0 | 56.9 | 29.9 | 56.6 | 26.4 | 59.4 | 32.3 | 42.7 | 14.1 |
| Large Open-source Omni-LLMs | ||||||||||||
| Gemma4 12B | 50.6 | 25.2 | 69.3 | 45.3 | 46.5 | 19.1 | 55.2 | 29.2 | 54.9 | 26.0 | 28.1 | 6.0 |
| Nemotron3-Omni 30B-A3B | 62.7 | 44.5 | 86.5 | 60.9 | 65.3 | 52.4 | 67.4 | 47.2 | 64.9 | 51.4 | 31.8 | 15.1 |
| Qwen3-Omni-Instruct 30B-A3B | 66.6 | 43.8 | 88.0 | 63.0 | 66.3 | 53.8 | 77.4 | 38.2 | 60.8 | 48.6 | 41.7 | 17.4 |
| Qwen3-Omni-Think 30B-A3B | 65.9 | 48.4 | 87.8 | 67.4 | 66.0 | 56.6 | 79.9 | 56.9 | 72.2 | 55.2 | 28.6 | 11.7 |
| MiMo-V2.5 310B-A15B | 78.4 | 60.2 | 93.5 | 73.7 | 72.2 | 63.5 | 81.9 | 60.8 | 84.7 | 69.4 | 60.7 | 37.0 |
| Proprietary Omni-LLMs | ||||||||||||
| Gemini-2.5-Flash | 73.7 | 56.2 | 94.0 | 76.8 | 69.4 | 60.1 | 74.3 | 54.9 | 78.1 | 63.2 | 52.6 | 28.6 |
| Gemini-2.5-Pro | 81.0 | 63.1 | 95.1 | 80.5 | 74.7 | 66.7 | 81.9 | 50.3 | 83.0 | 69.8 | 69.5 | 47.7 |
| Gemini-3.1-Pro (Preview) | 85.6 | 73.4 | 95.3 | 81.0 | 76.4 | 71.5 | 97.6 | 80.9 | 87.5 | 79.2 | 72.4 | 57.3 |
| Qwen3.8-Omni-Flash | 81.7 | 63.1 | 93.2 | 76.8 | 76.7 | 64.2 | 92.7 | 66.0 | 85.4 | 72.2 | 63.0 | 39.6 |
Main results. Small open-source models show generally low grounding accuracy across all grounding types, with MiniCPM-o-4.5 performing best in this group. Their goal accuracy is often substantially lower than their grounding accuracy, in some cases by more than a factor of two: even when these models identify the target device, they still struggle to complete the requested goal. Larger open-source models achieve higher overall accuracy with a smaller gap between grounding and goal performance. Their grounding is strongest on G0, where identifying the target device requires little beyond speech recognition, but drops on G1–G4, which require integrating additional multimodal evidence. Proprietary models generally outperform open-source models, yet show a similar gap between G0 and G1–G4. For these stronger models, the contrast between nearly saturated grounding and goal performance on G0 and lower performance on G1–G4 highlights multimodal grounding as the primary challenge of OmniSmartHome.
| Model | All | G0 Speech | G1 Gesture | G2 Sound | G3 Space | G4 User identity | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gr. | Goal | Gr. | Goal | Gr. | Goal | Gr. | Goal | Gr. | Goal | Gr. | Goal | |
| Qwen2.5-Omni + PROME | 38.2 | 22.4 | 61.5 | 34.6 | 41.0 | 28.1 | 35.1 | 14.2 | 28.1 | 22.9 | 22.9 | 11.7 |
| +5.2 | +4.8 | +6.0 | +9.1 | +10.1 | +8.7 | +7.0 | +0.0 | +0.9 | +5.5 | +2.6 | +0.8 | |
| Gemma4 + PROME | 74.0 | 59.3 | 86.2 | 73.4 | 65.3 | 58.3 | 89.6 | 70.8 | 75.7 | 58.3 | 55.5 | 37.8 |
| +23.4 | +34.1 | +16.9 | +28.1 | +18.8 | +39.2 | +34.4 | +41.6 | +20.8 | +32.3 | +27.4 | +31.8 | |
| Nemotron + PROME | 65.2 | 49.1 | 88.8 | 66.7 | 64.9 | 55.6 | 70.1 | 53.8 | 70.8 | 57.6 | 33.9 | 16.9 |
| +2.5 | +4.6 | +2.3 | +5.8 | -0.4 | +3.2 | +2.7 | +6.6 | +5.9 | +6.2 | +2.1 | +1.8 | |
| Qwen3-Think + PROME | 70.2 | 53.9 | 87.5 | 69.8 | 70.1 | 58.3 | 86.5 | 68.8 | 77.1 | 62.2 | 35.4 | 17.2 |
| +4.3 | +5.5 | -0.3 | +2.4 | +4.1 | +1.7 | +6.6 | +11.9 | +4.9 | +7.0 | +6.8 | +5.5 | |
| Gemini-2.5-Flash + PROME | 78.2 | 62.3 | 91.9 | 76.6 | 72.2 | 62.2 | 81.2 | 63.5 | 77.8 | 65.6 | 67.2 | 44.5 |
| +4.5 | +6.1 | -2.1 | -0.2 | +2.8 | +2.1 | +6.9 | +8.6 | -0.3 | +2.4 | +14.6 | +15.9 | |
| Gemini-2.5-Pro + PROME | 85.0 | 69.2 | 96.6 | 79.4 | 74.3 | 66.3 | 86.8 | 68.4 | 87.8 | 73.6 | 78.1 | 58.3 |
| +4.0 | +6.1 | +1.5 | -1.1 | -0.4 | -0.4 | +4.9 | +18.1 | +4.8 | +3.8 | +8.6 | +10.6 | |
PROME improves overall grounding and goal accuracy across six Omni-LLMs. Notably, Gemma4, the weakest larger open-source model, gains 23.4 / 34.1 percentage points in grounding / goal accuracy, respectively, surpassing the other open-source base agents. PROME also improves the strong Gemini models by 4.5 / 6.1 points on Gemini-2.5-Flash and 4.0 / 6.1 points on Gemini-2.5-Pro. Across grounding types, as expected, gains are limited on G0 but larger on G1–G4.
We examine four representative grounding failures across Qwen3-Omni-Think, Gemma4, Gemini-2.5-Flash, and Gemini-2.5-Pro.
G1 Gesture. In around 50% of the episodes where the agent declares an incorrect device, the chosen device is the one nearest to the speaker’s hand or body (a). This suggests that agents often fail to trace the pointing direction and instead fall back to selecting a device near the gesture itself.
G2 Sound. Panel (b) reports, among the episodes where the agent declares an incorrect device, the fraction in which the agent declares the sound source prematurely, that is, without checking the state of any candidate device, such as whether it is running or its alarm is active. Instead, the agent often relies on semantic priors about which devices typically produce a given sound: an alarm from a refrigerator is attributed to a microwave, since the agent assumes that microwaves commonly emit alarm sounds.
G3 Space. Panel (c) shows failure rates by speaker orientation. When the speaker faces the camera, their left–right directions are reversed relative to the image, and three models fail on over 70% of such cases, whereas error rates remain low when the speaker faces away. This suggests that models struggle to interpret directions from the speaker’s egocentric view.
G4 User identity. The left of panel (d) shows, among failed episodes, the fraction in which the agent does not inspect any user-related information. Rather than querying user information to identify the device associated with the speaker’s room, most of these cases end up either declaring no valid device (hatched bars) or selecting a device in the room where the user is currently located (dotted bars), as shown on the right.