OmniSmartHome: A Multimodal Reasoning
Benchmark for Smart-Home Agents

Anonymous Authors
Figure 1: The same air-conditioner request expressed through pointing, humming sounds, egocentric direction, and room ownership.

Figure 1: Overview of OmniSmartHome. In real-world interactions, a user’s spoken request is often underspecified in language alone, with the intended target device resolved from multimodal cues distributed across the surrounding audio-visual context.

Demo Videos

Abstract

Smart-home assistants are expected to handle diverse, realistic requests that arise in daily life. In such interactions, users often rely on the surrounding multimodal context—pointing at objects or referring to what they see or hear, leaving their requests underspecified in language alone. Existing smart-home benchmarks, however, express user requests solely through language, leaving context-dependent real-world requests underexplored. To bridge this gap, we introduce OmniSmartHome, a multimodal smart-home benchmark where each spoken request is paired with the surrounding visual and spatial-audio context, providing complementary cues to disambiguate underspecified requests. OmniSmartHome comprises 1,360 synthetic and 272 real-world episodes. We evaluate 16 omnimodal large language models (Omni-LLMs) and reveal that, while they perform strongly when speech alone sufficiently conveys the user’s intent, performance drops substantially when resolving it requires reasoning over multimodal contextual cues. As a simple agent baseline, we provide PROME (PROcedural Memory for multimodal Evidence gathering), which equips agents with specialized audio-visual perception tools and procedural memory for orchestrating their use. PROME generally improves performance across six Omni-LLMs.

Benchmark

OmniSmartHome defines an episode profile along three dimensions: a grounding type, indicating which multimodal cue the user relies on to specify the target device; a query type, indicating what the user asks for; and a feasibility type, indicating whether the request can be fulfilled or not. Every episode provides a visual observation of the room where the speaker is, a first-order Ambisonics recording of the spoken request together with the sounds of active devices, and the initial home state: rooms, devices, and household members.

1,360

Synthetic episodes

272

Real-world episodes

5

Grounding types

4

Query types

2

Feasibility types

Figure 2: Taxonomy of OmniSmartHome, showing G0-G4 and their query types.

Figure 2: Taxonomy of OmniSmartHome. Each episode has a grounding, query, and feasibility type.

Five grounding types

One explicit baseline (G0) and four implicit types (G1–G4) that rely on multimodal contextual cues.

G0 Speech The user explicitly names the target device in the spoken utterance, so grounding reduces to speech recognition alone.
G1 Gesture The user refers to the target device deictically (“Turn that off”) while pointing toward it, requiring the agent to interpret the gesture and ground it in the visual scene.
G2 Sound The user refers to the target device by the sound it emits (“Silence that humming noise”), requiring the agent to detect and localize the sound source and associate it with the corresponding device.
G3 Space The user refers to the target device using an egocentric direction (“Turn off the light on my right”), requiring the agent to reason about that direction from the user’s perspective and identify the target located there.
G4 User identity The user refers to the target device through a possessive reference (“Turn off the AC in my room”), requiring the agent to identify the speaker by face or voice and resolve the reference against that person’s household profile.

Benchmark samples

Please use your headphones 🎧 for the best spatial audio experience.

The released real-world videos are not anonymized; faces in the samples shown here are blurred for the anonymous submission.

Simple agentic baseline

Before operating a device, an agent first resolves which device the user refers to from the audio-visual context. This involves actively gathering fine-grained multimodal cues and integrating them differently depending on how the request is grounded in the environment, which recent Omni-LLMs often struggle with. To support this process, we introduce PROME (PROcedural Memory for multimodal Evidence gathering) as a simple agent baseline. PROME equips the agent with a set of audio-visual perception tools — speech recognition, pointing estimation, sound localization, head-pose recognition, face recognition and speaker recognition — and augments its reasoning with a procedural memory that guides when such tools should be invoked and in what order.

Figure 4: PROME adds perception tools, a grounding ambiguity table, a state update, and a procedure policy to the base ReAct agent loop.

Figure 4: Illustration of our agent framework. Components highlighted in orange are used only by PROME. The agent receives the visual and acoustic observations and iteratively gathers evidence to identify the target device using simulator-interaction and perception tools. PROME additionally provides procedural guidance. Once the target device is identified, the agent declares it and either queries its information (Q1) or executes the requested operation (Q2–Q4) via simulator-interaction tools. Grounding accuracy evaluates the declared target device, while goal accuracy is evaluated from the simulator state or by an LLM judge.

Main results

We evaluate 16 Omni-LLMs: seven open-source models with fewer than 10B parameters, five larger open-source models, and four proprietary models. For models that do not support spatial audio input, the first-order Ambisonics recordings are downmixed to mono. PROME is applied to six of the models, each with a memory built from the 680 synthetic training episodes.

Results on OmniSmartHome across grounding types (G0–G4). For each grounding type, we report grounding accuracy (Gr.) and goal accuracy (Goal), averaged over query types.
ModelAllG0 SpeechG1 GestureG2 SoundG3 SpaceG4 User identity
Gr.GoalGr.GoalGr.GoalGr.GoalGr.GoalGr.Goal
Human Level94.882.397.882.197.389.493.191.793.484.191.780.4
Open-source Omni-LLMs (<10B)
Video-LLaMA2 7B5.84.96.83.46.26.96.26.27.66.22.92.9
Qwen2.5-Omni 7B33.017.655.525.530.919.428.114.227.217.420.310.9
Spatial-Omni 7B20.011.539.114.114.911.516.310.416.011.810.49.4
video-SALMONN2+ 7B27.013.748.220.620.111.823.613.524.714.915.47.6
video-SALMONN-o1 7B24.510.244.513.524.710.417.49.425.311.89.16.5
OmniVinci 9B17.88.831.29.99.06.914.99.717.09.713.57.6
MiniCPM-o-4.5 9B60.630.985.251.056.929.956.626.459.432.342.714.1
Large Open-source Omni-LLMs
Gemma4 12B50.625.269.345.346.519.155.229.254.926.028.16.0
Nemotron3-Omni 30B-A3B62.744.586.560.965.352.467.447.264.951.431.815.1
Qwen3-Omni-Instruct 30B-A3B66.643.888.063.066.353.877.438.260.848.641.717.4
Qwen3-Omni-Think 30B-A3B65.948.487.867.466.056.679.956.972.255.228.611.7
MiMo-V2.5 310B-A15B78.460.293.573.772.263.581.960.884.769.460.737.0
Proprietary Omni-LLMs
Gemini-2.5-Flash73.756.294.076.869.460.174.354.978.163.252.628.6
Gemini-2.5-Pro81.063.195.180.574.766.781.950.383.069.869.547.7
Gemini-3.1-Pro (Preview)85.673.495.381.076.471.597.680.987.579.272.457.3
Qwen3.8-Omni-Flash81.763.193.276.876.764.292.766.085.472.263.039.6

Main results. Small open-source models show generally low grounding accuracy across all grounding types, with MiniCPM-o-4.5 performing best in this group. Their goal accuracy is often substantially lower than their grounding accuracy, in some cases by more than a factor of two: even when these models identify the target device, they still struggle to complete the requested goal. Larger open-source models achieve higher overall accuracy with a smaller gap between grounding and goal performance. Their grounding is strongest on G0, where identifying the target device requires little beyond speech recognition, but drops on G1–G4, which require integrating additional multimodal evidence. Proprietary models generally outperform open-source models, yet show a similar gap between G0 and G1–G4. For these stronger models, the contrast between nearly saturated grounding and goal performance on G0 and lower performance on G1–G4 highlights multimodal grounding as the primary challenge of OmniSmartHome.

Results on PROME

Results on PROME across grounding types (G0–G4). Small numbers denote the change in percentage points relative to the base agent.
ModelAllG0 SpeechG1 GestureG2 SoundG3 SpaceG4 User identity
Gr.GoalGr.GoalGr.GoalGr.GoalGr.GoalGr.Goal
Qwen2.5-Omni + PROME38.222.461.534.641.028.135.114.228.122.922.911.7
+5.2+4.8+6.0+9.1+10.1+8.7+7.0+0.0+0.9+5.5+2.6+0.8
Gemma4 + PROME74.059.386.273.465.358.389.670.875.758.355.537.8
+23.4+34.1+16.9+28.1+18.8+39.2+34.4+41.6+20.8+32.3+27.4+31.8
Nemotron + PROME65.249.188.866.764.955.670.153.870.857.633.916.9
+2.5+4.6+2.3+5.8-0.4+3.2+2.7+6.6+5.9+6.2+2.1+1.8
Qwen3-Think + PROME70.253.987.569.870.158.386.568.877.162.235.417.2
+4.3+5.5-0.3+2.4+4.1+1.7+6.6+11.9+4.9+7.0+6.8+5.5
Gemini-2.5-Flash + PROME78.262.391.976.672.262.281.263.577.865.667.244.5
+4.5+6.1-2.1-0.2+2.8+2.1+6.9+8.6-0.3+2.4+14.6+15.9
Gemini-2.5-Pro + PROME85.069.296.679.474.366.386.868.487.873.678.158.3
+4.0+6.1+1.5-1.1-0.4-0.4+4.9+18.1+4.8+3.8+8.6+10.6

PROME improves overall grounding and goal accuracy across six Omni-LLMs. Notably, Gemma4, the weakest larger open-source model, gains 23.4 / 34.1 percentage points in grounding / goal accuracy, respectively, surpassing the other open-source base agents. PROME also improves the strong Gemini models by 4.5 / 6.1 points on Gemini-2.5-Flash and 4.0 / 6.1 points on Gemini-2.5-Pro. Across grounding types, as expected, gains are limited on G0 but larger on G1–G4.

Error analysis

We examine four representative grounding failures across Qwen3-Omni-Think, Gemma4, Gemini-2.5-Flash, and Gemini-2.5-Pro.

Representative grounding failures of the base agents for Qwen3-Think, Gemma4, Gemini-2.5-Flash and Gemini-2.5-Pro: (a) G1 devices nearest to the speaker, (b) G2 premature sound-source declarations, (c) G3 failure rate split by speaker facing the camera or facing away, and (d) G4 episodes without any user lookup, broken down into no device, the user-location device, and others.

Figure 5: Representative causes of grounding failures of the base agents. For each grounding type (G1–G4), we show a representative failure pattern.

G1 Gesture. In around 50% of the episodes where the agent declares an incorrect device, the chosen device is the one nearest to the speaker’s hand or body (a). This suggests that agents often fail to trace the pointing direction and instead fall back to selecting a device near the gesture itself.

G2 Sound. Panel (b) reports, among the episodes where the agent declares an incorrect device, the fraction in which the agent declares the sound source prematurely, that is, without checking the state of any candidate device, such as whether it is running or its alarm is active. Instead, the agent often relies on semantic priors about which devices typically produce a given sound: an alarm from a refrigerator is attributed to a microwave, since the agent assumes that microwaves commonly emit alarm sounds.

G3 Space. Panel (c) shows failure rates by speaker orientation. When the speaker faces the camera, their left–right directions are reversed relative to the image, and three models fail on over 70% of such cases, whereas error rates remain low when the speaker faces away. This suggests that models struggle to interpret directions from the speaker’s egocentric view.

G4 User identity. The left of panel (d) shows, among failed episodes, the fraction in which the agent does not inspect any user-related information. Rather than querying user information to identify the device associated with the speaker’s room, most of these cases end up either declaring no valid device (hatched bars) or selecting a device in the room where the user is currently located (dotted bars), as shown on the right.