Skip to content
AI Atlas
PaperActive

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

arxiv.org/abs/2609.10895

quality89

Updated 2 h ago · first seen 12 Sept 2026

paper_01M29X34T3GJYTNWWP4WRK3839

Published
12 Sept 2026
T1 · 2 h ago
arXiv
2609.10895
T1 · 2 h ago
Category
cs.RO
T1 · 2 h ago

Abstract

Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision coreof household robots. Existing evaluations, however, probe intuitive physics passively through question answering over videos, or target deliberate, long-horizon tasks such as navigation and rearrangement; none measure whether a model can turn physical understanding into immediate, safety-critical action. We introduce ReactHuman, the first physics-grounded benchmark for human-like reactive decision-making, in which the evaluated MLLM acts as the brain of a simulated humanoid facing sudden household hazards; it spans 17 event families and over 1,000 bit-for-bit reproducible scenes with exact, annotation-free ground truth derived from 240 Hz rigid-body simulation, including adversarial objects whose appearance contradicts their physics (a foam anvil, a steel apple). We further design a five-metric suite that scores each reaction along three axes: reasonable, safe, and physically grounded. We physically execute every committed plan so that decisions have observable consequences. With this harness we evaluate seven representative MLLMs. Results show that reactive safety is far from solved: models mishandle roughly one hazard in three, act from fixed dispositions rather than the observed scene, trust appearance over motion, and miss interception points at meter scale even when the chosen action is correct; none of these failures shrink with model scale. ReactHuman thus offers both a fine-grained diagnosis and a scalable training signal toward physically grounded, safety-aware embodied agents. The benchmark can be found here: https://huggingface.co/datasets/Alan123/reacthuman-benchmark-scaled

Authors 8

Bang Liu, Dekun Wu, Dongqing Zhang, Jianxin You, Mengyang Xiong, Yinhuan Chen, Yizhan Li, Zicheng Zhao

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

arXiv id
2609.10895

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Categories
cs.AI, cs.RO

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Primary category
cs.RO

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

2 h ago

Conflicts

None