Spatial and Temporal Grounding Limits in Multimodal Models
L. Mbeki, S. Varga — Independent
Abstract
We separate perceptual competence from grounded reasoning in multimodal evaluation. Systems approach human baseline on scene understanding while failing at counting, temporal ordering and precise localisation. We argue the residual gap is not perceptual and discuss the consequences for embodied deployment, where these are exactly the operations that matter.
Key findings
- Scene understanding approaches human baseline
- Counting, ordering and localisation fail persistently
- The gap is not explained by perceptual resolution
- These operations are load-bearing for embodied tasks
Limitations
Published at equal prominence to the findings. A paper’s limitations are usually the part that determines how much its result should move your beliefs.
- Synthetic scenes used for controlled localisation tests
- Human baseline collected on a limited annotator pool
- No embodied evaluation is reported