Abstract:
Objectives: Spatial cognition is essential for multimodal large language models (MLLMs) to progress toward spatial intelligence and embodied agency. Existing evaluations often emphasize isolated tasks and end-point accuracy, providing limited evidence of whether models can maintain spatial representations, transform reference frames, integrate local observations, reason across scales, or adapt to dynamic environments. This study develops a brain-inspired framework for evaluating the geospatial cognitive abilities of large models.
Methods: Research on human spatial ability, environmental navigation, geospatial cognition, multimodal reasoning, cognitive maps, and embodied intelligence was synthesized. Existing geospatial tasks were reorganized according to cognitive requirements. A hierarchical ability system and evaluation scheme were constructed, covering spatial knowledge type, reference frame, input modality, spatial scale, environmental condition, and interaction mode. Functional principles from human spatial cognition were incorporated without assuming structural equivalence between artificial models and the brain.
Results: The framework comprises five interconnected levels: geospatial perception, representation, reasoning, decision-making, and interaction. These levels form a closed loop of perception, representation, reasoning, action, and updating. For each level, representative tasks and indicators are specified, including landmark and boundary recognition, cognitive-map construction, reference-frame transformation, path integration, route planning, active information acquisition, environmental change detection, replanning, and error recovery. The framework recommends combining static tests, representation probing, controlled interventions, interactive tasks, and real-world validation. Key indicators include localization error, topological accuracy, spatial consistency, reasoning stability, planning efficiency, information gain, confidence calibration, recovery rate, and long-horizon stability.
Conclusions: Evaluating large-model spatial cognition should move beyond benchmark accuracy toward evidence of persistent representation, structured spatial memory, predictive reasoning, active exploration, and closed-loop adaptation. Brain-inspired spatial intelligence should be treated as the computational abstraction of functional principles rather than direct replication of neural structures. Future research should emphasize hierarchical and ecologically valid assessment, explicit cognitive maps, human–AI collaboration, uncertainty communication, reproducibility, and trustworthy behavior in geographic environments.