面向类脑空间智能的多模态大模型空间认知能力评估框架

A Framework for Evaluating Geospatial Cognitive Abilities of Multimodal Large Language Models for Brain-Inspired Spatial Intelligence

  • 摘要: 空间认知是多模态大模型由语言理解迈向空间智能和具身智能的关键能力。现有模型虽已应用于地理问答、地图理解、遥感分析和代码生成,但相关评估多依赖静态任务与结果准确率,难以揭示模型在空间表征、参考框架转换、局部与整体整合及动态环境适应方面的能力边界。视觉认知与脑认知机制的不足,是限制模型空间认知能力的重要原因。基于空间认知、地理空间思维和类脑智能理论,构建由地理空间感知、表征、推理、决策和交互组成的五层能力体系,并从空间知识类型、参考框架、输入模态、空间尺度、环境条件和交互方式等维度提出综合评估方法。评估应结合静态测评、表征探测、控制实验、交互任务和真实场景验证,重点考察空间一致性、推理稳定性、决策效率、错误恢复及环境适应能力。未来研究需进一步推进分层与生态化评估、主动探索与闭环学习、人类—AI 协同以及可信空间智能,为发展具有稳定表征、迁移推理和动态交互能力的多模态大模型提供理论依据。

     

    Abstract: Objectives: Spatial cognition is essential for multimodal large language models (MLLMs) to progress toward spatial intelligence and embodied agency. Existing evaluations often emphasize isolated tasks and end-point accuracy, providing limited evidence of whether models can maintain spatial representations, transform reference frames, integrate local observations, reason across scales, or adapt to dynamic environments. This study develops a brain-inspired framework for evaluating the geospatial cognitive abilities of large models. Methods: Research on human spatial ability, environmental navigation, geospatial cognition, multimodal reasoning, cognitive maps, and embodied intelligence was synthesized. Existing geospatial tasks were reorganized according to cognitive requirements. A hierarchical ability system and evaluation scheme were constructed, covering spatial knowledge type, reference frame, input modality, spatial scale, environmental condition, and interaction mode. Functional principles from human spatial cognition were incorporated without assuming structural equivalence between artificial models and the brain. Results: The framework comprises five interconnected levels: geospatial perception, representation, reasoning, decision-making, and interaction. These levels form a closed loop of perception, representation, reasoning, action, and updating. For each level, representative tasks and indicators are specified, including landmark and boundary recognition, cognitive-map construction, reference-frame transformation, path integration, route planning, active information acquisition, environmental change detection, replanning, and error recovery. The framework recommends combining static tests, representation probing, controlled interventions, interactive tasks, and real-world validation. Key indicators include localization error, topological accuracy, spatial consistency, reasoning stability, planning efficiency, information gain, confidence calibration, recovery rate, and long-horizon stability. Conclusions: Evaluating large-model spatial cognition should move beyond benchmark accuracy toward evidence of persistent representation, structured spatial memory, predictive reasoning, active exploration, and closed-loop adaptation. Brain-inspired spatial intelligence should be treated as the computational abstraction of functional principles rather than direct replication of neural structures. Future research should emphasize hierarchical and ecologically valid assessment, explicit cognitive maps, human–AI collaboration, uncertainty communication, reproducibility, and trustworthy behavior in geographic environments.

     

/

返回文章
返回