SpatialCORE Improves Confidence-Aware Spatial Reasoning in Vision-Language Models
SpatialCORE is a new post-training framework that uses model confidence in object localization as a learning signal to improve spatial reasoning in large vision-language models. It sets new state-of-the-art results on spatial reasoning benchmarks and generalizes well to unseen data.
SpatialCORE is a post-training framework designed to improve the spatial reasoning capabilities of large vision-language models (LVLMs). It addresses a key limitation in previous methods: they only optimize for getting the answer right, without ensuring that reasoning is based on accurate and confident localization of task-relevant objects.

What SpatialCORE Changes
- Introduces a self-regulating spatial reward that weighs each predicted bounding box's quality by the model's confidence in its coordinates.
- Uses the model's own grounding confidence as a learning signal during post-training.
- Adds an answer gate mechanism that ties improvements in object localization directly to answer correctness.
- Achieves state-of-the-art performance on multiple spatial reasoning benchmarks.
- Demonstrates strong zero-shot transfer to new data distributions.
Developer Guidance
- Developers working on spatial reasoning with LVLMs should consider integrating SpatialCORE to boost grounding reliability and overall accuracy.
- Evaluate the framework on your own spatial reasoning or grounded question-answering tasks, especially if generalization to new data is important.
