Sign InOpen Brain
arXivPaperNeeds Review

3D-Aware VLMs with Implicit and Explicit Geometries

VLM-IE3D adds implicit and reconstructed geometry tokens to an RGB-video VLM, offering an open approach for agents that must reason about spatial scenes without dedicated 3D input.

arXiv · Jul 23, 2026
Open Source Open MarkdownOpen JSON
Source Summary

VLM-IE3D derives **implicit geometry tokens** from video and **explicit geometry tokens** from reconstructed 3D attributes. A 3D-aware adapter combines both with ordinary 2D visual cues, while requiring only RGB video as input.

Practical Implication

Builders working on scene-aware agents can test this pattern when 2D frames alone lose object position or spatial relationships. The released code and models make the architecture inspectable rather than merely conceptual.

Agent-Ready Context
VLM-IE3D derives **implicit geometry tokens** from video and **explicit geometry tokens** from reconstructed 3D attributes. A 3D-aware adapter combines both with ordinary 2D visual cues, while requiring only RGB video as input.

Builders working on scene-aware agents can test this pattern when 2D frames alone lose object position or spatial relationships. The released code and models make the architecture inspectable rather than merely conceptual.

The material reports gains across detection, grounding, dense captioning, and spatial reasoning, but provides no scores or deployment costs. Its usefulness outside the evaluated 3D tasks remains open.
Context Map
modelimagevideo#reasoning#open-models
Uncertainty
The material reports gains across detection, grounding, dense captioning, and spatial reasoning, but provides no scores or deployment costs. Its usefulness outside the evaluated 3D tasks remains open.