Abstract the Layout, Focus the Detail: A Dual-Granularity Representation Framework for Zero-Shot 3D Visual Grounding
Official repository for: Abstract the Layout, Focus the Detail: A Dual-Granularity Representation Framework for Zero-Shot 3D Visual Grounding, Accepted at ECCV 2026.
We present a dual-granularity framework for zero-shot 3D visual grounding. A Semantic Spatial Layout abstracts global scene structure for spatial reasoning, while Visual Detail Patches preserve local appearance cues for visual verification. Together with query-guided scene filtering, the two forms of evidence support joint spatial-visual reasoning for target grounding.
@inproceedings{lin2026abstract,
title={Abstract the Layout, Focus the Detail: A Dual-Granularity Representation Framework for Zero-Shot 3D Visual Grounding},
author={Lin, Zeyuan and Li, Hanxuan and He, Chen and Wang, Ruiping and Liu, Zhaoxiang and Chen, Xilin},
booktitle={European Conference on Computer Vision},
year={2026}
}