Back to Home

SpecBridge: Spectral Graph Bridging for 3D-2D-Text Pre-training

Dong Wang , Jie Jiang , Weidong Min , Lixin Zhan , Xinpeng Zhao , Ze Zhang

3D Understanding Vision-Language Learning Multi-Modal Learning CCF-B International Conference

Accepted By

International Joint Conference on Artificial Intelligence 2026 (IJCAI2026)

Accepted Date

May 1, 2026

Venue Type

Conference

Venue Level

CCF-B

SpecBridge: Spectral Graph Bridging for 3D-2D-Text Pre-training teaser

Abstract

Open-vocabulary 3D understanding aims to align 3D representations with a unified vision-language semantic space. However, existing methods suffer from the challenge of structural asymmetry caused by sparse observations and holistic geometries. Additionally, the inherent semantic chasm between discrete coordinates and abstract natural language hinders cross-modal mapping. This work introduces SpecBridge, a 3D-2D-Text pre-training framework that leverages CLIP priors as a foundational bridge to connect three modalities by synergizing spectral graph theory with transitive semantic learning. The method consists of two main sub-modules. First, Spectral Eigen-Modal Alignment is presented to construct cross-modal correlations within the intrinsic eigen-spectral space via Laplacian eigendecomposition. By aligning low-frequency geometric harmonics with CLIP-guided observation inference, it maintains structural consistency between 3D geometries and 2D projections. Second, a Transitive Spectral-Semantic Alignment method is developed to establish a 3D→2D→Text propagation chain across the CLIP priors bridge. It distills dense CLIP priors into 3D representations through transitive distillation, effectively mitigating the semantic chasm between geometry and language. Extensive experiments confirm that the proposed SpecBridge demonstrates state-of-the-art performance by overcoming challenges triggered by modality discrepancies.

Coming soon.