Vision-Language Constraint Graph Representation Learning for Unsupervised Vehicle Re-identification
Dong Wang , Qi Wang , Zhiwei Tu , Weidong Min , Xin Xiong , Yuling Zhong , Di Gai
Accepted By
Expert Systems With Applications (ESWA)
Accepted Date
June 11, 2024
Venue Type
Journal
Venue Level
SCI-Q1
Abstract
This paper proposes a vision-language constraint graph representation learning method for unsupervised vehicle re-identification. Unlike existing methods that mainly rely on visual features, the proposed framework introduces textual descriptions generated by conditional prompts to enhance the semantic understanding of vehicle images. A vision-language constraint graph topology is constructed by treating each training sample as a graph node and jointly modeling visual and textual feature correlations, enabling more reliable positive and negative sample relationship mining. To further reduce pseudo-label noise caused by visual feature clustering, the paper introduces neighboring node label smoothing, which combines clustering results with graph-neighbor relationships to generate more robust soft pseudo-labels. Experiments on VeRi-776 and VehicleID demonstrate that the proposed method effectively integrates visual and textual semantic information and achieves competitive performance in unsupervised vehicle Re-ID.