RWKV-CLIP: a robust vision-language representation learner
File(s) 2024.emnlp-main.276.pdf (1.9 MB)
Published version
Author(s)
Type
Conference Paper
Abstract
Contrastive Language-Image Pre-training (CLIP) has significantly improved performance in various vision-language tasks by expanding the dataset with image-text pairs obtained from the web. This paper further explores CLIP from the perspectives of data and model architecture. To mitigate the impact of the noise data and enhance the quality of large-scale image-text data crawled from the internet, we introduce a diverse description generation framework that can leverage Large Language Models (LLMs) to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags. Additionally, we propose RWKV-CLIP, the first RWKV-driven vision-language representation learning model that combines the effective parallel training of transformers with the efficient inference of RNNs. Extensive experiments across different model scales and pre-training datasets demonstrate that RWKV-CLIP is a robust vision-language representation learner and it achieves state-of-the-art performance across multiple downstream tasks, including linear probing, zero-shot classification, and zero-shot image-text retrieval. To facilitate future research, the code and pre-trained models are released at https://github.com/deepglint/RWKV-CLIP.
Date Issued
2024-11-01
Date Acceptance
2024-11-01
Citation
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp.4799-4812
Publisher
Association for Computational Linguistics
Start Page
4799
End Page
4812
Journal / Book Title
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Copyright Statement
©2024 Association for Computational Linguistics.
Source
2024 Conference on Empirical Methods in Natural Language Processing
Publication Status
Published
Start Date
2024-11-12
Finish Date
2024-11-16
Coverage Spatial
Miami, Florida, USA
