Skip to content

Latest commit

 

History

History
64 lines (64 loc) · 2.52 KB

File metadata and controls

64 lines (64 loc) · 2.52 KB
title Prompting vision-language fusion for Zero-Shot Composed Image Retrieval
booktitle Proceedings of the 16th Asian Conference on Machine Learning
year 2025
volume 260
series Proceedings of Machine Learning Research
month 0
publisher PMLR
pdf https://raw.githubusercontent.com/mlresearch/v260/main/assets/wang25d/wang25d.pdf
url https://proceedings.mlr.press/v260/wang25d.html
openreview 8NM3qmWjkt
abstract The composed image retrieval (CIR) aims to retrieve target image given the combination of an image and a textual description as a query. Recently, benefiting from vision-language pretrained (VLP) models and large language models (LLM), the use of textual inversion or generating large-scale datasets has become a novel approach for zero-shot CIR task (ZS-CIR). However, the existing ZS-CIR models overlook one case where the textual description is often too brief or inherently inaccurate, making it challenging to effectively integrate the reference image into the query for retrieving the target image. To address this problem, we propose a simple yet effective method—prompting vision-language fusion (PVLF), which adapts representations in VLP models to dynamically fuse the vision and language (V&L) representation spaces. In addition, by injecting the context learnable prompt tokens in Transformer fusion encoder, the PVLF promotes the comprehensive coupling between V&L modalities, enriching the semantic representation of the query. We evaluate the effectiveness and robustness of our method on various VLP backbones, and the experimental results show that the proposed PVLF outperforms previous methods and achieves the state-of-the-art on two public ZS-CIR benchmarks (CIRR and FashionIQ).
layout inproceedings
issn 2640-3498
id wang25d
tex_title Prompting vision-language fusion for Zero-Shot Composed Image Retrieval
firstpage 671
lastpage 686
page 671-686
order 671
cycles false
bibtex_editor Nguyen, Vu and Lin, Hsuan-Tien
editor
given family
Vu
Nguyen
given family
Hsuan-Tien
Lin
bibtex_author Wang, Peng and Chen, Zining and Zhao, Zhicheng and Su, Fei
author
given family
Peng
Wang
given family
Zining
Chen
given family
Zhicheng
Zhao
given family
Fei
Su
date 2025-01-14
address
container-title Proceedings of the 16th Asian Conference on Machine Learning
genre inproceedings
issued
date-parts
2025
1
14
extras