Skip to content

Latest commit

 

History

History
83 lines (83 loc) · 3.12 KB

File metadata and controls

83 lines (83 loc) · 3.12 KB
title Dude: Dual Distribution-Aware Context Prompt Learning For Large Vision-Language Model
booktitle Proceedings of the 16th Asian Conference on Machine Learning
year 2025
volume 260
series Proceedings of Machine Learning Research
month 0
publisher PMLR
pdf https://raw.githubusercontent.com/mlresearch/v260/main/assets/nguyen25c/nguyen25c.pdf
url https://proceedings.mlr.press/v260/nguyen25c.html
openreview SZ4LNs0zfE
abstract Prompt learning methods are gaining increasing attention due to their ability to customize large vision-language models to new domains using pre-trained contextual knowledge and minimal training data. However, existing works typically rely on optimizing unified prompt inputs, often struggling with fine-grained classification tasks due to insufficient discriminative attributes. To tackle this, we consider a new framework based on a dual context of both domain-shared and class-specific contexts, where the latter is generated by Large Language Models (LLMs) such as GPTs. Such dual prompt methods enhance the model’s feature representation by joining implicit and explicit factors encoded in LLM knowledge. Moreover, we formulate the Unbalanced Optimal Transport (UOT) theory to quantify the relationships between constructed prompts and visual tokens. Through partial matching, UOT can properly align discrete sets of visual tokens and prompt embeddings under different mass distributions, which is particularly valuable for handling irrelevant or noisy elements, ensuring that the preservation of mass does not restrict transport solutions. Furthermore, UOT’s characteristics integrate seamlessly with image augmentation, expanding the training sample pool while maintaining a reasonable distance between perturbed images and prompt inputs. Extensive experiments across few-shot classification and adapter settings substantiate the superiority of our model over current state-of-the-art baselines.
layout inproceedings
issn 2640-3498
id nguyen25c
tex_title {Dude}: {D}ual Distribution-Aware Context Prompt Learning For Large Vision-Language Model
firstpage 687
lastpage 702
page 687-702
order 687
cycles false
bibtex_editor Nguyen, Vu and Lin, Hsuan-Tien
editor
given family
Vu
Nguyen
given family
Hsuan-Tien
Lin
bibtex_author Nguyen, Duy Minh Ho and Le, An Thai and Nguyen, Trung Quoc and Diep, Nghiem Tuong and Nguyen, Tai and Duong-Tran, Duy and Peters, Jan and Shen, Li and Niepert, Mathias and Sonntag, Daniel
author
given family
Duy Minh Ho
Nguyen
given family
An Thai
Le
given family
Trung Quoc
Nguyen
given family
Nghiem Tuong
Diep
given family
Tai
Nguyen
given family
Duy
Duong-Tran
given family
Jan
Peters
given family
Li
Shen
given family
Mathias
Niepert
given family
Daniel
Sonntag
date 2025-01-14
address
container-title Proceedings of the 16th Asian Conference on Machine Learning
genre inproceedings
issued
date-parts
2025
1
14
extras