Skip to content
AI Atlas
PaperActive

Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models

arxiv.org/abs/2609.11310

Updated 22 min ago · first seen 11 Sept 2026

paper_01M294FQVVJZMH8FB1S0DWMF2R

Published
11 Sept 2026
T1 · 23 min ago
arXiv
2609.11310
T1 · 23 min ago
Category
cs.CV
T1 · 23 min ago

Abstract

We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting, where a small number of continuous prompt tokens are optimized while the pretrained backbone remains frozen. We identify two key design choices. First, placing prompt tokens at the cross-modal boundary between visual and text tokens outperforms other placements (10.0 vs. 8.4 mAP). Second, initializing prompts from the empty space token outperforms semantic and random initialization. With these choices, one to three learned tokens (7,168 parameters on average) match the best LoRA configuration on Roboflow20-VL (14.2 mAP, 10-shot) while training over 20,000x fewer parameters. Soft prompting remains harder to optimize, exhibiting higher variance across random seeds. Unlike LoRA, however, it causes no forgetting: the LoRA rank matching our accuracy reduces NaturalBench VQA accuracy by 35% relative, rising to 56% at the largest rank, whereas soft prompting leaves pretrained performance unchanged. The learned tokens behave like prompts rather than weights. They transfer to a newer model without retraining (+0.8 mAP on Qwen3.5-9B) and can be verbalized into readable prompts competitive with prompt-search methods (matching DetPO and outperforming GEPA). The approach also extends beyond detection. On RoboCasa manipulation tasks, the frozen $\pi_{0.5}$ vision-language-action policy benefits from soft prompting, matching the LoRA baseline on two of three tasks when tokens are placed at the gradient bottleneck. These results suggest modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask.

Authors 8

Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao, Wolfgang M. Pauli, John Galeotti, Deva Ramanan

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 22 min agohigh

arXiv id
2609.11310

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Categories
cs.CV, cs.AI, cs.LG, eess.IV, stat.ML

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Primary category
cs.CV

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 23 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

22 min ago

Conflicts

None