Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
Published 16 Sept 2026arXiv:2609.15671
Updated 7 h ago · first seen 15 Sept 2026
paper_01M2JK19E586V8G9WF0GVQM1AH
Abstract
Visual Question Answering (VQA) with Vision-Language Models (VLMs) is increasingly used in privacy-sensitive and bandwidth-constrained settings. Federated Learning (FL), Split Learning (SL), and U-Shaped Split Learning (USL) keep raw data local, but transmitting all visual tokens across a model partition remains costly and can expose private information. We propose QPriv-VL, a question-guided, privacy-aware token-pruning framework for FL, SL, and USL that prunes visual tokens before transmission based on task utility and privacy sensitivity. Its core component is a lightweight Dynamic Threshold Predictor (DTP) that jointly estimates a sample-specific pruning ratio and a token-level retention mask in one forward pass. DTP combines question relevance, computed from cross-modal similarity between visual patches and the pooled question embedding, with a sensitivity signal derived from frozen DINOv2 features. This allows the model to suppress potentially sensitive regions while preserving patches useful for answering the question, without requiring sensitivity labels. We evaluate QPriv-VL on GQA, OK-VQA, VQAv2, SLAKE, VQA-RAD, and PathVQA against four privacy attack families: FSHA, FORA, iDLG, and attribute-inference membership inference attacks. DTP matches or outperforms fixed-ratio pruning while using substantially fewer transmitted tokens. On VQA-RAD, it reduces membership-inference attack success from 0.99 to 0.76-0.79, lowers FSHA and FORA reconstruction PSNR relative to fixed-ratio pruning, and preserves competitive VQA accuracy using about 40% of the original visual-token budget. A sensitivity exclusion ratio of 1.20 +/- 0.18 indicates preferential removal of privacy-sensitive patches, while explainability analysis shows that retention adapts to question semantics rather than generic visual saliency.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 4
- Property changedPaperDon't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - Property changedPaperDon't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxiv - Property changedPaperDon't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperDon't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
New paper: Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.