RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
Published 16 Sept 2026arXiv:2609.14100
Updated 12 h ago · first seen 15 Sept 2026
paper_01M2JK0CG9XD437B2GSF5W5PZE
Abstract
Fashion Image Captioning (FIC) plays a vital role in enhancing user experience and product search in e-commerce platforms. Unlike natural scene image captioning, FIC requires fine-grained visual reasoning and knowledge of domain-specific terminology to capture subtle attributes such as neckline and closure types, graphic patterns, and dress silhouettes. Moreover, as fashion inventories evolve rapidly with new trends, styles, and frequently emerging vocabulary, developing training-free captioning solution becomes essential for scalability and real-world adaptability. Instruction-tuned vision-language models (VLMs) offer a promising solution to fashion image captioning dueto their strong zero-shot capabilities and natural language fluency. However, these general-purpose models often lack attribute-level coverage and precision, and tend to hallucinate or misidentify fine-grained fashion details, making them less suitable for high-fidelity applications like product cataloging or personalized recommendations. To address this, we propose RA-CoA (Retrieval-Augmented Chain-of-Attributes), a novel, training-free framework that disentangles fashion image captioning into two interpretable stages: (i) retrieval of relevant attribute sets from a product knowledge base, and (ii) attribute-level reasoning to generate the final caption. RA-CoA is a model-agnostic approach that works with frozen VLMs to improve fine-grained attribute precision in product captions without the need for fine-tuning. Extensive evaluations across diverse VLM model families under different prompting paradigms demonstrate that RA-CoA significantly improves caption quality, achieving an average gain of 26.3% METEOR score over zero-shot captioning. We make our code publicly available.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 4
- Property changedPaperRA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - Property changedPaperRA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxiv - Property changedPaperRA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes: arxiv announce type changed from cross to new
Arxiv announce typecross→newarxiv - New paperPaperRA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
New paper: RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
arxiv
Sources
Sources 3
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.