RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
Researchers propose RA-CoA, a training-free method for fashion image captioning that combines retrieval with a chain-of-attributes reasoning approach. The work targets e-commerce use cases, where captioning demands fine-grained visual analysis and correct domain-specific fashion terminology rather than generic scene description. It is published as an arXiv preprint in the cross-listed machine learning category.