papersSEP 10 04:00 UTC
On-Policy Distillation Proposed for Vision-Language Model Adaptation on Low-Quality Data
A new arXiv paper introduces an on-policy distillation approach for adapting compact vision-language models from a larger task-trained teacher. Rather than relying solely on teacher predictions as training targets, the method lets the student learn from its own outputs, which the authors report makes it especially effective when multimodal training data is noisy or low quality.