papersSEP 12 04:00 UTC
X-AuT Compresses Speech LLM Audio Encoders via Cross-Scale Distillation
Researchers propose X-AuT, a method that progressively compresses the audio encoder of speech large language models rather than deleting whole blocks at once. Because abrupt block removal distorts the embeddings the decoder receives and leads to word deletion and premature end-of-sequence errors, the approach uses cross-scale distillation to shrink encoder depth while preserving output quality. The aim is to cut inference cost without the accuracy loss typical of standard pruning.