CAL-MOS Uses Layer Adapters to Improve Speech Quality Prediction Across Foundation Models
A new arXiv paper introduces CAL-MOS, a method that uses adapters to combine representations from multiple layers of speech foundation models for non-intrusive speech quality assessment. The approach aims to make mean opinion score (MOS) prediction more robust when the underlying foundation model varies. The abstract notes that selecting which layer's representations to use remains an open question in this area.