papersSEP 10 04:00 UTC
StreamAlign: New Streaming Text-Aligned Speech Tokenization Approach for LLMs
Researchers introduce StreamAlign, a speech tokenization method that maps audio into tokens aligned with large language model token spaces while operating in a streaming, low-latency manner. Unlike existing text-aligned tokenizers that depend on offline automatic speech recognition, the approach addresses the latency and alignment limitations that offline processing imposes, enabling more efficient use of pretrained LLMs for speech tasks.