papersSEP 10 04:00 UTC
PRISM-Bench: Audio-Centric Benchmark for Evaluating Text-to-Audio-Video Generation
Researchers have introduced PRISM-Bench, a diagnostic benchmark for text-to-audio-video generation models that places the audio modality at the center of evaluation. The paper argues that prior benchmarks tend to treat sound as a minor add-on to video quality metrics or test it separately from the combined audiovisual output. The new benchmark is intended to give a fuller picture of how generative systems handle audio together with visuals.