papersTODAY 04:00 UTC
Mizan benchmark evaluates LLMs on Iraqi Arabic and civic context
Researchers introduced Mizan, a benchmark designed to test large language models on Iraqi Arabic and on civic topics relevant to Iraq. Existing Arabic evaluation efforts have largely centered on Modern Standard Arabic, leaving regional dialects and country-specific knowledge thinly covered. The work aims to give a national-level measure of model performance beyond aggregated MSA leaderboards.