papersSEP 10 04:00 UTC
KernelGenBench Tests Whether LLMs and Agents Can Write Efficient Kernels Across Hardware
Researchers introduced KernelGenBench, a benchmark that evaluates how well large language models and agentic systems can produce specialized accelerator kernels. The benchmark assesses code generation across diverse operator sources and hardware platforms, aiming to fill a gap left by earlier evaluations of kernel-writing capability.