Skip to content
AI Atlas
PaperActive

KernelGenBench: Can LLMs and Agents Write Efficient Kernels Across Operator Sources and Hardware Platforms?

arxiv.org/abs/2607.27231

Updated 50 min ago · first seen 11 Sept 2026

paper_01M294GP1SRQ32T660HNS50TJF

Published
11 Sept 2026
T1 · 50 min ago
arXiv
2607.27231
T1 · 50 min ago
Category
cs.AI
T1 · 50 min ago

Abstract

Modern AI systems depend on specialized accelerator kernels, whose development is complicated by increasingly diverse operators and hardware. LLMs and agentic systems promise to automate this work, but existing evaluations do not show whether their performance transfers across operator sources and hardware platforms, or what such transfer costs. We present KernelGenBench, the first unified multi-source and multi-chip infrastructure for evaluating LLM- and agent-generated Triton kernels. With a common Triton target spanning six hardware platforms, it provides the broadest cross-vendor hardware coverage among existing kernel-generation benchmarks. We report two controlled analytical views: KernelGenBench-MS (Multi-Source) covers 210 operators from PyTorch ATen, production vLLM operators, and proprietary cuBLAS routines, while KernelGenBench-MC (Multi-Chip) evaluates a semantically stable 110-operator subset across six hardware platforms. Our evaluation consumed over 15 billion tokens. Agentic execution improved correctness, but no method dominated across sources and platforms: vLLM posed the strongest correctness challenge, cuBLAS set the highest performance ceiling, and AutoKernel accuracy fell from 87% on NVIDIA to 25% on Iluvatar CoreX. These improvements were costly: specialized agents averaged 4.99 million tokens per successful operator, rising to 6.25 million for CUDA Optimized Skill. The results establish operator source, hardware platform, and agentic scaffold as distinct dimensions of kernel-generation capability, and show that success in a familiar source-hardware setting is not a reliable proxy for deployment readiness.

Authors 7

Peiyu Zang, Jian Tao, Jialing Zhang, Yichen Yuan, Wentao Zhang, Guang Liu, Yonghua Lin

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

arXiv id
2607.27231

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Categories
cs.AI, cs.LG

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 50 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

50 min ago

Conflicts

None