KernelBench is a reproducible benchmark and toolkit that measures whether large language models can generate correct and high-performance GPU kernels by transpiling PyTorch operators into CUDA or other DSL kernels. Framed as an ICML ’25 project with dataset releases on Hugging Face and arXiv/blog references, the tasks are organized into four levels: Level 1 (100 single-kernel operators like convolutions and matmuls), Level 2 (100 simple fused kernels), Level 3 (50 full model architectures such as MobileNet and MiniGPT), and Level 4 (optimizing Hugging Face models). The repository supplies benchmark dataset files, baseline timing results across hardware, example notebooks, and unit tests so experiments can be reproduced and extended to other backends and GPU vendors.
Evaluation is explicit and automated: generated kernels are validated for correctness against reference PyTorch operators on randomized inputs and profiled for speed over repeated trials. The core metric fast_p counts the fraction of tasks that are both correct and faster than the reference by a factor p (fast_1, fast_2, fast_0 for correctness). The repo includes scripts to generate samples, run single or batched evaluations (including cloud Modal support), compute benchmark summaries (greedy_analysis.py, benchmark_eval_analysis.py), and utilities for setup (pyproject.toml with uv, litellm API use, .env, ROCm/hip and ThunderKittens notes). Configuration exposes GPU_ARCH, precision (fp32/fp16/bf16), and backend choices (cuda, triton, hip, etc.), enabling researchers to reproduce or adapt large-scale LLM-to-kernel experiments.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.