Speeding up GPU kernels by 38% with a multi-agent system
Blog post from Cursor
Researchers have developed a multi-agent system capable of autonomously building, maintaining, and deploying complex software, which they tested by optimizing CUDA kernels crucial for AI model training and inference on NVIDIA GPUs. In collaboration with NVIDIA, the system tackled 235 optimization problems and achieved a 38% geometric mean speedup by building and optimizing Blackwell GPU kernels from scratch, demonstrating the potential to significantly enhance GPU performance, reduce energy consumption, and lower costs. This accomplishment, typically requiring months or years from experienced kernel engineers, was achieved in weeks, indicating the system's ability to address complex, open-ended optimization problems by exploring a broader solution space beyond traditional manual methods. Using SOL-ExecBench for problem generation and benchmarking, the system exceeded baseline performance on 63% of problems and delivered over 2x improvements on 19% of them. The experiment highlighted the system's adaptability in employing distinct optimization strategies across various real-world constraints, suggesting that multi-agent architectures could soon become the standard in software development to address novel challenges that exceed current training data distributions.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Multi-agent systems | 24 | 460 | 170 | 68 | -20% |
| LLM | 3 | 5,932 | 1,046 | 223 | -2% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.