Transferring AI Computing Expertise from CUDA to Apple Silicon
Researchers created a system that automatically adapts high-performance AI code from NVIDIA's CUDA ecosystem to run efficiently on Apple Silicon chips. This method matched expert-level performance on attention tasks and sped up Mamba prefill tasks.
In short: Researchers created a system that automatically adapts high-performance AI code from NVIDIA's CUDA ecosystem to run efficiently on Apple Silicon chips. This method matched expert-level performance on attention tasks and sped up Mamba prefill tasks.
Writing low-level code to make graphics chips run artificial intelligence efficiently is notoriously difficult. Now, a new artificial intelligence tool uses past computing knowledge to bridge the gap between different computer chips without requiring years of manual work.
What happened, in plain words
Researchers from IBM Research and UC Berkeley Sky Lab updated an artificial intelligence optimization tool called K-Search to work with Apple's machine-learning framework, called MLX. Because older systems like NVIDIA's CUDA have decades of hand-tuned code expertise, the team built a structured translation layer. This layer maps out hardware rules and memory limits, allowing an AI model to take existing code and rewrite it into high-performing GPU kernels for Apple Silicon chips. When tested, the new method achieved near-expert performance on attention kernels and provided up to a 20 times prefill speedup on the Mamba state-space model kernel compared to a community implementation.
Key points
- GPU kernels are crucial for AI GPU kernels are low-level programs that run inside graphics processors to handle heavy AI workloads, but writing efficient ones takes years of expertise.
- Adapting K-Search for Apple Silicon Researchers extended K-Search, an evolutionary optimization framework using a large language model, by adding a backend that compiles and executes code on Apple Silicon.
- Bridging the gap with a translation layer A structured translation layer provides the AI with concept mapping tables and architectural rules, helping it turn older NVIDIA code into valid, high-quality Apple Metal code.
- Strong speed gains on real hardware The newly generated attention kernel reached 0.97x the speed of Apple's native kernel, and the Mamba SSM kernel achieved up to a 20x prefill speedup over the community mlx-lm package.
Terms explained
- GPU kernel — A specialized low-level program that tells a graphics processing unit how to perform a specific math operation for artificial intelligence. Example: Like a recipe that tells a kitchen crew exactly how to chop vegetables simultaneously instead of one by one.
- Unified memory architecture — A system design where the processor and the graphics chip share the same pool of memory, removing the need to copy data back and forth. Example: Like two chefs sharing a single large kitchen counter instead of constantly walking data across the room to separate prep stations.
- Parallel scan — An algorithm that breaks a step-by-step math problem into pieces so multiple computing lanes can solve them at the same time. Example: Like an entire row of people adding up separate parts of a long list simultaneously instead of waiting for one person to do it sequentially.
Why it matters
As different companies build specialized computer chips for artificial intelligence, rewriting software for every single chip is a major bottleneck. Automated tools that transfer old performance knowledge to new hardware can help developers run local AI models faster and more efficiently on consumer devices like MacBooks.
What we still don't know
The research only tested this approach on a couple of specific kernels, such as attention and Mamba SSM, and it remains unknown how well the method will generalize to other hardware targets and broader software ecosystems.
Based on reporting from Berkeley AI Research Blog. This is an independent explainer, written in our own words with AI assistance; Berkeley AI Research Blog has not reviewed or endorsed it. Read the original for the full details.