1,062 stars | 91 forks | Cuda
FlashKDA: high-performance Kimi Delta Attention kernels
What it does
FlashKDA is a high-performance CUDA kernel for Kimi Delta Attention, optimized for SM90 and above GPUs with CUDA 12.9+. It integrates seamlessly with flash-linear-attention to enhance performance in AI applications.
Why it matters: 🚀 Dive into the future of GPU-accelerated AI with FlashKDA, a cutting-edge CUDA kernel for Kimi Delta Attention. #AI #CUDA
Trending today with 91 new stars
Want to create content about this repo? Use Nemati AI tools to generate articles, tutorials, and social posts.





