MoonshotAI/FlashKDA
FlashKDA: high-performance Kimi Delta Attention kernels
Cuda1.2K stars118 forks+91 today
What it does
FlashKDA is a high-performance CUDA kernel for Kimi Delta Attention, optimized for SM90 and above GPUs with CUDA 12.9+. It integrates seamlessly with flash-linear-attention to enhance performance in AI applications.
Star history
Not enough history yet — 1 day(s) recorded. The daily snapshot builds this up.
Tracking
- Last trending
- 2026-08-02
Creator kit
Hook
🚀 Dive into the future of GPU-accelerated AI with FlashKDA, a cutting-edge CUDA kernel for Kimi Delta Attention. #AI #CUDA
Content angles
- Create a tutorial on how to integrate FlashKDA into existing PyTorch projects.
- Analyze performance benchmarks and discuss real-world applications in machine learning research.
- Develop a comparative study of different GPU optimization techniques, including the benefits of using FlashKDA.
Who should care
AI researchers, developers working with CUDA and PyTorch, and anyone interested in high-performance computing for deep learning.