deepseek-ai/DeepGEMM
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
Cuda7.7K stars1.2K forks+109 today
What it does
DeepGEMM is a high-performance CUDA library for tensor core kernels, offering efficient FP8 GEMM operations and more. It simplifies GPU kernel optimization techniques while matching or surpassing expert-tuned libraries.
Star history
Not enough history yet — 1 day(s) recorded. The daily snapshot builds this up.
Tracking
- Last trending
- 2026-04-25
Creator kit
Hook
Revolutionize your AI workloads with DeepGEMM, a high-performance CUDA library for tensor core kernels. #AI #CUDA
Content angles
- Exploring the performance benefits of DeepGEMM in large language model training.
- A deep dive into how DeepGEMM optimizes FP8 GEMM operations on modern GPUs.
- Comparing DeepGEMM's efficiency and ease-of-use against traditional CUDA libraries.
Who should care
Researchers, developers, and engineers working with AI models and GPU optimization.