[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
GitHub Trending Cuda This Month
Discover GitHub Trending Cuda repositories gaining attention this month.
Trending Repositories
thu-ml/SageAttention
164 stars this month
deepseek-ai/DeepEP
DeepEP: an efficient expert-parallel communication library
166 stars this month
deepseek-ai/DeepGEMM
DeepGEMM: clean and efficient BLAS kernel library on GPU
179 stars this month
alibaba/rtp-llm
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
32 stars this month
NVIDIA/cuvs
cuVS - a library for vector search and clustering on the GPU
19 stars this month
NVIDIA/nccl-tests
NCCL Tests
32 stars this month
NVIDIA/cuopt
GPU accelerated decision optimization
35 stars this month
mirage-project/mirage
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
59 stars this month
BBuf/how-to-optim-algorithm-in-cuda
how to optimize some algorithm in cuda.
67 stars this month
NVlabs/instant-ngp
Instant neural graphics primitives: lightning fast NeRF and more
42 stars this month
karpathy/llm.c
LLM training in simple, raw C/CUDA
288 stars this month
HazyResearch/ThunderKittens
Tile primitives for speedy kernels
86 stars this month