LLM training in simple, raw C/CUDA
- Stars
- 30.8K
- Forks
- 3.7K
- Pushed
- Unknown
- Seen
- Today
Browse repositories that have appeared in the GitStar Trending catalog, ranked by adoption and recent activity.
LLM training in simple, raw C/CUDA
Instant neural graphics primitives: lightning fast NeRF and more
DeepEP: an efficient expert-parallel communication library
DeepGEMM: clean and efficient BLAS kernel library on GPU
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Tile primitives for speedy kernels
how to optimize some algorithm in cuda.
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
NCCL Tests
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
GPU accelerated decision optimization
cuVS - a library for vector search and clustering on the GPU