GitStar
GitHub TrendingTopicsLanguages
/

Blackwell GitHub Repositories

Explore popular GitHub repositories tagged “blackwell”.

Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.

RepositoriesGitHub topic
Ranked by:Stars

Trending Repositories

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

Python89,34620,852
sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

Python32,0217,982
NVIDIA/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Python14,4032,668
openlake-project/openlake

OpenLake is a high performance storage engine for efficient LLM inference and GPU Training

Rust2,329413
lightseekorg/tokenspeed

TokenSpeed is a speed-of-light LLM inference engine.

Python1,926244
GradientHQ/parallax

Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere

Python1,367144
NVIDIA/cudnn-frontend

cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

Python907252
AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-DFlash

Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.

Python44546
patrick-toulme/pyptx

A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch

Python37431
avifenesh/memra

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

Rust31435
IST-DASLab/qutlass

QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning

C++19825
0xSero/glm-5.2-sm120

GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode

Shell15411
5p00kyy/club-5060ti

Practical local LLM recipes and benchmarks for RTX 5060 Ti setups

Python1179
eelbaz/dgx-spark-vllm-setup

One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)

Shell10517
Saganaki22/ComfyUI-sol-attn

NVIDIA Sol-Attn for ComfyUI / Triton kernel on SM89 - SM121, with zero-copy MiniMax H3 nodes: memory-efficient attention, scheduled tau with graph preview, and feed-forward chunking. Measured 1.14–1.44× vs SageAttention and −37% MLP peak VRAM on H3

Python9412
dougeeai/llama-cpp-python-wheels

Pre-built wheels for llama-cpp-python across platforms and CUDA versions

833
AEON-7/comfyui-aeon-spark

Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs + Flux 2 Dev / LTX 2.3 22B / ACE-Step v1.5 XL Turbo pre-bundled with abliterated text-encoder paths.

Shell7917
AEON-7/vllm-dflash

DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding

Python549
6Morpheus6/deepspeed-windows-wheels

Prebuilt DeepSpeed wheels for Windows with NVIDIA GPU support. Supports GTX 10 - RTX 50 series. Compiled with pytorch 2.7, 2.8 and cuda 12.8

505
CosmicRaisins/glm-5.2-gb10

GLM-5.2 (744B/40B MoE) on a 4× DGX Spark / GB10 (sm_121) cluster: portable Triton sparse-MLA kernels, a data-free expert prune, MTP draft, and a one-script bootstrap.

Python387
kekzl/imp

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents. Decode beats llama.cpp b9976 by 42-48% on dense GGUF (measured 2026-07-12), at-or-ahead of vLLM on NVFP4. 100% written by Claude Code.

Cuda352
bidual/awesome-dgx-spark

A curated list of tools, guides, playbooks, and resources for the NVIDIA DGX Spark (GB10 Grace Blackwell personal AI supercomputer).

Shell344
hiroki-abe-58/ComfyUI-Win-Blackwell

No repository description provided.

PowerShell322
croll83/llama.cpp-dgx

llama.cpp fork optimized for NVIDIA DGX Spark / GB10 (Blackwell, SM 12.1) — TurboQuant weights + KV, NVFP4, DFlash MTP

C++313
Mekopa/whisperx-blackwell

GPU-accelerated WhisperX on NVIDIA Blackwell (SM_121) - DGX Spark compatible

Python309
egaoharu-kensei/flash-attention-triton

Cross-platform FlashAttention-2 Triton implementation for Turing+ GPUs with custom configuration mode

Python290
actypedef/ARCQuant

[ACL 2026 Main] Code for the paper "ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs"

Cuda298
MerkyorLynn/lynn-engine

Lynn 原生 LLM 推理引擎 · W4A8/NVFP4 量化 · 自写 CUDA/Triton kernel · MoE · 投机解码 | Lynn-native LLM inference engine for NVIDIA Blackwell

Python260
Sggin1/DGX-SPARK

DGX Spark research and tests - containers, benchmarks, and investigation notes for running models on GB10 (SM 12.1)

Python244
lna-lab/blackwell-geforce-nvfp4-gemm

NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.

Python212
GitStar

See what the GitStar community is most excited about today.

Trending

GitHub Trending TodayGitHub Trending WeeklyGitHub Trending Monthly

Languages

Browse all languagesTrending PythonTrending JavaScript

Explore

Browse GitHub topicsAI repositoriesDeveloper tools
© 2026 GitStarGitHub Trending source