Master AI inference, AI agent harness systems, and hardware engineering — then design a physical AI chip. That is the goal.
- Stars
- 255
- Forks
- 37
- Pushed
- Today
Current GitHub adoption and maintenance signals.
Master AI inference, AI agent harness systems, and hardware engineering — then design a physical AI chip. That is the goal.
12 Weeks, 24 Lessons, IoT for All!
AI accelerators, edge inference devices, compilers, runtimes, benchmarks, and research for building and evaluating machine-learning systems.
Vitis AI is Xilinx’s development stack for AI inference on Xilinx hardware platforms, including both edge devices and Alveo cards.
A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do
Fastest Miner for Neptune Cash
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
Privacy-first OpenAI-compatible AI gateway: anonymize agent prompts & tool I/O locally, rent reasoning from OpenRouter free models.
Interactive 3D visualization of dense decoder-only LLM inference. Companion to the AI Inference Engineer 2026 course.
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
Semi-automated research assistant for academic research and software development. Supports Claude Code, OpenCode, and Codex CLI across ideation, coding, experiments, writing, and publication.
😎 Awesome lists about all kinds of interesting topics
No repository description provided.
Hosted Solution (Jetson Linux) with ESP32 (Wi-Fi + BT + BLE)
Deep learning for dummies. All the practical details and useful utilities that go into working with real models.
The agent that grows with you
No repository description provided.
Arena for LLM Inference Server Optimization
slime is an LLM post-training framework for RL Scaling.
Autonomous nsys profiling dataset generator for Qwythos fine-tuning (CUDA-L1 + KernelBench)
An open-source AI coding agent that lives in your terminal.
Ralph recipe
Confidential VM Image
No repository description provided.
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
No repository description provided.
An Optimizer for Nvidia Compilers.
TokenSpeed is a speed-of-light LLM inference engine.
No repository description provided.
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
Open Source Continuous Inference Benchmarking Qwen3.5, DeepSeek, GPTOSS - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 vs H100 & soon™ TPUv6e/v7/Trainium2/3
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Machine Learning Systems
No repository description provided.
No repository description provided.
compiler learning resources collect.
HomeKit smart home control via MCP — lights, locks, thermostats, and scenes for Claude Desktop, Claude Code, and OpenClaw
DFlash: Block Diffusion for Flash Speculative Decoding
OpenMeow dogfood architecture for the OpenClaw app SDK
No repository description provided.
AI agent toolkit: coding agent CLI, unified LLM API, TUI & web UI libraries, Slack bot, vLLM pods
Arduino core for the ESP32
🧃 Token weight loss. Lean output compaction for terminal-heavy agent workflows. Works as a native CLI tool or as an extension to popular coding and agent frameworks.
I design and build at the intersection of hardware and AI — from bare-metal firmware and real-time Linux kernels to ML compilers and on-device inference.
Real-Time VLAs via Future-state-aware Asynchronous Inference.
ESP32 Audio Developent boards: HiFi-ESP32, Loud-ESP32, Amped-ESP32, Louder-ESP32
Set of tools to assess and improve LLM security.
A model family accelerating the development of useful quantum computers
A Home Assistant integration & Model to control your smart home using a Local LLM
Learn it. Build it. Ship it for others.
anonymous peer-to-peer cash
oneAPI Deep Neural Network Library (oneDNN)
No repository description provided.
No repository description provided.
verl: Volcano Engine Reinforcement Learning for LLMs
Open standard for machine learning interoperability
FlashInfer: Kernel Library for LLM Serving
Universal LLM Deployment Engine with ML Compilation
Low-code framework for building custom LLMs, neural networks, and other AI models
My learning notes for ML SYS.
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
🧪 Default models for ⚗️ Instill Model
Tile primitives for speedy kernels
Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
A community collection of OpenClaw use cases for making life easier.
Official Problem Sets / Reference Kernels for the GPU MODE Leaderboard!
Device tree and kernel module for running TI TAS5805M DAC on Raspberry Pi
An optimized JPEG decoder suitable for microcontrollers and PCs.
VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
tutorials for IterX
ESP32 Music streaming based on Squeezelite, with support for multi-room sync, AirPlay, Bluetooth, Hardware buttons, display and more
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
Use AirPlay to stream to UPnP/Sonos & Chromecast devices
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
[NeurIPS2025] "AI-Researcher: Autonomous Scientific Innovation" -- A production-ready version: https://novix.science/chat
This repository includes the official implementation of OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs.
🚀🚀🚀 This repository lists some awesome public CUDA, cuda-python, cuBLAS, cuDNN, CUTLASS, TensorRT, TensorRT-LLM, Triton, TVM, MLIR, PTX and High Performance Computing (HPC) projects.
LLM training in simple, raw C/CUDA
🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object detection datasets.
📽 Capture and develop clips of openpilot. UI optional. Already deployed on Replicate.com for YOUR immediate use!
🤖 Assemble, configure, and deploy autonomous AI Agents in your browser.
A list of tutorials, paper, talks, and open-source projects for emerging compiler and architecture
A reference application for a local AI assistant with LLM and RAG
Running large language models on a single GPU for throughput-oriented scenarios.
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models, multimodal agents, speech, vector DB, and RAG.
OpenResearcher, an advanced Scientific Research Assistant
[ECCV 2022] This is the official implementation of BEVFormer, a camera-only framework for autonomous driving perception, e.g., 3D object detection and semantic map segmentation.
How to deploy open source models using DeepStream and Triton Inference Server
Serving multiple LoRA finetuned LLM as one
Style transfer, deep learning, feature transform
High frequency trading (HFT) framework built for futures using machine learning and deep learning techniques
Real-time pose estimation accelerated with NVIDIA TensorRT
Collection of undergraduate course homework and projects
This is a list of useful libraries and resources for CUDA development.