GitStar
GitHub TrendingTopicsLanguages
/

Inference Server GitHub Repositories

Explore popular GitHub repositories tagged “inference-server”.

Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.

RepositoriesGitHub topic
Ranked by:Stars

Trending Repositories

jundot/omlx

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Python19,1751,650
Michael-A-Kuykendall/shimmy

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

Rust5,760555
containers/ramalama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

Python3,001354
superlinked/sie

Open-source inference server and production cluster for all the models your agent needs.

Python2,795272
roboflow/inference

Turn any computer or edge device into a command center for your computer vision projects.

Python2,420304
waybarrios/vllm-mlx

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

Python1,516211
basetenlabs/truss

The simplest way to serve AI/ML models in production

Python1,192120
aiptimizer/TurboOCR

TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC

C++1,00997
pipeless-ai/pipeless

An open-source computer vision framework to build and deploy apps in minutes

Rust85252
underneathall/pinferencia

Python + Inference - Model Deployment library in Python. Simplest model inference server ever.

Python54383
Epistates/pmetal

PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.

Rust30824
containers/podman-desktop-extension-ai-lab

Work with LLMs on a local environment using containers

TypeScript29684
BMW-InnovationLab/BMW-YOLOv4-Inference-API-GPU

This is a repository for an nocode object detection inference API using the Yolov3 and Yolov4 Darknet framework.

Python27668
raketenkater/ggrun

llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.

Go26415
BMW-InnovationLab/BMW-YOLOv4-Inference-API-CPU

This is a repository for an nocode object detection inference API using the Yolov4 and Yolov3 Opencv.

Python21859
kibae/onnxruntime-server

ONNX Runtime Server: The ONNX Runtime Server is a server that provides TCP and HTTP/HTTPS REST APIs for ONNX inference.

C++19518
BMW-InnovationLab/BMW-TensorFlow-Inference-API-CPU

This is a repository for an object detection inference API using the Tensorflow framework.

Python17848
autodeployai/ai-serving

Serving AI/ML models in the open standard formats PMML and ONNX with both HTTP (REST API) and gRPC endpoints

Scala16631
vertexclique/orkhon

Orkhon: ML Inference Framework and Server Runtime

Rust1534
notAI-tech/fastDeploy

Deploy DL/ ML inference pipelines with minimal extra code.

Python10517
RubixML/Server

A standalone inference server for trained Rubix ML estimators.

PHP6313
curtisgray/wingman

Wingman is the fastest and easiest way to run Llama models on your PC or Mac.

TypeScript472
TommyLemon/CVAuto

👁 零代码零标注 CV AI 自动化测试工具 🚀 免除大量人工画框和打标签等,直接零代码快速自动化测试 CV 计算机视觉 AI 人工智能图像识别算法:行人检测、动植物分类、人脸识别、OCR 车牌识别、旋转校正、舞蹈姿态、抠图分割 等,还可一键 下载测试报告、导出训练和测试数据集

JavaScript424
modelship-ai/modelship

Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and image models — many models sharing your GPUs, one gateway. Powered by Ray Serve.

Python404
k9ele7en/Triton-TensorRT-Inference-CRAFT-pytorch

Advanced inference pipeline using NVIDIA Triton Inference Server for CRAFT Text detection (Pytorch), included converter from Pytorch -> ONNX -> TensorRT, Inference pipelines (TensorRT, Triton server - multi-format). Supported model format for Triton inference: TensorRT engine, Torchscript, ONNX

Python337
haicheviet/fullstack-machine-learning-inference

Fullstack machine learning inference template

Jupyter Notebook3112
tensorchord/inference-benchmark

Benchmark for machine learning model online serving (LLM, embedding, Stable-Diffusion, Whisper)

Python283
leimao/Simple-Inference-Server

Inference Server Implementation from Scratch for Machine Learning Models

Python241
raspoli/mlx-serve

Local inference server for Apple Silicon — hot-swaps MLX models (LLM, vision, embeddings, TTS, STT) via OpenAI API

Python182
hec-ovi/vllm-qwen

vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.

Python173
GitStar

See what the GitStar community is most excited about today.

Trending

GitHub Trending TodayGitHub Trending WeeklyGitHub Trending Monthly

Languages

Browse all languagesTrending PythonTrending JavaScript

Explore

Browse GitHub topicsAI repositoriesDeveloper tools
© 2026 GitStarGitHub Trending source