LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Inference Server GitHub Repositories
Explore popular GitHub repositories tagged “inference-server”.
Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.
Trending Repositories
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
Open-source inference server and production cluster for all the models your agent needs.
Turn any computer or edge device into a command center for your computer vision projects.
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
The simplest way to serve AI/ML models in production
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
An open-source computer vision framework to build and deploy apps in minutes
Python + Inference - Model Deployment library in Python. Simplest model inference server ever.
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
Work with LLMs on a local environment using containers
This is a repository for an nocode object detection inference API using the Yolov3 and Yolov4 Darknet framework.
llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
This is a repository for an nocode object detection inference API using the Yolov4 and Yolov3 Opencv.
ONNX Runtime Server: The ONNX Runtime Server is a server that provides TCP and HTTP/HTTPS REST APIs for ONNX inference.
This is a repository for an object detection inference API using the Tensorflow framework.
Serving AI/ML models in the open standard formats PMML and ONNX with both HTTP (REST API) and gRPC endpoints
Orkhon: ML Inference Framework and Server Runtime
Deploy DL/ ML inference pipelines with minimal extra code.
A standalone inference server for trained Rubix ML estimators.
Wingman is the fastest and easiest way to run Llama models on your PC or Mac.
👁 零代码零标注 CV AI 自动化测试工具 🚀 免除大量人工画框和打标签等,直接零代码快速自动化测试 CV 计算机视觉 AI 人工智能图像识别算法:行人检测、动植物分类、人脸识别、OCR 车牌识别、旋转校正、舞蹈姿态、抠图分割 等,还可一键 下载测试报告、导出训练和测试数据集
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and image models — many models sharing your GPUs, one gateway. Powered by Ray Serve.
Advanced inference pipeline using NVIDIA Triton Inference Server for CRAFT Text detection (Pytorch), included converter from Pytorch -> ONNX -> TensorRT, Inference pipelines (TensorRT, Triton server - multi-format). Supported model format for Triton inference: TensorRT engine, Torchscript, ONNX
Fullstack machine learning inference template
Benchmark for machine learning model online serving (LLM, embedding, Stable-Diffusion, Whisper)
Inference Server Implementation from Scratch for Machine Learning Models
Local inference server for Apple Silicon — hot-swaps MLX models (LLM, vision, embeddings, TTS, STT) via OpenAI API
vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.