GitStar
GitHub TrendingTopicsLanguages
/

Gguf GitHub Repositories

Explore popular GitHub repositories tagged “gguf”.

Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.

RepositoriesGitHub topic
Ranked by:Stars

Trending Repositories

AlexsJones/llmfit

Hundreds of models & providers. One command to find what runs on your hardware.

Rust32,5952,014
mozilla-ai/llamafile

Distribute and run LLMs with a single file.

C++25,6261,558
Andyyyy64/whichllm

Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

Python6,335341
Michael-A-Kuykendall/shimmy

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

Rust5,760555
off-grid-ai/OGAM

The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.

TypeScript2,941280
Mobile-Artificial-Intelligence/maid

Maid is a free and open source application for interfacing with llama.cpp models locally, and with Anthropic, DeepSeek, Ollama, Mistral and OpenAI models remotely.

TypeScript2,636279
datawhalechina/handy-ollama

动手学Ollama,CPU玩转大模型部署,在线阅读地址:https://datawhalechina.github.io/handy-ollama/

Jupyter Notebook2,500314
heshengtao/comfyui_LLM_party

LLM Agent Framework in ComfyUI includes MCP sever, Omost,GPT-sovits, ChatTTS,GOT-OCR2.0, and FLUX prompt nodes,access to Feishu,discord,and adapts to all llms with similar openai / aisuite interfaces, such as o1,ollama, gemini, grok, qwen, GLM, deepseek, kimi,doubao. Adapted to local llms, vlm, gguf such as llama-3.3 Janus-Pro, Linkage graphRAG

Python2,337200
AtomicBot-ai/atomic-agent

Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.

TypeScript2,301216
MakazhanAlpamys/Soup

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

Python2,218327
withcatai/node-llama-cpp

Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level

TypeScript2,155212
sammcj/gollama

Go manage your Ollama models

Go1,835109
handy-computer/transcribe.cpp

ggml speech-to-text inference for 16+ model families

C++1,80284
intel/auto-round

A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.

Python1,569166
QwenAudio/Fun-ASR

Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.

C1,479146
edwko/OuteTTS

Interface for OuteTTS models.

Python1,435118
kitops-ml/kitops

An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.

Go1,401183
AtomicBot-ai/Atomic-Chat

Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V

TypeScript1,313144
PurpleDoubleD/locally-uncensored

Plug-and-play local AI studio: uncensored chat, image & video generation, coding agent. Runs abliterated LLMs + ComfyUI 100% offline. One installer, no Docker, no cloud.

TypeScript1,087178
alvarobartt/hf-mem

A CLI to estimate inference memory requirements for Hugging Face models, written in Python.

Python93884
eastriverlee/LLM.swift

LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.

Swift870124
mukel/llama3.java

Llama 3+ inference in pure Java

Java81694
jegly/Box

The most advanced, fully offline client-side AI suite on Android today.

Kotlin76245
ddalcu/mlx-serve

Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.

Zig72552
gokayfem/ComfyUI_VLM_nodes

ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.

Python58362
ciddwd/overlay-translator

无需 ROOT 的开源 Android 屏幕实时翻译工具,适合游戏、视觉小说和漫画。支持端侧与云端 OCR、离线 LLM、多种翻译服务和文字朗读(TTS),译文可直接显示在画面上。Open-source no-root Android real-time screen translator for games, visual novels, and manga. Supports on-device and cloud OCR, offline LLMs, multiple translation services, on-screen translations, and text-to-speech (TTS).

Kotlin57622
kelindar/search

Go library for embedded vector search and semantic embeddings using llama.cpp

Go55924
hybridgroup/yzma

Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.

Go55922
withcatai/catai

Run AI ✨ assistant locally! with simple API for Node.js 🚀

TypeScript49839
AudarAI/Audar-ASR-V1

Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.

Python4943
GitStar

See what the GitStar community is most excited about today.

Trending

GitHub Trending TodayGitHub Trending WeeklyGitHub Trending Monthly

Languages

Browse all languagesTrending PythonTrending JavaScript

Explore

Browse GitHub topicsAI repositoriesDeveloper tools
© 2026 GitStarGitHub Trending source