Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Benchmarks GitHub Repositories
Explore popular GitHub repositories tagged “benchmarks”.
Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.
Trending Repositories
No repository description provided.
Prime number projects in 100+ programming languages, to compare their speed - and their programmer's cleverness
Some benchmarks of different languages
VPS Fusion Monster Server Test GO Version Aiming to be the most comprehensive server testing project, implemented in Go with zero environment dependencies. VPS融合怪服务器测评项目 GO版本 尽量成为最全能的服务器测评项目,使用 Go 实现,无需任何环境依赖。
Avalanche: an End-to-End Library for Continual Learning based on PyTorch.
C++14 concurrent lock-free low-latency queue.
🏃♂️🏃♀️🏃 JS minification benchmarks: babel-minify, esbuild, terser, uglify-js, swc, google closure compiler, tdewolff/minify, oxc-minify
Therapeutics Commons (TDC): Multimodal Foundation for Therapeutic Science
Classifies GPUs based on their 3D rendering benchmark score allowing the developer to provide sensible default settings for graphically intensive applications.
Evaluate and improve models and agents using environments
Jeff Geerling's SBC review data - Raspberry Pi, Radxa, Orange Pi, etc.
Repo for AI Agents The Definitive Guide
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
📊 Benchmark Comparison of Packages with Runtime Validation and TypeScript Support
A curated list of 3D Vision papers relating to Robotics domain in the era of large models i.e. LLMs/VLMs, inspired by awesome-computer-vision, including papers, codes, and related websites
Yet another implementation of computer language benchmarks game
HammerDB: The industry standard open-source database benchmark
XBOW Validation Benchmarks
Benchmarks for Low Latency (Streaming) solutions including Apache Storm, Apache Spark, Apache Flink, ...
The wise choice for Ruby memoization
A unified framework for robot learning
[NeurIPS D&B '25] The one-stop repository for LLM unlearning
Deliver safe & effective language models
CodaLab Competitions
An in-the-wild benchmark for AI agents in the OpenClaw Environment.
TUI and CLI for browsing AI models, benchmarks, coding agents, and statuses for AI providers.
Swift benchmark runner with many performance metrics and great CI support
📊 Comparing deno, node and bun HTTP frameworks
NeurIPS 2023: Safe Policy Optimization: A benchmark repository for safe reinforcement learning algorithms