GitStar
GitHub TrendingTopicsLanguages
/

AI Benchmark GitHub Repositories

Explore popular GitHub repositories tagged “ai-benchmark”.

Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.

RepositoriesGitHub topic
Ranked by:Stars

Trending Repositories

microsoft/WindowsAgentArena

Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.

Python887101
TheAgentCompany/TheAgentCompany

An agent benchmark with tasks in a simulated software company.

Python766121
GoodStartLabs/AI_Diplomacy

Frontier Models playing the board game Diplomacy.

Python69696
rungalileo/agent-leaderboard

Ranking LLMs on agentic tasks

Jupyter Notebook22528
LeoYeAI/myclaw-bench

The definitive benchmark for AI agents on OpenClaw. 45 tasks across 4 tiers. Powered by MyClaw.ai

Python22336
seantins9/aiclothswap-showcase

AI clothes swap prompt cookbook with 50 virtual try-on prompts, KIE GPT Image before/after examples, failure fixes, and a benchmark rubric for AIClothSwap.

HTML1054
EastForce/worker-ai-spark

开源劳动者AI研究与评测项目:建设劳动知识、公开评测基准与共同治理机制|Open worker-centered AI research and evaluation project.

Python907
AnkitNayak-dev/llmBench

llmBench is a high-depth benchmarking tool designed to measure the raw performance of local LLM runtimes (Ollama, llama.cpp) while providing deep hardware intelligence.

Python474
chaosync-org/awesome-ai-agent-testing

🤖 A curated list of resources for testing AI agents - frameworks, methodologies, benchmarks, tools, and best practices for ensuring reliable, safe, and effective autonomous AI systems

4616
cowboycodr/sitegeist

A visual benchmark testing whether leading coding agents repeat the same design patterns across 100 neutral website briefs.

HTML416
qcai0427/TrustGraphBench

Benchmarking and evaluation resources for trust, robustness, and reliability analysis in graph-based AI systems.

Python310
oolong-tea-2026/arena-ai-leaderboards

📊 Daily auto-updated snapshots of all Arena AI (LMSYS Chatbot Arena) leaderboards — LLM, Vision, Code, Video, Image & more. Structured JSON with historical tracking.

Python298
ktwu01/benchmark-radar

Daily evidence-first radar for AI benchmarks, evaluations, datasets, and data quality.

Python249
petmal/MindTrial

MindTrial: Evaluate and compare AI language models (LLMs) on text-based tasks with optional file/image attachments and tool use. Supports multiple providers (OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, OpenRouter), custom tasks in YAML, and HTML/CSV/JSON reports.

Go175
tenurehq/precisionMemBench

Precision-aware retrieval benchmark for LLM memory systems.

TypeScript133
brandonhimpfen/awesome-ai-benchmarks-evaluation

A curated list of evaluation tools, benchmark datasets, leaderboards, frameworks, and resources for assessing model performance.

Python1010
lica-world/GDB

GDB: GraphicDesignBench - A real-world benchmark for evaluating AI on graphic design tasks

Python102
BennettSchwartz/ERR-EVAL

Benchmark for evaluating AI epistemic reliability - testing how well LLMs handle uncertainty, avoid hallucinations, and acknowledge what they don't know.

Python100
ctala/ai-benchmarks-alternativos

Benchmark abierto en español de 170 modelos de IA (118 con 20+ runs, 69 rankeados, juez Phi-4 independiente). Calidad, costo, velocidad, long-context y fuga de credenciales como dimensiones separadas. Alternativas a Claude, GPT y Gemini para agentes n8n/Hermes. Calculadora interactiva con tus propios pesos.

Python96
Habitante/gta-benchmark

GTA (Guess The Algorithm) Benchmark - A tool for testing AI reasoning capabilities

Python80
mlcommons/storage_results_v2.0

This repository contains the results and code for the MLPerf™ Storage v2.0 benchmark.

Python71
yasarshaikh/SF-bench

The first comprehensive benchmark for evaluating AI coding agents on Salesforce development tasks. Tests Apex, LWC, Flows, and more.

Python61
MetriLLM/metrillm

Benchmark local LLM models: speed, quality, and hardware fitness scoring. CLI, MCP server, and IDE plugins.

TypeScript50
vvsotnikov/astro-bench

Can AI agents do real science? Benchmarking AI agents on KASCADE cosmic ray classification

Python50
GitStar

See what the GitStar community is most excited about today.

Trending

GitHub Trending TodayGitHub Trending WeeklyGitHub Trending Monthly

Languages

Browse all languagesTrending PythonTrending JavaScript

Explore

Browse GitHub topicsAI repositoriesDeveloper tools
© 2026 GitStarGitHub Trending source