The fastest path to AI-powered full stack observability, even for lean teams.
Observability GitHub Repositories
Explore popular GitHub repositories tagged “observability”.
Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.
Trending Repositories
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
SigNoz is an open-source, OpenTelemetry-native observability platform for your team and their AI agents. Get logs, metrics, and traces in one tool with features like APM, distributed tracing, log management, infra monitoring, etc. Combined with SigNoz MCP and a native AI teammate (in SigNoz Cloud) it helps you build more resilient apps.
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
eBPF-based Networking, Security, and Observability
APM, Application Performance Monitoring System
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
CNCF Jaeger, a Distributed Tracing Platform
A high-performance observability data pipeline.
Self-Hosting Guide. Learn all about locally hosting (on premises & private web servers) and managing software applications by yourself or your organization. Including Cloud, LLMs, WireGuard, Automation, Home Assistant, and Networking.
Why is this running? Trace any process, port, container, or file back to what started it - CLI + TUI.
End-to-end, code-first tutorials for building production-grade GenAI agents. From prototype to enterprise deployment.
Your window into all of your data
Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage costs and single binary deployment.
VictoriaMetrics: fast, cost-effective monitoring solution and time series database
Zipkin is a distributed tracing system
The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
Apache Doris is a real-time analytics and hybrid search database for AI agents.
Build production-ready applications in TypeScript
Highly available Prometheus setup with long term storage capabilities. A CNCF Incubating project.
Nightingale is to monitoring and alerting what Grafana is to visualization.
eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.
Continuous Profiling Platform. Debug performance issues down to a single line of code
Build your own AI SRE agents. The open source toolkit for the AI era.
AI Agent Engineering Platform built on an Open Source TypeScript AI Agent Framework
Resolve production issues, fast. An open source observability platform unifying session replays, logs, metrics, traces and errors powered by ClickHouse and OpenTelemetry.
A curated collection of publicly available resources on how technology and tech-savvy organizations around the world practice Site Reliability Engineering (SRE)
Free, local tool to track AI coding token usage and cost across 37 tools and agents (Claude Code, Cursor, Codex, Gemini and more), by model, project, and task. npx codeburn
highlight.io: The open source, full-stack monitoring platform. Error monitoring, session replay, logging, distributed tracing, and more.
🫖 Status page with uptime monitoring & API monitoring as code 🫖