awesome-agent-observability
Health Pass
- License — License: CC0-1.0
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Community trust — 37 GitHub stars
Code Pass
- Code scan — Scanned 1 files during light audit, no dangerous patterns found
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
Curated list of tools, standards, and platforms for LLM and AI-agent observability: OpenTelemetry GenAI conventions, tracing, evals, guardrails, gateways, and MCP tooling.
Awesome Agent Observability 
Curated tools, standards, and platforms for tracing, evaluating, and governing LLM and AI-agent applications.
Agents fail in ways ordinary services do not: a run is non-deterministic, spans a dozen model and tool calls, and "wrong" is a quality judgement rather than a status code. The projects below cover the resulting stack — OpenTelemetry conventions for GenAI, tracing backends, evaluation harnesses, guardrails, gateways, and Model Context Protocol tooling.
Every entry was checked to resolve and to have been updated within the last 12 months at the time of the last audit (2026-10-01). Descriptions say what a project does, not what it markets.
Contents
- Standards and instrumentation
- Tracing and observability platforms
- Trace backends
- Evaluation and testing
- Guardrails and security
- Gateways and proxies
- Model Context Protocol
- Agent frameworks with built-in tracing
Standards and instrumentation
- OpenTelemetry GenAI Semantic Conventions - Repository that now owns the GenAI spans, metrics, and events, moved out of the main semantic-conventions repo.
- OpenTelemetry Semantic Conventions - Parent repo for all OTel conventions, including the HTTP, RPC, and database spans an agent stack also emits.
- OpenTelemetry GenAI Observability SIG - Charter and scope of the OTel project group driving GenAI telemetry; the place to track where the conventions are heading.
- OpenTelemetry Python GenAI instrumentation - Upstream instrumentation packages for OpenAI, Anthropic, Google GenAI, VertexAI, LangChain, and the Claude Agent SDK.
- OpenLLMetry - OpenTelemetry-based auto-instrumentation for Python LLM apps, exporting to any OTel backend.
- OpenLLMetry-JS - The same instrumentation for TypeScript and Node.js applications.
- OpenInference - OTel-compatible instrumentation spec and libraries from the Phoenix team, spanning Python, JS, and Java.
- OpenLIT - OTel-native auto-instrumentation SDK for LLM, vector-database, and GPU calls, with a self-hosted UI to read the resulting traces.
Tracing and observability platforms
- Langfuse - Self-hostable tracing, prompt management, datasets, and evals; ingests OTel as well as its own SDKs.
- Arize Phoenix - Self-hostable trace viewer and eval workbench that runs locally in a notebook or as a server.
- LangSmith - Hosted tracing and eval product from LangChain; the client SDKs are open source, the backend is not.
- Helicone - Proxy-based observability: point your base URL at it and get request logs, cost, and caching without code changes.
- Opik - Comet's open-source tracing and evaluation platform for LLM and agent workflows.
- LangWatch - OTel-based platform pairing production monitoring with pre-deploy agent simulation.
- Laminar - Rust-backed open-source tracing and eval platform aimed specifically at agent runs.
- Langtrace - OpenTelemetry-based, self-hostable tracing and evaluation for LLM calls, vector-DB queries, and framework usage, with Python and TypeScript SDKs.
- Agenta - Self-hostable workspace for authoring and versioning prompts and agents, with evaluation runs and tracing of each model and tool call.
- Pydantic Logfire - OTel-based observability with first-class Python, Pydantic, and agent instrumentation.
- W&B Weave - Weights & Biases toolkit for logging, comparing, and evaluating LLM app versions; docs.
- MLflow Tracing - GenAI tracing built into MLflow, so agent traces sit next to model runs and registries.
- Braintrust - Hosted eval and tracing platform; the scoring library autoevals is open source.
- AgentOps - Python SDK for session replay, cost tracking, and benchmarking across CrewAI, OpenAI Agents, LangChain, and AG2.
- AgentsView - Local-first session search, analytics, and token-use statistics across Claude Code, Codex, and other coding-agent session archives kept on your machine.
Trace backends
General-purpose OpenTelemetry backends that GenAI spans can be sent to when you do not want an LLM-specific product.
- OpenTelemetry Collector - Vendor-neutral pipeline to receive, process, filter, and fan out traces before they reach a backend.
- SigNoz - OTel-native, self-hostable traces, metrics, and logs in one UI.
- Jaeger - CNCF distributed tracing backend that stores spans and serves trace search and a waterfall UI over them.
- Grafana Tempo - High-volume trace store backed by object storage, queried from Grafana.
Evaluation and testing
- Ragas - Metrics and test-set generation for RAG and agent pipelines (repo moved from
explodinggradients/ragas). - DeepEval - Pytest-style eval framework with G-Eval, hallucination, and relevancy metrics.
- Promptfoo - Declarative CLI for prompt/model comparison plus LLM red-teaming and vulnerability scanning in CI.
- TruLens - Instrumentation plus feedback functions that score app internals, not just final output.
- Evidently - Python framework for evals, drift detection, and monitoring across tabular, text, and GenAI systems.
- Inspect - UK AI Security Institute's eval framework, built for agentic tasks, tool use, and sandboxed execution.
- OpenAI Evals - Eval framework and benchmark registry from OpenAI; broadly used, but the repo moves slowly.
- lm-evaluation-harness - EleutherAI's standard harness for few-shot academic benchmarks across model backends.
- Lighteval - Hugging Face evaluation runner supporting transformers, vLLM, and API backends.
- Giskard - Scans LLM agents for hallucination, prompt injection, and bias, and turns findings into test suites.
- Scenario - Simulates multi-turn users against an agent so conversations, not single calls, can be asserted on.
- Promptflow - Microsoft's flow authoring, batch evaluation, and tracing toolkit for LLM apps.
- Autoevals - Standalone library of model-graded and heuristic scorers usable outside Braintrust.
- ClawBench - Live-web benchmark for evaluating browser and computer-use agents on 283 everyday tasks across 144 websites, with request interception and five execution-evidence layers.
- Agent QA - Runs natural-language web/mobile application flows and records structured step evidence plus pass/fail results; it does not trace or score model or agent calls.
- HermesGate - Receipt-bound completion rail that runs a repository's own checks, records which content and tool versions were checked, and reuses a PASS only while those bytes still match, so an agent's own PR gets a repeatable, deterministic pre-review result.
- YYLO - Command-line orchestrator for coding agents that retains declared receipt hashes, session IDs, and terminal manifests per workflow run, keeps content-addressed task evidence tied to exact inputs, and treats missing or malformed terminal evidence as not success.
Guardrails and security
- Guardrails AI - Runs input/output validators around a model call and enforces structured output.
- NeMo Guardrails - NVIDIA toolkit for programmable dialogue rails defined in Colang.
- garak - LLM vulnerability scanner probing for jailbreaks, prompt injection, and data leakage.
- DeepTeam - Red-teaming framework that simulates jailbreaks, prompt injection, and related attacks against LLM and agent apps, built on DeepEval.
- Presidio - PII detection, redaction, and anonymisation for text and images; useful for scrubbing traces before export.
- Invariant Guardrails - Rule-based guardrail layer deployed between an app and its MCP servers or LLM provider; rules are Python-like matchers over tool-call sequences, so it can block a call pattern (e.g. "read one tool's output, then send it to an untrusted address") rather than just a single message.
Gateways and proxies
- LiteLLM - SDK and proxy exposing 100+ providers behind the OpenAI API, with cost tracking, logging, and callbacks to most platforms above.
- Portkey AI Gateway - Routing gateway with retries, fallbacks, caching, and inline guardrails.
- Bifrost - Go gateway that fronts multiple LLM providers behind one OpenAI-compatible API, with key load balancing, fallbacks, and plugin hooks for guardrails and telemetry.
- Agentgateway - Open-source proxy for agent-to-LLM, agent-to-MCP, and agent-to-agent (A2A) traffic with routing, guardrails, RBAC, and OpenTelemetry.
- Traceloop Hub - Small Rust LLM gateway from the OpenLLMetry authors, with OTel emission built in.
- Kong - API gateway whose AI plugins add LLM routing, token metrics, and request logging to an existing gateway deployment.
- Tuskira - Open-source gateway that secures, governs, and observes AI agents' MCP tool calls and LLM traffic.
Model Context Protocol
- Model Context Protocol - The specification itself; the source of truth for what a compliant client, server, and transport must do.
- MCP Inspector - Official UI for calling an MCP server's tools by hand and reading the raw JSON-RPC traffic.
- MCP Registry - Official community registry service for publishing and discovering MCP servers, run by the same working group as the spec.
- mcp-trace - CLI and Docker proxy for MCP servers speaking stdio, Streamable HTTP, or HTTP+SSE; emits an OpenTelemetry span per tool call with duration, status, and propagated trace context. v2.0.3
- ToolHive - Runs MCP servers in containers with permission policies, secrets handling, and audit logging.
- MCPJungle - Self-hosted registry and single proxy endpoint for the MCP servers an organisation runs.
- SandBase Harness - Self-hosted MCP agent runtime that keeps sessions, approvals, tool execution, and audit/replay records for governed runs.
- Docker MCP Gateway - Docker CLI plugin that fronts multiple MCP servers behind one gateway with container isolation.
- Snyk agent-scan - Security scanner for MCP servers, agents, and skills; formerly Invariant Labs'
mcp-scan.
Agent frameworks with built-in tracing
- OpenAI Agents SDK - Ships a built-in tracing layer with exporters to third-party platforms.
- Google ADK - Agent toolkit with built-in evaluation and OpenTelemetry-based tracing.
- Pydantic AI - Agent framework instrumented with OpenTelemetry out of the box, viewable in Logfire or any OTel backend.
Contributing
See CONTRIBUTING.md. Suggestions and corrections welcome, especially for anything on this list that has gone stale.
To the extent possible under law, Angel Hermon has waived all copyright and related or neighboring rights to this work. CC0 1.0, 2026.
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found