llm-latency-tracker

mcp
Security Audit
Warn
Health Warn
  • License — License: NOASSERTION
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 6 GitHub stars
Code Pass
  • Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Regional AI API latency and uptime measurements. Open datasets, JSON API and MCP access; edge TTFB and inference TTFT reported separately.

README.md

📊 LLM Latency Tracker

Measure AI API latency and uptime across regions, and publish the results as open data.

🌐 Live rankings · JSON API · MCP endpoint

⚡ What it does

  • 🌍 Regional measurements — compare provider response times and availability from distributed probe nodes.
  • ⏱️ Two probe types — network time to first byte (TTFB) and optional inference time to first token (TTFT).
  • 📈 Open dataset — daily rankings, historical snapshots and regional p50 / p95 statistics.
  • 🤖 Agent access — JSON API, MCP, Markdown pages and an llms.txt index.
  • 🗓️ Model lifecycle — a deprecation calendar with links to provider announcements.

🔧 Built for measurement

  • Python standard-library core for edge probes; provider keys are needed for inference probes.
  • SQLite storage and regional aggregation, with remote nodes shipping measurements to a central node.
  • Static publishing for the website and JSON datasets.
  • CI checks for tests and linting; published snapshots include their measurement date.

Reading the results: edge TTFB measures network/API responsiveness, not model generation speed. Compare it separately from inference TTFT; results depend on region, probe type and time window.

🚀 Quick start

Use the public data without installing anything:

curl https://llmlatency.dev/api/rankings.json

Or run edge probes locally with Python 3.12+:

git clone https://github.com/mazamaka/llm-latency-tracker.git
cd llm-latency-tracker
REGION=local python3 run.py
python3 aggregate.py --region local

For a local stdio MCP server over the published API:

python3 mcp_server.py

Inference probes, API access & site generation → · Deployment →

🧪 Contributing

Add a provider, bring a new probe region, or improve the measurement code. See CONTRIBUTING.md for setup and checks.


Python · SQLite · JSON API · MCP · Cloudflare Pages

Code: MIT · Data: CC-BY-4.0, with attribution to llmlatency.dev.

📈 Latest daily snapshot

Daily snapshot — 2026-10-05

Measured latency across 46 AI inference providers in 4 regions. Method: distributed edge (DNS→TCP→TLS→TTFB) + inference (TTFT) probes, last 24h. License: CC-BY-4.0.

Region Fastest provider (p50) p50 p95 Uptime
Asia (Tokyo) fireworks 18 ms 64 ms 100%
Europe (Germany) fireworks 98 ms 201 ms 100%
South America (São Paulo) openrouter 58 ms 91 ms 100%
US (Central) fireworks 28 ms 77 ms 100%

Snapshot generated 2026-10-05T07:21:23Z — this table is regenerated daily.

Reviews (0)

No results found