Prometheus

agent
Security Audit
Warn
Health Warn
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Pass
  • Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

OAra Prometheus — an open-source AI agent daemon that runs on hardware you own.

README.md

Prometheus

GitHub stars

A fire you own doesn't go out.
Not a local model. A local agent.

Pre-1.0 · known limits: oara.ai/docs/limits/

Prometheus is an AI agent that runs as a daemon on hardware you own, and keeps working after you close the laptop. It keeps its own schedule, runs background work and can message you on Telegram when a task finishes, remembers you in files you can read, and checks the file edits it claims against the disk. It talks to the inference servers you already run (llama.cpp, Ollama, LM Studio, vLLM), or to a cloud model on your own key.

If Prometheus is useful to you, or any of this code helps your own project, a ⭐ on the repo helps other people find it.

Beacon's Mission home — the Armilla telemetry sphere beside Mission Control: agent state, scheduled jobs, local backends, and tool-call telemetry

The counters in this screenshot are one rig's, not a benchmark. SENTINEL shows as active because that rig turned it on; it ships off.

Prometheus is two pieces that pair with a 6-digit code:

  • The daemon — an always-on Python agent runtime that you can run as a systemd user service on Linux: cron and background tasks stored on disk, chat gateways (Telegram, plus Slack and Discord gateways on the same command layer, not yet run in production; all off until a token is added), lossless memory, a registry of every local inference box you own, opt-in coding runs in a sandboxed clone, a security gate, a Model Adapter Layer that repairs open models' tool calls, and a bearer-token REST + WebSocket control plane — plus an OpenAI-compatible /v1 surface so anything that already speaks OpenAI can talk to it.
  • Beacon — free to use, closed source, in beta. Its native desktop cockpit (macOS / Linux): chat with live tool timelines and @-references, Mission Control, a Loop Manager for coding runs, per-turn file checkpoints you can restore, a documents editor with AI redlines, Kanban, telemetry feeds, and per-provider key management. Beacon for iOS rides the same control plane (per-device tokens, pagination); it is a private beta — request access at [email protected] with the subject "Beacon iOS".
pip install 'oara-prometheus[full]'   # Python 3.11+
# or, on Apple Silicon Macs only:
brew install oaralabs/tap/oara

oara setup          # auto-detects llama.cpp / Ollama / LM Studio / vLLM
oara                # chat in the CLI
oara daemon         # always-on: web API + gateways + cron + background layer

Use the full names. pip install oara installs a 0.0.0 placeholder that only reserves the name, and a bare brew install oara fails — the tap prefix is required.

To read, change, or test the code, install from a checkout instead:

git clone https://github.com/OAraLabs/Prometheus.git && cd Prometheus
pip install -e '.[full]'

(prometheus … still works as an alias for now; only the command name changed — the package and its paths did not.)

Running the tests from a uv-managed venv? Sync it with the same extras CI uses first:

uv sync --extra web --extra anthropic --extra mcp --group dev

That is the environment the CI test job runs in, so a local run exercises the same tests (pip install -e '.[full]' above already covers everything).

oara setup probes for a running local inference server, generates your agent's identity, writes a working config with the web API enabled, and smoke-tests the loop. In a hurry? oara setup --fast (or --noninteractive) is the three-question version. On first daemon start a web API token is minted and printed once — oara token show re-prints it. If anything misbehaves: oara doctor.

oara-prometheus has been on PyPI since 0.9.0. Every way to install, uv tool / pipx from Git included, is under Install.

What it gives you:

  • Always on, on your box — run it as a systemd user service (oara install-service, Linux) and it keeps going after you walk away. Background tasks are stored on disk: a watch resumes after a daemon restart, and a run that a restart cut off is marked failed instead of left "running". With the Telegram gateway set up, it messages you when a task finishes or fails, and pings every 10 minutes while a long one is still running. Cron jobs take cron syntax or plain English ("every weekday at noon"), and pass the security gate when they're created and again when they run.
  • Always-on gateways — Telegram, plus Slack and Discord gateways on the same command layer (not yet run in production). All off until a token is added. Mid-turn /steer and /queue let you redirect or line up work while the agent is mid-task.
  • Tool-call repair for open models — a Model Adapter Layer validates every call and auto-repairs common errors (fuzzy names, JSON inside markdown fences, type coercion). On llama.cpp, when the server isn't parsing a model's tool calls natively, Prometheus supplies a GBNF grammar, so tool calls are constrained while the model is still generating. For models that call tools natively (per the model registry or the server's chat template), the daemon picks a lighter tier and mostly stays out of the way.
  • Every GPU box you own, one table — name your inference boxes under backends: and each becomes a slash command and a row in the model picker: /4090, /mini, /mini qwen2.5:7b-instruct. The daemon probes each box (served model, reported context window, detected vision, latency), refuses to switch a chat to a box that is down and says why, budgets the conversation at that box's window, and remembers the choice across restarts. /backends shows the table; /local brings a chat home.
  • Visible memory that rides every prompt — MEMORY.md and USER.md you can read, structured facts mined from conversations every 30 minutes, and passive recall that FTS-matches each message against the memory store and injects what's relevant.
  • Lossless context — older turns are summarized in the live window; the originals stay in SQLite, full-text searchable, and the stored summaries expand back to them, until you purge the session.
  • Coding Mode (opt-in) — point it at a repo and an acceptance command; it iterates to green in a sandboxed clone and hands you a reviewable branch. Never merges, never pushes. Off until you set coding.enabled: true.
  • A working directory per conversation, with undo — /workspace <path> points a chat at a repo; the security gate follows it, project instructions (PROMETHEUS.md, CLAUDE.md, AGENTS.md, GEMINI.md, .cursorrules…) load from it, and every turn takes a file checkpoint first so a bad edit is one REST call (or one Beacon click) from undone.
  • Sees what you show it — images reach a vision-capable model as images, not paraphrases; stored by reference so the bytes never bloat the transcript. @file, @diff and @url in the composer resolve on the daemon, scoped to the conversation's workspace.
  • Telemetry that stays home — every tool call, repair, and token count logged to SQLite on your own disk and sent nowhere. It's the raw material for tuning the adapter, and the capture/export half of a fine-tuning loop (the training half is not built): the data big labs keep for themselves, kept by you instead.
  • A desktop cockpit — Beacon (free to use, closed source, in beta) pairs to the daemon over your LAN or tailnet and gives every subsystem a native surface; the model picker groups your own boxes above the cloud presets with each box's health.

Status: Active development. Expect rough edges. Fixes land weekly. Feedback welcome.

Deeper docs — the guide pages under docs/guide/:
Install & first flight · Feature reference · Beacon desktop · Coding Mode & Loop Manager · Record a Skill · Memory & knowledge · Models & providers · HTTP & WebSocket API


What we built from scratch

The interesting work is original:

  • Model Adapter Layer — the gap between Claude-quality tool-calling and what open models actually produce. Validates, auto-repairs, enforces output schemas, retries with specific error context.
  • SENTINEL — a proactive layer that watches for idle time and acts, instead of only reacting to prompts. Nudges, dreams, synthesizes. Opt-in: it ships off.
  • Wiki Knowledge System — turns every conversation into a compounding knowledge base that cross-references itself over time.
  • The coding engine's iterate-to-green policy — "done" is a verdict, not a claim: sandboxed rounds until the acceptance command exits 0, with failure-fingerprint step-back and zero-progress aborts.
  • The fine-tuning gym — frozen task-sets, dual scoring (raw emission vs post-repair execution), and a refusal to declare winners it can't statistically back.
  • SYMBIOTE — code assimilation and self-modification. Scout researches GitHub, Harvest clones and vets (license gate, AST-level security scan), Graft adapts and integrates with provenance headers and a full test run, and MorphEngine hot-swaps the result via blue-green deploy with automatic rollback, backed by a backup vault. Experimental, off by default — powerful and unstable; treat it as a research feature.

The guards

A passing test proves the code runs. It does not prove anything calls it. Every guard below exists because a feature was built, tested, green — and never wired; each one is a test that fails the build rather than a convention someone has to remember.

The invariant Enforced by
No source file may resolve the wiki root independently — one resolver, or the build fails test_no_site_resolves_the_wiki_root_independently (#131)
Every key in the live config must exist in the shipped template — live ⊆ template, with no allowlist. (Compares against a real install, so it skips on a fresh clone where there is no live config to compare) test_every_live_config_key_exists_in_the_default_template (#133)
A register of config keys with no reader that can only shrink — an unregistered key with no reader fails, and a registered key that gains one fails as stale test_no_new_config_key_without_a_reader · test_known_unread_register_is_not_stale (#136)
Every tool's example_call must validate against its own input model — the adapter puts one in the error a model gets back after a failed call, so a wrong one teaches the model a parameter that does not exist test_every_example_call_validates_against_its_input_model · test_the_walk_reaches_every_class_that_declares_an_example (#670) · test_example_call_uses_real_param_names (#134)
Every media and rate check declares fail-closed or fail-open at construction — the registry is built at module level, so an undeclared guard is an import error, not a test failure test_every_guard_declares_its_enforcement · test_a_guard_without_enforcement_cannot_be_constructed (#140)
An acceptance test that terminates in a registered test double fails — registration is what makes a double detectable, wildcard exemptions are refused, and only individually-named doubles can be let through. An unregistered substitute is the known gap, and raises a loud warning rather than passing silently test_tripwire_end_to_end — spawns a real inner pytest against the actual enforcement hook (TRIPWIRE, #75)
Every allowlisted file type must be provably admitted, not merely "not refused" — breach tests prove the door closes; only admission tests prove it opens test_every_advertised_document_extension_is_admitted (#141)

The last one is there because its absence shipped: a control suite whose every case asked "does disabling this let something bad through?" and none asked "does this let the permitted things through?" went green while the document surface silently degraded to PDF-only — 19 of 20 advertised types refused, including two the allowlist explicitly permitted. Over-refusal looks exactly like the control working.

Why a Model Adapter Layer

Open models are getting good at conversation. They're still uneven at doing things. Ask one to call a tool and it may invent the tool name. Ask for JSON and it may wrap it in markdown. Ask it to chain three tool calls and it may drop a required parameter on the second one.

Prometheus handles this with a Model Adapter Layer that sits between your agent loop and whatever LLM you're running. Every tool call gets validated before execution, common errors get auto-repaired (fuzzy name matching, JSON extraction from markdown fences, type coercion), and when something still fails, the model gets specific error feedback with the actual schema — not a generic "try again." On llama.cpp, when the server isn't parsing a model's tool calls natively, Prometheus supplies a GBNF grammar, so tool calls are constrained while the model is still generating rather than validated after the fact.

The design philosophy behind it: The model is the agent. The harness is the vehicle.

The result: open models that reliably call tools, chain multi-step tasks, and run autonomously — without you babysitting every interaction.

A finished reply from a local model in Beacon — the session list, the answer, and the composer's Agent / Chat / Voice toggle and model picker

What Makes This Different

Prometheus isn't a wrapper around ollama.chat(). These are the systems built around the loop:

Lossless Context Management keeps the originals. Every message is persisted to SQLite. Older turns are summarized in the live window; the originals stay in SQLite, and the stored summaries (a DAG the agent can walk with lcm_expand) expand back to them on demand, until you purge the session. Full-text search covers your entire conversation history. And memory isn't just storage: extracted facts ride back into each turn via passive recall, matched against what you just said.

SENTINEL makes the agent proactive (opt-in; it ships off). Most agents sit idle until you talk to them. Prometheus has a background intelligence layer that watches tool performance patterns, consolidates memory, lints its own knowledge base, and discovers cross-entity insights — all while you're away. Three of four phases use zero LLM calls. The fourth is budget-capped at 2,000 tokens. It nudges you via Telegram when it finds something interesting but never acts without permission.

The Model Adapter Layer is what makes local models work in a tool loop. Four cascading extraction strategies handle whatever mess the model produces. A retry engine feeds specific schema errors back to the model. On llama.cpp, when the server isn't parsing a model's tool calls natively, Prometheus supplies a GBNF grammar, so tool calls are constrained while the model is still generating. Telemetry tracks success rates per model per tool so you know exactly where your model struggles. It is tier-selected: strictness: NONE switches it off for cloud APIs that already emit clean tool calls, so the repair path is exercised exactly where it is needed.

A compounding knowledge base inspired by Karpathy's LLM Wiki concept. Every 30 minutes, a Memory Extractor pulls structured facts from your conversations, and the wiki recompiles the affected entity pages automatically — extraction and compilation are one loop, not a manual step. The wiki then maintains itself: SENTINEL's zero-LLM linter patrols for orphans, broken links, and stale pages; dedup gates keep repeated insights from piling up; and passive recall feeds the accumulated facts back into every turn. Point Obsidian at the markdown files and the graph view lights up.

Infrastructure self-awareness via the AnatomyScanner. At startup, Prometheus scans your hardware (CPU, RAM, GPU VRAM), detects the loaded model and its quantization, maps your Tailscale network peers, checks disk usage, and generates ANATOMY.md with Mermaid architecture diagrams of your entire setup. The agent knows exactly what machine it's running on, what model is loaded, and what resources are available.

An evaluation framework with a local LLM judge. Most agent evals require API calls to GPT-4. Prometheus uses constrained-decoding on your local model to judge task completion, tool usage accuracy, and hallucination — zero API cost. Failure classification (model vs harness vs unclear) and trend tracking across models and runs. The shipped config points the judge at a separate endpoint and pins the model that grades (judge_base_url + evals.judge_model); blank judge_base_url and the judge falls back to the endpoint under test, which means the model grades itself. Every score records which judge produced it either way.

LSP integration for compiler-grade code intelligence. Instead of grepping for function names, the agent queries language servers for real symbol definitions, type errors, and references. After every file edit, a diagnostics hook automatically checks for type errors and feeds them back to the model in the same turn. Off by default — one config flag turns it on.

Open Models First, APIs Welcome

Prometheus is built for local inference. That's the whole point — sovereignty, privacy, no subscriptions. But it's not religious about it. If you want to use cloud models, the same daemon works with:

  • OpenAI (GPT-4o, o3-mini)
  • Anthropic (Claude)
  • Google Gemini (Flash, Pro)
  • xAI (Grok) — via API key or by signing in with a SuperGrok subscription (OAuth device flow; no key needed)
  • DeepSeek, Kimi (Moonshot), GLM (Z.ai), MiMo (Xiaomi) (provider classes ship; not yet exercised)
  • Any OpenAI-compatible endpoint (vLLM, LiteLLM, Together, etc.)

Switch any single chat with a slash command — /claude, /gpt, /gemini, /xai, /deepseek, /kimi, /glm, /mimo — and /local to come home. Keys are managed from Beacon's Models tab (paste once, live immediately, no restart) or the env file. The adapter layer adjusts its strictness automatically: full validation for open models, passthrough for APIs that already handle tool calling well.

Beacon's Models tab — per-provider keys, auth-mode badges, and SuperGrok subscription sign-in

The architecture doesn't care where the tokens come from. It cares that the tools get called correctly.

Features

The feature reference covers everything below in depth — including which subsystems are on by default and which are opt-in.

Three commitments run through everything below:

Refuses to lie to you. Untrusted input — cron payloads, task output, file contents — carries provenance and is fenced off as data, not instructions, before the model sees it. A file-mutation verifier stat-checks every claimed edit and tells the model "CLAIMED but NO CHANGE ON DISK" when nothing actually changed. Telemetry keeps honest denominators (a failing test run is not a failed tool call), the daemon self-reports when it's running stale code, every swallowed exception lands in a queryable silent-failure ledger, "I'll let you know when it's done" must be backed by a real registered task, and a coding run's "done" is a verdict — the acceptance command re-run — not a claim.

Survives you, restarts, and itself. Sessions, background tasks, and message ids are durable contracts that outlive daemon restarts. The memory store heals itself — search-index rebuilds, snapshot backups before every migration, writes that either commit or raise. Signals persist before they broadcast, so a crash can never have told you something the disk doesn't know.

Guards in code, not in policy. Keys are stripped from the agent's shell environment, and the key file is on a hard deny list for file tools. The audit log redacts secrets before writing, outbound fetches resolve-and-block private address space, cron passes the same security gate at creation and at execution and fails closed, self-modification is gated behind a dangerous-code scanner, and agent deliverables are served by content id rather than by path.

Model Independence

  • Runs any model llama.cpp or Ollama can serve — Qwen, Gemma, Llama, Mistral, Phi, DeepSeek, Command-R
  • Optimized formatters for Qwen and Gemma, default formatter works with everything else
  • Auto-detects whatever model is loaded — swap the GGUF, restart, done
  • 10+ providers: llama.cpp, Ollama, OpenAI-compatible (OpenAI/Gemini/xAI/DeepSeek/Kimi/GLM/MiMo), Anthropic
  • Configurable adapter strictness: STRICT (small models), MEDIUM (Qwen/Gemma), NONE (cloud APIs)
  • Per-session model override via slash command, REST, or Beacon's model switcher
  • Deferred tool loading (tri-state, default auto): cloud models get the full tool catalog; local models get a compact deferred catalog that hands back roughly 8K tokens of context on a 32K window
  • Cache-shaped context: the tool catalog is frozen at run start, any history rewrite is flagged, and mid-run compaction stays off on cloud providers — your prompt prefix is treated as a cache asset
  • Fallback chains that degrade loudly: when a provider fails terminally the turn moves to the configured fallback and says so — a provider_degraded frame on the wire, a line in history, and a named reason when a fallback declines (the context doesn't fit, the key is missing) instead of silence
  • The loop stops itself: 3 identical calls in a 10-call window that return nothing new, or 6 adjacent calls to the same read-only tool that return nothing new, halts the turn; a 500-call ceiling is the backstop (the same for local and cloud, the live value reported on /api/status) — so a flailing model ends the turn instead of burning the budget
  • Thinking blocks persist and round-trip across providers with one wire vocabulary, so a reasoning model's traces survive a restart and a client switch
  • A backend registry owns "which local boxes exist and what each is serving" — the model picker, the per-chat override, the fallback chain, Anatomy and /api/status all read it, so no two surfaces can disagree about a box. It reports; it never restarts a server (llama-server is one model per process, and swapping it is that box's business)

50+ Builtin Tools

bash, read_file, write_file, edit_file, grep, glob, web_search, web_fetch, youtube_transcript, download_file, browser (Playwright), image_generate, video_generate, tts, message, dashboard, notebook_edit, cron_create/delete/list, task_create/get/list/update/stop/output, todo_write, skill, agent (subagent spawning), ask_user, sessions_list/send/spawn, lcm_grep/expand/describe/expand_query, wiki_compile/query/lint, sentinel_status, audit_query, anatomy, lsp (7 actions; when enabled), plus dynamic MCP tools (mcp__{server}__{tool}).

Coding Mode — iterate to green

Opt-in. coding.enabled ships false: until you set it to true, oara code exits non-zero and POST /api/code returns 403. Coding runs execute model-authored commands against a clone of your repo, so you turn them on deliberately.

Point the agent at a repo, a task, and an acceptance command. It clones the repo into a sandbox (cwd jail, env-scrubbed so your provider keys never reach the subprocess), works in rounds until the acceptance command exits 0, and leaves a reviewable branch. "Done" is a verdict, not a claim — the session re-runs your acceptance command itself and rejects no-evidence turns. Mid-run supervision (pause / inject / resume) rides a control channel the run polls between episodes (experimental — not yet working reliably). Rounds stream live to Beacon.

  • Supervision is fail-safe by construction: a corrupt or missing control file reads as "not paused", and a run with no control channel is byte-identical to an unsupervised one
  • Repeated failures are caught by fingerprint: the failure output is normalized (timings, addresses) and hashed, so hitting the same wall twice triggers an explicit step-back — and zero-progress runs abort instead of burning rounds

A finished coding run in Beacon — Converted ✓, acceptance exit 0, and the reviewable diff

Beacon's Loop Manager turns this into a PM cockpit: register repos, keep a TASKS.md board, edit the LOOP.md run contract, and fire — Autonomous, Composed, or Supervised (the Beacon UI doesn't call pause/inject/resume yet). Kanban stories can be dispatched straight into coding runs. See the Coding Mode guide.

Skills

The agent writes skills for itself: the SkillCreator turns successful multi-step traces into markdown skill files, the SkillRefiner updates them when better executions come along, and a weekly Curator pass consolidates and prunes (pinned skills are protected; nothing is hard-deleted). Three core skills ship in the package (commit, debug, plan), and the repo carries a 103-file skill library in skills/ you can drop into ~/.prometheus/skills/ selectively — it's deliberately not auto-loaded, to keep prompts lean.

Record a Skill

Show the agent a workflow instead of describing it. Two capture paths, two trust levels:

  • Live DOM recording — record a browser workflow; a deterministic pipeline (no model calls) turns the event trace into a skill, runs it through a five-check quality gate, and auto-persists it to skills/auto/
  • Video / YouTube ingestion — screen recordings and videos are transcribed and vision-digested into skill drafts that never auto-persist: they wait in Beacon for human accept or reject

Ground-truth DOM traces earn autonomy; lossy vision output stays human-reviewed. What ships in this repo is the daemon side — the DOM-trace pipeline, its upload endpoint, and video ingestion (off by default). The browser recorder extension that captures a live DOM trace is not distributed yet. Full walkthrough: Record a Skill guide.

Why this only works here: a screen recording of you doing your job is among the most revealing data you own, and every frame of it stays on your disk. A hosted product can't offer the same feature, because shipping your screen to someone else's server is the feature — and the problem.

MCP Integration

  • Dynamic tool discovery from any MCP server, in the daemon and the CLI alike
  • Collision-free naming (mcp__{server}__{tool}), stdio transport today (HTTP/SSE planned), config fingerprinting
  • Scoped and gated: per-server allowed_tools allowlists (enforced at discovery and at call time); every MCP tool call requires confirmation before it runs — a server's readOnlyHint is recorded and shown, never trusted to skip the prompt
  • Managed over REST as well as config: /api/mcp/servers adds, edits and removes servers live (Beacon's Connectors tab is its client); a server added at runtime is advertised to the model and reaches the llama.cpp grammar on the next turn, or the add fails loudly
  • Context7 is a two-line config away for up-to-date library documentation
  • Packs — the extension contract at the boundary: a pack declares its tools, skills and panels; skills arrive quarantined, panels are discoverable, and nothing a pack ships is trusted before it is declared (FOUNDATION Part 2)

Identity System

  • SOUL.md — persistent identity loaded into every prompt. Survives /reset. Generated at setup — no hardcoded names.
  • AGENTS.md — agent registry with specializations for subagent spawning
  • ANATOMY.md — live infrastructure snapshot with Mermaid diagrams (hardware, VRAM, model + quant, Tailscale peers), queryable via the anatomy tool
  • MEMORY.md + USER.md — the agent learns who you are over time (bounded: 12K + 8K chars)
  • Agent Profiles — full, coder, research, assistant, minimal via /profile to trade tool breadth for context budget — and a profile can belong to one conversation rather than the whole daemon
  • Project instructions — walking up from the daemon's cwd, or from a conversation's own /workspace, the assembler stacks PROMETHEUS.md, HERMES.md, CLAUDE.md, AGENTS.md, GEMINI.md, .cursorrules, .windsurfrules, .github/copilot-instructions.md and .prometheus/rules/*.md, one per directory level, under a per-file and an aggregate cap that names what it had to omit
  • Node & instance identity — a local Ed25519 node keypair plus a vault-resident instance UUID (oara vault adopt / vault status); see Identity, and what does not phone home and docs/FOUNDATION.md

Security

  • 4-level trust model (BLOCKED → APPROVE → AUTO → AUTONOMOUS), origin-aware: background work (SENTINEL, cron, gym) faces stricter gates than what you ask for directly
  • A small hardcoded floor no mode can waive — a recursive rm aimed at /, ~, ~/.prometheus or a workspace root is refused, and file tools can never touch ~/.ssh, ~/.gnupg or a ~/.config/*/…env file — plus configurable deny lists (security.denied_commands, security.denied_paths), bash intent analysis, and a workspace boundary on write_file / edit_file — a speed bump, not confinement: bash is gated on the command string, not on the paths a command writes to, so a shell redirect goes anywhere. denied_paths is the hard stop, a turn-end check detects writes that landed outside the permitted area — after the fact — and for a conversation with a workspace every turn takes a file checkpoint first, so its writes can be undone over REST or from Beacon. With bubblewrap available, the bash write floor makes the boundary real in the kernel: the filesystem is mounted read-only except the workspace, and the gate follows the conversation's /workspace. What each one actually catches
  • Untrusted-input fencing: every message carries a provenance tag, and content from cron jobs, task output, and files is wrapped as data — not instructions — before it reaches the model
  • Secrets structurally absent: tool sandboxes strip key/token/secret variables from the environment (the agent can't env its own keys), key updates reject control characters, and key reads return booleans — never values
  • Secrets do not survive into logs or kept data: every log handler redacts token shapes (Telegram bot<token>, Bearer headers, ?token= query values, private-key blocks, and the GitHub/sk-/AWS/GitLab/Hugging Face/Stripe/JWT/Slack/Discord and other vendor key families) before a line is written — including the web server's own request log — and daemon.log/cli.log rotate. The same redactor runs before anything is kept (conversation history and its summaries, memory facts, telemetry, training pairs, trajectory exports) and on every prompt the learning loop sends to a model; the live conversation is left alone, so a token you hand the agent still works in that session. oara scrub cleans what was written before it existed (a dry run unless --apply). Not config-gated, on purpose
  • SSRF-hardened outbound: web_fetch and download_file resolve DNS and refuse private, loopback, and link-local address space before any request leaves the machine
  • Rate limiting on the public chat surface: a per-chat budget with a global ceiling above it, so one peer can't exhaust the daemon and the aggregate is capped even when every chat is individually under. Messages and media carry separate budgets, a refusal doesn't consume budget, and the sender is warned once per window rather than per message
  • Inbound media is checked cheapest-and-earliest-first: declared MIME before any transfer, then the size cap before download — then the download itself runs under a hard byte ceiling, because file_size is supplied by the peer and the pre-check believed it. Magic bytes are sniffed after, and a declared type that disagrees with the sniffed one is refused; that's the renamed-extension case
  • Media cache is LRU-bounded with a free-disk floor, and classified as a convenience: an unwritable or full cache declines to cache, never to serve
  • The honest limit: signature-less text formats have no magic bytes to verify, so they are admitted on their declared type — trusted, but bounded by the allowlist. That's strictly weaker than verifying bytes and strictly stronger than refusing the type outright, which is what it replaced (and which protected nothing while breaking everything)
  • Cron is not a bypass: jobs pass the security gate at creation and again at execution; blocked runs are recorded and reported — fail closed
  • Audit logging (SQLite + JSONL, queryable via /audit), exfiltration detection, prompt-injection defense
  • Approval queue — /approve, /deny, /pending via Telegram, or one-click Approve/Deny cards in Beacon
  • Authenticated control plane — bearer-token REST plus first-frame token auth on the WebSocket bridge
  • Secrets live in ~/.config/prometheus/env, never in the yaml; a pre-commit hook scans staged blobs for provider keys and this project's own opaque tokens before anything lands in the repo

Always-On

  • Telegram gateway with photo (real vision when the served model can see; captioning when it can't), voice (Whisper STT), document (20+ formats), and sticker handling — and control-plane commands (/steer, /approve, /status) answer mid-turn, because the data plane no longer holds the fetcher
  • Slack and Discord gateways on the same command layer as Telegram (not yet run in production): Slack over Socket Mode with thread-based long replies and channel whitelists; Discord with /prometheus app commands and DM + guild/channel whitelists. All three gateways are off until a token is added
  • Cron scheduler (natural-language scheduling supported), heartbeat monitoring, systemd service
  • Durable background tasks (tasks.db survives restarts) with an honesty check: "I'll let you know when it's done" must be backed by a real registered task — and tasks orphaned by a restart are marked failed instead of pretending to still run
  • 40+ slash commands on Telegram — including mid-turn /steer, /queue, and per-chat provider overrides
  • Paperclip fleet gateway (experimental, off by default) — Prometheus as a hireable agent: a fleet manager wakes it over HTTP, it checks out an issue, works a turn, reports back, and bills from real token usage
  • A mobile-ready control plane — per-device tokens (mint, verify, revoke), APNs push support for the iOS client (not yet verified on a physical iPhone), message pagination (?limit= / ?before=), a WebSocket subscribe that is a real session filter, and auto-titled sessions; the iOS client is built on nothing the desktop doesn't also use

Sessions that behave like sessions

  • They survive restarts — the session list is durably indexed, so a daemon restart restores every conversation; "Forget" is a durable tombstone that new activity revives
  • You can stop a running turn — interrupt over WebSocket or REST; completed rounds persist and the partial reply is kept as a real assistant turn, not discarded
  • Liveness you can trust — a progress pulse (phase / tool / round / elapsed) every few seconds distinguishes a long turn from a dead daemon, and failures arrive classified — a billing error says it's a billing error and how to fix it, instead of a blank timeout
  • An outbox for deliverables — files the agent produces land in ~/.prometheus/files and are served by content id (renames don't break links, no path-traversal surface); Beacon shows them as download chips
  • A real sync contract — every message has a durable id that doubles as a cursor, so clients resync incrementally after any disconnect
  • Rehydrate, pin, fork, purge — a restarted daemon restores a session's working context, not just its list entry; pins are daemon-side so every client agrees; a conversation can be forked at any point (edit-and-branch without branching the model); purge deletes for real, with a retention window for the tombstones
  • Search across everything — /api/search spans messages and summaries, and a summary hit carries the durable row ids to jump to
  • Honest context accounting — /api/lcm reports against the real detected context window (one detection, one resolution, the same number /context shows), splits fresh from summarized tokens, and /api/status says how much of the store is hidden behind summaries

Documents & Board

  • Documents editor — a confined documents folder served over the API; Beacon gives it a calm writing surface with auto-save and Ask AI redlines: describe a change, get {find, replace, reason} edits as inline tracked-changes, accept or reject each one. Nothing touches disk until you accept.
  • Kanban board — projects and stories over REST, drag-and-drop in Beacon, and stories dispatchable into coding runs.

Image & Video Generation

  • image_generate: Pollinations (free, hosted), ComfyUI (free, local GPU), or WAN 2.5 via DashScope (paid). auto never selects the paid backend.
  • video_generate: Kling 3.0 text/image-to-video (paid, dormant until keyed).
  • Details in the providers guide.

Fine-Tuning Flywheel (in progress)

  • Successful tool-call traces and adapter repair-pairs are captured, stored, and mined into an exportable dataset (capture → store → miner → export); browse with /pairs
  • Using the big model is collecting the corpus: first-try cloud successes are flagged golden at write time and banked as training examples for the local model
  • A gym runs frozen task-sets against live models with deterministic dual scoring (raw emission vs post-repair execution) and refuses to declare winners below sample-size thresholds
  • The gym also refuses untrustable results by construction: it probes the backend before starting, rejects two-variable experiments outright, and pins task-set and manifest hashes into every run
  • Corpus harvesting is Goodhart-proofed: pairs are classified by transition type with per-type caps, after one harvest came back 97% a single pattern
  • This is the data-collection half of a LoRA loop for the local model; the training step itself is still on the roadmap

Observability

Telemetry that stays home. Local-first people rightly refuse telemetry — because it usually means someone else's server. Prometheus inverts it: every tool call, repair, token count, and failure is recorded to SQLite on your own disk and sent nowhere — the telemetry module contains no network code at all, and the optional tracing exporter is off by default and points at localhost. It exists so you hold the data the big labs keep for themselves: the per-model success rates that tune the adapter, the golden traces a fine-tuning loop would learn from (capture and export exist; training is not built), and the receipts for what your model actually did. And it's neither slow nor bloated — WAL-mode appends cost sub-milliseconds next to tool calls that take seconds, and months of heavy daily use produce a database around 11 MB. Don't want it anyway? infrastructure.telemetry_enabled: false turns it off (from v0.9.7), and prometheus --reset-telemetry wipes the slate whenever you like.

  • Tool-call telemetry (SQLite) — success rates per model per tool, surfaced in Beacon's Tool Feed and /health — with honest denominators: a correctly executed command whose task fails (pytest exit 1) is not counted against the model
  • Every claimed file mutation is verified against disk — created / modified / deleted / no change — and "CLAIMED but NO CHANGE ON DISK" goes back to the model on its next turn
  • Every model call wrapped in an LLMCallEnvelope — per-round token accounting, and every swallowed exception lands in a queryable silent-failure ledger before any failure policy runs
  • Per-round prompt-cache stats (cached vs cache-write tokens, hit ratio) across providers — a provider that doesn't report is recorded as unknown, never as a fake zero
  • The daemon knows when it's stale on both axes: /api/status reports the running SHA against the checkout's HEAD and the checkout against origin/main, so neither merged-but-not-restarted nor merged-but-not-pulled can masquerade as deployed. The first axis alone reported stale: false while two merged commits sat undeployed — both operands came from the same clone. When freshness cannot be determined the answer is unknown, never current
  • The ledger is readable over the API: GET /api/tools/recent (the per-call tool history, not the loop's echo), GET /api/usage (token cost that states what it does not know), GET /api/tasks (background work, all of it — not just since the last restart), the wiki's pages and search, and the approval lifecycle pushed over the WebSocket
  • Configuration honesty: a config read that substitutes a default has to say so, a missing key has one of three rulings (absent / null / present-but-empty), config pins correct the running config and leave the file alone, and /api/status reports which checkout booted
  • Phoenix/OpenTelemetry tracing — env-gated, zero-cost no-ops when off
  • Failure classification in evals (model vs harness vs unclear) with trend tracking — the judge runs on your local endpoint; judge_base_url chooses it and evals.judge_model pins which model grades
  • Every score records which judge produced it — base URL, model, and whether that model was pinned or auto-detected from whatever the endpoint had loaded. pinned is the field that matters and can't be inferred from the model name: an auto-detected judge that resolves to qwen2.5:7b-instruct records the same name as one pinned to it, but only the pinned run is reproducible. Result files written before this carry no judge key at all — that means unknown, permanently, and they must not be compared across paths. Backfilling them would manufacture exactly the false provenance the change exists to prevent

On by default vs opt-in

Chat, tools, adapter, memory + LCM + passive recall, security gate, telemetry, and the web API are on out of the box. Coding Mode and the chat gateways ship off (oara setup turns a gateway on when you add its token). The bigger autonomous subsystems — SENTINEL dreaming, the router's autonomous half (task classification and fallback chains — its per-chat /claude-style overrides ship on), LSP, GEPA, escalation-to-teacher, SYMBIOTE (GitHub research → license gate → AST scan → safe graft → blue-green hot swap with auto-rollback; experimental), the Paperclip gateway, and Record-a-Skill's video ingestion — ship off by default and are one config flag away when you want them. The feature reference marks every subsystem's default.

Quick Start

Prerequisites

  • Python 3.11+
  • llama.cpp or Ollama running with any model loaded (or a cloud API key)
  • A Telegram bot token (from @BotFather) — optional, CLI works without it
  • bubblewrap or Docker — optional, only for coding mode's stronger sandbox backends (coding.sandbox_type: bwrap | docker; the Docker backend is experimental, and no test runs a coding task on it). The default process backend needs neither; oara doctor reports which backend you have configured and whether it can actually start.

Install

Four ways to get it:

One command, from Git — an isolated install of the oara command straight from the main branch, with uv or pipx:

uv tool install 'oara-prometheus[full] @ git+https://github.com/OAraLabs/Prometheus'
# or: pipx install 'oara-prometheus[full] @ git+https://github.com/OAraLabs/Prometheus'

From source — clone and install editable. Best if you want to read or change the code, or run the test suite:

git clone https://github.com/OAraLabs/Prometheus.git && cd Prometheus
pip install -e '.[full]'

From PyPI — the packaged release, oara-prometheus:

pip install 'oara-prometheus[full]'

A plain install, with no extras, runs oara daemon and pairs with Beacon: the web API is part of the base package. [full] adds the Slack and Discord gateways, MCP, the browser tool, the Anthropic provider and voice output.

With Homebrew — Apple Silicon Macs (M1 and later) only; Intel Macs aren't verified yet:

brew install oaralabs/tap/oara

Use the full name: a bare brew install oara fails because the tap isn't trusted. The install takes a few minutes, since every Homebrew dependency comes prebuilt, and like a plain pip install it runs oara daemon and pairs with Beacon.

Then run the setup wizard:

oara setup

The setup wizard generates your personalized identity, detects your inference server (llama.cpp:8080, Ollama:11434, LM Studio:1234, vLLM:8000), writes the config with the web API enabled, and runs a smoke test. No server running? The wizard offers a remote URL, a cloud provider, or copy-paste install instructions — it never writes a config it knows is broken.

Variants:

oara setup --fast            # quick path: probe → yaml → env, 3 questions
oara setup --noninteractive  # zero questions (first detected server, CLI gateway)
oara setup --gateway-only    # add/change Telegram, Slack, or Discord later
oara setup --provider anthropic --api-key-env ANTHROPIC_API_KEY --model claude-sonnet-5   # no local GPU yet: start on a cloud model, switch later

Prefer doing setup from a couch? Skip oara setup, run oara daemon bare, and it boots in setup mode — a pairing-only API that prints a one-time 6-digit code. Beacon's wizard takes it from there (detects backends, names the agent, configures gateways) and the daemon wakes fully configured:

Setup-mode pairing banner

Run

prometheus                                        # interactive CLI
oara --once "List the Python files here"    # one-shot
oara daemon                                 # always-on

On the first daemon start, Prometheus mints a secure PROMETHEUS_API_TOKEN, saves it to ~/.config/prometheus/env, and prints it once:

oara token show     # re-print the token
oara token rotate   # invalidate + mint a new one
curl -H "Authorization: Bearer $(oara token show | head -1)" http://localhost:8005/api/status

Run as a service (Linux and macOS)

oara install-service          # Linux: writes ~/.config/systemd/user/prometheus.service
systemctl --user start prometheus
journalctl --user -u prometheus -f

On macOS the same command writes a LaunchAgent (~/Library/LaunchAgents/com.oaralabs.prometheus.daemon.plist) instead; add --now to load it immediately, otherwise launchd loads it at your next login. It exits non-zero when it could not install or enable the service (a differing file refused without --force, another supervisor already owning the job, or systemctl / launchctl missing). The install guide has the details.

When something is off

oara doctor

oara doctor output — every subsystem checked with a fix hint per failure

Exit code is nonzero when anything is broken, so it also works in scripts.

Get Beacon

Beacon — free to use, closed source, in beta. Builds — macOS dmg, Linux AppImage/deb — are published when Beacon leaves beta, and until then early users get draft builds (oara.ai/beacon). Beacon for iOS: private beta — request access at [email protected] with the subject "Beacon iOS" (oara.ai/beacon-ios). First launch walks you through pairing — the full flow with screenshots is in the install guide, and the app tour is in the Beacon guide.

Beacon's setup wizard pairing with a daemon

Where the config lives (search order)

  1. an explicit --config path
  2. config/prometheus.yaml — repo-local (checkout installs; gitignored)
  3. $PROMETHEUS_CONFIG_DIR/prometheus.yaml — default ~/.prometheus/prometheus.yaml

Secrets never go in the yaml — they live in the env file ~/.config/prometheus/env, which both oara daemon and the systemd unit load.

Multi-Machine Setup

Run the agent on one machine, point it at a GPU machine for inference:

model:
  provider: "llama_cpp"
  base_url: "http://gpu-machine:8080"
  fallback:
    - provider: "ollama"
      base_url: "http://gpu-machine:11434"
    - provider: "anthropic"
      api_key_env: "ANTHROPIC_API_KEY"
      model: "claude-haiku-4-5-20251001"

Connect via Tailscale, WireGuard, or any network — Beacon pairs over the same address. Prometheus talks HTTP; localhost or remote, it doesn't care.

What About Smaller GPUs?

16GB VRAM runs Gemma 2 9B or Qwen 2.5 14B (Q4 quantized). Set strictness: STRICT — the adapter compensates with more validation and retries. No GPU at all? Use a cloud provider and you still get the whole daemon: memory, wiki, SENTINEL, security, profiles, all of it.

Architecture

┌──────────────────────────────────────────────────────────┐
│                    INTERFACE LAYER                        │
│  Telegram │ Slack │ Discord │ CLI │ Beacon desktop │ iOS  │
│  (REST :8005 + WebSocket :8010, token-authed; /v1 OpenAI) │
└────────────────────────┬─────────────────────────────────┘
                         │
┌────────────────────────┴─────────────────────────────────┐
│                  ALWAYS-ON LAYER                          │
│  Heartbeat │ Cron │ SENTINEL │ Memory Extractor │ Tasks   │
└────────────────────────┬─────────────────────────────────┘
                         │
┌────────────────────────┴─────────────────────────────────┐
│                 ORCHESTRATION LAYER                       │
│  Agent Loop → Model Adapter → Tool Dispatch               │
│  ┌──────────────────────────────────────────────────┐     │
│  │  MODEL ADAPTER LAYER                             │     │
│  │  Validator │ Formatter │ Enforcer │ Retry │ Telem │     │
│  └──────────────────────────────────────────────────┘     │
│  Model Router │ Coding Mode │ LSP │ MCP │ Subagents       │
└────────────────────────┬─────────────────────────────────┘
                         │
┌────────────────────────┴─────────────────────────────────┐
│               IDENTITY & KNOWLEDGE LAYER                  │
│  SOUL.md │ AGENTS.md │ ANATOMY.md │ Profiles              │
│  LCM (DAG compression) │ Wiki │ MEMORY.md │ Passive recall│
└────────────────────────┬─────────────────────────────────┘
                         │
┌────────────────────────┴─────────────────────────────────┐
│                 MODEL PROVIDER LAYER                      │
│  llama.cpp │ Ollama │ OpenAI │ Anthropic │ Gemini │ xAI   │
│  DeepSeek │ Kimi (Moonshot) │ GLM (Z.ai) │ MiMo (Xiaomi)  │
│  (xAI: API key or SuperGrok subscription OAuth)           │
└──────────────────────────────────────────────────────────┘

Identity, and what does not phone home

At first run Prometheus generates a node keypair under ~/.prometheus/node/
(Ed25519; the private key is mode 0600 and never leaves the machine). This is
an identity, nothing more:

  • Nothing is transmitted anywhere. The key exists locally and is inert
    until you deliberately turn on something that uses it. There is no
    telemetry-home, no registration, no callback. The public half is shown on
    your own daemon's /api/status and that is all.
  • It never encrypts your data. Losing the key can never make local files
    unreadable.
  • Your vault carries a separate instance_id (a UUID in
    .prometheus-vault) that travels with the data when you move machines;
    the node key deliberately does not. Every file Prometheus writes stays
    plain markdown or plain SQLite — leaving takes your data with you.

Configuration

model:
  provider: "llama_cpp"              # or ollama, openai, anthropic, gemini, xai,
                                     #    deepseek, kimi, glm, mimo
  base_url: "http://localhost:8080"
  # model auto-detected from llama.cpp on startup

context:
  effective_limit: 24000
  compression_trigger: 0.75

security:
  permission_mode: "default"
  workspace_root: "~/.prometheus/workspace"   # or a list of roots
  # A path outside every root makes write_file/edit_file ask first; it does
  # not make the write impossible (see Security). A conversation can override
  # this with /workspace <path> — the gate, project instructions and per-turn
  # file checkpoints all follow the conversation.

gateway:
  telegram_enabled: true   # default: false — setup enables it when you add a token
  # token via env: PROMETHEUS_TELEGRAM_TOKEN

memory:
  recall:
    enabled: true      # passive recall — stored facts ride each turn

sentinel:
  enabled: false       # opt-in: idle-time dreaming, wiki lint, synthesis
  dream_budget_tokens: 2000

wiki:
  root: ~/.prometheus/wiki   # the wiki is relocatable — every consumer
                             # resolves through this one key

heartbeat:
  maintenance_db: ""   # path to a SQLite file with a maintenance(until_ts) row.
                       # While the window is open, the "merged-but-dark" drift
                       # nudge is suppressed — the merge-to-restart gap is
                       # exactly when drift is expected. Empty = off; fails OPEN.

profile:
  active: "full"       # full | coder | research | assistant | minimal

Several GPU boxes? Name them once and every surface gets them — slash commands, the model picker, the fallback chain, /api/status:

backends:                       # the primary in `model:` is registered automatically as `local`
  4090:
    provider: llama_cpp
    base_url: http://gpu-box:8080
  mini:
    provider: ollama
    base_url: http://localhost:11434
    model: qwen2.5:14b-instruct                          # what this box serves for you
    models: [qwen2.5:14b-instruct, qwen2.5:7b-instruct]  # vetted choices (ollama swaps per request)
backend_probe:
  ttl_s: 60                     # a read older than this re-probes; /backends refresh forces it
  timeout_s: 5.0                # per box — a dead box costs one timeout, never a hang

/mini points a chat at that box after a live probe (a down box is a named refusal, not a failed turn); the chat is budgeted at the window the box reported, the choice survives a restart, and /local brings it home. model.fallback.backend: mini makes the same entry the fallback target, so the switcher and the fallback chain cannot disagree about what mini is.

Gateways

Telegram, plus Slack and Discord gateways on the same command layer (not yet run in production). All off until a token is added: every onboarding surface (oara setup, the fast path, the remote setup API, and Beacon's wizard) can enable any subset, and oara doctor reports each one's state.

Gateway What you need Env vars Extra
Telegram A bot token from @BotFather (/newbot) PROMETHEUS_TELEGRAM_TOKEN built-in
Slack A Slack app (api.slack.com/apps) with Socket Mode — both tokens: bot (xoxb-…) + app-level (xapp-…) PROMETHEUS_SLACK_BOT_TOKEN, PROMETHEUS_SLACK_APP_TOKEN pip install 'oara-prometheus[slack]'
Discord A bot from the developer portal with Message Content Intent, invited with bot + applications.commands scopes PROMETHEUS_DISCORD_TOKEN pip install 'oara-prometheus[discord]'

Tokens live in the env file, never in the yaml. The easiest way to configure any of them is oara setup --gateway-only.

Commands

The full 40+ command surface is in the feature reference; the daily drivers:

Command Description
/status /health /context Model, uptime, subsystem health, token budget
/steer Inject a mid-turn course-correction while the agent is working
/queue /unqueue Line up follow-up messages while it's busy
/wiki /note /memory Knowledge base stats, quick capture, memory files
/skills /profile /anatomy Skills, agent profiles, infrastructure snapshot
/workspace <path> Point this conversation at a repo — gate, project files and checkpoints follow
/backends /<backend> What every local box is serving; /4090, /mini [model] point this chat at a box (probed first, refused if down)
/approve /deny /pending Human-in-the-loop approval queue
/claude /gpt /gemini /xai /deepseek /kimi /glm /mimo Per-chat cloud override
/local /route Back to the local model · show this chat's routing
/sentinel /gepa /curator /symbiote /audit The opt-in autonomous layers

Cloud slash-commands are configurable per command (provider, key env, model) in prometheus.yaml — see the providers guide.

Benchmarks

Measured results — per-model task success with 95% intervals, and what each adapter tier adds or costs on the same model — are in docs/MODEL-LADDER.md, with every rung pinned by revision and SHA-256 and the per-rung reports committed under gym/results/ladder/.

All evaluation runs locally — the LLM judge uses constrained decoding on your own hardware. The shipped config gives the judge its own endpoint and pins the grading model (judge_base_url + evals.judge_model); leave judge_base_url blank and the judge falls back to the endpoint under test, i.e. the model grades itself. Each score carries its judge's provenance, so two runs can be compared — or refused comparison.

Project Structure

prometheus/
├── src/prometheus/
│   ├── engine/          # Agent loop, sessions, streaming, honesty check
│   ├── adapter/         # Model Adapter Layer (validator, formatter, enforcer, retry)
│   ├── providers/       # llama_cpp, ollama, openai_compat, anthropic, xai_oauth, registry
│   ├── tools/builtin/   # 50+ builtin tools
│   ├── coding/          # Sandboxed iterate-to-green runs + supervision + livestream
│   ├── hooks/           # PreToolUse / PostToolUse + hot reload + LSP diagnostics
│   ├── permissions/     # Security gate + audit + exfiltration + approval queue
│   ├── security/        # Log + capture redaction, path guard, dangerous-code scanner
│   ├── checkpoints/     # Per-turn workspace snapshots + restore
│   ├── memory/          # LCM engine, wiki compiler, extractor, passive recall
│   ├── context/         # Token budget, compression, prompt assembly
│   ├── gateway/         # Telegram, Slack, Discord, cron, heartbeat, paperclip
│   ├── web/             # REST API, /v1 OpenAI-compatible surface, WebSocket bridge, setup-mode server, @-references
│   ├── documents/       # Confined documents service + AI redline suggestions
│   ├── kanban/          # Projects + stories store
│   ├── sentinel/        # Observer, AutoDream, wiki lint, consolidation, digest
│   ├── mcp/  lsp/       # MCP runtime · language-server client
│   ├── evals/  gym/     # Local-judge evals · fine-tuning gym (dual scoring)
│   ├── coordinator/     # Subagent spawning, divergence detection
│   ├── learning/        # Skill creator/refiner, curator, GEPA, pair capture,
│   │                    #   live recorder + video ingest (Record a Skill)
│   ├── symbiote/        # Code assimilation + self-modification (experimental, off by default)
│   ├── infra/           # AnatomyScanner, project configs
│   ├── telemetry/       # Tool-call tracking + cost
│   └── config/          # Settings, paths, env overrides, profiles
├── templates/           # Identity templates (no personal data)
├── skills/              # 103-file skill library (.md, opt-in)
├── tests/               # pytest suite (see Tests below)
├── docs/                # Guides, architecture, sprint reports
│   └── guide/           # Install · features · Beacon · coding · skills · memory · providers · API
├── gym/                 # Frozen task-sets, harvest corpus
├── packaging/           # systemd unit
└── PROMETHEUS.md        # Agent instructions (CLAUDE.md, AGENTS.md, GEMINI.md… also read)

Tests

9,786 tests passed in CI on the v0.9.5 release commit (run 36461450849, Linux, Python 3.11–3.13).

If you build something with Prometheus, we'd love to hear about it. And if it helped, a star goes a long way.

License

MIT — see LICENSE. Upstream copyright notices for adapted code are collected in NOTICE.

Credits

Built by OAra Labs, starting with early scaffolding from OpenHarness. Design informed by Andrej Karpathy's LLM Wiki concept, Lossless-Claw, and Sigrid Jin's analysis of Claude Code's agent-loop patterns. Full lineage and upstream licenses: NOTICE.

Reviews (0)

No results found