aesop
Health Warn
- License — License: NOASSERTION
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 5 GitHub stars
Code Pass
- Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
PolyForm Noncommercial-licensed (≤0.7.0 remains MIT) multi-agent orchestration harness for autonomous software development. Swap orchestrator + worker models from one config. Event-sourced state via SQLite; web dashboard; built-in secret-scan gates.

Crash-only multi-agent orchestration — restart IS recovery. Stateless workers, durable filesystem state, no central server to lose.
What It Is
Aesop is a multi-agent orchestration harness for autonomous software development. It runs fleets of LLM coding agents across ranked backlog items, verifies their output locally, and ships merge-ready code to CI. Crash-only by design: workers are stateless, state lives in git + SQLite + durable files, and restart is the only recovery path.
Credibility: The Trap Tests
Aesop does not trust itself. Every agent writes code claiming to be correct, and the trap tests deliberately reproduce patterns of agent deception:
- Fake-green trap: Tests that skip real validation (caught incident #464: playwright browser-proofs reported green without actually running). Prevention: counts are computed live from git ls-files (no stored artifact to drift) and
python tools/verify_test_suite_count.py --checkverifies the CI shard matrix actually covers every tracked test file. - Gate-activation trap: Forbidden flags (--admin, --no-verify, --auto) in dispatch templates (caught 7+ incidents). Pre-push secret scan blocks leaks with fail-closed exit on read errors.
- Doc-invented trap: Documentation claims not backed by facts (caught hallucinated 0.3.0 CHANGELOG entries). Prevention: statistics gate verifies README matches git.
- Test-pollution trap: Test state leaking across isolation boundaries (caught 6 incidents including sys.modules mock pollution). Prevention: all tests use isolated temp directories and subprocess cwd= isolation.
- CI-drift trap: Workflow state out of sync (caught #450: pytest missing from main-full.yml). Prevention: YAML validation and required-tool assertion.
See docs/INCIDENTS.md for the 41-incident log backing each trap class.
Credibility: The Cancelled Architecture
A measured design decision, not a feature.
In wave-11, a hierarchical dispatch architecture was built: orchestrator → Sonnet specialists → Haiku fleets. A/B testing measured it as 4.3x weighted cost (docs/ab-cost-dataset.md) producing identical quality (100% test pass , zero repair rounds) on the same fixture. The architecture was cancelled on 2026-07-14. docs/archive/spikes/tiered-cognition/ keeps it for reference; nothing is wired into live settings.
The lesson: flat Haiku-first dispatch is not a limitation—it is the measured optimum for this problem. Architectural changes require A/B proof or they stay off.
Quick Try (2 min, no API keys)
Just Node.js >= 18 and git. No API keys needed.
git clone https://github.com/matt82198/aesop.git && cd aesop
npm install && npm run test:all # Run full test suite
npx . --help # See CLI
Explore the CLI, verify the test infrastructure, and study examples/first-wave-baseline — a replay kit of a complete autonomous wave cycle.
Full Setup (LLM orchestration)
To run the actual multi-agent orchestration loop you need:
- Claude Code CLI (or another supported backend) with a valid API key
- Python 3.10+ for guardrails, secret scanning, and the dashboard
- Bash 4+ (or Git Bash on Windows) for daemon scripts
# Scaffolds the fleet AND installs /power, /buildsystem, /fleet, /dashboard,
# /healthcheck into ~/.claude/skills/ — the only path Claude Code scans.
# Add --install-deps to also run npm install and pip install -r requirements.txt.
npx @matt82198/aesop my-fleet --name "api" --repos "/path/to/repo" --install-deps
cd my-fleet
# Restart Claude Code — skills are enumerated at startup.
npx . doctor # Confirms skills, hooks, config, and interpreters
Existing skills of yours are never clobbered: an installed skill that differs
from the shipped one is preserved and reported, and --force overwrites it.
Use --no-skills to scaffold without touching ~/.claude/.
See the Learn More section for setup and architecture guides.
How It Works
Three layers: An orchestrator (Opus/Fable, main thread) dispatches parallel Haiku workers in isolated worktrees—1/5 the per-token cost of Opus. Durable state lives in SQLite WAL + git-committed files (STATE.md, BUILDLOG.md). On crash, the orchestrator re-reads from disk. No special recovery.
Verification gates: Fail-closed secret scan (pre-push), re-run of the exact CI gate (no mock), adversarial review (default-on). These are not optional best-practices—they are the default architecture.
Observable machinery: Heartbeats detect stalls; watchdog auto-restarts; every decision is readable in logs and git history. See docs/ARCHITECTURE.md for the full picture.
Known Limitations
- Benchmark is curated, not sampled. The 39-task judgment set (Haiku vs Sonnet vs Opus) scored 39/39 on all tiers, hitting the pre-declared ceiling rule (discrimination failure). It measures a sufficiency floor for seam-level tasks, not tier equivalence. See bench/METHODOLOGY.md.
- Seam-level only. This repo's agents operate within code review, test triage, severity assessment, and orchestration seams. Frontier reasoning (architecture redesign, novel algorithms) is out of scope.
- MCP server is read-only by design. State-store projections reflect the last successful run state. Real-time multi-agent coordination is not implemented.
Learn More
- docs/INSTALL.md — Setup and first wave
- docs/FIRST-WAVE.md — Walkthrough of a complete autonomous wave cycle
- docs/MICROKERNEL.md — Two-seat architecture (worker + orchestrator), model swapping
- docs/HOW-THE-LOOP-WORKS.md — Concrete walkthrough of one wave cycle
- THE-AESOP-HYPOTHESIS.md — Design philosophy, trade-offs, and why crash-only
- DISPATCH-MODEL.md — Cost analysis, Haiku sufficiency, scaling properties
- docs/how-i-built-aesop.md — Technical narrative of the orchestration design
- CARDINAL-RULES.md — 10 foundational principles for safe autonomous code
- autonomous-swe.md — What "autonomous" means and doesn't. Honest limits.
- docs/INCIDENTS.md — All 41 incidents: fake-green, gate-activation, test-pollution, stalls, conflicts, flakes, CI drift, hallucinations.
- RELEASE-NOTES.md — Version 0.7.2: bug fixes & compliance, encoding hardening, fake-green gate fixes.
Contributing
Aesop is source-available under the PolyForm Noncommercial License 1.0.0. Patches and contributions are welcome.
- Issues and bug reports — tell us what's broken or confusing.
- Discussion and ideas — feature requests, design critiques, use-case questions.
- Code contributions — fork, commit, and open a PR; we'll review and merge.
See CONTRIBUTING.md for details. The repo develops itself via its own /buildsystem loop.
License
Aesop is licensed under the PolyForm Noncommercial License 1.0.0 from this version forward: free for personal, research, and noncommercial use; commercial use is prohibited without a separate agreement.
Versions 0.7.0 and earlier were released under the MIT License and remain available under those terms — the relicense is not retroactive.
For commercial licensing, contact the author (Matt Culliton).
Copyright 2026 Matt Culliton.
Citing
If you use Aesop, please cite it using the metadata in CITATION.cff (GitHub
renders a "Cite this repository" button from this file). It also lists the preferred research
citation, The Receipts Loop (manuscript draft, arXiv submission pending).
References
Aesop: Autonomous developer for any repository, built by Aesop itself. May your orchestrator be wise and your subagents swift.
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found