medula
Health Uyari
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 6 GitHub stars
Code Basarisiz
- rm -rf — Recursive force deletion command in bench/check_openrouter.sh
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
Several coding agents, one repo, at the same time: a reproducible lab to measure how they coordinate, and Médula, a kernel that decides who waits for whom.
Médula
Several coding agents, one repository, at the same time. Who should wait for whom, and when?
A small, fully reproducible lab for measuring how coding agents coordinate, and Médula, a kernel
that coordinates them: before every write it asks a fast decider whether the change collides with
what the other agents are doing.
Six tasks on a room-booking API, written to collide in known ways. Six coordination modes. The same
agents (Claude Code, anthropic/claude-sonnet-5) in every mode, and every run published raw: agent
sessions, diffs, evaluations and each decision the kernel took.
What the first matrix of runs says:
- Branches hide semantic conflicts. Git merges the six branches without a single conflict on the
code, and the result is broken:/exportstill callslogin(username, password)afterlogin()
started requiring a second factor. 6 red tests in every mode B run. - With Médula, both real conflicts are caught in every run with no unnecessary blocks (mode D),
at the same total cost as classic per-file locks, which block things that don't collide. - The fast decider is almost free but unsure on real states. Jev decides in about 0.3 s for
$0.00006, yet 61 % of its decisions on real write requests fall in the "not sure" band and go to
the slow path (Sonnet, then Opus). That is the biggest open problem, and where you can help.
Try it in two minutes (no API key)
Everything below runs offline and costs nothing: the tests mock the model provider.
git clone https://github.com/JoaquinRuiz/medula.git && cd medula
uv sync && (cd demo-app && uv sync) && (cd medula && uv sync)
(cd medula && uv run pytest) # the kernel: 69 tests, simulated provider, a real server and hook script
bench/self_check.sh # the collision design: every pair merges cleanly in git,
# yet merging all six breaks exactly T2, T4 and integration
Then look at a real decision: results/runs/d1/medula.db is a SQLite file with every question Médula
asked, the answer, the probability, the latency and the cost.
Help wanted: the open problems
These come straight out of the published runs. Each one is concrete, has data behind it and a
place in the code to start from. Comment on its issue to claim it or to discuss an approach
(all open issues, good first issues).
- Human labels for the calibration set (#1) · no code, highest value. The 100 calibration pairs
and their labels were written by a model of the same family as two of the deciders, which likely
flatters them. Label them yourself, blind, in your own file:uv run bench/e04_etiquetar.py --salida calibration/etiquetas_humanas/<your-github-user>.yaml
(about 20–30 minutes;ccollides,ndoesn't,dunsure).uv run bench/acuerdo_etiquetas.py
shows how you agree with the model and with other people. - Calibration pairs that look like real runs (#2) · no code or light code. In calibration, 13 % of
Jev's decisions were unsure; in real runs, 61 %. Real states have several agents and long
intentions; the calibration pairs have one of each. New pairs incalibration/candidatas.yaml,
or a script that extracts them from themedula.dbof past runs, would close that gap. - Signature changes on the fast path (#3) · code. Jev ranks the real T1–T2 conflict (
login()gaining a
required parameter) below a false one, and it's the slow path that catches it. Two ideas were measured and rejected
(bench/e04b_firma.py: a shared-symbol rule and a directional question); better ones are welcome.
Start atmedula/src/medula/intencion.py(how the intent is built) anddecisores/jev.py. - A new decider (#4) · code. Any model or rule that answers "does this collide?" with a
probability fits the same interface (medula/src/medula/decisores/): a local model, another
provider, a static analyser. Measure it on the calibration set and it can run the whole matrix. - A robust slow path (#5) · light code. In the published runs, Sonnet sometimes answered with JSON
the kernel couldn't parse, or hit the 30 s timeout, and the decision escalated for nothing. Seemedula/src/medula/camino_lento.py. - More scenarios (#6) · no code, mostly tests. The six tasks cover a changed signature and a
renamed field. A changed behaviour with the same signature, a schema migration, a dependency
bump: each one needs a task, acceptance tests and a line inground_truth.yaml. - More runs (#7) · costs money, no code. Modes E and F have few runs, and F's only run had agents
messaging each other (see the caveats). If you have credit,bench/run_matrix.shresumes the
matrix and records everything; send the evidence in a PR.
Good first contributions
| Difficulty | What | Where |
|---|---|---|
| 🟢 No code | Label the calibration pairs (blind, in your own file) | calibration/etiquetas_humanas/ |
| 🟢 No code | Argue a label you disagree with, in an issue | calibration/etiquetas.yaml, ground_truth.yaml |
| 🟢 No code | Propose new calibration pairs | calibration/candidatas.yaml |
| 🟡 No code | Add a scenario: task, acceptance tests, ground truth | tasks/, acceptance/, ground_truth.yaml |
| 🟡 Light | Make the slow path robust to unparseable answers and timeouts | medula/src/medula/camino_lento.py |
| 🔴 Code | Extract calibration pairs from real runs | results/runs/*/medula.db → calibration/ |
| 🔴 Code | A new decider, or a better intent for signature changes | medula/src/medula/decisores/, intencion.py |
Disagree with a label or with how a conflict is defined? That's a contribution, not a complaint:
the ground truth is opinion made inspectable. Open an issue with your reasoning.
Before you send a PR
(cd medula && uv run pytest) # kernel; the model provider is simulated
(cd demo-app && uv run pytest) # the demo app
bench/self_check.sh # the collision design still holds
uv run python bench/lib/check_ground_truth.py # the ground truth covers the 15 pairs
See CONTRIBUTING.md for the data formats, how to add a decider or a scenario,
and how to contribute runs.
What won't be merged
- Tests that need an API key or money. Anyone must be able to check a change for free.
- Hand edits to published results. A run's evidence is what the run produced. Fixing a test and
re-evaluating is fine (bench/reevaluar.shrecords why); editingsummary.csvor a log is not. - Hiding a caveat. If a run is flawed, it gets flagged, not quietly dropped.
How it works
The six tasks
| Task | What it does | Collides with | Kind |
|---|---|---|---|
| T1 | Adds a TOTP second factor; login() now requires otp |
T2 | semantic: different file, same contract |
| T2 | GET /export in CSV, authenticating via login(username, password) |
T1 | semantic |
| T3 | Renames the booking field fecha → inicio in model, DB and API |
T4 | semantic |
| T4 | GET /reservas?desde=YYYY-MM-DD, filtering on fecha in SQL |
T3 | semantic |
| T5 | Rewrites the error messages of validar_sala() |
— | false conflict: same file as T6 |
| T6 | Adds the 8:00–20:00 rule to validar_horario() |
— | false conflict |
Acceptance tests are always evaluated on the final state with all six tasks applied, so they target
the final contract: every authenticated call sends the OTP, and bookings use inicio.
The modes
| Mode | What it is |
|---|---|
| A | Sequential: one agent after another in the same directory |
| B | One branch per task, merge at the end; textual conflicts are resolved by an agent |
| C | Shared directory, classic per-file locks (Médula with the locks decider) |
| D | Shared directory, Médula with Jev as the fast decider |
| E | Shared directory, Médula with Haiku as the fast decider |
| F | Shared directory, Médula with Sonnet as the fast decider |
Same spec, same agents (Claude Code, anthropic/claude-sonnet-5, effort high) and same model in every
mode. Rounds: T1–T4 in parallel, then T5–T6.
What Médula does
agent wants to Edit / Write / Bash
│ PreToolUse hook
▼
Médula builds the intent: the agent's task + the file, diff or command
│ reads never wait
▼
fast decider: "does this collide with what each other agent is doing?" → probability p
│
├── p below the low threshold → go ahead
├── p above the high threshold → wait until the other agent finishes
└── in between → slow path: Sonnet proposes a way out, Opus if it can't
(never one that breaks a task's acceptance criteria)
after every write (PostToolUse): "does this change invalidate another agent's plan?" → notice in its mailbox
The kernel's design, decisions and trade-offs are in medula/SPEC.md andmedula/README.md.
Results
Per-run data is in results/summary.csv; everything behind it (agent sessions, final diff,
evaluation, Médula's SQLite with every decision) is in results/runs/<run-id>/. Charts inresults/graficos/. Averages over the valid runs of the matrix:
| Mode | Runs | Tests green / red | Real conflicts detected (of 2) | Unnecessary blocks | Total time | Agents cost | Decision cost | Decisions | Slow-path escalations |
|---|---|---|---|---|---|---|---|---|---|
| A | a1, a2 | 30 / 7 | — | — | 341 s | $0.86 | — | — | — |
| B | b1–b5 | 31 / 6 | — (git: 2 textual conflicts per run, resolved) | — | 329 s | $1.20 | $0.40 (conflict resolution) | — | — |
| C | c1–c3 | 37 / 0 | 1.7 | 1.7 | 452 s | $1.65 | $0 | 88 | 0 |
| D | d2–d4 | 37 / 0 | 2.0 | 0 | 415 s | $1.36 | $0.29 | 118 | 20 |
| E | e1–e3 | 37 / 0 | 2.0 | 1.0 | 824 s | $1.74 | $0.55 | 149 | 22 |
| F | f1 | 37 / 0 | 2.0 | 0 | 376 s | $1.40 | $0.59 | 114 | 9 |
Times use only runs that did not overlap with another run (tiempo_fiable in summary.csv).
- A shared directory changes the game. In modes C–F agents see each other's code and adapt, so
all of them end green. What Médula changes is how they get there: fewer unnecessary blocks than
locks, and conflicts caught before the code is written. - In sequential mode A, the T2 agent detects the conflict and refuses to implement a login
without the second factor, asking for a confirmation that never comes in non-interactive mode. - Most of mode D's decision cost is the slow path: about 20 escalations per run, one in six of
all decisions. Only write requests can escalate, and 61 % of Jev's decisions on write requests
fell in the uncertain band.
Calibration
The kernel asks each decider for a collision probability. On the 100 calibration pairs
(calibration/, 2 repetitions): Jev 92.5 %, Haiku 91.5 %, Sonnet 98.5 % accuracy at p > 0.5. Jev
ranks cases very well but its probabilities are compressed towards the middle (chart:results/e04b/calibracion/). Thresholds per decider (calibration/umbrales.yaml) were chosen with one
rule: maximum fast-path coverage with ≥ 95 % accuracy, no more missed conflicts than with 0.2/0.8,
and at most one extra false alarm.
Caveats
- Unreliable times: a1, b2, b3. These runs overlapped in time with another run, so they count for
tests and cost but not for time (tiempo_fiable = Falseinsummary.csv). Mode A's time comes
from a2; mode B's from b1, b4 and b5. - Agent-to-agent channels in e1 and f1. Some agents used Claude Code's
ListAgents/SendMessage
tools to reach each other outside Médula: in f1, five messages (T3 told T4 about thefecha→iniciorename); in e1, a single listing with no message. Both runs are kept and flagged
(canal_entre_agentesinsummary.csv). Those tools are now disabled for every agent. - Thresholds fitted on the same pairs they are measured on. There is no separate validation set.
In real runs, 61 % of Jev's decisions on write requests fell in the uncertain band and went to the
slow path, versus 13 % of the calibration pairs. - Origin of the calibration labels. The 100 pairs and their labels were written by Claude, a
model of the same family as Haiku and Sonnet, with no human review (calibration/README.md,calibration/etiquetas.yaml). This likely favours Haiku and Sonnet in the accuracy comparison
(Sonnet 98.5 %). Human labels are the first open problem above. - d1 and d2-sin-calibrar are mode D runs with the initial, uncalibrated thresholds; e08-demo is a demo
run with a short wait. None of them is in the averages above.results/runs/_invalidas/holds runs
lost to an exhausted API credit. - Agent cost is Claude Code's own estimate at the provider's list prices; decision cost is what
OpenRouter reports.
Reproduce
Python is managed with uv; there are three uv projects (repo root for the
bench tooling, demo-app/, medula/). Agents and deciders go through OpenRouter, so paid experiments
need an OpenRouter key in .env.
cp .env.example .env # add OPENROUTER_API_KEY; MEDULA_AGENT_MODEL and MEDULA_AGENT_EFFORT are preset
export MEDULA_RUNS_DIR=/tmp/runs # where workspaces go (outside the repo)
Costs are approximate, at September 2026 prices.
| Experiment | What it produces | Command | Cost |
|---|---|---|---|
| Collision design | untouched app red; each task alone green; the 15 pairs merge cleanly in git; merging all six leaves T2, T4 and integration red; the correct integration green | bench/self_check.sh |
free |
| Kernel | Médula's tests against a simulated provider, plus an end-to-end test with a real server and the hook script | (cd medula && uv run pytest) |
free |
| Effort check | confirms --effort high reaches the model through OpenRouter (required by the runners) |
bench/check_openrouter.sh |
cents |
| Decider micro-benchmark | latency and cost per decision for Jev, Haiku and Sonnet (results/e03/) |
uv run --env-file .env bench/e03_decisores.py |
~$1 |
| Calibration | accuracy of each decider on the labelled pairs, and the fast-path thresholds | uv run --project medula --env-file .env python bench/e04b_firma.py --variantes base,haiku,sonnet |
~$2.5 |
| Calibration chart | declared collision probability vs. observed frequency | uv run --project medula --with matplotlib python bench/grafico_calibracion.py |
free |
| Signature experiments | shared-symbol rule and directional question, each measured separately | uv run --project medula --env-file .env python bench/e04b_firma.py --variantes base,regla,direccional |
cents |
| One run of mode B | branches, merge, agent conflict resolution, evaluation | bench/run_mode_b.sh --run-id b1 (live panel: uv run python bench/modo_b_en_directo.py --run-id b1) |
~$1.6 |
| One run of modes A, C–F | shared directory; C–F coordinated by Médula | bench/run_mode.sh --modo D --run-id d1 --umbral-bajo 0.25 --umbral-alto 0.5 (with Médula's terminal UI: bench/en_directo.sh …) |
~$0.9–2.3 |
| The run matrix | A×1, B×3, C×3, D×3, E×3, F×1 into results/summary.csv, each mode with its decider's thresholds; resumable, stops on API errors |
bench/run_matrix.sh (estimate first: bench/run_matrix.sh --estimar) |
~$20, ~1 h |
| Wait demo | a run where an agent is blocked with a reason it can see, outside the statistics | bench/en_directo.sh --modo D --run-id demo --demo --espera-max 30 --umbral-bajo 0.25 --umbral-alto 0.5, and uv run python bench/ver_agente.py --run-id demo T2 |
~$1.5–3 |
| Replay a mode B merge | the merge and the acceptance tests of a finished run, without agents | bench/repetir_merge.sh b1 |
free |
| Re-evaluate a run | acceptance tests on a finished run's final state (e.g. after fixing a test) | bench/reevaluar.sh d1 "reason" |
free |
Agents are stochastic: expect the same pattern, not identical numbers.
Isolation of agent runs
- Workspaces live outside the repo, in
$MEDULA_RUNS_DIR(default$TMPDIR/medula-runs). The scripts
refuse a path inside the repo, or one with aCLAUDE.mdor.claude/in any parent directory. - Each run uses an empty
CLAUDE_CONFIG_DIR, so agents don't inherit the user's plugins, skills, MCP
servers, hooks or memory. - Agents cannot use tools that end the session waiting for a wake-up, talk to other sessions or leave
the working directory (bench/lib/agent.sh). - In mode B each task gets its own clone, not a worktree, so an agent can't see other tasks' branches
before the merge. - Paths in the published logs are anonymised (
<repo>,<runs>,<home>,<tmp>).
Layout
demo-app/ Room-booking API (FastAPI, SQLite, pytest) in its initial state; its own uv project
tasks/T1.md … T6.md The six tasks, written as an agent receives them
acceptance/ Acceptance tests per task + integration test (never copied into agent workspaces)
ground_truth.yaml The 15 task pairs, labelled conflict / no conflict
calibration/ Calibration pairs, their labels, human labels and the calibrated thresholds
reference/ Reference solutions: each task done in isolation and all six integrated correctly
medula/ The coordination kernel (its own uv project; see medula/README.md)
bench/ Runners, evaluation, calibration and analysis scripts
results/ Raw evidence of every run, summary.csv and charts
SPEC.md Spec of the lab (in Spanish); medula/SPEC.md is the kernel's spec
Code comments and specs are in Spanish; issues and PRs are welcome in English or Spanish.
About the author
Joaquín Ruiz — jokiruiz.com ·
youtube.com/@jokioki
📗 Del vibe coding al Spec-Driven Development
📙 El motor de la Inteligencia Artificial
📘 Programar con Inteligencia Artificial
📙 Explora la Inteligencia Artificial
License
MIT © Joaquín Ruiz
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi