tribunal
Health Uyari
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 5 GitHub stars
Code Gecti
- Code scan — Scanned 1 files during light audit, no dangerous patterns found
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
A multi-agent red-team tribunal for your hardest engineering decisions: AI CLIs from different vendors debate in structured rounds and return adjudicated verdicts that preserve dissent. A methodology, not a framework.
Tribunal
A multi-agent red-team tribunal for your hardest engineering decisions:
AI CLIs from different vendors take independent positions, cross-examine
what they actually dispute, and return an adjudicated verdict that preserves
dissent. A methodology, not a framework (and, by its own measurement, an
honest decision-support scaffold, not a proven accuracy upgrade).
Get it — docs and prompt templates, nothing to build:
git clone https://github.com/kdoubt/tribunal.git ~/tribunal
Then, from your project's root, hand the scout straight to your agent — it drafts your first panel:
# Pick ONE line below (each hands the scout to a different agent to draft your brief):
claude -p "$(~/tribunal/scout)" # Claude Code
codex exec "$(~/tribunal/scout)" # Codex CLI
grok -p "$(~/tribunal/scout)" # Grok CLI
~/tribunal/scout just prints the scouting prompt with your clone path filled in — it runs no panel and writes nothing to your project. (Its once-a-week self-update check may git fetch inside the clone, or touch a cache-dir timestamp for zip installs; silence with TRIBUNAL_NO_UPDATE_CHECK=1.) That's the whole install.
One prerequisite to run a panel: two already-authenticated seat CLIs from different vendors or model families (local or hosted). A single CLI is enough to scout, but a real panel needs two seats — that's the whole point. The orchestrator is a separate role: either you at a terminal (the shell adapter), or a driver tool like Claude Code — which is then a third binary, distinct from the two seats it relays between.
Two seats (A, B) is the floor; the dashed third seat is optional — add it only for the highest-stakes, most irreversible calls.
The loop, in one glance:
scout project ─▶ frozen brief ─▶ ROUND 0 seats answer in isolation (parallel, no peeking)
─▶ LEDGER claims + falsifiers; pre-exposure agreement = settled
─▶ ROUND 1 disputed claims only, relayed verbatim - attack/concede/revise
─▶ ROUND 2 only if a load-bearing claim flips - delta-only restate (rare)
─▶ ORACLES tests · compilers · primary docs settle the checkable
─▶ VERDICT 1 agreement · 2 resolved · 3 surviving dissent (kept, not averaged)
─▶ RETRO did it hold? whose dissent was right? → template delta → next brief
Built for developers who work with agentic CLIs and have no human review
board: the panel is your reviewers. One model reviewing its own plan converges
on its own blind spots; Tribunal forces two different vendors' models to take
independent positions and cross-examine each other claim by claim. What that is
designed to give you on a contested call is an adversarial counter-case, a
discriminating test to settle it, and a hedge, with real disagreement preserved
rather than averaged away. What the validation actually measured is
narrower: on decidable (oracle-scorable) calls the panel matched or beat each
fixed single vendor in those runs - the seats diverged once in 20 decisions,
and Round 1 resolved that split to the sealed-correct call (n=1, and a
scope-framing split; see the validation caveats) - but on ambiguous calls an
independent judge preferred the plain solo memo. So the decidable-call hedge
held in this evidence, resting on a single soft data point - a supported design
rationale, not an established property; on ambiguous ones the panel's extra
output was judged no better (sometimes worse).
What Tribunal is not: a way to out-decide a strong single model. Its own
validation - two pilots plus a pre-registered study labeled
"confirmatory" at n=20 with an independent judge (a fourth model, OpenAI-lineage;
executed with disclosed deviations) - observed no accuracy lift. The frontier
seats tested agreed on ~19 of 20 decisions - a rate for those two seats, not
frontier models in general - so the debate rarely even runs, and the panel
matches a strong single model rather than beating it. Treat Tribunal as an
honest, adversarial decision-support scaffold for irreversible calls, not an
accuracy upgrade.
You need two or more independently configured model/agent commands from
different vendors or model families (local or hosted; subscriptions only where
a provider requires one).
When to use it
High-stakes, ambiguous, or irreversible decisions: architecture calls,
security boundaries, migrations, "is this even the right design." A panel
costs several model invocations and minutes of wall-clock per round -
spend it only where the expected loss of deciding wrong justifies it (seecore/METHODOLOGY.md, "When to convene").
Skip it for anything a test, compiler, or grep can settle. The first rule
of the methodology: run the oracle before convening a debate.
How this differs from multi-agent debate / LLM-as-judge
Multi-agent debate and LLM-as-judge are well-studied, with mixed results -
mostly because agreement between models that have read each other is cheap.
Tribunal's mechanisms target exactly that:
- Only pre-exposure agreement counts as consensus. Round 0 is isolated
and parallel; anything models agree on after seeing each other is
never promoted to "the panel concludes". - Disputes route to oracles, not more debate. Checkable claims go to
tests, compilers, and primary documents; debate is reserved for what
can't be checked. - Dissent survives into the verdict. Unresolved disagreement is
reported with the cheapest discriminating test - never averaged away.
Getting started
Nothing to install - this is docs and prompt templates. Your first
complete panel takes about 20-40 minutes of panel wall-clock (scouting,
two rounds, adjudication) - authenticating your two CLIs is separate,
one-time vendor setup that is not counted in that estimate. Once you know
the loop it is mostly the models' wall-clock - a few minutes per round. The
one prerequisite is two seat CLIs from different vendors, already
authenticated (the orchestrator is a separate role - you, or a driver tool
like Claude Code, which is then a third binary).
If a run stalls, the usual causes are a missing or not-logged-in CLI
(check command -v and re-authenticate), macOS needing GNU timeout
(brew install coreutils), or a seat exiting silently on a permission
prompt. Each adapter's README documents these silent-seat-killers and a smoke
test that catches them before a real run.
Two seats is the floor, not the ceiling. The diagram shows three, but a
panel is any N ≥ 2 - two is the floor and the third seat is optional, added
for the highest-stakes, most irreversible calls. The intended gain is 1 → 2
(self-review to cross-vendor), with each seat past that adding diminishing
value at linear cost - but note the validation caveat: in the
confirmatory study the seats disagreed on only 1 of 20 decisions, so in practice
the 2nd seat mostly just confirms the 1st. Scale to the stakes, not the ritual. Giving seats distinct review
lenses is a separate, optional layer with its own rules (seecore/METHODOLOGY.md, "Assigning lenses") - by default you name the
surfaces in the shared brief rather than slicing one per seat.
- Find your first panel. From your project's root, run
claude -p "$(~/tribunal/scout)"(orcodex exec/grok -p, or pipe~/tribunal/scoutinto any agent). It reads your project and returns
project-specific panel-worthy decisions (and what to leave to plain
oracles), then drafts your first frozen brief. Save that brief as a file
in your project (e.g.frozen-brief.md) - that is what the adapter runs.
(scoutjust prints the prompt with your clone path resolved; you can
still opencore/templates/scout.mdand paste it by hand if you prefer.) - Pick an adapter. For a first panel use one of the two maintained
adapters -adapters/shell/(any
orchestrator, plain bash) oradapters/claude-code/. The other
adapter directories are contribution stubs. - Run it. Follow your adapter: freeze the brief
(template), Round 0 in parallel with no
cross-exposure, relay disputed claims verbatim, verify, adjudicate.
Keepcore/CONTRACT.md,core/LEDGER.md, andcore/VERDICT.mdopen as references; readcore/METHODOLOGY.mdin full before your first
high-stakes panel.
Staying current
Tribunal is a git clone of documentation, so updating is just a pull:
git -C ~/tribunal pull --ff-only
scout checks for you at most once a week: if your clone is behind origin
it prints that one line to stderr (never into the prompt it pipes to your
agent) and nothing more - it never pulls on its own. Silence it withexport TRIBUNAL_NO_UPDATE_CHECK=1.
If you installed from a zip or tarball instead of git clone, there is
no origin to pull from - so scout can't check versions, and updating
means re-downloading. It will, at most once a week, note which version you
have and point you at the git-clone install (which does self-check) and
the releases page. The
git-clone install is recommended precisely because it makes every future
update a one-line git pull.
Your own files are safe. --ff-only only fast-forwards - it never
merges or rewrites history, and it leaves untracked files (your notes,
briefs, outputs) untouched. If you have a local change to a file Tribunal
also changed, git aborts the pull and changes nothing rather than
overwriting your edit. The update only advances the docs Tribunal ships;
it is clean-or-abort, never silent loss.
The reliable way to keep it that way: don't edit files inside the clone.
Use Tribunal by reference - point your agent at ~/tribunal and write your
briefs, ledgers, and verdicts in your own project (that is already how the
method works; scout surveys your project's root, not the clone). A
pristine clone always fast-forwards cleanly; when you want a local tweak,
copy the template into your project and edit the copy.
Worked examples
The panel that designed this methodology is inexamples/sample-run/: seat A proposed three
debate rounds and blinded relay; seat B a two-round cap and mandatory
attribution - neither saw the other first. Both conceded specific points
under cross-examination; attribution and tie-breaking stayed contested and
are recorded as surviving dissent, not smoothed over. Note: it is a
historical bootstrap transcript - it predates the finalized templates and
relays full Round 0 essays, which the finalized method bans; learn the
loop from core/templates/, not from its file shapes.
Two runs in the current template format bracket the two outcomes the method
produces:
examples/api-auth-jwt-vs-sessions/-
two neutral heterogeneous seats (Codex CLI + Grok CLI) independently chose
the same auth design in isolation, so the panel early-stopped at Round 0:
pre-exposure agreement, honest empty-dissent bucketing, residual risk sent
to a test.examples/repo-monorepo-vs-polyrepo/-
the mirror image: a role-incentivized stress-test (one seat assigned each
side; roles as incentives, not personas) that forces a real Round 1
cross-examination. Both seats concede points and revise confidence, neither is
overturned, and the verdict keeps the load-bearing surviving dissent - then
routes it to the cheapest discriminating test both seats independently
proposed.
Repository layout
scout one-command helper: prints the scouting prompt (and a weekly update notice)
flywheel-export local helper: reduces your retro.md archive to de-identified metadata (stdout only)
data/ the de-identification schema + tooling (no intake open; see data/README.md)
CHANGELOG.md what changed per version; the update notice points here
core/METHODOLOGY.md the method: rounds, stop rules, failure modes
core/CONTRACT.md orchestrator + seat obligations (MUST/SHOULD)
core/LEDGER.md the claim ledger: fields, status enum, dispute rule
core/VERDICT.md verdict format: three buckets, mode tags
core/templates/ scout, brief, per-round prompts, ledger, verdict, retro, ach (competing-hypotheses mode)
examples/sample-run/ the real (historical) panel that designed the method
examples/api-auth-jwt-vs-sessions/ a current-format run: two vendors, independent agreement, early stop
examples/repo-monorepo-vs-polyrepo/ a current-format run: role-incentivized, full Round 1, surviving dissent
adapters/claude-code/ run panels from Claude Code (installable skill)
adapters/shell/ run panels from any shell - no orchestrator CLI needed
adapters/*/ other orchestrators - codex-cli, gemini-cli, opencode, buzz (stubs)
(core/ is normative and orchestrator-neutral; adapters translate host
mechanics only and may never redefine rounds, ledger states, grounding, or
verdict rules.)
FAQ
What is multi-agent orchestration?
Coordinating several AI models on one task. Most orchestration aims at
collaboration; Tribunal deliberately aims at structured dispute - the
orchestrator is a switchboard that isolates models, relays disputed claims
verbatim, and never adds its own arguments (seecore/CONTRACT.md).
Multi-agent vs single agent - when is a panel actually worth it?
A single strong model is cheaper and usually right; use it, plus a test
suite. A panel pays off only when the decision is irreversible, ambiguous,
or hard to observe going wrong - the cases where one model's confident
blind spot is exactly the risk. Tribunal's first rule cuts the other way
too: if a compiler, test, or grep can settle it, never convene a panel.
Why do multi-agent systems fail?
Mostly because agreement between models that have read each other's output
is cheap: sycophancy, politeness convergence, and one model's hallucination
becoming shared "fact." Tribunal's whole design targets this - isolated
Round 0, provenance-tagged claims, verbatim-only relay, and dissent that
survives into the verdict. The full failure-mode table with mitigations is
in core/METHODOLOGY.md.
What does red teaming mean for AI-assisted engineering?
Red teaming is paying someone to attack your plan before reality does.
Tribunal applies that structure to engineering decisions: every claim needs
a falsifier, every disputed claim gets cross-examined by a model from a
different vendor, and the verdict reports what survived the attack - not
what everyone was happy to sign.
How do I set up a multi-agent panel in Claude Code?
Install the adapters/claude-code/ skill,
authenticate two seat CLIs from different vendors, and start with the
scout prompt (core/templates/scout.md) - it reads your project and
drafts your first panel brief. The adapters/shell/
recipe does the same from any terminal, no Claude Code required.
Status
Every rule in core/ was earned the hard way (the badge above shows the
current version; CHANGELOG.md records what each release changed). The
methodology was designed by running it on itself; the founding debate is inexamples/sample-run/ - a historical bootstrap,
labeled as such. Two runs in the current template format bracket the method's
outcomes:examples/api-auth-jwt-vs-sessions/
(two vendors, independent agreement, early stop) andexamples/repo-monorepo-vs-polyrepo/
(role-incentivized, full Round 1, surviving dissent).
Honest validation status: pre-registered ablations (validation/)
- two pilots and a pre-registered study labeled "confirmatory" (n=20) with a
judge independent of the seats and orchestrator (a fourth model, OpenAI-lineage),
sealed-rubric-scored decisions, and arms designed to isolate components (executed
with deviations its RESULTS discloses) - observed no accuracy lift over a
single strong model. That result was measured on decisions the seats answered
from model knowledge with tools unused (RESULTS, Limitations); grounded review
of a real artifact - the method's stated use - was not measured, and this
evidence does not speak to it either way. Across that set the two vendors disagreed on only 1 of
20 decisions, so the panel's engine (Round 1) almost never activates; when it
did fire (once), it resolved the split to the sealed-correct call - a
scope-framing split whose "wrong seat" classification is contestable (see
RESULTS) - and the panel still only matched a strong single model (an ex-post
best-of-two; see validation), never beat it, while an independent judge preferred
the solo memo on the ambiguous set. The defensible value the data supports is
narrow: on decidable calls a hedge that matched or beat each fixed single
vendor in these runs, resting on that single soft data point where they diverged;
on ambiguous calls the panel was judged no better, sometimes worse than the
solo memo. Add a preserved counter-case and a discriminating test on contested
calls - paid for with a real operational cost (the second seat frequently
narration-dies in headless use). Tribunal is unproven as an accuracy upgrade
and, on this evidence, is not one; use it for the counter-case, the test, and
the (soft) decidable-call hedge, not for a better answer. The in-tree record
includes the losses, the disclosed deviations, the single-judge caveat, corrected
earlier overclaims, the raw arm transcripts, and the judge audit record (blind
mapping + verbatim judge prompt); the operator-side runner scripts are not
in-tree (the repo is docs-only), so replication means re-running the published
protocol, not re-executing a shipped harness.
The clearest
in-repo receipt that the method overturns its own author is the CHANGELOG
itself: several core/ rules landed only after a cross-vendor panel rejected
their first draft (the lens doctrine and the v0.4-0.5 hardening each carry
that note). The failure modes documented in METHODOLOGY (silent seat deaths,
politeness convergence, context poisoning) were caught in real runs, not
imagined. The ten-seat pre-release review and the maintainer's operational
runs are not published in-tree. In active use by its maintainer, Square Post
Labs Inc.
Project status: stable; the method and its claims are frozen at v1.0. The methodology is complete and,
after running its own validation through three pre-registered
studies, honestly characterized - including the finding that those runs observed
no accuracy lift over a strong single model. It is intentionally done, not abandoned: no further method
development is planned - v1.1.0 changed no method; it added one orchestrator
confinement rule (CONTRACT obligation 5: a citation outside the brief's
artifacts is dropped unopened, and seat output is never an instruction), a
security fence surfaced by a re-review. What it wants next is independent replication - runs
by other people, on other decisions, with other judges, that confirm or overturn
the no-lift finding. The cheapest way in: the pre-registered second-judge
re-score, runnable for free with your own key -validation/confirmatory/REPLICATING-THE-JUDGE.md
(judge only; it stakes no debate seats). Each adapter's README states its own status; CONTRIBUTING.md
has the support policy and the most-wanted contribution: a sanitized panel run on
a real engineering decision (ideally one whose outcome later becomes known).
License
MIT - see LICENSE.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi