orka
Health Gecti
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Community trust — 11 GitHub stars
Code Basarisiz
- rm -rf — Recursive force deletion command in install.sh
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
Make your best coding agent the CTO: it plans, reviews and merges while Codex, Grok, Claude Code, Copilot & co. do the work in parallel git worktrees. An Agent Skill + a few hundred lines of bash.
Orka
Make your best coding agent the CTO. Let cheaper agent CLIs do the work.
One agent plans, writes task cards, reviews and merges. Codex, Grok, Claude Code, Copilot, Cursor, Gemini or OpenCode workers
do the coding in parallel, each in its own git worktree. A different model reviews every critical change before it lands.
Your strongest agent is also your most expensive, and its context and quota run out first. Orka is an Agent Skill plus a few hundred lines of bash that turn it into a manager. It talks with you, keeps the backlog and writes a task card for each piece of work. Headless coding-agent CLIs implement the cards, each in its own git worktree.
"Prompt + git worktree" is the easy part. What Orka adds are the rules that made it work on real projects:
- A different model reviews every critical change before merge.
- The orchestrator re-runs the acceptance itself. A worker's report is a claim, and an implementation card with
commits: 0delivered nothing. - QA checks against your original words, not the developer's checklist.
- Every run is logged and scored, so the next assignment is based on how each worker actually did.
These came from 1,035 worker runs (477 cards, 242 worker-hours) on eight projects between August and October 2026. In all of them Claude Code orchestrated Codex, Grok and its own sub-agents. The anonymized numbers are in docs/evidence.md.
A session, shortened (a card is one task file for one worker; c03r is the lead's review of card c03):
you › Use Orka. Fix the five registration-form bugs we just listed, then the payment page.
orka › Team from .orka/orka.json: lead codex/astra · senior codex/sol · mid grok · junior codex/luna
Wrote 3 cards. Predicted: c03 senior (hard point: old stored forms must still load) …
Say "go" when the backlog is right.
you › go
orka › c03r (lead): REQUEST CHANGES, old stored forms vanish from the admin page.
Fix list appended to c03, attempt 2 running. c04 accepted: build, tests, browser check OK.
$ .orka/bin/orka status
CARD STATE TRY RC MIN GIT WORKER
c03-registration-form running 2 - 12 -
c03r-review done 1 0 2 0 codex/gpt-6-astra/high
c04-payment-page done 1 0 24 3 grok/grok-4.7/medium
c03q-qa todo
Why
- Tokens and quota. One strong model managing four cheaper ones gets through a lot more work before anything runs out. The orchestrator reads diffs and reports, not the whole codebase.
- A second model finds what the author can't. 115 of 127 first-round lead reviews came back REQUEST CHANGES, for real problems that had got past the author's own passing tests: a data-loss regression, a session-replay window, a double-order race. The median review took about 6 minutes (details).
- Parallel, isolated work. Each card runs in its own git worktree and branch, with its own database, ports and browser session if you want them. Workers never step on each other or on your checkout.
- A ledger instead of a gut feeling. Every run gets a predicted and an actual score (smart / dumb / speed / cost) in
ledger.jsonl. The skill tells the orchestrator to read the per-worker averages before it assigns the next card. Nothing is trained; it's a log the next session actually reads, and the orchestrator's own mistakes go into it too.
How it works
you ──── talk, decide, say "go"
│
┌───────▼────────┐ backlog.md · decisions.md · ledger.jsonl
│ orchestrator │ (Claude Code, Codex, Grok, Copilot, Gemini, OpenCode, Cursor …)
│ = the CTO │
└───────┬────────┘
card.md │ .orka/bin/orka queue "c03-form senior" "c04-pay mid" …
┌───────┼──────────────────────┬─────────────────────┐
▼ ▼ ▼ ▼
senior mid junior lead
worktree worktree QA in a headless reviews the diff
orka/c03 orka/c04 browser, screenshots REQUEST CHANGES / APPROVE
│ │ │ │
└───────┴──── report.txt · meta.json (exit code, seconds, commits, usage) ─┘
│
orchestrator re-runs the acceptance itself → merge → score → retro
Roles are abstract. lead reviews and judges; senior takes the hard cards and may split them into sub-cards for cheaper workers; mid does normal development; junior does research, browser, mobile and CLI QA, and simple fixes. Hands-on work stays with the junior: the orchestrator and the lead are the most expensive models on the team, so they don't browse, click through pages or dig through the codebase themselves. The lead can send the junior its own sub-cards to reproduce a bug. You map each role to any CLI and model you have: four vendors, or one vendor at four price points (for example Opus / Sonnet / Haiku).
Works with
As the orchestrator: any agent harness that can load an Agent Skill (SKILL.md) and run shell commands: Claude Code, Codex CLI, Grok CLI, GitHub Copilot CLI, Gemini CLI, OpenCode, Cursor, Amp and others. In harnesses without skill support, point the agent at SKILL.md.
As workers: each worker CLI is a small adapter in .orka/bin/workers/<cli>.sh.
| Worker CLI | Adapter | Verified end to end¹ |
|---|---|---|
| OpenAI Codex CLI | codex.sh |
✅ codex-cli 0.157 (580+ production runs) |
| Grok CLI | grok.sh |
✅ grok 1.0.41 (140+ production runs) |
| Claude Code | claude.sh |
✅ 2.1.284 |
Your orchestrator's own sub-agents ("cli": "subagent") |
built in (orka run + orka finish) |
✅ Claude Code 2.1.284 Agent tool (298 production runs) |
| GitHub Copilot CLI | copilot.sh |
✅ 1.0.88 |
| Cursor Agent CLI | cursor-agent.sh |
⚠️ flags checked against --help; not run (no login on the test machine) |
| Gemini CLI | gemini.sh |
⚠️ flags checked against --help; not run (no login on the test machine) |
| OpenCode | opencode.sh |
⚠️ flags checked against --help; not run (no provider on the test machine) |
| Anything else | ~10 lines |
¹ A real worker received a card, edited a file in its worktree, committed it and returned its report. Tried one on the ⚠️ ones? A PR that flips it to ✅ is very welcome.
Install
Before you install: workers run headless with their CLI's approval prompts turned off, as your user, with your credentials. A worktree keeps them from colliding; it is not a sandbox. Use a container or VM on machines that matter. More
Pick one:
# Every harness on this machine: installs to ~/.agents/skills/orka, links it for Claude Code (and Cursor, if installed)
curl -fsSL https://raw.githubusercontent.com/ugorur/orka/master/install.sh | bash
# Claude Code plugin
/plugin marketplace add ugorur/orka
/plugin install orka@orka
# With the skills CLI (asks which agents to install for)
npx skills add ugorur/orka
# Or by hand: copy skills/orka to wherever your harness reads skills
git clone https://github.com/ugorur/orka && cp -R orka/skills/orka ~/.agents/skills/
Install at user scope, not inside your project. A copy inside the repo would be checked out into every worker's worktree.
Runtime needs bash, git, jq and timeout (on macOS: brew install coreutils jq), plus the worker CLIs you want, each logged in once interactively.
Quick start
In a git project, tell your agent:
Use Orka. Let's plan the next piece of work.
Bootstrap. It runs
init.sh, sees which CLIs are installed, proposes a team and asks you once. Your answer goes to.orka/orka.json. Model ids are whatever your CLI accepts; these are the ones we used in September 2026:{ "team": { "lead": { "cli": "codex", "model": "gpt-6-astra", "effort": "high" }, "senior": { "cli": "codex", "model": "gpt-5.6-sol", "effort": "high" }, "mid": { "cli": "grok", "model": "grok-4.7", "effort": "medium" }, "junior": { "cli": "codex", "model": "gpt-6-luna", "effort": "medium" } }, "parallel": 2, "timeoutMinutes": 120, "worktreeDir": "../myproject-wt", "autoMerge": false }Smallest setup: one CLI you already use, at different price points. In Claude Code on a subscription, all four roles can use
"cli": "subagent", withopusas lead,sonnetas senior and mid, andhaikuas junior. Or use one headless CLI for every role. You can add other vendors later; a lead from a different vendor than the author gives the best reviews. Headlessclaude -pworkers are for API-key or gateway setups.Talk. Describe bugs and features the way you would to a team lead. They go into
.orka/backlog.mdin your words. Nothing runs yet.Say "go". Cards are written, scores predicted, workers dispatched in parallel.
Watch or sleep. The orchestrator accepts each branch itself, gets critical ones reviewed by the lead, has the junior test the UI, merges (or asks first), scores everyone and reports back: what's merged, how it was verified, what is still open.
Everything lives in .orka/ in your project, which is git-excluded and never committed:
.orka/
orka.json the team backlog.md your requests, your words
COMMON.md project rules decisions.md locked decisions + lessons
tasks/ the cards ledger.jsonl every score, predicted vs actual
runs/<card>/attempt-N/ prompt.md · report.txt · meta.json · logs
env/<slot>.env per-card ports/DBs bin/ runner + worker adapters
The commands
The orchestrator calls these; you rarely need to, but they are plain bash and easy to read.
| Command | What it does |
|---|---|
orka run <card> <role> [worktree] |
Run one card on the worker for that role; creates worktree + branch orka/<card> if needed |
orka run <card> <cli> <model> <effort> [worktree] |
Same, with a one-off worker |
orka finish <card> [exit-code] |
Finish a prepared native sub-agent attempt; use --abandon after a crash |
orka queue "<card> <role>" … |
Run many cards, at most parallel at once (no spaces in worktree paths) |
orka status |
Every card: todo · running · done · failed · timeout · stopped, minutes, commits, worker |
orka score <card> predicted|actual … |
Log a score; actual also copies worker, time and usage (tokens / $) from the run |
orka score --summary |
Average scores per worker/model, which the orchestrator reads before assigning |
(orka = .orka/bin/orka.)
The scores
For each worker run, the orchestrator records four numbers from 0 to 100: before the run what it predicts for the worker it picked, after the run what actually happened. A card that needed three attempts has three actual lines. The gap between the two is how it learns whom to give the next card.
| Score | Question it answers | Better | Example |
|---|---|---|---|
| Smart | Did the worker clear the card's hard point: the real root cause, everywhere it applies, honestly verified? | higher | 95: found the actual cause and proved it with a probe · 50: fixed the easy part, missed the risky part |
| Dumb | How silly was the worst mistake it made? One bad mistake is enough; it is not an average. | lower | 0: none · 30: tested against a hand-made object instead of real data · 80: faked a success or broke something unrelated |
| Speed | How fast was it for the size of the card? | higher | a 2-minute review scores high; 50 minutes for a small fix scores low |
| Cost | How cheap was it (tokens or dollars) for the value delivered? | higher | a $0.04 critique scores high; a $9 run for a one-line change scores low |
Smart and dumb are separate on purpose. A worker can crack a hard problem (smart 90) and still do something careless on the way (dumb 30). You want to know both before you hand it the next card.
$ .orka/bin/orka score c03-registration-form predicted 85 10 50 45 "senior: contract → API → admin → mobile"
$ .orka/bin/orka score c03-registration-form actual 78 35 35 35 "legacy rows vanished; tests used a fake object twice"
(each prints the ledger line it appended, as JSON)
$ .orka/bin/orka score --summary # example from a longer ledger: averages of "actual" lines per worker
WORKER RUNS SMART DUMB SPEED COST AVG_MIN
codex/gpt-6-astra/high 5 94 3 86 90 4
grok/grok-4.7/medium 2 87 7 86 92 3
codex/gpt-6-luna/medium 1 50 35 90 92 2
The scores are the orchestrator's honest judgement, not a benchmark. In the retro after each piece of work, independent reviewers (including the lead) score the orchestrator itself on the same scale, as retro lines.
What 1,035 runs taught us
These are written into the skill as rules. Each one is here because we paid for it once:
- "Done" means the user's words, everywhere they apply. When QA checked against the developers' own checklists, the owner's re-check found only about 10% of the items truly done. QA now checks against the original request.
- Author and reviewer must be different models. It is the cheapest insurance in the whole setup.
- A worker's "done" is a claim. The orchestrator re-runs the acceptance in the worktree. An implementation card with
commits: 0delivered nothing, whatever the report says. - Tests must use real data shapes. One fix passed its unit test twice against a hand-made object and still failed on the real stored row. A junior's browser QA caught it.
- UI changed = seen in a headless browser. A main button shipped broken (blocked by CSP) while every test was green.
- Isolate everything: worktree, database, ports, cache index, browser session. Never the user's browser, never a shared container. Kill processes by PID only.
- Every card names its hard point. Otherwise you get the easy 80% done and the risky 20% quietly skipped.
- Weak QA goes straight to the next level. If the junior cannot run the checks, re-run the QA on mid immediately instead of retrying the junior.
- Research cards need web access, or they will quote from memory as if they had fetched it.
- The orchestrator and the lead don't do the junior's work. Their browsing, screenshots, manual QA and inline fixes were the waste users flagged most often. They write a card and read the report instead.
- The orchestrator is scored too. An independent reviewer and the lead grade it after each piece of work. In our runs its logged mistakes included a stdin hang, editing the runner while jobs were running, and a
pkillthat killed its own shell.
Add a worker CLI
An adapter gets five environment variables and must leave the worker's final message in report.txt:
#!/usr/bin/env bash
# .orka/bin/workers/mycli.sh: ORKA_WT (worktree) ORKA_MODEL ORKA_EFFORT ORKA_PROMPT (file) ORKA_OUT (dir)
args=(--non-interactive --yes --cwd "$ORKA_WT")
[ -n "$ORKA_MODEL" ] && args+=(--model "$ORKA_MODEL")
mycli "${args[@]}" < "$ORKA_PROMPT" > "$ORKA_OUT/report.txt"
Then use "cli": "mycli" in orka.json. The runner handles the worktree, prompt assembly, timeout, secret masking in reports, meta.json and the ledger. Optionally write $ORKA_OUT/usage.json (for example {"usd": 0.42}) and it will show up in the ledger. Please send adapters upstream.
Safety: read this
- Workers run with approvals bypassed (
--dangerously-bypass-approvals-and-sandbox,--always-approve,--dangerously-skip-permissions, …). Nobody is there to click "allow" for a headless worker. They run as your user, with your credentials. - A git worktree is not a sandbox. It keeps workers from colliding; it does not keep them out of the rest of your disk. On any machine that matters, run Orka inside a container or VM.
- Worker reports are pattern-masked for obvious secrets (
*_KEY=…,Bearer …,sk-…,ghp_…) before the orchestrator reads them. This is a seatbelt, not DLP. Keep real secrets out of.orka/env/. - Orka never pushes. Merging to your main branch asks you first unless you set
"autoMerge": true.
FAQ
Can I use my harness's built-in sub-agents? Yes. Set a role's CLI to subagent; Orka prepares its worktree and prompt, your orchestrator runs it, and orka finish records its report and ledger metadata. Native sub-agents usually share the orchestrator's vendor and quota, so separate CLIs still help when you want different quotas or blind spots.
Do I need Claude Code? No. Any harness that can read SKILL.md and run bash can be the orchestrator, and any CLI with an adapter can be a worker. Our production runs happened to use Claude Code → Codex + Grok.
What does it cost? Only what your worker CLIs cost. Real examples from the ledger: a lead review, about 6 minutes (median of 348) and a few hundred thousand mostly-cached tokens; a senior feature card spanning contract → API → admin → mobile, 20–50+ minutes; a Grok run, median $3.22 and at most $10.40 (83 runs). orka score --summary shows your own numbers.
Is this a framework? No. It's a skill (a markdown playbook) plus about 300 lines of bash. No daemon, no server, no database, no Node/Python runtime.
Status
v0.2. Used on eight projects so far: two private production repos before packaging, then six projects on this runner (818 runs, 28 Sep to 10 Oct 2026). CI runs the plumbing tests on Linux and macOS (including the stock bash 3.2). Issues and PRs are welcome, especially adapters and ✅ verifications for more CLIs.
Run the plumbing tests with bash tests/smoke.sh (fake worker, no network, about a minute).
License
MIT © Umurcan Gorur
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi