crux

agent
Guvenlik Denetimi
Basarisiz
Health Gecti
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Community trust — 186 GitHub stars
Code Basarisiz
  • process.env — Environment variable access in .github/workflows/approve-contributor.yml
  • fs module — File system access in .github/workflows/approve-contributor.yml
  • rm -rf — Recursive force deletion command in .github/workflows/build-binaries.yml
  • child_process — Shell command execution capability in .github/workflows/issue-analysis.yml
  • process.env — Environment variable access in .github/workflows/issue-analysis.yml
  • fs module — File system access in .github/workflows/issue-analysis.yml
  • network request — Outbound network request in .github/workflows/issue-analysis.yml
Permissions Gecti
  • Permissions — No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

Taking the pi coding agent to Claude Code-level performance: 0.539 → 0.773 pass@1 on Terminal-Bench 2.1 with the same self-hosted 27B model, by re-engineering the agent.

README.md

Crux

Taking the open-source pi coding agent to Claude Code-level performance on Terminal-Bench — same model weights, re-engineered agent.

pass@1 0.773 pi 0.539 to Crux 0.773 Claude Code on the same model 0.730 4399 tests self-hosted model built on pi

English · 简体中文


Crux takes pi, an open-source coding agent, and re-engineers
its agent layer for long-horizon terminal work. On
Terminal-Bench 2.1 — 89 real tasks spanning compilers,
emulators, cryptanalysis, ML training and systems administration, each graded by
a hidden test suite inside a container — it lifts pi from 0.539 to 0.773
pass@1, to the level of Claude Code running the same model (0.730).

The model weights never change. Crux's work is in the agent layer: the loop and
its recovery paths, context management and compaction, tool contracts, and the
runtime that keeps multi-hour runs alive. Every change ships with the
measurement that justified it.

Highlights

  • pi → Claude Code-level, same weights. pass@1 from 0.539 to 0.773
    (+23.4 points) on Terminal-Bench 2.1 with a self-hosted 27B model; Claude Code
    on the same model scores 0.730. Reproduced at 0.793 on an independent run.
  • A context engine that holds up under pressure. Compaction success taken
    from 14% to 100%; output truncation cut from 18.0% to 0.7%; single-task
    sessions summarized in full, with the task's exact wording carried through
    every compaction.
  • Hang-free long runs. Timed-out trials had been idle for 94% of their
    wall clock; three root causes found and fixed, taking idle time to 3%.
  • Statistically rigorous evaluation. Paired sign tests on concurrently run
    arms, a measured noise floor, and a failure-attribution pipeline that separates
    agent failures from infrastructure failures automatically.
  • Designs proven against data, not adopted on reputation. Mechanisms from
    Claude Code, Codex, opencode, hermes-agent and grok-build were evaluated
    against 356 trajectories; 34 candidate designs were ruled out by the
    numbers before any code was written.

Results

Self-hosted Qwen3.8-27B · the 89 Terminal-Bench 2.1 tasks that run without a GPU
inside the container · 8× agent budget · pass@1 over whole runs.

Configuration pass@1
Crux 0.773  (0.793 on a second run)
Claude Code, same model 0.730  (32K window)
Crux — agent loop and tool contracts only 0.591
pi, upstream 0.539

Reproducibility. Two independent runs of the final configuration scored
0.773 and 0.793. The benchmark's run-to-run variance was measured directly, and
every comparison in this repository is paired task by task to account for it.

Architecture

flowchart TB
    subgraph agent["Agent · packages/"]
        direction LR
        ctx["Context engine<br/>window-fitted compaction<br/>verbatim task"] <--> loop["Agent loop<br/>state machine · recovery<br/>budget pacing"] <--> tools["Tools<br/>bash · read · edit · write<br/>stream watchdog"]
    end
    subgraph evals["Evaluation system · benchmark/"]
        direction LR
        pre["Preflight"] --> run["Containerized runs"] --> attr["Failure attribution"] --> stats["Paired sign tests<br/>→ next change"]
    end
    agent ==> evals

Core engineering

Each capability below is tied to the trajectory evidence that motivated it.

Context engine

Capability Evidence
Summarization requests sized to the window they live in History and summary share one window and nothing checked their sum: 2,115 of 2,370 compactions had been rejected outright. Now 100% succeed
Adaptive retry on rejection Characters per token ranges from 1.95 to 3.99 across real trial text, so a rejected request halves and retries rather than trusting a constant
Partial summaries kept A summary cut at the output cap still carries the work; discarding it disabled compaction on smaller windows
Cut-point search that always makes progress A large trailing tool result no longer leaves the search with nowhere to cut
Full summaries for single-task sessions An agent run on one task is one conversational turn, and every compaction had taken a short turn-fragment path that kept 2,302–4,443 characters of ~250,000 tokens. Now the full structured summary
The task's exact words survive compaction The original request is carried verbatim and re-read from the session each time, never paraphrased — the approach Codex (openai/codex#48115) and hermes-agent take

Liveness and robustness

Capability Evidence
Stream watchdog SDK timeouts cover getting a response, not keeping one: 25 timed-out trials had been silent for a median 112 of their 121 minutes
Default command timeout Ten minutes, chosen from data: 161 commands legitimately ran 2–10 minutes, against 60 runaway ones
Process-tree reaping A detached child writing to an inherited pipe once held a container for eight hours after 48 seconds of work
Resume after OOM kill A cgroup OOM kill ends every process in the group; the run now resumes instead of scoring zero

Agent loop

Capability Evidence
Explicit loop state machine Every recovery path — truncation, budget, escalation — lives in one immutable LoopState with explicit limits and a recorded transition reason per turn
Two-phase truncation recovery Raise the output ceiling and retry silently first; speak to the model only if it truncates again — Claude Code's design
Reasoning carried across a cut Reasoning is not replayed between turns, and 292 turns in 60 trials had ended mid-thought and restarted from scratch. The tail of the reasoning is now handed back
Stall detection A turn that stops inside its reasoning with no answer is treated as a stall, not as completion
Escalation for truncated tool calls All 42 cut tool calls had stopped at the initial 16K with the model's own 32K unused; they now get the raised ceiling, as in Claude Code and hermes-agent
Budget-aware pacing The loop sees its deadline, warns before it, and questions a stop while most of the budget is unspent; each task's budget follows its own time limit (80–1,600 minutes)

Tools and prompting

Capability Evidence
Edit failures that show the file Instead of "must match exactly", a failing edit returns the nearest region of the file by bigram similarity, so the next attempt targets real text
Head-and-tail command output Long output keeps its first lines as well as its last, so the first compiler error survives truncation
Visible working notes The model is told its reasoning is not kept and writes short notes beside tool calls; the median visible text beside a tool call had been 0 characters
Calibrated output ceiling A per-response default sized from this deployment's own p50/p95/p99 output distribution

Evaluation methodology

  • Paired, concurrent comparisons. Arms are compared task by task with a
    sign test, and only when run in the same time window — run-to-run variance on
    this benchmark is large enough that totals alone mislead.
  • Mechanism health before tuning. A mechanism's success rate is counted
    before any of its parameters are touched; this is how the compaction failure
    rate was found and fixed.
  • Failure attribution. A seven-category taxonomy classifies every trial
    automatically and separates the agent's failures from the environment's,
    including verifiers whose own test runners failed to install.
  • Within-task analysis. For tasks solved in some runs and not others, the
    solved and failed runs are compared directly, which controls for task
    difficulty.
  • Preflight gating. Every launch is checked for endpoint, launch settings
    and harness checksum before it spends GPU hours.

Repository layout

packages/          the agent: core loop, model layer, tools, CLI (built on pi)
benchmark/         the evaluation system
  src/crux/        harbor agent, prompt sections, failure analysis
  scripts/         launchers, preflight, paired comparison, health checks
  tests/           486 tests

Quick start

cd benchmark && uv sync
harbor run --dataset terminal-bench/terminal-bench-2-1 \
  --agent crux.pi_agent:CruxPiAgent --model openai/<model>

benchmark/scripts/preflight.sh <launcher> validates the endpoint, the launch
settings and the harness checksum before a run. See
benchmark/README.md for the evaluation system.

Acknowledgements

Crux is built on pi by Earendil Works. Upstream issues and
pull requests belong at
earendil-works/pi.

Package Description
@earendil-works/pi-coding-agent Coding agent CLI (installs as crux and pi)
@earendil-works/pi-agent-core Agent runtime with tool calling and state management
@earendil-works/pi-ai Unified multi-provider LLM API
@earendil-works/pi-tui Terminal UI components
@earendil-works/chord Application-composition runtime for services, RPC and plugins

Licensed as upstream; see LICENSE.

Yorumlar (0)

Sonuc bulunamadi