OpenRUA
Health Warn
- License — License: Apache-2.0
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 5 GitHub stars
Code Fail
- rm -rf — Recursive force deletion command in examples/plugins/kimi/agent.yaml
- rm -rf — Recursive force deletion command in examples/plugins/zcode/agent.yaml
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
Let Your Claude Code or Codex Control Any Robot, Real or Simulated
OpenRUA
Let Your Claude Code or Codex Control Any Robot, Real or Simulated
Through the standard ROS 2 CLI and client library, without relying on any VLA model.
https://github.com/user-attachments/assets/3b134c51-a949-44dd-9474-5249c3879aa0
Codex (GPT-6 Astra) on LIBERO-10, "put the yellow and white mug on the left plate and put the white mug on the right plate": the commands it typed on the left, the robot's cameras on the right, task success. More in docs/demos.md.
A robot-use agent (RUA)
uses a robot just as a computer-use agent uses a computer.
OpenRUA connects off-the-shelf coding agents to robots through their native
ROS 2 interfaces. The agent works in a workspace containing robot
documentation and starter tools, writes perception and control programs, and
uses execution feedback to continue the task.
You can chat with the agent through OpenRUA's terminal UI or browser, use
its original terminal, or run recorded benchmark experiments. Shared
chat lets you send follow-up instructions, inspect tool activity, manage the
queue, and revisit saved observations without starting a new robot each turn.
- Start a shared session: OpenRUA TUI, browser, or plain CLI.
- Use the agent's original terminal: launch a task with
openrua run. - Connect your robot: describe its ROS 2 interfaces in a profile.
- Run experiments: fresh trials, benchmark scoring, and recorded artifacts.
OpenRUA turns your robot into a coding project: your agent explores it
like a live codebase, pulls sensor streams into files for reading, and
runs commands and programs to move it.Start playing with your robot like you code a project :)
Quick start
Chat with a robot
Use a Linux host with Docker or a supported Podman setup and an existing
Claude Code or Codex login (installation).
pip install -U openrua
openrua
The keyboard-first TUI guides environment selection and prepares missing images.
Each launch creates a new session; no name is required. Type an instruction and
press Enter, for example:
Inspect the workspace documentation and describe the scene without moving the robot.
Use /files to browse saved observations and code, /resume to find earlier
sessions, and /end to stop execution while keeping your work. Ctrl+D on an
empty input detaches while the session continues running.
See the walkthrough for setup and continuous tasks,
terminal guide for controls, and
validation scope for what has been tested.
Shared chat remains experimental.
Native Claude Code or Codex login is the default. To use an API key, explicitly
select it in openrua --setup or through the
API configuration guide.
OpenRUA never switches to API billing after a login error.
Run a benchmark and make a demo
For researchers, the CLI runs fresh benchmark trials and saves the code,
observations, transcripts, scores, and configuration needed to inspect a run.
The following example runs one CaP-Bench Lift trial, records its cameras,
and renders a video with terminal commands beside the robot views.
Install the demo dependencies and prepare the environment once (Docker must
be running; an existing native Claude Code login is reused):
pip install -U 'openrua[demo]'
claude login # only if not already signed in
openrua build --bench capbench
openrua build sandbox --distro humble --agent claude-code
openrua build proxy --agent claude-code
openrua doctor panda --sim robosuite --bench capbench --agent claude-code
Then run and render:
openrua bench --config capbench --run-id lift-demo \
--task-suite capbench_lift --task-ids 0 --seeds 0 \
--operator agent --record --record-every 4 --wall-clock-min 10
openrua demo runs/capbench/lift-demo/trials/capbench_lift-0/seed0 --speed 4
The video is saved as demo.mp4 in that trial directory. Its result.json
reports the benchmark verdict, and provenance.json records the resolved
configuration and software versions. This small, recorded run demonstrates the
method; reproducing the paper's aggregate results requires its full task/seed
sets, model settings, and evaluation budgets. Recording also adds rendering cost.
The bundled CaP-Bench configuration selects Claude Opus 5 with high reasoning
effort; --agent and --model can select another available configuration.
Native login requires the corresponding host CLI. See installation
for prerequisites and credentials. To check the benchmark setup without a model
call, use the same bench command with --operator none and omit recording;
this checks startup and interfaces, not task success. See
running experiments for full runs and recorded
replays of existing trials, and the paper
for the experimental protocol.
Choose your interface
| Interface | Entry | When you leave |
|---|---|---|
| OpenRUA TUI | openrua --name chat-demo |
Ctrl+D detaches; the shared session continues |
| Browser | openrua --gui --name chat-demo |
Closing the page leaves the shared session running |
| Plain CLI | openrua --cli --name chat-demo |
Leaving the client keeps the shared session running |
| Native agent terminal | openrua run panda --sim robosuite --bench capbench --agent codex --name native-demo |
Exiting the agent stops the resources owned by run |
The first three use one shared queue per session. These examples explicitly
name a session; omitting --name creates a new one with an automatic ID.
To reconnect to a running session, pass only its name and interface choice;
startup options cannot change an active conversation. Existing chat andsession commands remain available for explicit client operations.
Native run starts a separate session using the agent's original TUI; it cannot
yet take over that shared conversation. For a native terminal with a robot that
stays up between agent visits, use up, agent, and down as described in
First task in simulation.
In shared chat, new instructions queue behind the active turn. Interrupting
pauses the queue for review; Resume continues it. Ending the session stops
its resources while preserving files; deleting them is a separate operation.
For development, openrua serve still runs the service in the foreground.
Stopping that process ends its session, unlike closing one of its clients.
The default terminal uses Enter to send and Ctrl+J for a newline.
Use /tools, /files, and /queue to open details on demand. Configuration,
history, and confirmations use arrow keys, Enter, and Escape; no mouse is needed.--tui pi remains an explicit alias. See the keyboard guide.
How it works
Treat the robot as an interactive software project. Give an off-the-shelf
coding agent a task, terminal access to the robot's native ROS 2 interface,
and a workspace. The agent explores the machine, writes and tests programs,
and uses execution feedback to refine its actions.
your terminal the robot (real or simulated)
┌──────────────────────────────────┐ ┌────────────────────────────────┐
│ openrua │ │ ROS 2 graph │
│ └─ Claude Code / Codex │ DDS │ /joint_states · /tf · /camera │
│ in a sandbox workspace │◄────────►│ FollowJointTrajectory │
│ ros2 · rclpy · docs │ │ GripperCommand · MoveIt │
│ starter tools · saved files │ │ Velocity control │
└──────────────────────────────────┘ └────────────────────────────────┘
- Workspace as harness. Documentation and readable starter tools give the
agent a starting point. Files preserve observations, measurements, programs,
and notes as the task progresses. OpenRUA prescribes no robot-task workflow. - Perception as file I/O. The agent pulls sensor data on demand, inspects
saved images, and can write code to process images and compute measurements. - Manipulation as coding. The agent writes and executes commands or programs
against native ROS 2 interfaces, incorporating sensor feedback when needed.
It decides when to observe, how to act, and how to verify completion.
Task-specific perception and control strategies emerge from the agent's own
coding during execution. The paper analyzes
these behaviors, including image processing, metric measurement, and feedback
controllers, without a learned robot policy or a task-specific primitive library.
For the software modules, plugin boundaries, and shared-session implementation,
see Architecture.
Choosing what to run
- Robot and scene.
--simselects the simulator;--bench,--task-suite,
and--task-idselect a benchmark scene. Without a benchmark selection or saved
benchmark default, robosuite uses its nativeLiftscene. - Agent and model. Choose
--agentand optionally--modelwhen starting
the session. Switching them inside a running conversation is not implemented. - Custom profiles. Pass a file with
--robot,--sim, or--benchto use
your own environment in the TUI. It is validated and prepared without being
replaced by a guided choice. See your own robot. - Defaults. Save common choices with
openrua config set; explicit command
arguments override them. Configuration documents the fields. - Available extensions.
openrua robots,openrua simulators,openrua benchmarks, andopenrua agentslist the bundled and local choices.openrua benchmarks libero_prolists its suites and tasks.
Supported robots
A robot is what is true of it wherever it runs: joints, limits, frames,
gripper, ports, planner. Which simulator embodies it, and what surrounds
it, come from the other two kinds of file.
| Robot | Model | Embodied by |
|---|---|---|
panda |
Franka Emika Panda | robosuite, maniskill, calvin, vlabench |
panda-omron |
Panda on an Omron mobile base | robosuite through robocasa / robocasa365's assets |
widowx |
Trossen WidowX 250S | maniskill (the Bridge dataset's arm, through simpler) |
aloha-agilex |
AgileX Cobot Magic with two ARX X5 arms | robotwin |
| your robot | any ROS 2 arm or mobile manipulator, real | its own file; see docs/your-own-robot.md |
openrua robots prints this list from the files on disk, yours included.
Supported simulators
| Simulator | Engine | Robots | Native scene |
|---|---|---|---|
robosuite |
robosuite 1.5 on MuJoCo | panda |
Lift: a table and a cube |
maniskill |
ManiSkill 3 on SAPIEN 3 (PhysX, CPU) | panda, widowx |
PickCube-v1: a table, a cube and a goal marker |
robotwin |
RoboTwin 2.0's harness on SAPIEN 3 (PhysX, CPU) | aloha-agilex |
none: name a benchmark |
calvin |
calvin_env on PyBullet (TinyRenderer, CPU) | panda |
none: name the benchmark |
vlabench |
VLABench's dm_control environments on MuJoCo 3.2 | panda |
none: name the benchmark |
A simulator file knows the engine and how it drives each robot it
embodies; it knows no benchmark. Its install, and the benchmarks' own
over it, is what openrua build renders into one image per declaration
(openrua-sim-<name>: ROS 2, the checkouts, the assets, the Python
environment); docs/simulation.md describes them.
Supported benchmarks
| Benchmark | Robot | Simulator | Brings |
|---|---|---|---|
LIBERO (libero) |
panda |
robosuite |
the four standard suites and LIBERO-90, on LIBERO's robosuite 1.4 fork, ROS 2 Jazzy |
LIBERO-PRO (libero_pro) |
panda |
robosuite |
LIBERO's scenes under five perturbation axes, same fork as libero |
LIBERO-Plus (libero_plus) |
panda |
robosuite |
~10,000 perturbed variants of the four suites, its own fork and assets, ROS 2 Jazzy |
LIBERO-Mem (libero_mem) |
panda |
robosuite |
ten non-Markovian tasks with subgoal sequences, its own fork, ROS 2 Jazzy |
RoboCerebra (robocerebra) |
panda |
robosuite |
long-horizon tabletop cases on its LIBERO fork, the Ideal protocol, ROS 2 Jazzy |
CaP-Bench (capbench) |
panda |
robosuite |
CaP-X's tabletop scenes on robosuite 1.5, ROS 2 Humble |
RoboCasa (robocasa) |
panda-omron |
robosuite |
the original release's 24 atomic kitchen tasks (v0.2 on robosuite 1.5.0), ROS 2 Humble |
RoboCasa365 (robocasa365) |
panda-omron |
robosuite |
the 365-task release's kitchens and the Panda-Omron body, ROS 2 Humble |
ManiSkill (maniskill) |
panda |
maniskill |
the eleven table-top Panda tasks that ship with ManiSkill 3, seeded resets, ROS 2 Jazzy |
SimplerEnv (simpler) |
widowx |
maniskill |
the four WidowX Bridge tasks as their authors ported them to ManiSkill 3 (the SAPIEN 2 original needs a GPU; its Google Robot tasks are not ported), the visual-matching placement grid, ROS 2 Jazzy |
MIKASA-Robo (mikasa) |
panda |
maniskill |
the 90 language-conditioned memory tasks (remember, shell game, intercept, ...), its own image on ManiSkill 3.0.1, ROS 2 Jazzy |
RoboTwin 2.0 (robotwin) |
aloha-agilex |
robotwin |
the fifty dual-arm tasks under the Easy protocol (demo_clean); the Hard protocol needs its 11 GB textures and is not declared; ROS 2 Jazzy |
CALVIN (calvin) |
panda |
calvin |
the 1000 five-subtask chains of the long-horizon evaluation on play table D, each with its fixed initial condition and the benchmark's task oracle, ROS 2 Humble |
VLABench (vlabench) |
panda |
vlabench |
every task registered in the pinned checkout (5 GB of objects and scenes), seeded resets, the task's own termination as success, ROS 2 Humble |
A benchmark names its robot and simulator and brings its own world:install: (its image contents and ROS distro) and scenes: (scene
cameras, and robot embodiments its assets add). openrua benchmarks prints this
list; openrua bench --config <name> runs one.
Supported agents
| Agent | Status |
|---|---|
| Claude Code | supported (--agent claude-code) |
| Codex | supported (--agent codex) |
Experimental external plugins are available for
Kimi Code and
ZCode. Both have passed small real-provider
file tasks and two-turn shared-chat checks. Their API paths have real-provider checks; native Kimi OAuth and ZCode
Coding Plan key profile staging have offline checks only. ZCode OAuth reuse
remains unsupported, and robot tasks remain unvalidated. They are not default setup choices.
Bring your own agent. An agent is a manifest (how to install its CLI in the sandbox, which
hosts it talks to, how it logs in) and a small hooks class (how to
launch it); everything else is optional. Pass yours as a path
(--agent ./my-agent.yaml) or send a pull request;openrua agents lists what is available and what each can do. See
docs/agents.md.
Results
Results below are from our paper, evaluated
zero-shot in simulation without privileged simulator state or access to the
benchmark's success predicate during execution.
Main benchmark results use Claude Code with Claude Opus 5 at high reasoning effort:
| Benchmark | Evaluation scope | Success |
|---|---|---|
| CaP-Bench | Seven manipulation tasks | 99.0% |
| LIBERO-PRO | All four suites, both perturbation types | 87.0% |
The model comparison uses the LIBERO-10 suite of LIBERO-PRO, with both
perturbation types and 200 trials per configuration:
| Agent | Model (reasoning effort) | Success |
|---|---|---|
| Claude Code | Claude Sonnet 5 (high) | 24.0% |
| Claude Code | Claude Opus 5 (high) | 72.5% |
| Claude Code | Claude Fable 5.1 (high) | 75.0% |
| Codex | GPT-6 Astra (medium) | 62.5% |
| Codex | GPT-6 Astra (high) | 69.0% |
See the paper for baselines, evaluation protocols and detailed analysis.
Use your own robot
Draft a profile from the robot's live graph, finish the TODO lines,
and pass it to run:
openrua probe --host > my-ur5.yaml # joints, limits, frames, ports, cameras from the graph
openrua run ./my-ur5.yaml "..." # or: openrua config set --robot ./my-ur5.yaml
The profile's machine: section is what the agent's machine.yaml is
generated from: model, joint names and limits, frames,
gripper, and the ports (trajectory, gripper, twist, wrench) the
robot serves. Details in docs/your-own-robot.md.
Simulation and benchmarks
The simulated robots run the community benchmark scenes unchanged;
their original success predicates score the trial in place.
openrua bench --config libero_pro --run-id demo \
--task-suite libero_goal_task --task-ids 0,1 --seeds 0 --operator agent
Every trial writes result.json (verdict, preflight, termination,
token accounting), provenance.json (code and simulator commits, image
digests, config and prompt hashes), the agent's full transcript, and
the workspace it left behind; the run directory keeps a SUMMARY.md
regenerated from those files after every trial. Building the simulator checkouts:
docs/simulation.md.
A trial replays from its own commands.sh, and a replay with--record renders as a video, terminal on the left, cameras on the
right (openrua demo <trial>; see
running-experiments.md).
Architecture
The implementation separates robot resources, native agent adapters, shared
conversations, and presentation. runner/ composes resources; sessions/
coordinates messages and events; terminal/, ui/terminal/ and web/ provide client interfaces.
Robot, sandbox, proxy, and agent knowledge stay in their respective modules.
Benchmark execution and recorded demos have separate entry points.
Import-linter and architecture tests enforce these boundaries. See the
module map, agent extension guide,
and contribution guide.
Documentation
| page | read when |
|---|---|
| docs/install.md | setting a machine up: images, logins, doctor |
| examples/first-task.md | your first task on the simulated Panda in the original agent terminal |
| examples/shared-session.md | TUI, browser and CLI chat, queue checks, reconnection, saved images and shutdown |
| docs/terminal.md | TUI shortcuts, expandable panels, and current limitations |
| docs/sessions.md | shared-session operations, API, architecture and validation scope |
| docs/your-own-robot.md | describing your robot in one profile |
| examples/real-robot.md | the same flow on a real ROS 2 arm |
| docs/simulation.md | the simulator checkouts and GPU rendering |
| docs/podman.md | machines without Docker |
| docs/demos.md | recorded trials rendered as videos, one per task type |
| docs/running-experiments.md | openrua bench, the runs/ layout, every record field, replays and demo videos |
| docs/cli.md | every verb and flag, exit codes (generated) |
| docs/config.md | every config key (generated) |
| docs/agents.md | adding a coding agent |
| docs/architecture.md | the units and the layering contract |
| CONTRIBUTING.md | conventions for code, names and docs |
| CHANGELOG.md | release changes and upgrade notes |
Citation
If you use OpenRUA in your research, please cite our paper:
@misc{chu2026openrua,
title = {{OpenRUA}: Robot-Use Agents Are Zero-Shot Visuomotor Policies},
author = {Zhaoyang Chu and Earl T. Barr and Claire Le Goues and Peter O'Hearn and Mark Harman and Federica Sarro and He Ye},
year = {2026},
eprint = {2610.02459},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2610.02459}
}
License
Apache-2.0
Feedback
Try OpenRUA with your coding agent and robot, and tell us how it goes.
Bug reports, confusing steps, and ideas for improving the experience are welcome
in GitHub Issues.
If something fails, include your OpenRUA version, agent, environment, and the
command or steps that led to it. Please remove credentials and private data
from any logs you share.
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found