agents_control

agent
Security Audit
Warn
Health Warn
  • License — License: Apache-2.0
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Warn
  • network request — Outbound network request in lib/agents_control/channels/telegram/api.rb
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Control iTerm2 tabs including AI agents from Telegram

README.md

agents_control

Gem Version

A remote for iTerm2 and its AI agents (Claude Code, Codex) from Telegram:
tab list, commands, screen, new sessions. For Claude Code — also the
agent's questions and permission requests: it stops, buttons show up in
Telegram, you answer, the session unblocks.

Everyone runs their own bot — the token lives in the Keychain or
libsecret, never in files or in git. The daemon is invisible from
outside: the port only binds to 127.0.0.1, there are no incoming connections.

Status: terminal control and Claude Code work in full. Codex
sessions are visible and controllable through the terminal (like any
tab), but without the question relay: its hooks don't give us
anything to hook into for that yet.

Platforms and requirements

  • macOS — terminal backend: iTerm2 (via AppleScript) or tmux;
    secrets: Keychain; autostart: launchd.
  • Linux — terminal backend: tmux; secrets: libsecret; autostart: systemd.
  • Windows is not supported.
  • Ruby 3.1+. The only external dependency is thor (pure Ruby, no
    C extensions) — installing it compiles nothing.

On macOS, terminal control (sending commands, reading the screen,
creating tabs) works through iTerm2 or tmux — Terminal.app isn't
supported. Notifications about questions and permissions aren't tied to
a terminal at all: their source is Claude Code's hooks, which work
everywhere, including sessions with no terminal (a VS Code session, for
instance).

Why

You need to run a command in a session, check the screen, or switch to
a tab, and you're not at the computer — now you can do that from
Telegram. If the session is Claude Code, there's a bonus too: it
stops and waits for an answer — the question and its buttons arrive in
Telegram, no need to go home just to say "continue."

How it works

Hooks, not screen scraping. Claude Code itself calls agents_control
when it stops — the hook waits for an answer and passes the decision
back into the session. There's currently one adapter, for Claude Code;
a new agent needs a file with the same interface. Codex didn't fit:
its hooks only see shell commands and only understand deny — there's
no "agent stopped" event to hook into at all.

The terminal is iTerm2 or tmux, and each session picks its own backend.
Terminalless sessions (VS Code) are visible and answer hooks too — they
just have nothing to type into and nowhere to read a screen from.

Installation

git clone https://github.com/saparjohnick/agents_control
cd agents_control
bundle install

Or as a RubyGem:

gem install agents_control

Or straight from Claude Code, as a plugin — the repo doubles as a marketplace:

/plugin marketplace add saparjohnick/agents_control
/plugin install agents-control@agents-control

The plugin doesn't replace the install above — it's just a way to find
the tool and get install instructions without leaving Claude Code.

Telegram

Create a bot with @BotFather and run the wizard:

agents_control setup    # asks for the token, waits for your /start
agents_control           # after that, just open the console

The wizard catches your chat_id from your own message and adds it to
the allowed list — no need to type in a long number by hand.

At the end it offers to set a passphrase for destructive commands.
A few commands are held back when they arrive from Telegram — rm -rf ~,
mkfs, dd writing to a raw device, a fork bomb. With a passphrase set
you're asked to type it before one of those is sent; without one they
ask for a tap on a button instead. It's optional and skippable, and
agents_control passphrase set adds one later. See
Destructive commands ask first.

Bot commands:

Command What it does
/agents sessions with a live agent
/tabs all terminal tabs
/screen N show a tab's screen
/focus N switch to a tab
/run N command run a command in a tab and show the result
/new [directory] create a tab
/away intercept agent questions
/status current status

The number N comes from the last list shown.

Against a plain tab, /run treats the command itself as the reference
point: the shell echoes back whatever's typed, so it looks for that
exact text on screen and shows from there — sharper than diffing
screenshots, and it still works even if something else wrote to the
same tab in between, since it doesn't need the screen from right
before typing to relate to the screen after at all. A short reply
("y", "n" mid git add -p) isn't a safe anchor on its own — too
likely to match something unrelated — so those fall back to a
before/after diff instead, and when even that can't cleanly tell
what's new, to the current screen outright: seeing the result,
possibly with a little stale context around it, beats not seeing it
at all. How much gets captured either way is a /settings option
(terminal.run_result_lines, 200 by default) — a command whose output
runs longer just gets its last N lines, same as the default gets cut
by a screen that's too tall.

Against an agent, /run just confirms the send — an agent isn't a
shell command that finishes in a couple of seconds, so there's no
"result" to capture yet by the time it would look. Hooks already own
telling Telegram when it's actually done or needs something, the same
as replying to one of its own questions.

A tab stuck inside less, vim, a REPL, or anything else that reads
keystrokes as its own input rather than a line to submit gets a
confirmation first instead of a blind send — the text would go to
whatever's actually running there, not run as a command. A pager not
on that recognized list still gets caught: a bare : as the entire
last line is the one thing practically every pager agrees on for
"waiting on you," and a real shell prompt never looks like that.

The result stays a live target: replying to it — "y", "n", anything —
types straight into that same pane and shows what came back, so
something like git add -p's hunk-by-hunk prompts works as an actual
back-and-forth over Telegram, not a one-shot fire-and-forget.

Both /run and /screen head their reply with the tab's label and
tty (valkyrie · ttys017) — several tabs can share a label when
they're open on the same project, and the tty is what actually tells
them apart.

This list also populates Telegram's own / command menu automatically
— setup and every daemon start publish it via the Bot API, no manual
BotFather step needed. If the menu still shows only /start after
that, it's Telegram's client caching the old list, not a missing step
on your end: close and reopen the chat, or restart the Telegram app,
to force it to refresh.

Sending files in

Photos, videos, voice notes and documents — logs, a CSV, a spreadsheet,
a screenshot — can be sent to the bot and handed to a session. The file
is downloaded to ~/.local/state/agents_control/inbox/ and what reaches
the pane is its local path, so the agent opens it the way it opens
anything else:

you  › [screenshot.png]  caption: /run 3 why does this render wrong?

pane › why does this render wrong? /Users/you/.local/state/agents_control/inbox/20260907-171530-screenshot.png

Three ways to say where it goes:

caption /run N … that tab, with the rest of the caption as the message
reply with a file to a session's own message that session
neither saved, and the path comes back so you can use it

Guessing a target for a file is how it ends up pasted into the wrong
project, so with nothing to go on it isn't guessed. However it gets
there, a file's path is checked and escaped like anything else typed
into a pane — a caption is a way of saying /run, so it's guarded like
one.

A path and never a link, deliberately. Telegram serves file bytes
from https://api.telegram.org/file/bot<TOKEN>/… — the download URL
is the bot token. Pasting that into a pane would put the token in the
terminal, in the agent's context, and in the shell history, and a
leaked token is remote code execution on this machine. The URL is built
inside the API client and never leaves it.

The name comes from whoever sent the file, so it doesn't get to be a
path: no separators, no .., no leading dot, and cut to a length any
filesystem takes. Files are written 0600 inside a 0700 directory —
logs and screenshots are exactly what carries a token or a customer's
name through in passing — and swept after telegram.inbox_keep_days
(14 by default), since an agent reads them within minutes and the
directory would otherwise only grow. Telegram caps what a bot can
download at 20MB; anything larger is refused before the round trip.

Two modes

This whole section is about Claude Code specifically — its hooks are
what makes any of it possible. While you're at the keyboard,
intercepting its questions is counterproductive: you'll answer in the
terminal faster than you can reach for your phone, and a blocked hook
keeps the dialog from ever appearing on screen. So there are two modes:

  • present (default) — questions are mirrored to Telegram but stay
    in the terminal;
  • away (/away) — a question arrives with buttons and waits for a
    reply; Claude Code stands by until you answer or time runs out.

A question can also be answered by replying directly to the message —
it goes to the right session, even with several tabs open.

Silence is treated as a refusal. If nobody answered while you were
out, the action doesn't happen, and the session just keeps waiting in the terminal.

AskUserQuestion is the one exception to all of this: its answer never
travels back through the hook at all, so there's nothing to block on —
it always arrives with real buttons for each option, in both modes.
Tapping one types that choice straight into the terminal, the same
keystroke you'd type by hand. An open-ended option ("something else,"
"explain what you mean") has no button — just reply to the message
with your own words instead. With more than one question in a single
batch, or several tabs sharing the same directory so the target pane
is ambiguous, buttons are skipped in favor of a plain reply, since
guessing at the terminal's exact sequencing there risks typing into
the wrong place.

A "continue" reply is sent automatically, but tool permissions aren't.
These are two independent settings on purpose: merged into one, they'd
let it approve itself everything while nobody's watching. A question
that offers a choice ("rewrite it or leave it?") is never answered
automatically, even if it contains the word "continue."

Hooks

The daemon connects hooks on start and removes them on stop — otherwise
Claude Code prints a warning about an unreachable address in every
session. If the daemon crashed and the hooks are still there:

agents_control hooks          # check status
agents_control hooks uninstall

Entries in ~/.claude/settings.json are tagged, and other settings
aren't touched: installing and removing return the file to exactly its original shape.

Usage

agents_control          # opens the console and stays in the tab

The tool lives in a tab: while it's open, it listens to Telegram and
receives agent events. Commands inside start with a slash, same as the bot's:

> /sessions          sessions with a live agent
> /tabs              all terminal tabs
> /away              intercept agent questions (before stepping out)
> /settings          settings; /settings away — toggle
> /doctor            check that everything is in place
> /quit              quit

State icons: ⏳ working · ▸ at a shell prompt · 🖥 no terminal
(VS Code) · · everything else.

One-off commands exist too — agents_control sessions, doctor,
daemon — but the normal way to run it is an open console.

Rate-limit anchors

A five-hour window starts at the minute of the first message and
expires exactly three hundred minutes later. An anchor doesn't add a
single extra token — it moves window boundaries to where they're
convenient: the difference between "the window reset at 2:37pm,
mid-work" and "windows at exactly 7am, noon, and 5pm."

The ping uses a cheap model, and that's not economizing for its own
sake. The five-hour window is shared across the account, but weekly
limits are tracked per model family: an anchor on opus would spend the
scarcest bucket for an effect haiku gives for free.

Turned on in /settings. If you were working recently and a window is
already open, the ping is skipped — the daemon sees every agent event
and knows this without polling anything.

On macOS, a 7am anchor won't fire if the laptop is asleep: doctor
catches this and suggests pmset repeat wakeorpoweron.

Watchers

Hooks see Claude Code's own decisions, but not everything: the CLI's own
local menus (model switch, folder trust) and text on screen (a
rate-limit message) aren't covered by hooks at all — these events never
produce a single hook call. Two independent watchers handle them,
working over the screen rather than over Claude Code. Both go through
the full Registry — they see bare iTerm2 tabs and tmux panes alike.

CLI menus (terminal.watch_menus, on by default, polled every 20
seconds — terminal.menu_poll_interval). Notices the "❯ 1. … / 2. …"
pattern Claude Code (or a skill's own wizard) uses to draw any choice —
options can carry a couple of lines of description each, or a divider
before a trailing one — and sends it to Telegram as buttons; pressing
one types the option's number straight into the pane.

Limit reset (answers.auto_resume_after_limit, on by default,
polled once a minute — terminal.rate_limit_poll_interval). Notices a
message like "resets 3pm (UTC)" / "resets Oct 9, 10am" and types the
continuation itself once the time comes (with a minute of headroom).

Checking and autostart

agents_control doctor           # is everything in place
agents_control service install  # autostart (launchd / systemd)

doctor checks the environment the daemon will actually get, not
the current one: an interactive shell can show a different Ruby and a
different PATH than the process a service manager launches separately
— for instance, if a Ruby version manager puts a broken shim on PATH
ahead of the working interpreter.

That's why the service always starts via an absolute path to the
interpreter, never through PATH.

Why a tab's title can't be trusted

An agent sets a tab's title via an OSC sequence, and the title stays up
after it exits — a tab with an agent icon isn't necessarily still
running anything. That's why an agent's presence is confirmed by the
process tree, and the title is only ever used as a label.

Security

This tool lets you run commands on your machine from Telegram. That comes with some rules:

  • There's no shared service. Everyone sets up their own bot with
    @BotFather and connects it with their own token — the daemon talks to
    a bot only you own, not to any third-party infrastructure.
  • The daemon's port only listens on 127.0.0.1. Reaching it from
    outside the machine isn't just blocked by a password, it's physically
    impossible. The daemon talks to Telegram itself via long polling —
    outgoing connections only; no incoming port is opened for this either.
  • An allowed chat_id list is required. While it's empty, the bot
    answers nobody, even if someone learns its name. The bot's token
    isn't an access secret; the access secret is the chat_id filter.
  • A leaked bot token is equivalent to remote code execution.
  • The token is stored in the Keychain (macOS) or libsecret (Linux) —
    the same place as website and Wi-Fi passwords — and never lands in
    the config. There's deliberately no command-line argument for the
    token: it would leak into ps and shell history.
  • This protects against leaks through git, config files, ps, and
    shell history — not against malicious code already running as the
    same user: like any CLI tool without its own signed .app, the
    Keychain entry trusts the security utility itself, not
    agents_control specifically, so any process on the same account that
    knows the service name can read the token.
  • Auto-replying "continue" is on by default; automatic tool approval is
    off. These are separate settings on purpose.

Destructive commands ask first

A few commands are held back when they arrive from Telegram: a
recursive force delete aimed at / or $HOME, a fork bomb, mkfs,
dd writing to a raw device. Each comes back saying what it actually
does — "this recursively deletes ~ — there is no undo" — rather than
which rule caught it. rm -rf /tmp/build goes straight through, as it
should: the list is a seatbelt for a thumb on the wrong line in a tab
listing, not a policy engine, and it stays short enough to hold in
your head.

By default that's a confirmation button, and the first times one of
these turns up the bot offers a passphrase alongside it — that's the
one moment the reason for it is on screen. Don't ask again makes
the refusal stick; /settings keeps the offer reachable either way.
The bot also says so once, on the first daemon start with no
passphrase set, because someone who never types a destructive command
themselves would otherwise never learn the option exists — and the
first one that does turn up might not be theirs.

A passphrase can be set from the chat itself (two entries, both
deleted) or with agents_control passphrase set at the machine. Set
from the chat, it guards against whoever gets to that chat later,
not whoever has it right now — the bot says so when it saves. That's
the honest trade: the alternative is that most people never set one,
since you only think about this while away from the keyboard.

Set a passphrase — during agents_control setup, from the bot, or any
time after — and the button becomes a question instead:

agents_control passphrase set      # entered without echo, stored in the Keychain
you  › /run 3 rm -rf ~

bot  › this recursively deletes ~ — there is no undo

         rm -rf ~

       Reply to this message with the passphrase to send it to
       mobile-app. Anything else cancels.

you  › ················            ← deleted from the chat once read

A tap answers "did you mean this?" — which covers the wrong tab
number. A passphrase answers "are you you?", which is the other half:
an unlocked phone in someone else's hand taps just as well as its
owner does. It replaces the button rather than joining it — two steps
for one decision is how people learn to hurry through both — and the
message carrying it is deleted from the chat the moment it's read,
right or wrong, since a chat history syncs to every device the account
is signed in on.

The passphrase is stored derived (PBKDF2-HMAC-SHA256, salted) next to
the bot token, and is never accepted as a command-line argument.
agents_control passphrase clear goes back to the button.

Changes between versions

CHANGELOG.md — what changed and why, newest first. Read
the current entry before upgrading if you rely on /run: while the
major is 0, a minor bump may change behaviour, and the entry says so
where it does.

Development

rake test

Tests run on minitest, with external commands stubbed via
FakeExecutor — neither iTerm2 nor tmux is needed to run them.
Fixtures live in test/fixtures.rb — recorded output from real commands.

License

Apache 2.0.

Reviews (0)

No results found