talktome
Health Uyari
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 5 GitHub stars
Code Basarisiz
- spawnSync — Synchronous process spawning in desktop/main.cjs
- process.env — Environment variable access in desktop/main.cjs
- fs module — File system access in desktop/main.cjs
- network request — Outbound network request in desktop/main.cjs
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
Voice calls with the coding agent you already use. A macOS menu-bar app.
talktome
TalkToMe is a macOS menu-bar app for voice calls with a coding agent.
The agent rings you from the session that it already runs. You answer, and you talk.
The agent keeps its model, tools, files, and history. TalkToMe does not supply a language model.
The app records the microphone, turns speech into text, sends the text to the agent, and speaks the reply.
How a call works
- You ask the agent in its session to call you, for example "Call me with TalkToMe."
- The agent runs
talktome callwith its session ID. - TalkToMe shows a ring on the screen. You select Answer.
- The agent speaks a short greeting.
- You speak. The agent receives your words as a message in its session and replies.
- You select End, or the agent runs
talktome end.
The app has no window that starts a conversation.
After setup, TalkToMe stays in the menu bar and waits for a call.
The menu-bar icon opens Settings and gives controls to answer, decline, mute, and end a call.
Requirements
- A Mac with Apple silicon (arm64) and macOS 14 Sonoma or later.
- One agent host: Codex, Claude Code, Hermes Agent, OpenClaw, or another host that can run shell commands.
Install
Disk image
- Download
TalkToMe-<version>-arm64.dmgfrom Releases. - Open the disk image and drag TalkToMe to Applications.
- Open TalkToMe. macOS tells you that it cannot verify the app. Close the message.
- Open System Settings > Privacy & Security. Below the message about TalkToMe, select Open Anyway.
- Open TalkToMe again, then select Open Anyway. macOS can ask for your password.
You do steps 3 to 5 only once.
On macOS 15 Sequoia and later, Control-click and Open does not skip this check.
To skip the check from Terminal, run xattr -dr com.apple.quarantine /Applications/TalkToMe.app.
macOS asks for these steps because the app is not notarized.
Notarization needs a paid Apple Developer ID.
The app has an ad-hoc signature. The signature lets macOS find a change to the app files.
Homebrew (coming soon)
This command works after the tap rohanprichard/tap exists:
brew install --cask rohanprichard/tap/talktome
The first start needs the same Open Anyway step as the disk image.
Build from source
You need Node.js 22 or later, uv, and the Xcode command-line tools.
uv installs Python 3.11 to 3.13 when necessary.
git clone https://github.com/rohanprichard/talktome.git
cd talktome
npm ci
uv sync --frozen
npm start
npm start runs uv sync --frozen if the Python environment is missing. Then it starts Electron.
To build the disk image, see Build the disk image.
Server install for remote calls
An agent on another server can ring the app through the remote bridge.
On that server, install only the command. It has no speech libraries:
uv tool install "talktome-local @ git+https://github.com/rohanprichard/talktome"
First-run setup
The first start opens a setup window with six steps:
- Welcome. The window explains the call flow.
- Microphone. Select Allow microphone. macOS asks for permission.
- ElevenLabs. Enter an ElevenLabs API key, or select Later to use local speech.
- Agent connection. Select Install. This installs the agent skill and the
talktomecommand. - Glow color. Select the color that the call surface shows during a live call.
- All set. Ask your agent to call.
You can skip a step with Later. Settings contains the same options.
To run setup again, quit the app and run npm run reset.
This command clears the saved speech settings and the setup progress. It keeps the downloaded models and the local token.
Connect an agent
Install writes the skill file to ~/.codex/skills/talktome/SKILL.md.
It also writes the skill to ~/.hermes/skills/ and ~/.openclaw/skills/ if those directories exist.
It installs the talktome command in ~/.local/bin, /opt/homebrew/bin, or /usr/local/bin.
The skill tells the agent which commands to run. See the skill.
For Claude Code, or for another host, copy the skill into the host's skill directory:
mkdir -p ~/.claude/skills/talktome
talktome skill > ~/.claude/skills/talktome/SKILL.md
| Host | Command | How replies reach the call |
|---|---|---|
| Codex | talktome call --agent codex --thread "$CODEX_THREAD_ID" |
TalkToMe reads the session's public replies automatically. |
| Claude Code | talktome call --agent claude --thread SESSION_ID |
The agent runs talktome listen and talktome reply. |
| Hermes Agent chat | talktome call --agent hermes --connection cooperative --thread ID |
The agent runs talktome listen and talktome reply. |
| OpenClaw chat | talktome call --agent openclaw --connection cooperative --thread ID |
The agent runs talktome listen and talktome reply. |
| Other hosts | talktome call --agent generic --thread ID |
The agent runs talktome listen and talktome reply. |
The listen and reply commands are the "cooperative" connection.
They work with any host that can run shell commands on the Mac that runs TalkToMe.
The commands exchange private files with the app. Thus, they work when a sandbox blocks local network access.
Hermes and OpenClaw also have experimental adapters for an API session or a Gateway session.
Run talktome providers to see the connection methods that are ready.
Agent support gives the setup and the limits.
Agent protocol gives the commands, the ring flow, and the local HTTP interface.
Remote bridge (experimental)
The remote bridge lets an agent on another server ring the laptop.
Only text and call events cross the bridge. Microphone audio stays on the laptop.
The setup uses SSH:
- Install talktome on the server:
uv tool install "talktome-local @ git+https://github.com/rohanprichard/talktome". - On the laptop, run
talktome remote-connect user@server --install-service. - Restart TalkToMe.
The app opens an SSH tunnel to the server itself, so SSH must log in with a key and no password prompt.
This feature is experimental. See remote bridge.
During a call
After the agent finishes its reply, speak to start the next turn.
TalkToMe uses Smart Turn, a small local model, to decide when you finished speaking.
Settings can select a fixed pause instead. See Smart Turn.
Call latency and call timing describe the delays in a call.
Settings also sets the position of the call surface: Bottom or Top center.
Speech providers
| Function | Local option | ElevenLabs option |
|---|---|---|
| Speech recognition | Whisper Small (484 MB) or Whisper Base English (145 MB) | Scribe v2 |
| Agent voice | System voice or Kokoro | Flash v2.5 |
Whisper is the default for recognition. The system voice is the default voice.
Download a Whisper model in Settings before you use local recognition.
Kokoro downloads a 114 MB model and a 28 MB voice file. The app examines their SHA-256 hashes before use.
Local models run offline after the download.
The ElevenLabs key needs access to the voice list and to each selected speech service.
Select Remember key to keep the key in the macOS keychain. If you do not, the key stays in server memory until the app closes.
The app never returns the key to the interface or writes it to its settings file.
Remove key removes the key from the app session and from the keychain.
See speech providers for the exact interfaces.
Data and network access
The server listens only on 127.0.0.1:8765. A generated local token protects its interface.
The desktop windows use an HTTP-only session cookie. Other web origins cannot use the interface.
The app keeps its token, settings, and models in ~/Library/Application Support/talktome.
The app keeps up to 200 transcript messages and 512 events in memory. Closing the app clears them.
The app connects to the network for these purposes only:
- Hugging Face, to download Whisper and the Smart Turn model.
- GitHub, to download the Kokoro model and voice file.
- ElevenLabs, only if you select an ElevenLabs service. ElevenLabs recognition sends microphone audio. ElevenLabs voice sends reply text. Service charges and the provider's retention rules apply.
- A Hermes or OpenClaw host, or a relay, only if you configure one.
The agent receives the text of what you say. The agent's provider and tools have their own data rules.
The cooperative command files contain conversation text.
Settings for Hermes and OpenClaw are in agent-hosts.json in the data directory. This file holds a plaintext token with mode 0600.
Environment settings:
| Name | Purpose |
|---|---|
TALKTOME_DATA_DIR |
Change the local data directory |
TALKTOME_PORT |
Change the desktop server port |
TALKTOME_URL |
Set the server address for external clients |
TALKTOME_TOKEN |
Supply an existing shared token |
TALKTOME_RELOAD |
Restart the server when Python files change. Development only. |
TALKTOME_FLOATING_CALL |
Set to 0 to turn off the call window. The call then has no controls on screen. |
TALKTOME_ALLOW_REMOTE_AGENTS |
Set to 1 to permit a Hermes host that is not on this Mac. The URL must use HTTPS. |
Build the disk image
npm run build:app
This command freezes the Python server into one binary, draws the icon, compiles the notch helper, and runs electron-builder.
The result is dist/app/TalkToMe-<version>-arm64.dmg and a zip of the app for the updater.
The script mounts the disk image after the build and examines its contents and its signature.
To reuse the last frozen server when only the desktop code changed, run npm run build:dmg.
To install the app, open the disk image and drag TalkToMe to Applications.
The build has an ad-hoc signature. It opens on the Mac that built it without a prompt.
On another Mac, it needs the Open Anyway step from Disk image.
The bundle does not include the speech models. The app downloads them at first use.
Development
uv sync --frozen # install the Python environment
npm ci # install the Node packages
npm start # start the app from source
uv run pytest -q # Python tests
uv run ruff check # Python lint
node --test tests/ # JavaScript tests
The dev group includes the speech extra, so uv sync --frozen installs the full voice server.npm test runs the Python tests and the JavaScript tests together.
These commands start the real app for end-to-end checks:
| Command | What it examines |
|---|---|
npm run test:call |
The call window: position, stacking, and growth of the transcript |
npm run test:attach-call |
A full call against a real Codex session. It sends a few short model requests. |
npm run test:stream |
Time to first audio for ElevenLabs. It spends credits on two short replies. |
The call test turns off the Chromium sandbox.
A normal start keeps the sandbox on.
The development notes hold plans, research, and a work log. They can be out of date.
Security
To report a security problem, read SECURITY.md.
Contributing
Read CONTRIBUTING.md before you open a pull request.
License
MIT. See LICENSE.
TalkToMe uses Electron, FastAPI, faster-whisper, kokoro-onnx, Pipecat Smart Turn, and optional ElevenLabs services.
It does not contain copied SpeakType or AgentCall code. NOTICE lists the third-party references.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi