jarvis
Health Warn
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 6 GitHub stars
Code Fail
- exec() — Shell command execution in bridge/chrome.mjs
- exec() — Shell command execution in bridge/net.mjs
- network request — Outbound network request in bridge/net.mjs
- exec() — Shell command execution in bridge/page.mjs
- process.env — Environment variable access in bridge/server.mjs
- network request — Outbound network request in bridge/server.mjs
- spawnSync — Synchronous process spawning in scripts/setup.mjs
- process.env — Environment variable access in scripts/setup.mjs
- process.env — Environment variable access in scripts/start.mjs
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
J.A.R.V.I.S. voice assistant with an Iron Man holographic UI, driven by Claude Code. Built on adewaskar/jarvis, with Windows support and a loopback-only bridge.
J.A.R.V.I.S.
A browser voice assistant with an Iron Man holographic interface. Say
"Hey Jarvis", he wakes, listens, and does real things through your tools —
searches the web, generates images, drives your phone, reads your mail. The face
is a web page (React + Vite + Three.js + custom GLSL). The brain is Claude Code,
run headless as a library.
The only subscription you need is Claude Code. No API keys, no OpenAI
account, no cloud bill — the brain runs on your existing Claude Code login, and
the heavy work (the model itself) runs on Anthropic's servers, so even a low-end
laptop only has to draw the interface. ElevenLabs is an optional add-on that
gives JARVIS a much better voice and sharper hearing; without it he speaks and
listens through the browser's own speech, and everything still works.
What this fork adds
- Bridge bound to loopback. The bridge used to listen on every network
interface, and itsOrigincheck is a browser convention, not authentication
— any client on the same network could set the header and drive an agent that
may have shell access. It now binds to127.0.0.1only (bridge/server.mjs). - Windows support. The
chrome_*extension bridge does not run on Windows,
so onwin32the system prompt points JARVIS at the Chrome DevTools MCP server
instead and teaches it the Git Bash quirks of Windows commands (taskkill //IM,winget). macOS keeps the original behaviour. The prompt also requires a
spoken confirmation before shutdown, restart, sign-out or deleting files. - "Jarvis, logout". A voice dismissal that acknowledges and returns to
standby. The pattern must match the whole utterance, so "log out of my
account" is still treated as a task rather than a dismissal (src/App.tsx). - Unattended mode.
?auto=1powers up on load, with no INITIALISE click or
clap, so JARVIS can run as a background assistant. - Windows launchers. Start JARVIS in a Chrome window parked off-screen (kept
un-throttled so the wake word still works), bring it back into view, and stop
exactly the processes it started. See Running on Windows.
Requirements
In one line: a Claude Code subscription, plus two free things every computer
can have — Node.js and Chrome. That's the whole list.
- Claude Code, installed and logged in — this is the only account you need.
Install it with the official method —npm install -g @anthropic-ai/claude-code,
or the platform installer at https://docs.claude.com/en/docs/claude-code —
then runclaudeonce and complete login. The bridge reuses that login. No
API key, and usage is billed to your existing Claude account. - Node.js 20 or newer — free, one installer from https://nodejs.org. This
is a Node web app, so it is the one unavoidable tool. - Google Chrome or Microsoft Edge, in a real browser window — not an
embedded preview pane. Preview panes (including the one inside editors and
Claude Code) block microphone access, so the page loads and looks right but
never hears you. JARVIS also needs WebGL, which these browsers provide. - Optional: an ElevenLabs API key — a good add-on, not a requirement. It
gives a better voice and sharper transcription; the free tier is plenty for a
demo. Without it, everything runs on the browser's own speech.
Run npm run setup after cloning and it checks all of this for you, in plain
language.
Quick start
First, install, then start it:
npm install
npm start # runs the brain and the face together
Then open the URL it prints (http://localhost:5173) in Chrome, click INITIALISE, and say “Hey Jarvis”.
Prefer two terminals? Run them separately instead:
npm install
Terminal 1 — the brain:
npm run bridge
Terminal 2 — the face:
npm run dev
Then open the app in a real Chrome or Edge window:
open http://localhost:5173
Click INITIALISE, allow the microphone when asked, and say "Hey Jarvis".
It has to be a real browser window. Embedded preview panes block the
microphone, so JARVIS will look perfectly alive and simply never respond.
Running on Windows
Double-click launchers live in the repo root:
| File | Does |
|---|---|
start-jarvis.cmd |
Visible mode: starts the bridge and the interface, then opens Chrome |
start-jarvis.vbs |
Hidden mode: runs jarvis.ps1, which starts everything and parks the interface in a Chrome window off the edge of the desktop — you just talk |
show-jarvis.vbs |
Brings the hidden window back on screen |
stop-jarvis.vbs |
Stops only the processes JARVIS started |
Both start modes run with writes enabled — read
Enabling actions first.
The hidden launcher uses a Chrome profile of its own (%LOCALAPPDATA%\JarvisApp)
with the microphone pre-allowed for localhost:5173 and the camera denied,
because nobody can answer a permission prompt in a window they cannot see.
For the ElevenLabs voice, copy jarvis-secrets.example.cmd tojarvis-secrets.cmd and paste your key in. That file is git-ignored.
For browser control, add the
Chrome DevTools MCP server
to your user-level Claude Code config, where the bridge looks for it:
claude mcp add --scope user chrome-devtools -- npx chrome-devtools-mcp@latest
How it works
JARVIS is two processes. The browser is the face and the voice; the bridge is
the brain and the hands.
┌─ browser (the face) ───────────────┐ ┌─ bridge (the brain) ─────────────┐
│ "Hey Jarvis" wake word │ │ Node · bridge/server.mjs │
│ local VAD → speech to text │ ws │ Claude Agent SDK │
│ reactor UI (Three.js + GLSL) │◄─────► │ = Claude Code, headless │
│ text to speech │ 8787 │ spawns your MCP servers │
│ heads-up display │ │ permission gate (decideTool) │
└────────────────────────────────────┘ └──────────────────────────────────┘
Everything you see and hear happens in the browser. The bridge is a single Node
process (bridge/server.mjs) that runs the Claude Agent SDK
(@anthropic-ai/claude-agent-sdk) — this spawns the real claude CLI as a child
process, so the brain literally is Claude Code, headless. They talk over a
WebSocket (plus a few HTTP endpoints) on ws://localhost:8787.
Why a bridge at all? A browser tab cannot spawn the local stdio MCP servers —higgsfield, elevenlabs, android, playwright, exa, serper, and the
rest. The bridge can. And because it is the Agent SDK, it authenticates off your
existing Claude Code login: no API key, billed to that same Claude account.
The model. claude-opus-5 at effort medium by default. Override with theJARVIS_MODEL and JARVIS_EFFORT environment variables. On startup the bridge
prints its choice, e.g. [jarvis] model claude-opus-5 · effort medium.
The voice pipeline
The loop is designed so that nothing silently dies and barge-in feels natural.
- Detection is local. An energy-based voice-activity detector
(src/lib/vad.ts) decides when you are speaking. It is instant, cannot quietly
fail, and is what makes barge-in work — speak while JARVIS is talking and he
stops. - Transcription has two tiers, chosen automatically at boot. The browser asks
the bridge/healthand picks the best available:- ElevenLabs key present → ElevenLabs Scribe, via the bridge
/sttendpoint. - Nothing configured → the browser's own
SpeechRecognition(Chrome/Edge),
guarded by a heartbeat so it recovers when Chrome throttles it.
- ElevenLabs key present → ElevenLabs Scribe, via the bridge
- Speaking uses the ElevenLabs voice when a key is present, and the
browser'sspeechSynthesisotherwise. If a cloud call fails it falls back to
the browser voice, and if the OS voice itself is broken it latches over to the
cloud voice.
So it works with no keys and auto-upgrades when a key appears — there is no flag
to set. Capability detection lives in src/lib/capabilities.ts, which probes the
bridge's GET /health (returning { ok, tts, stt }, both tracking the
ElevenLabs key) once at boot and picks the engines.
What JARVIS can do
Beyond answering, JARVIS reaches every MCP server in your Claude Code
configuration, and can drive his own interface.
Your tools
Every server in your ~/.claude.json is handed to the SDK explicitly. Depending
on what you have installed, that is roughly:
- Web & search —
exa,serper,serpapi - Images & video —
higgsfield,openrouter-image,palmier-pro - Voice —
elevenlabs - Your phone —
android - The browser —
playwright
A few things you can say:
- "What's happening in AI this week?"
- "Generate an image of the Mark VII suit."
- "Take a screenshot of my phone."
- "Open my GitHub notifications."
Note on account connectors. Servers you added through your claude.ai
account are not stored on disk, so the bridge cannot see them — it works from
the servers in~/.claude.json(about 14), not the claude.ai ones.
JARVIS controls the interface
He drives the UI through MCP tools the bridge exposes:
ui_theme— accent, background, per-phase coloursui_reactor— colour, scale, intensity, spin, and style (ring|sphere|wire), visibilityui_orbit— put images in orbit around the reactorui_chrome— show or hide rails, transcript, badgesui_effect—glitch|pulse|scan|shake|flashui_screen— clearui_reset— back to defaults
So "make it red, hide the systems list, put that render in orbit" is a spoken
command.
The heads-up display
JARVIS authors panels with a display tool against a fixed .hud-* design
system. The browser sanitises the markup (DOMPurify, a class allowlist and a
strict CSP) before rendering. Rich media works — images, <video>, and
YouTube/Vimeo embeds. Remote images and video are fetched server-side through
the bridge (/img and /media, both SSRF-guarded), so hotlink-blocked news
thumbnails still appear and the page never beacons your IP to a host the model
chose.
Controls
| Key / phrase | Does |
|---|---|
| "Hey Jarvis" | Wake him |
| Space | Talk without the wake word |
| Just speak | Interrupt him mid-sentence (barge-in) |
| V | Cycle the browser voice |
| "Jarvis, logout" | Acknowledge and stand down (also "stand down", "goodbye", "that'll be all") |
| Escape | Stand down |
| D | Live diagnostics panel |
| T | One-line audio self-test |
The boot sequence
Power-up plays a four-beat Iron Man start-up (src/ui/Boot.tsx): an
"INITIATING SYSTEM" status bar with a segmented progress bar and boot log; then
concentric reticle rings resolving into "J.A.R.V.I.S"; then a suit schematic;
then the triangular arc reactor lighting up — with a start-up sound under it
(public/audio/boot-music.mp3).
Configuration
Everything is optional in bridge mode. Frontend settings live in .env.local
(copy .env.example); bridge settings are environment variables.
Bridge
| Variable | Default | Effect |
|---|---|---|
JARVIS_BRIDGE_PORT |
8787 |
Port for the WebSocket + HTTP endpoints |
JARVIS_MODEL |
claude-opus-5 |
Model to run |
JARVIS_EFFORT |
medium |
Reasoning effort |
JARVIS_ALLOW_WRITES |
off | 1 allows effectful tools (see below) |
JARVIS_ALLOWED_ORIGINS |
local dev | Extra WebSocket origins to accept |
JARVIS_ALLOW_NO_ORIGIN |
off | Accept connections with no Origin header |
JARVIS_FILE_ROOTS |
— | Roots the /file endpoint may serve from |
JARVIS_VOICE_ID |
— | ElevenLabs voice id |
ELEVENLABS_API_KEY |
— | Optional; enables the ElevenLabs voice + Scribe |
Frontend (.env.local)
| Variable | Effect |
|---|---|
VITE_BACKEND |
bridge (default) or direct |
VITE_BRIDGE_URL |
Where to reach the bridge |
VITE_TTS_ENGINE |
system or kokoro |
VITE_KOKORO_VOICE |
Voice for the Kokoro engine |
VITE_USE_ELEVENLABS |
Force the ElevenLabs voice on |
VITE_ANTHROPIC_API_KEY |
Direct mode only |
Adding an ElevenLabs key
You do not have to touch a flag. Either:
- Set
ELEVENLABS_API_KEYon the bridge before starting it, or - Add the key to your
elevenlabsMCP server's env in~/.claude.json— the
bridge reads it from there too.
Either way, /health starts reporting the capability, the browser picks it up on
the next boot, and both the voice and transcription upgrade automatically.
Enabling actions
The tool gate starts read-only. Search, generation and lookups run freely;
anything effectful — send, tap, delete, install, pay — is denied. Voice is a poor
interface for a confirmation dialog, so the decision is made ahead of time indecideTool() in bridge/server.mjs, not at the moment of use. The bridge setssettingSources: [], which makes its own gate the only authority — filesystem
settings and any global bypassPermissions cannot override it.
To allow effectful tools (phone, browser driving, sending), run the bridge this
way instead:
npm run bridge:writes
Read
decideTool()before you do. "Hey Jarvis, clean up my downloads folder"
means something rather different with writes enabled.
Troubleshooting
I can't hear him, or he can't hear me. Press D for the diagnostics panel
— it states plainly whether he is hearing you and whether he is producing sound.
Press T for a one-line audio self-test.
No voice at all. You must be in Chrome or Edge, in a real browser
window (not an embedded preview), and you must have allowed the microphone.
Bridge not reachable. Check that npm run bridge is still running in its
terminal, and that nothing else is holding port 8787.
Security
All of this lives in bridge/server.mjs:
- The bridge listens on
127.0.0.1only, so nothing else on the network can
reach it. The origin check below is a browser convention, not authentication. - The WebSocket accepts only local dev origins (add more with
JARVIS_ALLOWED_ORIGINS). /file,/imgand/mediavalidate the scheme, confine to allowed roots,
resolve the real path, and refuse private and loopback addresses (SSRF guard).- The tool gate (
decideTool) is default-deny for effectful MCP tools. - A strict CSP in
index.html; model-authored panel HTML is sanitised.
Credits & licence
MIT. The original project is by
Aditya Dewaskar —
adewaskar/jarvis. The additions in this
fork are released under the same licence.
Music by Kevin MacLeod (incompetech.com), licensed under Creative Commons
Attribution 4.0 — see public/audio/CREDITS.md.
The boot sound and any tracks in public/audio/ ship with the project for the
demo. If you go on to monetise something built on this, clearing the rights to
that audio is your responsibility.
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found