nekoro-browser

mcp
Guvenlik Denetimi
Uyari
Health Uyari
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Gecti
  • Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Gecti
  • Permissions — No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

Lightweight browser automation CLI + MCP server driving your everyday Chrome via the extension chrome.debugger API — keeps your login state, no --remote-debugging-port

README.md

nekoro-browser — browser automation CLI + MCP server

tests PyPI Python versions MIT License MCP supported

Let your AI coding tool drive your own Chrome — open pages, click, read them back, still logged in everywhere.

中文

Quick Start · Examples · Use Cases · MCP · API · Architecture · Site Knowledge · Limitations · Reference


nekoro-browser drives the Chrome you already use — logins, cookies, sessions, all intact.

An MV3 extension connects your existing Chrome profile to a local Python daemon.
Use CLI snippets or built-in MCP tools from your AI client. Python 3.12+, standard
library only, no bundled browser engine. MIT; extension source included.

Quick Start

1 — Install (Python 3.12+, zero third-party dependencies)

uv tool install nekoro-browser

No uv? pipx install nekoro-browser works too.

From source: git clone https://github.com/zeshuochen/nekoro-browser && cd nekoro-browser && uv pip install -e .

2 — Load the extension

nekoro-browser setup

Copies the extension directory to your clipboard and waits until it connects. Meanwhile:
chrome://extensions/ → Developer mode → Load unpacked → paste.

3 — Start in the background and check readiness

Open a regular webpage in Chrome, then run:

nekoro-browser --ensure
nekoro-browser --doctor

--ensure starts the daemon in the background and checks a real page response.
You can close this terminal. --doctor diagnoses; --stop stops the daemon.

4 — Drive the browser. Pick the way you actually work:

From your AI coding tool (MCP) — one command for Claude Code, other clients in
MCP:

claude mcp add nekoro-browser -- nekoro-browser-mcp

Restart the client and ask it to open a page.

From the terminal — run a snippet:

nekoro-browser -c "await page_info()"
# → {"ok": true, "result": {"title": "...", "url": "..."}}

Upgrading

uv tool upgrade nekoro-browser
nekoro-browser --ensure
nekoro-browser --doctor

For pipx, use pipx upgrade nekoro-browser. --ensure restarts an outdated daemon,
preserves its port and domain allowlist, and reloads the extension. Doctor must show
matching CLI, daemon and extension versions. If Chrome disabled the extension or its
path changed, re-enable/load the directory printed by nekoro-browser --extension-path
in chrome://extensions, then run --ensure again.


Examples

Send a multi-step flow in one shot. Every helper is already await-able at top level — no
asyncio boilerplate, no imports:

nekoro-browser <<'PY'
await new_tab("https://example.com")
print((await page_info())["title"])            # Example Domain
print((await get_markdown(max_chars=200))["result"])
print((await state(max_items=3))["result"])    # indexed interactive elements, model-ready
await close_tab()
PY
On Windows? <<'PY' is bash-only — PowerShell equivalent
@'
await new_tab("https://example.com")
print((await page_info())["title"])
'@ | nekoro-browser

The closing '@ must sit at the start of its own line. One-liners:
nekoro-browser -c "await navigate('https://example.com')".

Not cmd.exe — its echo keeps the quotes, so the snippet arrives as a string and comes back
{"ok": true, "result": "page_info()"} with the browser untouched.

state() numbers the elements and click_index(n) clicks by number — the model never has to guess a CSS selector:

nekoro-browser <<'PY'
await navigate("https://github.com/search?q=browser+automation&type=repositories")
await wait_for_load()
print((await state(max_items=40))["result"])   # every interactive element carries an index
PY

Then click the one you saw — indices shift with page content, so don't copy a fixed number:

nekoro-browser -c "await click_index(7)"

All helpers are documented in SKILL.md.


Where It Fits

Use it for personal browser workflows that need existing logins, a small Python
runtime, CLI/MCP access, and editable site scripts. Installation requires loading
an unpacked extension; screenshots can change the visible tab.

It does not provide isolated parallel browser contexts or multi-browser testing.
The full Chrome loop is validated on Windows; long-running stability remains unverified.


MCP (any MCP client)

MCP is how Claude Code, Cursor and friends call outside tools. Hook it up once and the
model gets navigate, click_index, get_markdown… as first-class tools.

Prerequisite: run nekoro-browser --ensure first. The daemon owns the Chrome
connection; the MCP server forwards calls to it.

The command to register is always nekoro-browser-mcp. Only the config shape differs:

Claude Code

claude mcp add nekoro-browser -- nekoro-browser-mcp

Claude Desktop (Settings → Developer → Edit Config) · Cursor (~/.cursor/mcp.json,
or .cursor/mcp.json for one project) · Cline (MCP Servers → Configure MCP Servers)

{ "mcpServers": { "nekoro-browser": { "command": "nekoro-browser-mcp" } } }

Claude Desktop config file: macOS ~/Library/Application Support/Claude/claude_desktop_config.json · Windows %APPDATA%\Claude\claude_desktop_config.json

opencode (opencode.json) — note command is an array, and the key is mcp

{ "mcp": { "nekoro-browser": { "type": "local", "command": ["nekoro-browser-mcp"], "enabled": true } } }

Codex (~/.codex/config.toml, or codex mcp add nekoro-browser -- nekoro-browser-mcp)

[mcp_servers.nekoro-browser]
command = "nekoro-browser-mcp"

VS Code / Copilot (.vscode/mcp.json, or MCP: Open User Configuration) — the key is
servers, not mcpServers

{ "servers": { "nekoro-browser": { "command": "nekoro-browser-mcp" } } }

Prefer not to install anything up front? Replace the command with uvx, which fetches and
runs on demand the way npx -y does — e.g. "command": "uvx", "args": ["--from", "nekoro-browser", "nekoro-browser-mcp"]. That only removes the install step for the MCP
server; the daemon still has to be installed and running.

Restart the client afterwards. If the tools don't show up, run nekoro-browser --doctor
first — a dead daemon looks exactly like a broken MCP config — then check the client's MCP
log (Claude Desktop keeps them in ~/Library/Logs/Claude on macOS, %APPDATA%\Claude\logs
on Windows).

Beyond the tool list:

  • cdp — raw CDP command, and exec_python — arbitrary Python in the daemon namespace, so
    a whole multi-step flow costs one round trip.
  • Screenshots return as image content; clients render them inline.
  • A helper failure ({"ok": false}) surfaces as isError, never dressed up as success.
  • Navigating to a site you have notes or scripts for ships them in the tool result — see
    Site Knowledge.

API

Category Commands
Navigation navigate(url), new_tab(url), ensure_tab(url), new_tab(url, reuse=True), list_tabs(), switch_tab(id), close_tab(id), close_tabs(ids), sweep_tabs()
Page info page_info(), page_html(), page_text(), get_markdown(), state(), refs(), find_text(t), iframe_target(url_substr)
JavaScript js(code), cdp(method, **p), cdp_batch(*cmds)
Interaction click(loc), click(loc, tab=id), click_selector(sel), click_ref(ref), click_index(n), click_at_xy(x,y), type_text(t), fill_input(sel,t), press_key(k), upload_file(sel,path)
Dialogs dialog_off(), get_last_dialog()
Waiting wait_for_load(), wait_selector(sel), wait_for_network_idle(), sleep(s)
Downloads wait_for_download()
Screenshots capture_screenshot(), capture_screenshot(scale="device"), capture_screenshot("jpeg", 90)

Helpers exposing tab= can target an already attached tab (default: the active tab).
capture_screenshot defaults to scale="css" — pixel size equals the CSS
viewport, so coordinates can be fed straight to click_at_xy; scale="device"
keeps physical pixels.

Screenshots bring the requested tab to the foreground before capture, so the visible
Chrome tab may change. A window that still has a zero-size viewport reports not_rendered.


Architecture

flowchart TD
    A["Chrome tab — your profile, your logins"]
    B["Extension background.js<br/>chrome.debugger / CDP"]
    C["Python daemon<br/>127.0.0.1:28417"]
    D["CLI<br/>nekoro-browser"]
    E["MCP server<br/>nekoro-browser-mcp"]

    A <-->|CDP| B
    B <-->|persistent WebSocket| C
    D -->|"HTTP /exec · token auth"| C
    E -->|"HTTP /exec · token auth"| C
Same diagram as plain text (for renderers without Mermaid, e.g. PyPI)
Chrome extension (background.js) —— chrome.debugger / CDP
        ↕ persistent WebSocket
Python daemon (127.0.0.1:28417)
        ↕ HTTP /exec (token auth)
CLI (nekoro-browser)  ·  MCP server (nekoro-browser-mcp)
  • helpers.py — general browser helpers, reflected into MCP tools.
  • lifecycle.py — pid file + process fingerprint (never kills a reused pid), stale-daemon
    self-heal (CDP probe fails → cleanup and restart), localhost bypasses the system proxy.
  • Extension, against MV3 service worker eviction — content_scripts heartbeat (wake vector
    living in the page, revives a killed SW) + onStartup (reconnects on Chrome cold start) +
    reattaches the last-driven tab instead of drifting to a blank one.

Self-Healing and Site Knowledge

When an agent hits a gap it writes the missing piece and uses it immediately — nothing is
recompiled, no daemon restart, no extension reload.

  • src/nekoro_browser/agent_helpers.py is scratch paper: reloaded on every /exec, good
    for a quick experiment. It lives inside the installed package, so an upgrade overwrites it.
  • Anything worth keeping goes in your own skills directory (NEKORO_DOMAIN_SKILLS, falling
    back to domain-skills/ in the repo), one folder per site holding both kinds of material:
    <site>/*.md for knowledge and <site>/*.py for workflows. Scripts are loaded into the
    /exec namespace on every call and can use the built-in helpers directly.

The point is that this material finds the agent instead of waiting to be discovered.
navigate() and new_tab() return two extra fields when the site has any:

{'ok': True, 'loaded': True,
 'notes':   ['example/search.md — Example — search results'],
 'actions': ['open_first_result(query) — search and open the top hit']}

notes lists titles only; actions lists functions that are already callable, so the agent
runs one instead of rebuilding the flow. list_site_actions() shows everything loaded,
failed files included. What to record — and what not to — is in
domain-skills/README.md.

Tabs work the same way: a tab left over from last time still holds its login and page state,
so new_tab() adds an existing field when the managed group already has tabs for that site:

{'ok': True, 'tabId': 42, 'loaded': True,
 'existing': {'hint': 'switch_tab(id) reuses an open tab, or new_tab(url, reuse=True)',
              'tabs': [{'tabId': 17, 'title': 'Example Domain'}]}}

The tab still opens — the field only makes reuse visible at the moment a duplicate is about
to appear; reuse=True navigates the existing one instead. Nothing is ever closed
automatically
: sweep_tabs() only reports candidates (same-site duplicates, stray
about:blank), sweep_tabs(dry_run=False) / close_tabs([...]) act on them, and the active
tab is never a candidate.


Platform Support

Platform Status
Windows Primary development platform, exercised end to end
Linux / macOS Platform branches + CI, full Chrome loop untested — reports welcome

Linux/macOS have the platform branches (~/.config / ~/Library/Application Support
data dirs, chmod 600 token, /proc + ps liveness probes) and CI runs unit tests on all
three — but the full "Chrome + extension" loop has never run on a real macOS/Linux box.

Known Limitations

  • Unpacked extensions get disabled by Chrome. An extension installed via "Load unpacked" may be switched off automatically after a Chrome update or restart, or hidden behind the "Disable developer mode extensions" prompt. When --doctor reports Extension/SW not responding, re-enable it in chrome://extensions/ first. This project is not published to the Chrome Web Store, so the limitation is not going away soon.
  • MV3 recovery depends on Chrome. Heartbeats, onStartup and reattachment help restore connections. Short browser regressions cover reload and restart recovery; multi-hour unattended stability remains unverified. Run --ensure before a workflow and handle failures.
  • One shared Chrome session. Helpers with tab=id can target another already attached tab; an unattached target is an error. Other helpers follow the active tab. Concurrent clients share this tab state and must coordinate their actions; separate browser contexts are not provided.
  • Downloads land wherever Chrome is configured to put them; the path cannot be changed from here. wait_for_download() returns {url, filename, bytes} — a filename, not a full path. Set the directory in Chrome's own settings. Both Browser.setDownloadBehavior (-32601) and the deprecated Page.setDownloadBehavior (-32000 "Cannot not access browser-level commands") are browser-level and get rejected under chrome.debugger, which only ever hands out a tab target.
  • The MCP server handles requests serially. During a wait_selector(timeout=90) every other request on that connection (including ping) queues behind it. Open separate client connections if you need concurrency.

Reference

CLI flags, configuration, troubleshooting, security — click to expand

CLI

Command What it does
nekoro-browser Start the daemon (foreground)
nekoro-browser setup Guided install: copies the extension path, then waits until the extension actually connects
nekoro-browser --ensure Launch Chrome / daemon, restart an outdated daemon, and reload an unresponsive or outdated extension. Exit 0 requires a real page response and matching component versions. It never starts a second daemon while the port is still held
nekoro-browser --doctor End-to-end diagnostic, including matching CLI / running daemon / extension versions — reports only, repairs nothing
nekoro-browser --stop Stop the daemon
nekoro-browser --restart Stop and restart (foreground)
nekoro-browser --reload-ext Reload the extension and wait for a new connection plus a real page response; disconnected or failed reloads exit nonzero. Use --ensure after upgrading
nekoro-browser --extension-path Print the extension directory (for "Load unpacked")
nekoro-browser --version Print the installed version (check it against the extension you loaded)
nekoro-browser --port N Run the daemon on port N (default 28417)
nekoro-browser -c "code" Run one snippet, print the result
nekoro-browser --timeout N Seconds to allow a snippet (default 120 — page loads are slow)
nekoro-browser --allow-domains "jd.com,*.taobao.com" Only allow these domains (comma-separated); unset = unrestricted
echo "code" | nekoro-browser Pipe mode (daemon must already be running)

Configuration

The daemon listens on 28417 by default. To change it:

Side How
Python (daemon + CLI + MCP) nekoro-browser --port 30500, or set NEKORO_PORT=30500
Extension Extension details → Extension options → set the port → Save (reconnects immediately, no reload)

Both sides must agree. Clients don't need the flag repeated: the daemon records its
actual port in <data dir>/port, so a plain echo ... | nekoro-browser finds a daemon
running on a non-default port. Precedence is --port > NEKORO_PORT > that file > default.

The data dir holding token / pid / port is %LOCALAPPDATA%\nekoro-browser on Windows,
~/Library/Application Support/nekoro-browser on macOS, $XDG_CONFIG_HOME/nekoro-browser
or ~/.config/nekoro-browser elsewhere. NEKORO_DATA_DIR overrides it on any platform — it replaces the parent of that path; a nekoro-browser/ directory is still created inside it. So with NEKORO_DATA_DIR=/my/dir the token lives at /my/dir/nekoro-browser/token, not /my/dir/token.

The same limit can be set via NEKORO_ALLOW_DOMAINS (comma-separated, same syntax).
Rule syntax: example.com matches exactly; *.example.com matches subdomains and the
bare domain; * allows everything. See Security below.

Troubleshooting

Symptom Cause Fix
Daemon not running Daemon not started Run nekoro-browser --ensure
CDP timeout Extension not connected / service worker asleep nekoro-browser --doctor to diagnose; try --reload-ext or manually reload in chrome://extensions
Extension disabled by Chrome Unpacked extension + Chrome update Re-enable it in chrome://extensions/, then re-run --doctor
Page unchanged Extension not attached to tab Open a regular (non-chrome://) page, restart daemon
Another debugger is already attached Another debugging extension owns that tab Only one debugger per tab. Use a different tab, or disable the other extension in chrome://extensions
Port in use Another process or busy daemon Let a busy daemon finish; use --stop for your daemon, or choose a different port and match it in extension options
Red Errors badge on the extension card in chrome://extensions Daemon isn't running; the extension keeps retrying Run --ensure, then clear old connection errors on the card

Nearly everyone hits the last one: between loading the extension and starting the daemon, every
reconnect logs WebSocket connection to 'ws://127.0.0.1:28417/ws' failed: ERR_CONNECTION_REFUSED.
Chrome's network stack emits that message below the JS layer — the extension's try/catch and
ws.onerror cannot suppress it, and probing with fetch first logs the same thing.
It can be explained, not silenced.

Security

The daemon listens on 127.0.0.1 and /exec runs arbitrary Python, so the transport is guarded:

  • CLI / MCP → daemon (/exec, /raw): a per-session token is written to a user-private file — %LOCALAPPDATA%\nekoro-browser\token on Windows, ~/Library/Application Support/nekoro-browser/token on macOS, $XDG_CONFIG_HOME or ~/.config/nekoro-browser/token elsewhere, chmod 600 on POSIX. Clients read it and send X-Nekoro-Token; missing/wrong token → 403. Web pages and remote hosts can't read local files, so they can't obtain it. /ping stays open.
  • Extension → daemon (/ws): the handshake Origin must be chrome-extension://…; a web page's WebSocket to localhost carries its own origin and is rejected.

Same-user local processes can read the token file — that boundary matches the OS user account, as with browser-harness's chmod 600.

  • Optional domain allowlist: --allow-domains "jd.com,*.taobao.com" (or NEKORO_ALLOW_DOMAINS) gates navigate / new_tab to listed hosts — anything else is refused before reaching CDP. Unset = unrestricted (fail-open): this tool drives your personal Chrome, so the default stays permissive.

Feedback

Hit a problem, or missing a helper you need? Open an
issue.
For bugs, include the output of nekoro-browser --doctor, your Chrome version and OS — saves a round trip.

PRs welcome. Run the tests first: for f in tests/test_*.py; do uv run python "$f"; done (CI runs them on all three platforms too).

For the real browser regression, use Chrome for Testing and Node.js (the test bridge
uses Node's standard library; runtime dependencies are unchanged):

uv run --with build python tests/browser_regression.py --chrome /path/to/chrome

Add --headful to exercise an ordinary browser window. The test builds and installs
the wheel into temporary venvs, enables Developer mode in an isolated profile, and
checks cold startup, 30 navigation/input/click cycles, background-tab screenshots,
three extension reloads, MCP success/error results, browser/daemon restarts, and an
in-place upgrade from PyPI 0.3.5 that retains the domain allowlist. Port 28417 must be
free; the test fails before launching if it is occupied. Temporary browser profiles
and runtime data are removed on exit. Windows CI runs this gate before publication;
this is a short regression, not a long-running stability test.


Acknowledgments

Core architecture derived from:

  • browser-harness — thin-wrapper philosophy (each function is a CDP alias, ≤10 lines), pipe mode, self-healing agent_helpers.py, domain-skills directory structure, cdp() raw access
  • browser-act — state() indexed element tree, *[N] change markers, waitSelector() state polling, getMarkdown() page extraction
  • Playwright — CDP Input.dispatchMouseEvent real mouse events (isTrusted:true), extension + daemon dual-path architecture

Ideas drawn from:

  • ego-lite — "code base, not CLI base" (agent writes a script, not a command loop), unified locator syntax (css: / text: / xpath= …) with transient/permanent element-resolution errors as a retry/abandon signal (→ click()), "name says the intent" openOrReuseTab ergonomics (→ ensure_tab()), and experience-accumulation as a first-class design goal (nekoro's domain-skills already chase this)

Yorumlar (0)

Sonuc bulunamadi