mcp-agy
Health Uyari
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 6 GitHub stars
Code Gecti
- Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
FastMCP server exposing Google Antigravity as an autonomous background coding worker for architect agents (Claude Code, Cursor, Cline, Roo Code)
mcp-agy — Google Antigravity FastMCP Server
Bridging High-Level AI Architect Agents with Autonomous Antigravity Coding Workers via Model Context Protocol (MCP)
🌟 1. Overview & High-Level Vision
In contemporary AI-assisted software engineering, single-agent workflows frequently suffer from cognitive overload, context exhaustion, and execution drift. Large Language Models forced to simultaneously handle macro-level system architecture, cross-module requirements, and low-level line-by-line code editing often lose sight of overarching design invariants.
mcp-agy establishes a paradigm shift: the Architect-Worker Division of Labor.
+-----------------------------------------------------------------------------------+
| THE ARCHITECT-WORKER PARADIGM |
+-----------------------------------------------------------------------------------+
| |
| STRATEGIC TIER (The Architect) |
| - Claude Desktop, Cursor Composer, Cline, Roo Code |
| - Responsibilities: System design, domain modeling, security audit, code review |
| |
| │ |
| │ Delegates via FastMCP JSON-RPC 2.0 |
| ▼ |
| |
| TACTICAL EXECUTION TIER (The Worker) |
| - Google Antigravity (AGY) |
| - Responsibilities: Multi-file edits, build runs, test executions, git diffs |
| |
+-----------------------------------------------------------------------------------+
Key Pillars:
- Separation of Concerns: The AI Architect focuses on reasoning, design trade-offs, and verification. Google Antigravity (AGY) operates as a high-speed execution sandbox performing atomic edits, shell commands, and automated test runs.
- Pristine Stdio Protocol Purity: Standard output (
sys.stdout) is strictly reserved for JSON-RPC 2.0 frames. All logging, startup banners, and diagnostic traces route exclusively tosys.stderr, preventing framing corruption in MCP clients. - Resilient 3-Tier Backend Fallback: Seamlessly transitions across Python SDK (
Tier 1), Subprocess CLI (Tier 2), and High-Fidelity Simulation Mock (Tier 3) for continuous offline testing and CI workflows. - Zero-Trust Workspace Isolation: Path canonicalization, system-critical directory protection, and per-workspace asynchronous mutex concurrency locking prevent race conditions and unintended file modifications.
- Pluggable Worker Reasoning Protocol: Optional delivery of the AI Deep Reasoning Toolkit into each workspace, constraining how the worker writes code while leaving the architect's own behaviour untouched. See §6.
🏗️ 2. System Architecture & Data Flow
+-------------------------------------------------------------------------------+
| External AI Architect Agent |
| (Claude Desktop, Cursor Composer, Cline, Roo Code, Custom) |
+---------------------------------------+---------------------------------------+
|
| JSON-RPC 2.0 over Stdio (Pure Stream)
| (Zero stdout pollution; logs to stderr)
v
+---------------------------------------+---------------------------------------+
| mcp_agy (FastMCP) |
| |
| +--------------------+ +------------+ +--------------+ +-------------+ |
| | agy_execute_task | | agy_chat | | agy_get_diff | |agy_run_tests| |
| +---------+----------+ +-----+------+ +-------+------+ +------+------+ |
| | | | | |
| v v | | |
| +---------+-------------------+------+ | | |
| | Backend Manager | | | |
| | - Tier 1: PythonSDKBackend | | | |
| | - Tier 2: SubprocessCLIBackend | | | |
| | - Tier 3: MockAGYBackend (CI) | | | |
| | - 3-Tier Fallback Hierarchy | | | |
| +-----------------+------------------+ | | |
| | | | |
| | v v |
| | +--------+----+ +--------+------+ |
| | | Diff Engine | | Test Runner | |
| | | - Git diff | | - 13 Tiers | |
| | | - Untracked | | - Auto-Detect | |
| | | - Stats | | - Diagnostics | |
| | +--------+----+ +--------+------+ |
| | | | |
| +-----------------------------+----------------+ |
| | |
| v |
| +------------+------------+ |
| | Workspace Safety Engine | |
| | - Path Canonicalization | |
| | - Boundary Validation | |
| | - Async Concurrency Lock| |
| +------------+------------+ |
+---------------------------------------------------|---------------------------+
|
v
+---------------------------------------------------+---------------------------+
| Target Workspace Directory |
| - Multi-file source code, git repository, virtual environments, tests |
+-------------------------------------------------------------------------------+
Core Subsystems:
- FastMCP Server (
server.py): High-performance asynchronous stdio MCP server exposing 4 rich tools with strict parameter validation and stderr logging. - Backend Manager (
backend_manager.py): Dynamic fallback coordinator orchestrating execution across Python SDK, Subprocess CLI, and offline Simulation Mock. - NDJSON Stream Telemetry Engine (
cli_backend.py): Real-time event parsing for tool calls, thought streams, and token consumption metrics. - Git Unified Diff Engine (
diff_engine.py): Comprehensive diff capturing tracked modifications, staged files, and newly created untracked files. - Multi-Ecosystem Test Runner (
test_runner.py): 13-tier test framework detector (pytest, unittest, npm, cargo, go, gradle, maven, dotnet, etc.) with ANSI color stripping and structured regex diagnostics. - Workspace Safety Engine (
workspace.py): Realpath canonicalization, boundary validation, system root blocking, and async mutex locking per workspace.
📦 3. Installation & Prerequisites
Prerequisites:
- Operating System: Windows 10/11, macOS 12+, Linux (Ubuntu 20.04+, Debian, Arch, Fedora)
- Python: Version 3.10, 3.11, or 3.12+
- Package Manager:
uv(strongly recommended) orpip - Git: Version 2.20+ (for unified diff engine and repository tracking)
- Google Antigravity: Antigravity CLI (
agy) installed and authenticated, orgoogle-antigravityPython SDK (offline simulation mock runs automatically if not present).
Quick Installation:
# Clone the repository
git clone https://github.com/google/mcp-agy.git
cd mcp-agy
# Create virtual environment and install in editable mode using uv
uv venv
uv pip install -e .
# Or using standard pip:
python -m venv .venv
# On Windows:
.venv\Scripts\activate
# On macOS/Linux:
source .venv/bin/activate
pip install -e .
🧰 4. Complete Tool Suite Reference
mcp-agy exposes 8 tools tailored for AI Architect agents: 4 synchronous ones that answer
within a call, and 4 that run AGY as a background job so a long task never has to fit inside
one.
Which to reach for. agy_execute_task and agy_chat block until AGY is done, so they only
work for runs shorter than your client's per-call timeout — 60 seconds on Claude Code. For
anything larger, and that is most real work, start a job:
agy_start_task(workspace, prompt) -> {"job_id": "...", "status": "running"} # returns at once
… the architect keeps working …
agy_job_status(job_id, wait_seconds=30) -> {"status": "completed", "result": {...}} # collect
agy_get_diff(workspace) -> review what AGY actually changed
agy_run_tests(workspace) -> verify it
agy_job_status with wait_seconds returns the instant the job finishes, so a short job needs
no polling loop and a long one costs one cheap call per check. The recommended waiting pattern is
monitoring done_marker_path on disk (e.g. until [ -f <path> ]; do sleep 5; done) and callingagy_job_status(job_id, wait_seconds=0). wait_seconds should always be kept strictly below your
client's per-call timeout: a call with wait_seconds=300 cut by a 60s client timeout will tear down
and restart the server process. Completed job results are atomically persisted to disk before .done
markers appear, so finished jobs survive server restarts and return with recovered_from_disk=True.
| Tool Name | Parameter | Type | Required | Default | Description |
|---|---|---|---|---|---|
agy_start_task |
workspace_path |
str |
Yes | — | Target workspace. Returns a job_id immediately; AGY keeps running in the server. |
prompt |
str |
Yes | — | Self-contained instructions for AGY. | |
auto_approve |
bool |
No | True |
Auto-approve AGY's tool executions and file edits. | |
mode |
Literal["accept-edits", "plan"] |
No | "accept-edits" |
"plan" is the read-only form — the long-running equivalent of agy_chat. |
|
timeout_seconds |
int |
No | 600 |
Bounds the run, not this call (1 to 3600). | |
rigor |
Literal["standard", "deep", "off"] |
No | "standard" |
Reasoning protocol for the run. Inert unless MCP_AGY_TOOLKIT_PATH is set — see §6. |
|
agy_job_status |
job_id |
str |
Yes | — | Id returned by agy_start_task. |
wait_seconds |
int |
No | 0 |
Block up to N seconds, returning early on completion. Capped at 45 so the call always fits inside a client timeout. | |
agy_cancel_job |
job_id |
str |
Yes | — | Stops the run and kills its process tree. Not a rollback — files already written stay written. |
agy_list_jobs |
— | — | — | — | Every known job, newest first, without their results. The recovery path when a job_id has fallen out of context; collect a result with agy_job_status. |
And the 4 synchronous tools:
| Tool Name | Parameter | Type | Required | Default | Description |
|---|---|---|---|---|---|
agy_execute_task |
workspace_path |
str |
Yes | — | Target workspace path where AGY will execute coding tasks. Must be an existing directory. |
prompt |
str |
Yes | — | Comprehensive coding instructions, file paths, requirements, and constraints. | |
auto_approve |
bool |
No | True |
Automatically approve tool executions and file edits without interactive confirmation. | |
mode |
Literal["accept-edits", "plan"] |
No | "accept-edits" |
"accept-edits" for full file modifications; "plan" for read-only analysis without disk writes. |
|
timeout_seconds |
int |
No | 600 |
Maximum execution duration in seconds (1 to 3600). Enforces process tree termination on timeout. | |
rigor |
Literal["standard", "deep", "off"] |
No | "standard" |
Reasoning protocol for the run. Inert unless MCP_AGY_TOOLKIT_PATH is set — see §6. |
|
agy_chat |
prompt |
str |
Yes | — | Analytical query, codebase question, architecture review, or planning consultation. |
workspace_path |
str |
No | "" |
Optional workspace directory to provide codebase context for analysis (strictly read-only). | |
conversation_id |
str |
No | "" |
UUID of a previous conversation session to maintain multi-turn dialogue context. | |
timeout_seconds |
int |
No | 300 |
Maximum consultation duration in seconds (1 to 1800). | |
agy_get_diff |
workspace_path |
str |
Yes | — | Target workspace repository directory to inspect for git and file modifications. |
agy_run_tests |
workspace_path |
str |
Yes | — | Target workspace directory where test suites will be discovered and executed. |
test_command |
str |
No | "" |
Optional explicit test command (e.g. pytest tests/ -v, cargo test). If empty, auto-detects from 13 supported tiers. |
|
timeout_seconds |
int |
No | 300 |
Maximum test execution duration in seconds (1 to 1800). |
Detailed Return Schemas
1. TaskExecutionResult (from agy_execute_task)
{
"status": "success",
"conversation_id": "9b1deb4d-3b7d-4bad-9bdd-2b0d7b3dcb6d",
"response": "Implemented JWT authentication module in src/auth.py and created unit tests in tests/test_auth.py.",
"modified_files": [
"src/auth.py",
"tests/test_auth.py"
],
"diff_summary": "2 file(s) modified (+54, -2)",
"duration_seconds": 14.82,
"tokens_used": {
"input_tokens": 2100,
"output_tokens": 920,
"thinking_tokens": 450,
"cache_read_tokens": 0,
"total_tokens": 3470
},
"backend_used": "sdk",
"error_details": null,
"reasoning": {
"rigor": "standard",
"toolkit_active": true,
"toolkit_source": "/home/you/src/ai-deep-reasoning-toolkit",
"toolkit_revision": "a7236c1",
"installed": ["GEMINI.md", ".agents/skills/deep-verify"],
"gate_line": "Simplicity gate: cut _fmt_row, _DEFAULTS; kept parse (LOAD-BEARING: contract rule 3).",
"deep_verify_declined": false,
"notes": ""
}
}
reasoning is null unless the deep reasoning toolkit is configured — see §6.
2. ChatResult (from agy_chat)
{
"status": "success",
"conversation_id": "4a1c5b8e-7e9a-4123-b123-abcdef012345",
"response": "Based on the codebase analysis, `src/services/billing.py` depends directly on `src/db/raw_queries.py`...",
"duration_seconds": 3.45,
"tokens_used": {
"input_tokens": 1200,
"output_tokens": 480,
"total_tokens": 1680
},
"backend_used": "cli",
"error_details": null
}
3. DiffResult (from agy_get_diff)
{
"status": "success",
"has_changes": true,
"unified_diff": "--- a/src/auth.py\n+++ b/src/auth.py\n@@ -10,6 +10,18 @@\n+class JWTManager:\n+ def generate_token(self, user_id: str) -> str:\n+ ...",
"changed_files": [
{
"path": "src/auth.py",
"status": "M",
"insertions": 48,
"deletions": 2
}
],
"untracked_files": [
"tests/test_auth.py"
],
"summary": "1 modified, 1 untracked file(s)",
"error_details": null
}
4. TestRunResult (from agy_run_tests)
{
"status": "passed",
"exit_code": 0,
"framework": "pytest",
"test_command_executed": "pytest tests/ -v",
"output": "============================= test session starts =============================\ntests/test_auth.py::test_generate_token PASSED\ntests/test_auth.py::test_verify_token PASSED\n============================== 2 passed in 0.45s ==============================",
"summary": {
"total": 2,
"passed": 2,
"failed": 0,
"skipped": 0,
"errors": 0
},
"failures": [],
"duration_seconds": 0.52,
"error_details": null
}
⚙️ 5. Step-by-Step Client Setup Guides
1. Claude Desktop Setup
Claude Desktop interacts with MCP servers via local stdio processes.
- Locate or create your Claude Desktop configuration file:
- Windows:
%APPDATA%\Claude\claude_desktop_config.json(e.g.C:\Users\<YourUser>\AppData\Roaming\Claude\claude_desktop_config.json) - macOS:
~/Library/Application Support/Claude/claude_desktop_config.json - Linux:
~/.config/Claude/claude_desktop_config.json
- Windows:
- Add the
mcp-agyconfiguration entry:
{
"mcpServers": {
"mcp-agy": {
"command": "uv",
"args": [
"--directory",
"C:\\path\\to\\mcp_agy",
"run",
"mcp-agy"
],
"env": {
"MCP_AGY_BACKEND": "cli",
"MCP_AGY_AUTO_FALLBACK": "0",
"MCP_AGY_LOG_LEVEL": "INFO",
"MCP_AGY_MODEL": "gemini-3.8-flash-high",
"AGY_BIN_PATH": "C:\\Users\\<you>\\AppData\\Local\\agy\\bin\\agy.EXE"
}
}
}
}
(Replace C:\\path\\to\\mcp_agy with the absolute path to your mcp_agy repository).
- Fully restart Claude Desktop.
- Click the 🔨 Hammer icon in the bottom right corner of Claude's prompt bar. Verify that all 8 tools (
agy_execute_task,agy_chat,agy_get_diff,agy_run_tests,agy_start_task,agy_job_status,agy_cancel_job,agy_list_jobs) appear with green indicators.
2. Cursor Setup
Cursor supports project-level MCP server definitions in .cursor/mcp.json.
- In your project's root directory, create
.cursor/mcp.json. - Add the following JSON content:
{
"mcpServers": {
"mcp-agy": {
"command": "uv",
"args": [
"--directory",
"C:\\path\\to\\mcp_agy",
"run",
"mcp-agy"
],
"env": {
"MCP_AGY_BACKEND": "cli",
"MCP_AGY_AUTO_FALLBACK": "0",
"MCP_AGY_LOG_LEVEL": "INFO",
"MCP_AGY_MODEL": "gemini-3.8-flash-high",
"AGY_BIN_PATH": "C:\\Users\\<you>\\AppData\\Local\\agy\\bin\\agy.EXE"
}
}
}
}
(Replace C:\\path\\to\\mcp_agy with the absolute path to your mcp_agy repository).
- Open Cursor Settings -> Features -> MCP Servers. Verify
mcp-agyis listed and connected. - Use Cursor Composer or Chat to instruct AGY directly.
3. Cline Setup (VS Code Extension)
Cline allows automated autonomous agent workflows with configurable auto-approval.
- In VS Code, open the Cline side panel.
- Click the ⚙️ Settings gear -> MCP Servers tab -> Configure MCP Servers.
- Edit
cline_mcp_settings.json:
{
"mcpServers": {
"mcp-agy": {
"command": "uv",
"args": [
"--directory",
"C:\\path\\to\\mcp_agy",
"run",
"mcp-agy"
],
"env": {
"MCP_AGY_BACKEND": "cli",
"MCP_AGY_AUTO_FALLBACK": "0",
"MCP_AGY_LOG_LEVEL": "INFO",
"MCP_AGY_MODEL": "gemini-3.8-flash-high",
"AGY_BIN_PATH": "C:\\Users\\<you>\\AppData\\Local\\agy\\bin\\agy.EXE"
},
"disabled": false,
"autoApprove": [
"agy_execute_task",
"agy_chat",
"agy_get_diff",
"agy_run_tests"
]
}
}
}
(Replace C:\\path\\to\\mcp_agy with the absolute path to your mcp_agy repository).
4. Roo Code Setup (VS Code Extension)
Roo Code allows specialized custom modes (Architect, Code, Test) delegating to AGY.
- Open Roo Code settings panel -> MCP Servers.
- Click Edit MCP Settings and insert the server definition:
{
"mcpServers": {
"mcp-agy": {
"command": "uv",
"args": [
"--directory",
"C:\\path\\to\\mcp_agy",
"run",
"mcp-agy"
],
"env": {
"MCP_AGY_BACKEND": "cli",
"MCP_AGY_AUTO_FALLBACK": "0",
"MCP_AGY_LOG_LEVEL": "INFO",
"MCP_AGY_MODEL": "gemini-3.8-flash-high",
"AGY_BIN_PATH": "C:\\Users\\<you>\\AppData\\Local\\agy\\bin\\agy.EXE"
},
"disabled": false,
"autoApprove": [
"agy_execute_task",
"agy_chat",
"agy_get_diff",
"agy_run_tests"
]
}
}
}
(Replace C:\\path\\to\\mcp_agy with the absolute path to your mcp_agy repository).
5. Claude Code Setup
Claude Code reads .mcp.json from the project root (see configs/claude_code_mcp.json):
{
"mcpServers": {
"mcp-agy": {
"type": "stdio",
"command": "uv",
"args": ["--directory", "C:\\path\\to\\mcp_agy", "run", "mcp-agy"],
"env": {
"MCP_AGY_BACKEND": "cli",
"MCP_AGY_AUTO_FALLBACK": "false",
"MCP_AGY_LOG_LEVEL": "INFO",
"MCP_AGY_MODEL": "gemini-3.8-flash-high"
},
"timeout": 1800000
}
}
}
(Replace C:\\path\\to\\mcp_agy with the absolute path to your mcp_agy repository).
timeout is not decoration — read the next section before your first real task.
⏱️ Raise the client's call timeout, or nothing here works
Every tool in this server is a wrapper around a call that takes minutes. A realagy_execute_task on a real repository runs for 2–20 minutes; even an agy_chat review of a
handful of files takes several. Every MCP client caps how long it will wait for a single tool
call, and the defaults are all far below that. When the cap fires the client reportsError: Request timed out and drops the call — the architect gets a bare failure, and the
work AGY did in those minutes is discarded.
Claude Code's default cap is 60 seconds — measured, not quoted: a client-side log of a liveagy_chat shows still running (30s elapsed), still running (60s elapsed), thenError: Request timed out. Two ways to raise it, either is enough:
| Where | Setting | Notes |
|---|---|---|
.mcp.json, per server |
"timeout": 1800000 |
Milliseconds. Overrides the env var for this server only. Values under 1000 are ignored. |
settings.json, env block |
"MCP_TOOL_TIMEOUT": "1800000" |
Milliseconds. Applies to every MCP server the client launches. |
Both are read when the client starts, so restart the client after changing either.
Progress notifications will not save you. Claude Code documents
timeoutas a hard
wall-clock limit per call, explicitly stating that progress notifications do not extend it.
A keepalive heartbeat from this server would be wasted effort against that client — raising
the cap is the only fix.
Other clients cap calls too — Claude Desktop, Cursor, Cline and Roo Code each expose their own
timeout setting. If a long task fails while the server log shows AGY still working, that cap is
the first thing to check.
🧠 6. Deep Reasoning Toolkit Integration
mcp-agy can deliver the AI Deep Reasoning Toolkit
into every workspace it hands to AGY. The toolkit is a rule file (GEMINI.md) plus an
on-demand skill (deep-verify) that constrain how Gemini/Antigravity writes code.
The integration is off unless you switch it on, and it changes nothing about the architect.
Why this belongs in the MCP server
Antigravity reads GEMINI.md and .agents/skills/ from the directory it is working in, and
this server already spawns agy with the workspace as its cwd. Claude Code, Cursor and Copilot
read different filenames, so installing the toolkit changes the worker's behaviour and
leaves the architect's untouched:
Claude Code (architect) ──▶ mcp-agy ──▶ agy --print (cwd = workspace)
does not read GEMINI.md reads GEMINI.md + .agents/skills/
Doing it here rather than by hand buys one thing that matters. The toolkit's own benchmark
measured deep-verify self-activating in 0 of 10 runs on a task built to exactly the shape
the skill describes — and 5 of 5 once the prompt named the trade-offs out loud. Discovery is
what fails, not the skill. The architect is the only party that knows whether a task has two
genuinely different designs, so rigor="deep" is how it says so.
Setup
git clone https://github.com/DuongNAD/ai-deep-reasoning-toolkit.git ~/src/ai-deep-reasoning-toolkit
Point the server at the checkout in your client config's env block:
{
"mcpServers": {
"mcp-agy": {
"command": "uv",
"args": ["--directory", "/path/to/mcp-agy", "run", "mcp-agy"],
"env": {
"MCP_AGY_BACKEND": "cli",
"MCP_AGY_AUTO_FALLBACK": "false",
"MCP_AGY_TOOLKIT_PATH": "/home/you/src/ai-deep-reasoning-toolkit"
},
"timeout": 1800000
}
}
}
A path that does not contain GEMINI.md logs a warning and disables the integration, rather
than looking like a working install.
The rigor parameter
Accepted by agy_execute_task and agy_start_task.
rigor |
What happens | When to use it |
|---|---|---|
"standard" (default) |
Installs the toolkit in the workspace. AGY loads the always-on rules; your prompt is passed through untouched. | Everything. This is the setting the measured gains come from. |
"deep" |
Also prepends a preamble naming the deep-verify skill and requiring its Comparative Matrix. |
Only when the task admits two or more genuinely different designs and choosing wrong is expensive to reverse. |
"off" |
Nothing is written, nothing is added. | A workspace you want left exactly as it is. |
deep is not a free upgrade, and this is measured. On a fully specified contract the
toolkit's benchmark recorded an identical judge score, 37% more code, and all five subagent
tournaments electing the same architecture. Reach for it when you want the alternatives
enumerated, not when you want the answer to be more correct.
deep also permits the skill to decline. The skill states that a matrix of one real design
against two strawmen "launders a foregone conclusion as deliberation", so the preamble tells it
to decline in one line rather than manufacture alternatives — and that declination comes back
as deep_verify_declined: true.
Asking for deep without a configured toolkit does not silently downgrade: the result
carries toolkit_active: false and a note saying why, because a plain answer read as the output
of a verification pipeline is the one failure mode worth being loud about.
What lands in the workspace
Exactly two paths, and neither is ever overwritten:
<workspace>/GEMINI.md
<workspace>/.agents/skills/deep-verify/
An existing GEMINI.md is left alone and reported in notes — it may be your own rules for
that repo, and there is no meaningful way to merge two rule files.
This happens in mode="plan" too. That mode promises AGY will not touch your code, and it
still does not — but the two toolkit files are written before the investigation starts, because
that is what makes the rules available to it. installed on the result says so every time.
AGENTS.md is never written. Antigravity reads it, but so does Claude Code, so writing it
would destroy the worker/architect isolation the whole design rests on.
Your repository stays clean. Both paths are registered in .git/info/exclude, which is
local to the clone and never committed. Because agy_get_diff and modified_files both readgit status, the toolkit disappears from them for free — no skip-list to keep in sync, and
nothing showing up in your own git status either:
$ git status --porcelain
M src/calc/stats.py # your change
$ agy_get_diff(workspace)
"1 file(s) changed: 1 modified (+8, -1)" -> ["src/calc/stats.py"]
What comes back
TaskExecutionResult.reasoning records the conditions the run was carried out under. Two runs
of one task under two rigor settings or two toolkit revisions are two different experiments; a
result that cannot name its arm cannot be compared against another.
"reasoning": {
"rigor": "standard",
"toolkit_active": true,
"toolkit_source": "/home/you/src/ai-deep-reasoning-toolkit",
"toolkit_revision": "a7236c1",
"installed": ["GEMINI.md", ".agents/skills/deep-verify"],
"gate_line": "Simplicity gate: cut _fmt_row, _DEFAULTS; kept parse (LOAD-BEARING: contract rule 3).",
"deep_verify_declined": false,
"notes": ""
}
gate_line is the load-bearing one. GEMINI.md §4.4 mandates that line on every response that
ships code, and states that a missing gate line means the gate did not run. Lifting it out of
the prose gives the architect a machine-checkable signal that the protocol actually executed —
the same evidence the toolkit's benchmark uses to conclude its rule file was loaded at all.
Only the two formats the toolkit genuinely mandates are parsed. The Stage 0 tier declaration is
required to be one line but its wording is left open, and skill activation has no declared
marker, so neither is guessed at: a field that is confidently wrong some of the time is worse
than no field.
What the toolkit does and does not buy
From the toolkit's own benchmark — one model, two tasks, ~48 runs, graded by a suite the agent
never sees:
| baseline | with toolkit | ||
|---|---|---|---|
safe-path judge score |
417/420 | 417/420 | unchanged |
safe-path AST statements |
52.3 | 33.2 | −36% |
event-bus judge score |
400/400 | 320/320 | unchanged (both perfect) |
event-bus AST statements |
55.5 | 49.1 | −12% |
Correctness never improved. The toolkit does not make the model think of a better solution;
it stops the model shipping things nobody asked for. Install it if you are tired of Gemini
inventing helpers, metrics and config options on its own. Do not install it expecting it to
catch bugs.
Those numbers come from one model on two single-file Python tasks with fully specified
contracts. Ambiguous requirements, multi-file refactors, and legacy codebases are untested —
and that is exactly the territory the rest of GEMINI.md aims at.
Verified end to end
Against a real agy run on a real repository, with MCP_AGY_AUTO_FALLBACK=false so a
simulated answer could not pass: 14/14 checks — both artefacts installed, AGENTS.md absent,git status free of them, agy_get_diff and modified_files reporting only the real change,
provenance stamped with the checkout's actual revision, and gate_line carrying aSimplicity gate: line that Gemini emitted because it had read the rules file.
🎯 7. High-Yield Prompt Templates for Architect Agents
These templates are battle-tested prompts designed for LLMs acting in the Architect Role to orchestrate Google Antigravity:
Template 0: The delegation loop (use this for anything that takes minutes)
The shape every other template should be run in. The architect never blocks on the worker, and
never takes the worker's word for what it did.
1. DELEGATE — hand over one self-contained unit of work:
agy_start_task(workspace_path="E:/Project/Thing",
prompt="<one job, exact files, explicit acceptance criteria>")
-> {"job_id": "…", "status": "running"} # returns in milliseconds
2. KEEP WORKING — read the code you are about to review, plan the next unit,
or start a second job in a *different* workspace. Do not sit on the job id.
3. COLLECT — agy_job_status(job_id, wait_seconds=30)
'running' -> ask again later, the run is unharmed
'completed' -> read result.status, result.response, result.modified_files
4. VERIFY — never accept the worker's own report as evidence:
agy_get_diff(workspace_path=…) # what actually changed on disk
agy_run_tests(workspace_path=…) # whether it still works
5. DECIDE — accept, or send back one corrective job naming exactly what was wrong.
agy_cancel_job(job_id) if a run is heading the wrong way; cancelling is not a
rollback, so follow it with agy_get_diff to see what already landed.
Writing the prompt is the architect's real work. A worker prompt earns its keep when it
carries: the exact files to touch and the ones not to, the defect stated as what happens
rather than what to change, an acceptance criterion the worker can check itself, and the
project's own constraints (test command, style gate, forbidden dependencies). Vague delegation
is what produces a confident report and an unusable diff.
Template 1: Greenfield Feature Implementation & Scaffolding
You are the High-Level Software Architect. I need you to implement a new feature in the workspace using the `agy_execute_task` tool.
**Workspace**: `/path/to/project`
**Task**: Implement a JWT-based authentication module with Redis token blacklisting.
**Prompt for `agy_execute_task`**:
"""
Implement a production-ready JWT authentication module in `src/auth/jwt.py` and `src/auth/redis_store.py`.
Requirements:
1. Create `JWTManager` class supporting token generation (`access_token` 15m expiry, `refresh_token` 7d expiry) using `PyJWT`.
2. Implement token revocation via Redis key-value expiration store with prefix `jwt_revoked:`.
3. Provide unit test suite in `tests/test_jwt_auth.py` covering token creation, validation, expiration, and revocation.
4. Run `pytest tests/test_jwt_auth.py` to ensure all tests pass with 100% assertions.
Constraints:
- Do not introduce breaking changes to existing `src/config.py`.
- Type annotations must be strict and pass `mypy`.
- Use async Redis client (`redis.asyncio`).
"""
After execution:
1. Call `agy_get_diff(workspace_path="/path/to/project")` to inspect the code changes.
2. Call `agy_run_tests(workspace_path="/path/to/project")` to confirm test suite integrity.
Template 2: Bug Investigation, Fix & Git Diff Review Loop
You are the AI Architect. A critical bug has been reported in our data processing pipeline.
**Workspace**: `/path/to/project`
**Bug Report**: `ZeroDivisionError` occurring in `src/analytics/metrics.py:84` when computing average latency over empty sample batches.
**Workflow**:
1. Invoke `agy_execute_task` with:
"""
Investigate and fix the ZeroDivisionError in `src/analytics/metrics.py`.
1. Check line 84 where batch latency division occurs. Guard against `len(samples) == 0` by returning 0.0 or `None` as appropriate.
2. Locate existing tests in `tests/test_metrics.py` and add regression test cases for empty batches, single-item batches, and large batches.
3. Verify the fix by running pytest.
"""
2. Inspect the exact diff using `agy_get_diff(workspace_path="/path/to/project")`.
3. Verify that only the intended lines and tests were modified.
Template 3: Safe Code Refactoring with Regression Testing
You are the AI Architect managing a legacy code refactoring.
**Workspace**: `/path/to/project`
**Goal**: Refactor synchronous database operations in `src/db/repository.py` to modern SQLAlchemy 2.0 AsyncSession.
**Workflow**:
1. Run baseline tests first: `agy_run_tests(workspace_path="/path/to/project")` to verify existing tests pass.
2. Invoke `agy_execute_task` with:
"""
Refactor `src/db/repository.py` from sync SQLAlchemy 1.4 syntax to SQLAlchemy 2.0 async syntax (`AsyncSession`, `select()`, `await session.execute()`).
- Update `tests/test_repository.py` with `pytest-asyncio` fixtures.
- Preserve all public method signatures and return models.
- Ensure all existing tests continue to pass.
"""
3. Run `agy_run_tests(workspace_path="/path/to/project")` to verify 0 regressions.
4. Call `agy_get_diff` to review the refactor diff.
Template 4: Test-Driven Development (TDD) Workflow
You are the AI Architect executing a Test-Driven Development workflow.
**Workspace**: `/path/to/project`
**Feature**: Implement a Rate Limiter using Leaky Bucket algorithm.
**Step 1: Write Failing Tests (Red Phase)**
Call `agy_execute_task`:
"""
Create `tests/test_rate_limiter.py` with comprehensive unit tests for a `LeakyBucketLimiter(capacity=10, leak_rate=2.0)` class.
Include tests for:
- Allowing burst requests up to capacity.
- Rejecting requests when capacity is exceeded.
- Leaking tokens over time.
- Thread-safety under concurrent access.
Do NOT create the implementation yet. Run pytest and confirm tests fail as expected.
"""
**Step 2: Implement Code (Green Phase)**
Call `agy_execute_task`:
"""
Implement `src/rate_limiter.py` containing `LeakyBucketLimiter` to satisfy all test cases in `tests/test_rate_limiter.py`.
Run `pytest tests/test_rate_limiter.py` until all tests pass.
"""
**Step 3: Verification & Review**
Call `agy_run_tests` and `agy_get_diff` to finalize.
Template 5: Analytical / Architectural Consultation (Read-Only)
You are the AI Architect consulting Google Antigravity on repository structure and design.
**Workspace**: `/path/to/project`
**Question**: "Analyze the project's dependency graph in `src/` and evaluate whether our microservices module boundaries violate domain-driven design principles."
**Action**:
Call `agy_chat(prompt="Analyze src/ dependencies and evaluate DDD module boundaries.", workspace_path="/path/to/project")`.
Review AGY's reasoning, architectural trade-offs, and refactoring plan before commissioning any disk modifications.
🔒 8. Security, Workspace Isolation & Concurrency Locking
When external AI agents manipulate files and run shell commands, filesystem safety is paramount. mcp-agy enforces multi-layer defenses:
Incoming Request (workspace_path)
│
▼
+─────────────────────────────────+
│ Path Canonicalization │
│ - Resolves symlinks/junctions │
│ - Strips null bytes (\x00) │
│ - Expands user tilde (~) │
+────────────────┬────────────────+
│
▼
+─────────────────────────────────+
│ Filesystem Root Protection │
│ - Blocks C:\, D:\, / roots │
│ - Blocks raw user home root │
+────────────────┬────────────────+
│
▼
+─────────────────────────────────+
│ System-Critical Directory Blk │
│ - Blocks C:\Windows, /etc, │
│ /sys, /bin, /usr, /var │
+────────────────┬────────────────+
│
▼
+─────────────────────────────────+
│ Allowed-Roots Whitelist Check │
│ - Optional boundary fence │
+────────────────┬────────────────+
│
▼
+─────────────────────────────────+
│ Async Mutex Concurrency Lock │
│ - Serializes ops per workspace│
│ - Prevents Git index collisions│
+────────────────┬────────────────+
│
▼
Safe Execution Sandbox
- Path Canonicalization (
canonicalize_workspace_path): Resolves all relative.and..segments, follows symlinks/junctions, and blocks null-byte string injection attacks. - System Root & OS Protection (
is_system_critical_path): Explicitly blocks requests targeting drive roots (C:\,/), operating system directories (C:\Windows,/etc,/sys,/usr,/bin), and raw user profile roots (C:\Users\Admin,/home/user). Project subdirectories within user homes (e.g.C:\Users\Admin\Projects\repo) are fully permitted. - Boundary Validation (
validate_workspace_path): Enforces optionalallowed_rootsconstraints to sandbox agent operations within approved directories. - Asynchronous Concurrency Locking (
WorkspaceLockManager): Prevents race conditions and.git/index.lockcontention by ensuring only one mutating operation executes per workspace directory at any given moment.
🔧 9. Environment Variables & Backend Configuration
MCP clients launch this server from a JSON config whose only tunable is the env block —
they cannot append CLI flags. Every variable below is therefore honored from the environment.
Where an equivalent CLI flag exists, the explicit flag wins; the environment is the fallback.
| Environment Variable | Allowed Values | Default | Description |
|---|---|---|---|
MCP_AGY_BACKEND |
auto, sdk, cli, mock |
auto |
Forces specific AGY execution backend or enables 3-tier probing hierarchy. |
MCP_AGY_AUTO_FALLBACK |
1, true, yes, on / 0, false, no, off |
false |
Falls back to the simulated mock engine when a real backend is unavailable or errors. Off by default on purpose — with it on, a failed run returns a fabricated status="success". |
AGY_BIN_PATH |
File path string | None (searches PATH) |
Absolute path to the agy binary. Checked before MCP_AGY_CLI_PATH and before the PATH search — set this when the client launches the server without a usable PATH. |
MCP_AGY_CLI_PATH |
File path string | None (searches PATH) |
Alternate spelling of AGY_BIN_PATH, checked second. |
MCP_AGY_MODEL |
Model ID string | None (agy CLI default) |
Default model for both CLI and SDK backends, e.g. gemini-3.8-flash-high. Run agy models for valid ids. A per-call model argument overrides it. |
MCP_AGY_DEFAULT_MODEL |
Model ID string | None |
Accepted alias for MCP_AGY_MODEL, which takes precedence when both are set. |
MCP_AGY_EFFORT |
low, medium, high |
None (agy CLI default) |
Default reasoning effort. A per-call effort argument overrides it. |
MCP_AGY_MAX_CONCURRENCY |
Positive integer | 0 (no limit) |
Most agy processes allowed to run at once; extra background jobs wait inside the server and timeout_seconds counts only the run, not the wait. Each agy run loads every MCP server in your Antigravity config - measured ~730 MB and ~18 child processes per run on one machine (20 at once: ~14.6 GB; capped at 5: ~2.3 GB). Blank, non-numeric and non-positive values mean no limit. |
MCP_AGY_KILL_CHILDREN_ON_EXIT |
1, true / 0, false, no, off |
on (Windows only) | Puts the server in a job object so the OS kills every agy run it started if the server dies, however it dies. Without it, a server killed by its client left the whole agy trees running: measured 43 surviving processes (3 agy.exe plus the MCP servers each loads) after a hard kill with 3 jobs in flight, 0 with this on. |
MCP_AGY_LOG_LEVEL |
DEBUG, INFO, WARNING, ERROR |
INFO |
Logging verbosity directed strictly to sys.stderr. Case-insensitive; an unrecognized value logs a warning and falls back to INFO rather than refusing to boot. |
MCP_AGY_DEBUG |
1, true, yes, on / 0, false, no, off |
false |
Enables debug mode and raises logging to DEBUG. An explicit MCP_AGY_LOG_LEVEL still wins over the level this implies. |
MCP_AGY_TOOLKIT_PATH |
Directory path string | None (integration off) |
Checkout of the AI Deep Reasoning Toolkit. When set, rigor on agy_execute_task / agy_start_task becomes live and results carry a reasoning profile. A path without a GEMINI.md in it logs a warning and stays off. See §6. |
MCP_AGY_JOB_MARKER_DIR |
Directory path string | System temp dir | Where background-job .done markers and persisted results are written. Redirect it to keep one machine's job records isolated (the test suite does). |
Running many tasks at once
- One workspace per write task. Edits to one workspace run one at a time;
mode='plan'jobs may share one. Give each write task its owngit worktree. - Cap the processes. Set
MCP_AGY_MAX_CONCURRENCY(16 is a sensible start). Queued jobs show asrunninginagy_list_jobs. - Keep replies short. Every
agy_job_statusresult lands in the architect's context. End each prompt withREPLY: at most 5 lines - what changed, test result, blockers, and verify withagy_get_diff/agy_run_testsrather than the reply. /teamwork-previewas the prompt. A prompt that starts with AGY's/teamwork-previewcommand makes oneagyprocess run a team of subagents. Measured once on four small independent utilities: 1 process / ~1.1 GB / 313 s, against 4 plain jobs at 4 processes / ~3.2 GB / 190 s. Prefer it for one large project; prefer plain jobs when latency matters or the tasks are unrelated.- Slim the worker. Disable the Antigravity MCP servers a worker does not need (
agy mcp disable <name>) - above allmcp-agyitself, so a worker cannot start workers of its own. This edits your global Antigravity config.
🧪 10. Development, Testing & Verification
Running the Full Test Suite:
# Run all unit, integration, and protocol tests
pytest -v
# Run with coverage report
pytest --cov=src/mcp_agy -v
# Run specific test modules
pytest tests/test_workspace_isolation.py -v
pytest tests/test_cli_packaging.py -v
pytest tests/test_server_protocol.py -v
pytest tests/test_diff_engine.py -v
pytest tests/test_test_runner.py -v
Verifying Stdio Stream Purity:
FastMCP stdio communication requires zero stdout corruption. All logs, diagnostics, and banners must route exclusively to stderr.
# Run CLI with debug logging and redirect stderr to a file
# Standard output must be empty until a JSON-RPC request is sent
python -m mcp_agy --debug 2> stderr.log
❓ 11. Troubleshooting & FAQ
- Q: Claude Desktop reports "Could not connect to MCP server"?
- A: Ensure
uvis installed and reachable in your systemPATH. Checkclaude_desktop_config.jsonto ensure JSON syntax is valid and all file paths use escaped backslashes\\on Windows.
- A: Ensure
- Q: What happens if Google Antigravity CLI is not installed?
- A:
mcp-agyautomatically falls back toMockAGYBackend, allowing complete local testing, offline development, and CI verification without requiring live credentials.
- A:
- Q: How does
agy_run_testsdiscover virtual environments?- A:
MultiEcosystemTestRunnerautomatically checks for local.venv,venv, orenvdirectories and injects their executable directory intoPATHduring test execution.
- A:
- Q: Can multiple MCP clients use
mcp-agysimultaneously?- A: Yes. FastMCP supports concurrent sessions, and
WorkspaceLockManagerserializes operations targeting the same workspace repository to prevent race conditions.
- A: Yes. FastMCP supports concurrent sessions, and
- Q: A call fails with
Error: Request timed outafter about a minute, every time.- A: That is the client giving up, not AGY. See "Raise the client's call timeout" at the end of §5. AGY keeps running for a moment after the client drops the call, then the server kills the process tree — no orphans, but the work is lost.
- Q:
error_detailssaysinvalid model selection ... conflicts with --effort=low.- A: Model ids ending in
-low/-highalready carry a reasoning effort, so passingeffortas well is a contradiction the CLI refuses before it starts. Either drop theeffortargument or pick a model id without the suffix. Runagy modelsfor the valid ids.
- A: Model ids ending in
- Q: A call returns
status: "error"andAGY returned an empty response.- A: AGY started, decided it needed a tool, and could not use it — most often a permission it cannot prompt for in headless mode. The rest of
error_detailscarries the tool that was refused and AGY's own stderr. Check that the config passesauto_approve(the default) so the CLI runs with--dangerously-skip-permissions.
- A: AGY started, decided it needed a tool, and could not use it — most often a permission it cannot prompt for in headless mode. The rest of
📜 12. License & Contributing
- Distributed under the MIT License. See
LICENSEfor details. - Contributions, bug reports, and feature requests are welcome via GitHub Pull Requests and Issues.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi