grom

mcp
Security Audit
Fail
Health Warn
  • License — License: NOASSERTION
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Fail
  • fs module — File system access in package.json
  • network request — Outbound network request in package.json
  • child_process — Shell command execution capability in src/builtin-tools.ts
  • exec() — Shell command execution in src/builtin-tools.ts
  • network request — Outbound network request in src/builtin-tools.ts
  • network request — Outbound network request in src/client.ts
  • exec() — Shell command execution in src/composer.ts
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

BYOLLM. AI coding assistant for VS Code that runs on your machine — agentic, private, zero friction.

README.md

Grom

Grom

Documentation · Marketplace · GitHub

BYOLLM

Bring Your Own LLM. Your model. Your machine. Your rules.


Grom is a local-first agent harness and AI coding assistant for VS Code. Bring your own LLM. Your code never leaves your machine.

No cloud. No account. No telemetry. No upsell.


Why Grom?

Other tools say they support local models. Try it and you'll find yourself three config screens deep, staring at a broken connection, wondering why the cloud path is suspiciously smooth.

Grom was built because local AI shouldn't require fighting your tools.

If Ollama is running, Grom works. That's the whole deal.


Meet Grom

Grom is the little robot who lives in your sidebar. He watches your cursor, thinks while you type, and goes to sleep when you're idle. He's not a feature He's the heart of the tool.

When Grom goes idle he does two useful things quietly in the background: for Ollama users, he sends a keep-alive ping so your model stays loaded in VRAM and is ready instantly when you return. And if the file you have open has errors, he shows the count in his thought bubble, a gentle nudge that doesn't interrupt your flow.

His antenna tells you what's happening before you read a word:

State What it means
PLAN Gold, antenna bent PLAN mode, thinking broadly
BUILD Blue, antenna straight BUILD mode, focused, ready to ship
Generating Antenna bouncing (animated) Model is generating a response
Disconnected Grey, unplugged Your server isn't running or isn't connected
Error Glitching Server returned an error
Idle Eyes closed, thought bubble (animated) Idle, keeps model in VRAM; shows error count if active file has issues
Float PLAN Gold, riding a cloud Floating window, PLAN mode detached from the sidebar
Float BUILD Blue, riding a cloud Floating window, BUILD mode detached from the sidebar
Mic PLAN Gold, mic icon Voice recording active, PLAN mode listening
Mic BUILD Blue, mic icon Voice recording active, BUILD mode listening

Two Modes, One Purpose

PLAN mode: warm honey gold. Grom thinks architecturally. Break down problems, plan features, talk through ideas before a single line is written.

BUILD mode: focused blue. Grom is direct and implementation-ready. Write code, fix bugs, ship things.

⚡ Tools: a toggle in the toolbar, only available in BUILD mode. Off by default so plain chat stays fast. When off, Grom skips codebase indexing and RAG queries entirely, so replies are instant regardless of workspace size. When on it fills solid with the accent colour, no ambiguity. Turn it on when you want Grom to read and write files, run terminal commands, call MCP tools, and search your codebase for context. Each session remembers its own Tools state.

Reasoning effort: a kettlebell icon appears in the toolbar when your active model has reliable thinking control. Click to adjust how hard the model thinks before answering. The icon only appears where Grom can actually enforce the level. Models where it would be guesswork show no icon at all.


Cloud models (claude-3-7-sonnet, claude-opus-4, OpenAI o-series, Gemini 2.5+) have four levels:

Level Icon What Grom sends
Off Off No thinking budget, direct answer
Low Low Small budget, brief reasoning
Medium Medium Balanced budget
High High Large budget, full extended reasoning

Maps directly to the provider's native API: output_config.effort with adaptive thinking for Claude 4.6+, thinking.budget_tokens for Claude 3.7 and 4.5, and reasoning_effort for OpenAI and Gemini. The provider enforces the level.


Qwen3 has two levels:

Level Icon What Grom sends
Off Off /no_think token, thinking suppressed, fast reply
High High /think token, full thinking enabled

Qwen3 exposes native control tokens that hard-switch thinking on or off. You will see a clear speed difference between Off and High.


Other local reasoning models (DeepSeek-R1, QwQ, Magistral, phi-4-reasoning, and others): no icon shown.

These models have no reliable mechanism to control how much they think. Grom could send a system prompt hint, but whether the model honours it is unpredictable and untestable. A control that might do nothing isn't worth showing. The kettlebell icon is hidden for these models. If you want them to think less or more, ask directly in your message.


Set a default with grom.reasoningEffort or use /effort off|high in chat to change it mid-session. To disable the feature entirely, set grom.showReasoningToggle: false.

Model not recognised as a reasoning model? Grom maintains a built-in list of known reasoning models but new ones appear faster than any list can keep up. If your model supports thinking but isn't flagged as a reasoning model, add a substring of its name to grom.reasoningModels in your settings:

"grom.reasoningModels": ["my-r1-finetune", "local-thinker"]

Any model whose name contains one of those strings will be treated as a reasoning model: it gets the reasoning capability icon and the "may not support thinking" warning is suppressed. The kettlebell icon still only appears for models Grom can actually control (Claude, OpenAI o-series, Gemini 2.5+, Qwen3), so listing an R1 fine-tune here will not add it. Matching is case-insensitive and strips provider prefixes (e.g. lmstudio/ or ollama/) before checking.

The UI colour shifts with the mode. So does Grom's personality.


What Grom Does

Providers

Grom works with local servers out of the box with no account or key. Cloud providers are supported optionally if you want to bring your own key.

Built-in:

Provider Notes
Ollama Local, 127.0.0.1:11434, recommended
LM Studio Local, 127.0.0.1:1234
Open Code api.opencode.ai, requires API key
OpenAI GPT-4o, o1, o3-mini
Anthropic Claude Sonnet, Claude Opus
Groq Llama 3, Mixtral, fast inference
Mistral Mistral Large, Small, Codestral
Gemini Gemini 2.5 Pro, Flash, Google AI API

Custom providers: add any OpenAI-compatible endpoint or Anthropic-compatible proxy via grom.customProviders. Gemini, OpenRouter, Together AI, and most other cloud APIs work out of the box. See Adding a Custom Provider below.

Switch providers and models without leaving the panel. Each session remembers its own model. Switching sessions restores it automatically. Grom detects model capabilities automatically: vision, tool use, and reasoning models each show their own icon. Vision models (llava, qwen2-vl, etc.) can receive images via the + button or paste.

Capability detection reads directly from your provider where possible (chat template inspection for Ollama, explicit capability fields from LM Studio), so the icons reflect what your model can actually do rather than guessing from its name. New models always arrive faster than any built-in list, so Grom gives you two ways to fix detection yourself without waiting for an update:

  • grom.reasoningModels: add model name substrings that should be treated as reasoning models. Simpler when you just need the model flagged as a reasoning model.
  • grom.modelCapabilities: per-model override for any combination of tools, vision, and reasoning. Use this when detection is wrong on more than one capability, or when you need to disable a capability the auto-detection got wrong.
"grom.modelCapabilities": {
  "my-vision-model": { "vision": true },
  "broken-tools-model": { "tools": false },
  "my-r1-finetune": { "reasoning": true }
}

Keys are partial model names matched case-insensitively. Values win over all auto-detection.

If you find a model that should be in the built-in list, PRs to src/model-caps.ts are welcome. That's the single file that owns all capability detection. Adding a keyword or pattern there means everyone benefits.

Knows What You're Working On

Use @ in any message to attach context:

Mention What it includes
@selection Currently selected text in the active editor
@filename Any workspace file, open tabs shown first
@problems All current VS Code errors and warnings
@git Your current uncommitted diff (git diff HEAD)
@terminal Recent output from the integrated terminal
@url:https://... Fetches a web page and includes its text
@docs Searches all indexed documentation sources (grom.docSources)
@docs:name Searches a specific doc source by name

Auto-context is on by default. Grom reads the file you have open automatically.

Inline Autocomplete

Ghost-text completions as you type, powered by FIM models.

  • Adaptive debounce: speeds up when you're accepting, slows down when you're not
  • Word-by-word accept: Tab accepts the next word; keep pressing for more
  • Dedicated model: set a fast FIM model (e.g. qwen2.5-coder:1.5b) separate from your chat model
  • Per-language routing: different models for different languages via grom.languageModels
  • Toggle: click ✦ Grom in the status bar to enable/disable instantly

Inline Edit

Select code, press Ctrl+Shift+I, describe what you want. Grom rewrites it and opens a diff. Accept or Reject.

Compose: Multi-file Edit

Press Ctrl+Shift+O or type /compose. Describe changes across your codebase. Review per-file or apply everything at once. Undo the whole run with one click.

Every code block in compose format gets a 💾 Save button that opens a syntax-highlighted diff showing exactly what will change.

Agentic Loop

Enable ⚡ Tools in the toolbar (BUILD mode only) and Grom doesn't just reply once. It works through tasks step by step, calling tools based on what the last one returned.

Tool What it does
read_file Read any file in your workspace
write_file Write or create a file, then open it in the editor
list_directory List files and folders at a path
delete_file Delete a file
search_files Search workspace files by regex pattern
run_terminal Run a shell command and return its output
browse_web Fetch a live web page and return its text content

Note on model size: Tool call accuracy scales with model size. 32B+ local models call tools reliably. Smaller models (1.5B–7B) occasionally write prose instead of a tool call. Grom handles this by re-prompting once and enabling structured JSON mode after the first tool use. For complex agentic tasks, 14B+ is significantly more reliable. Capable Ollama models (qwen2.5, llama3.1, mistral) also support native structured tool calls, Grom detects this automatically and uses it when available.

When the agent reaches its round limit (grom.agentMaxIterations, default 20), Grom tells you in the chat so you know to review and continue manually rather than wondering why it stopped.

Documentation Sources

Index web documentation so you can reference it with @docs in any message. Grom crawls the URLs you configure, strips the HTML, and builds a searchable index — no copy-pasting docs into context. Any HTTP/HTTPS URL works, including local dev servers (http://localhost:3000/docs).

"grom.docSources": [
  { "name": "react",   "url": "https://react.dev/reference" },
  { "name": "mdn",     "url": "https://developer.mozilla.org/en-US/docs/Web/API" },
  { "name": "mylib",   "url": "http://localhost:3000/docs" }
]

Each source needs a name (short identifier) and a url (root page to start crawling). Grom follows links within the configured path only, up to 40 pages per source. Indexing runs at startup and whenever grom.docSources changes.

Note: Grom fetches pages directly and does not run JavaScript. Sites that render content client-side (pure SPAs) will return little or no usable text. Use a server-side rendered URL, a static export, or a local dev server instead.

Once indexed, use @docs to search all sources or @docs:name to target one:

@docs how do I use useEffect?
@docs:react suspense boundaries

Voice Input

Speak your prompts instead of typing them. Grom captures audio locally and transcribes it on-device using Whisper. Nothing is ever sent to a server.

  • Mic button in the toolbar (enable in Settings → Voice Input on first use)
  • Push-to-talk: click to start recording, click again to stop and transcribe
  • Six Whisper models: from Tiny EN (~40 MB, fast) up to Small (~244 MB, best accuracy), downloaded on demand. English-only .en variants are faster and more accurate for English speakers
  • Model pre-warming: the selected model loads silently when Grom starts so the first utterance transcribes without delay
  • No account, no cloud: audio stays on your machine. The mic is optional and designed for those who want or need voice input as an accessibility tool.
  • Requires a one-time ffmpeg download (~50 MB), managed entirely within VS Code. Remove it anytime from Settings → Voice Input.

Platform support: Windows (DirectShow), macOS (avfoundation), Linux (PulseAudio/PipeWire, ALSA fallback).

Floating Panel

Pop Grom out of the sidebar into a standalone window — useful for multi-monitor setups where you want Grom on a second screen while keeping your file tree visible.

  • Click the expand arrows button in the header to detach Grom into its own window
  • The sidebar becomes a passive mirror: the "floating" pill and a banner confirm it's active
  • The floating window is fully functional: chat, voice input, model switching, all modes work
  • Grom's icon changes to the cloud variant in the floating window so you always know which panel is live
  • The mic changes Grom's face to the listening variant while recording
  • Close the floating window at any time. The sidebar restores automatically
  • The floating panel persists across VS Code restarts. It reopens exactly where you left it

The sidebar input is disabled while floating is active. Click Close floating in the banner to return to the sidebar.

MCP Tool Use

Connect any Model Context Protocol server and Grom's model can call its tools during chat. Tool calls stream live with a badge showing which tool is running. MCP tools are available as soon as the servers connect — no restart or provider switch needed. Configure via grom.mcpServers.

Grom Memory

Persistent memory injected into every new chat — like custom instructions, but yours.

Only use TypeScript.
Never push code directly, always explain changes first.
My stack is React 18 + Express.

Open it with the brain icon in the header.

Conversations

  • Multiple chat sessions, persistent across restarts
  • Session dates: the history list shows a relative last-modified timestamp (e.g. "5m ago", "yesterday") next to each session
  • Prompt history: up/down arrow in the input cycles your previously sent messages
  • /compact trims long histories, and a divider marks exactly where the cut was made
  • Export any conversation as .md, import it back to continue
  • Search through any conversation with live highlighting
  • Per-session system prompt override via the chat bubble icon
  • Context window indicator: the radial circle in the toolbar shows token usage; hover for exact counts and estimated cost (when pricing is configured in grom.modelPricing)

Custom Prompt Files

Create .grom/*.md files in your workspace. A file at .grom/deploy.md becomes /deploy, shareable with your whole team via git.


Slash Commands

Type / to open the command menu:

Command What it does
/explain Explain the active file
/refactor Refactor for clarity and best practices
/fix Find and fix bugs
/tests Write unit tests
/docs Write documentation
/review Full code review
/commit Draft a commit message from your actual git diff, no @git needed
/compose Multi-file edit mode
/effort <level> Set reasoning effort for the current session: off, low, medium, high
/search <query> Web search via DuckDuckGo
/<name> Any .grom/<name>.md file in your workspace

Keyboard Shortcuts

Shortcut Action
Ctrl+Shift+G / Cmd+Shift+G Open Grom
Ctrl+Shift+I / Cmd+Shift+I Inline edit (requires selection)
Ctrl+Shift+Y / Cmd+Shift+Y Accept inline diff
Ctrl+Shift+U / Cmd+Shift+U Reject inline diff
Ctrl+Shift+O / Cmd+Shift+O Open Compose mode
Ctrl+Shift+M / Cmd+Shift+M Toggle voice recording (when mic is enabled)
Enter Send message
Shift+Enter New line

Requirements

Grom works with local servers (no account or key needed) or cloud providers (bring your own key).

Editors:

Grom runs in VS Code and any VS Code-compatible editor:

Local (runs entirely on your machine):

  • Ollama: recommended, free, runs most open models
  • LM Studio: great UI for managing models

Cloud (optional, requires an API key from each provider):

Recommended local models:

Use Model
Chat qwen2.5-coder:32b, deepseek-coder-v2, llama3.1
Autocomplete qwen2.5-coder:1.5b, deepseek-coder:1.3b, starcoder2:3b
Embeddings (RAG) nomic-embed-text, mxbai-embed-large

Settings

Setting Description Default
grom.apiUrl Your local server URL http://127.0.0.1:11434
grom.model Chat model name qwen2.5-coder
grom.useOllamaFormat Use Ollama's chat format true
grom.autocomplete Enable inline completions true
grom.autocompleteModel Dedicated FIM model (chat model)
grom.languageModels Per-language model overrides {}
grom.ragEnabled Enable codebase indexing true
grom.embeddingModel Embedding model for semantic RAG, works with Ollama and LM Studio (blank)
grom.mcpServers MCP server definitions []
grom.customProviders Custom provider endpoints; keys stored securely in OS keychain []
grom.robotAnimations Enable Grom's animations true
grom.theme UI theme: Grom, Cyberpunk, Classic, High Contrast Grom
grom.agentEnabled Master switch, disables tools globally when off true
grom.agentMaxIterations Max tool-call rounds per task 20
grom.fontSize Chat panel font size: small, medium, large medium
grom.debugLogging Write diagnostics to the Grom Output channel false
grom.hints Show in-chat hint cards (e.g. context window full warning) true
grom.modelCapabilities Per-model capability overrides for tools, vision, and reasoning. Keys are partial model names (case-insensitive). Wins over all auto-detection, use to fix wrong detections or force-enable/disable a capability. {}
grom.modelPricing Per-model token pricing and context window size overrides {}
grom.docSources Documentation sources available via @docs []
grom.presets Custom prompt presets shown in the / menu []
grom.chatLanguageModels Per-language chat model overrides (takes priority over grom.languageModels) {}
grom.autocompleteLanguageModels Per-language autocomplete model overrides (takes priority over grom.languageModels) {}
grom.customGreeting Override the greeting shown in the empty chat state (blank)
grom.customLogo Override the chat logo, URL, data: URI, or emoji (blank)
grom.reasoningEffort Default reasoning effort for new sessions: off, low, medium, high off
grom.showReasoningToggle Show the reasoning effort (kettlebell) icon in the toolbar (hidden when false) true
grom.reasoningModels Extra model name substrings to treat as reasoning models, extend the built-in list for new or custom models. Example: ["my-r1-finetune"] []
grom.voiceInput Enable the mic button in the toolbar false
grom.voiceModel Whisper model: tiny.en, tiny, base.en, base, small.en, small tiny.en
grom.voiceSensitivity Mic energy gate (RMS threshold). Raise if phantom transcriptions appear; lower for quiet mics 0.010
grom.ffmpegPath Path to a custom ffmpeg binary (skips the built-in download) (blank)

The codebase index updates automatically as files change. If search results ever look stale, type /reindex (or pick Reindex from the / menu) to force a full rebuild.

Per-Language Model Routing

grom.languageModels sets a model for a language in both chat and autocomplete. Use grom.chatLanguageModels or grom.autocompleteLanguageModels to set them independently. These take priority when set.

{
  "python": "qwen2.5-coder:1.5b",
  "typescript": "qwen2.5-coder:32b",
  "rust": "deepseek-coder-v2"
}

Adding a Custom Provider

OpenAI and Anthropic are built-in. Select them from the provider dropdown. Use grom.customProviders for everything else.

API keys are never stored in settings files. Grom prompts for a key the first time you select a provider that needs one, then stores it securely in the OS keychain (Windows Credential Manager / macOS Keychain / libsecret on Linux). Click the lock icon next to the provider dropdown at any time to update or clear a key.

For custom providers you can also set an apiKey field directly in grom.customProviders. That's convenient if you manage settings declaratively, but keep that file out of source control.

[
  { "name": "OpenRouter", "url": "https://openrouter.ai/api" },
  { "name": "Together",   "url": "https://api.together.xyz" },
  { "name": "Local (no key)", "url": "http://127.0.0.1:8080", "authType": "none" },
  { "name": "Claude proxy",   "url": "https://my-proxy.example.com", "providerFormat": "anthropic" },
  { "name": "My Gemini",  "url": "https://generativelanguage.googleapis.com/v1beta/openai", "apiKey": "AIza..." }
]

For most cloud providers, name and url are all you need. Optional fields:

Field Values Default When to set
providerFormat openai, anthropic openai Only for a self-hosted Claude-compatible proxy
authType bearer, x-api-key, none bearer Set to none for keyless local servers
useOllamaFormat true, false false Only for servers using Ollama's /api/chat format
apiKey string (none) API key for cloud endpoints not in the built-in provider list

MCP Servers

[
  {
    "name": "filesystem",
    "command": "npx",
    "args": ["-y", "@modelcontextprotocol/server-filesystem", "/path/to/workspace"]
  }
]

Context Window

The radial circle in the toolbar shows how full your context window is. Hover it to see the exact token count and window size. When it fills up, /compact trims old messages. Grom marks the cut point so you always know what's been removed.

Grom reads the context window size directly from your model, including architecture-specific keys for Gemma, Mistral, Phi, and other non-Llama families, so the circle and auto-compact threshold are accurate without any manual configuration. On Ollama, it prefers the size your model is actually running with (respecting any custom context-length setting you've configured) over the model's raw architectural maximum, so the number reflects what's genuinely available. Set grom.modelPricing to pin a size for a specific model if needed.


License

PolyForm Shield 1.0.0, free to use for any purpose, including commercially. You may not redistribute it as a competing product.


Built in Ireland. Shipped with care.

Reviews (0)

No results found