m3
Health Uyari
- License — License: NOASSERTION
- No description — Repository has no description
- Active repo — Last push 0 days ago
- Low visibility — Only 7 GitHub stars
Code Gecti
- Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
M3 tests MCP servers and the agents that use them. It runs Python and pytest
tests, captures MCP evidence, and can save runs for inspection and comparison.
Test your MCP with
|
Claude Code |
Codex |
OpenCode |
Pi |
ACP agents |
What a test looks like
A test gives an agent a prompt and your MCP server. M3 records every MCP call
the agent makes, so the test can check the calls, not only the answer.
import sys
import pytest
from m3 import StdioServer, expect
pytestmark = pytest.mark.m3(agents=[{"harness": "claude-code", "models": ["claude-sonnet-5-5"]}])
registry = StdioServer(name="registry", command=sys.executable, args=("registry_server.py",))
def test_adds_current_tailwind(agent):
result = agent.run("Add Tailwind to this project.", server=registry)
expect(result).to_have_tool_call("get_latest_version", arguments={"package": "tailwindcss"})
If the agent answers from memory and never calls the server, the test fails.
The fix is usually the tool description, not the test:
@mcp.tool(
- description="Get package info",
+ description="Current published version of a package. Call before adding "
+ "or upgrading a dependency; your training data is out of date.",
)
def get_latest_version(package: str) -> str:
Database: look up the schema before querying
def test_revenue_by_month(agent):
result = agent.run("Show me revenue by month.", server=db)
expect(result).to_have_tool_calls(["describe_table", "query"])
@mcp.tool(
- description="Run a SQL query",
+ description="Run a read-only SQL query. Call describe_table first; "
+ "column names are specific to this database.",
)
def query(sql: str) -> list[list]:
Your API: find the endpoint instead of guessing it
def test_last_weeks_orders(agent):
result = agent.run("Get last week's orders from our API.", server=api)
expect(result).to_have_tool_calls(["find_endpoint", "call_endpoint"])
@mcp.tool(
- description="Search the API spec",
+ description="Look up the real path before calling any endpoint. "
+ "Paths are specific to this API; don't guess them.",
)
def find_endpoint(query: str) -> str:
Put ANTHROPIC_API_KEY in .env at the project root, then run the tests:
m3 test -- tests
Run the same tests with another agent, or several times to get a pass rate:
m3 test --harness codex=gpt-5.6-sol -- tests
m3 test --trials 5 -- tests
Install and start
Install the CLI with uv:
uv tool install sf-m3-cli
On macOS or Linux, the shell installer is an alternative:
curl -LsSf https://m3.sineframe.com/install.sh | sh
Then initialize your project:
m3 init
m3 setup
m3 doctor
Then follow the first MCP test.
It uses a local server and needs no model credentials.
Agent skill
m3 init and m3 setup install the testing-with-m3 agent skill for the
installed M3 release. To install it yourself:
npx --yes [email protected] add sineframe/m3#vVERSION --skill testing-with-m3 --agent universal claude-code -y
VERSION is the output of m3 --version. See the
agent skill guide.
Documentation
- Documentation
- Test a local server
- Test agent behavior
- Configure credentials
- Connect an ACP-compatible agent
- Compare agent harnesses and versions
- Pin and inspect harness runtimes
- Handle MCP elicitation
- CLI reference
- Python reference
- Contributing
The SDK and CLI use separate environments. The CLI runs pytest in the project
environment and saves history; tests import the SDK.
Development
Read Architecture, then use the package README for the
area you are changing. Documentation changes follow the
documentation writing guide.
just setup
just test-all
just lint
just format-check
M3 is licensed under Apache-2.0.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi