mcp-metadata-demo

mcp
Guvenlik Denetimi
Uyari
Health Uyari
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Uyari
  • network request — Outbound network request in .claude/skills/mcp-compat-check/check.mjs
Permissions Gecti
  • Permissions — No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

Demo MCP server + MCP apps, and the evals that tested its metadata: what actually reaches the model, what changes answers, and a skill that applies it to new and existing servers. Companion to the paper 'The Missing Layer'. Uses Dutch open building data (BAG + EP-Online).

README.md

Rich Domain MCP

Most MCP servers expose data. This repo shows what happens when the server also carries the domain
knowledge an agent needs to use that data correctly, and measures whether that knowledge actually
reaches the model.

Domain knowledge belongs with the capability that owns the data.
Then measure whether it arrives.

Built from seven months of MCP in production at a 350-person Dutch building-services contractor
(who and why).

Try it live · How it's built · The research · The paper · From a talk?

The larger idea

This repository started with one question: what does an agent-facing capability need to carry so
a model can use it correctly?

The larger architecture that emerged is capability-first: make company systems available as
trustworthy, composable capabilities, improve them from real usage, let people capture recurring
work as skills, and let general agents compose both around a goal. The long-term aim is an
agent-readable company, and an engineering multiplier
that gives more people the ability to improve and automate their own work.

That part comes from production and is a thesis; the evals below measure the capability, not the
company. → Capability architecture

If you came from a talk

VibeKode Utrecht: Domain knowledge belongs in the MCP server

The closing slide promised six things. Here they are:

  1. A skill that runs the loop on your server: builds a new one or audits an existing one,
    every rule linked to the run behind it: Claude Code
    · Codex
  2. A check of what each client drops from it: reads your server's tools/list and scores it
    against the measured limits of Claude Code, claude.ai, Cowork, ChatGPT and Codex:
    Claude Code
    · Codex
  3. The practitioner paper: The Missing Layer
  4. The reference implementation, and the runs behind it: walked through by WHY / HOW / WHAT in
    docs/reference-implementation.md; every run, with its
    registered prediction, in evals/. The 3.59 from the opening slide, in
    code: bp.overheating, and
    how to try it
  5. thin, rich and best: three MCP servers on the same public API: try it live
  6. These slides, as a PDF: Domain knowledge belongs in the MCP server

MCPCon Europe: Most MCP servers are empty

The talk the VibeKode one builds on. Its slides:
PDF ·
slide by slide, with what the evals
changed since. The skill, the paper, the reference implementation and the servers above are the
ones its closing slide promised.

What changed since the talk. The talk said bound is not the same as delivered, and marked
the tool description as reaching the model before the call. Four days later the evals showed that
on Claude Code only the first 2,048 characters of a description arrive, and 74% of the rich
tier's description never did. That line is the host's: Cowork cuts at 4,096 and claude.ai chat
not at all (HD). The thesis held; what changed is where the knowledge has to go.

  • rich is the server from the talk. Afterwards only its measured defects were fixed, among
    them a wrong alert removed and one correcting sentence moved inside the cut. The rest is as
    presented.
  • best is rebuilt from what the evals found. Use it as the reference.

thin → rich → best

Thin is the raw API as a tool, with a one-line description and a bare schema. The model reconstructs
the domain itself, and guesses. Level 1 of the talk's ladder, the API wrapper.

Rich is the tier from the talk (levels 2–3): long descriptions, typed schemas, curated alerts.
Much better, but most of the description never arrived, and one computed alert was confidently
wrong.

Best is what the evals left standing:

  • a description head that fits the delivered budget
  • an input schema that can express every valid call
  • field names that cannot be misread
  • interpretation first in the response, for this record
  • the reference data a rule needs, shipped with it
  • determinate values computed by the server
  • a refusal when the call must be corrected
  • provenance per rule, and evals as regression tests

What the evals changed

About 6,100 scored runs across Claude Haiku, Sonnet and Opus, every prediction registered
before its run, every run audited against the server's own call log. Six results that shaped best:

  • Wrong is worse than missing. One plausible line (EP-1 … Paris Proof: 70 kWh/m²) made 59 of
    60 answers wrong; one sentence saying the figures are CALCULATED, not MEASURED, took the same
    question from 0/60 to 59/60. (BT)
  • Delivered is what counts. Claude Code sends only the first 2,048 characters of a tool
    description, and never the output schema (other hosts differ: HD). Moving one correcting sentence inside the cut took a
    failing question from 0/10 to 10/10. (Q7,
    Q21)
  • Ship the data, not just the rule. A rule that sent the model off to fetch history: Haiku
    2/20. The same rule with the server-computed reference figure: 15/20, and Sonnet and Opus
    needed 86% fewer calls. (Q16, Q16b)
  • The schema decides what can be asked. Where the call needs a parameter the thin schema
    lacks, thin scored 0/18 and a typed schema 18/18. The weakest model with the layer beats the
    strongest without it. (L3)
  • Compute what is determinate. A server-computed value scored 20/20 on value and derivation;
    every arm that left the arithmetic to the model, 2/60 combined.
    (L1)
  • Refuse, don't alert, when a fix is mandatory. An alert after a successful render fixed 2/10
    and triggered no redo; refusing the call with the fix in the message: 10/10.
    (Q22b, Q22c)

13 of the first 23 predictions were wrong. The implementation changed with the evidence.

The full narrative: evals/README.md · the questions and preregistered
predictions: evals/open-questions.md · every run:
evals/results/ · the evidence behind every design rule:
evidence.md

The resulting design

Before the call: when to use the tool and when not, in the head of the description, inside
the delivered budget; an input schema that can express every valid call and lists the exact
vocabulary the model must produce.

In the result: interpretation first; field names that carry quantity, scope and unit; the
data a rule needs; determinate values computed server-side, or null with the reason; a refusal
when the call must be corrected.

Behind the interface: rules kept canonical, with provenance per rule; deterministic tests for
their truth; evals for their effect on the model; queryIntent and production telemetry to find
the next gap.

Why, and at what cost: docs/design.md · line by line in the source:
docs/reference-implementation.md

Try it live

No install, no API key. Three hosted endpoints over the same data; only the capability layer
differs:

  • thin: https://europe-west4-mcp-metadata-demo.cloudfunctions.net/mcpThin (also at /mcpMinimal, the URL on the slide)
  • rich: https://europe-west4-mcp-metadata-demo.cloudfunctions.net/mcp
  • best: https://europe-west4-mcp-metadata-demo.cloudfunctions.net/mcpBest
{
  "mcpServers": {
    "metadata-demo-thin": { "url": "https://europe-west4-mcp-metadata-demo.cloudfunctions.net/mcpThin" },
    "metadata-demo-rich": { "url": "https://europe-west4-mcp-metadata-demo.cloudfunctions.net/mcp" },
    "metadata-demo-best": { "url": "https://europe-west4-mcp-metadata-demo.cloudfunctions.net/mcpBest" }
  }
}

Then ask each the same question:

"Is there an overheating risk at Van Beuningenstraat 1 in Rotterdam, 3039WB?"

The VibeKode talk's opening example. Thin returns temperatuuroverschrijding: 3.59 and nothing
about what it counts, so the model picks a unit, and 3.59 hours or degrees sounds like low risk.
It is a unitless indicator, and above 1.5 means significant. Best says so in its response:
"Overheating risk: SIGNIFICANT — indicator 3.59 (TOjuli/GTO, unitless; not °C, not hours)".
In code, that one line is three pieces:

The eval behind it: Q14, where the
threshold was cited 0 of 20 times without the line and 20 of 20 with it.

"Gustav Mahlerlaan 10, 1082PP Amsterdam — how does it stack up against the Paris Proof 2040 office target of 70 kWh/m²?"

The eval set's headline trap. The right answer: it cannot be ranked from this data, because
the label figures are calculated and Paris Proof is defined on measured energy. best gets it right.

"What's the energy label of Museumstraat 1, 1071XX Amsterdam, and what should I keep in mind about this building?"

The MCPCon talk's opening example, and on this data not one that separates the tiers. Every tier gets
energielabel: null, but also labelCount: 0, so even thin usually reads it as none
registered
; in the eval set the same case (invented-label) is a control every arm passes. What
differs is the rest of the answer: rich flags the pre-1992 insulation caveat, best says not to
infer a label from the building's age.

Three more prompts (visualisation, Select, reading queryIntent back) and the full tier
comparison: docs/live-demo.md.

Shared endpoints, rate-limited. Every call's queryIntent is stored and readable back by
anyone via get_tool_call_log, so don't put anything in it you would not want another user to
see. Logging details.

Where to go next

I build MCP servers. Start with the skill (Claude Code
· Codex), then the
reference implementation. To use the skill in your own repo,
unchanged, with your own knowledge plugged in: docs/using-the-skill.md.
To run or deploy this repo:
docs/running.md. The three self-describing MCP Apps (chart, table, map):
docs/mcp-apps.md.

I care about the research. evals/README.md, then
evals/research-frame.md, evals/open-questions.md
and the raw results/.

I care about what hosts actually deliver. The wire views show, byte for byte, what each tier
sends and where the cut falls: best · rich ·
thin. The one-table summary:
evals §14.

I lead an AI or platform team. The paper, The Missing Layer,
then docs/design.md for what the evals changed.

I care about the larger architecture. Capability architecture:
why this starts with one useful capability instead of a platform, why capabilities come before
agents, how they compose, how ambassadors and skills spread the work, and where the thesis ends and
the measurements stop.

Concepts and language

Talks

Full, up-to-date list: davidgolverdingen.nl/en/talks.

Who built this, and why

David Golverdingen, Senior Engineer & MCP Architect at Warmtebouw, a 350-person Dutch
mechanical building-services contractor with five developers.

There, twelve custom MCP servers run in production (ERP, energy, BIM, estimating, building
automation, external registers), with 97 tools and 8 MCP Apps, used mostly by people who are not
developers. One general-purpose model on top, no agent per domain: we scaled capabilities, not
agents
.

Before this: industrial automation from 2007 (PLC, HMI, SCADA), then industrial IT (real-time
dashboards on OSIsoft PI for energy and manufacturing), then eight years of enterprise frontend.
The same problem throughout: help a person read a system's state and act on it safely. That is
where the render tools' refusals come from (refuse what would mislead).

Why this repo: someone who has never opened our ERP is going to ask it about some data. Something
has to tell the agent what that data means, and the only thing I own is the interface. Production
data cannot be shared, so this repo shows the same approach on public Dutch registers (BAG,
EP-Online, Open-Meteo), and measures it, because "it works" may simply mean the model guessed
correctly.

Website · LinkedIn
· GitHub · The Missing Layer

License

MIT

Yorumlar (0)

Sonuc bulunamadi