open-design-system-mcp-starter

An MCP server that stops agents inventing your components.

Coding agents invent component names and props because nothing in your repo can answer them. This template turns your design system into an MCP server they can ask, and generates the AGENTS.md section and llms.txt that go with it. Point it at your system and it runs in about an afternoon.

MIT · Node 22+ · no build step · 9 tools · 7 adapters · 252 offline tests

Without it, from memory

// the agent writes what most libraries call it
<Card>
  <Heading>Storage</Heading>
</Card>

Nothing in the repo told it otherwise, so it used the vocabulary of the libraries it saw most in training. The import fails at build time, long after the agent moved on.

With the server attached

// the agent asks first
resolve_component({ name: "Card" })

You wrote 'Card'; this system calls it acme-surface. There is no Heading component either.

// so it writes this instead
<acme-surface>
  Storage
</acme-surface>

When a name has no match at all, the answer says so and lists the nearest real components. It will not supply a substitute that merely looks right.

Both answers are real output from the example system bundled with the template. You can reproduce them a minute after cloning, before you configure anything of your own.

The same invented names come back across models and systems.

These counts come from 898 graded generations by six model configurations against two production design systems. The template ships that mined list and uses it: when an agent reaches for one of these names, the server can tell what the model meant and answer with the name your system uses.

Props invented that did not exist

as170
spacing85
color49
tone32
variant32
icon20
label18
fullWidth17

Components imported that did not exist

Heading29
Typography29
Box26
Card25
TextField21
cn17
Switch16

Each of these is a question the agent had no way to ask. Written instructions help, but they have to be in front of the model before the work starts, and they take up context whether or not the task needs them.

Each one answers a question that comes up mid-task.

Before you write anything

search_components
What do I use for a dismissible notice? Ranked matches by name, your team's aliases, the mined vocabulary, props, description and docs.
get_pattern
A labelled field with an error? A complete recipe your team wrote, composed from real components with the accessibility wiring done.
get_guidance
When should I not use a modal? The passages of your own docs that answer it, with the file they came from.

While you write it

resolve_component
Does TextField exist here? Exact, alias with the real name, or missing with the nearest candidates and the reason. Never a fabricated entry.
get_component
What are Button's real props? Typed props with defaults and descriptions, the import or tag line, events, slots, accessibility rules.
find_token
Which token is #1a1a1a? Nearest by colour distance with the theme named, exact or nearest length, or a word match. Returns what to write.
list_tokens
Show me the spacing scale. One category at a time, in a compact table, with per-theme values.

Before you call it done

check_usage
Is this file valid? Unknown components, invented props, invalid values, raw hex codes, missing accessible names, disallowed imports.
get_migration
This prop is deprecated, now what? The replacement, the version, the changelog lines that mention it, and any codemod you ship.

The same data is also exposed as MCP resources for clients that attach context up front, plus a build-ui prompt that loads the routine: search, resolve, read the props, use a token, check the file, never invent.

Read your system once, then serve the committed result.

Adapters are the only part that knows file formats. Everything downstream reads one committed, stamped data layer, so a system with an unusual setup needs one adapter and nothing else changes.

ConfigOne entry per kit in ds.config.json: where the components are, which adapters to use, where the docs live.
AdaptersReact source or a published package, custom elements manifests, CSS custom properties, design token JSON, markdown, Figma Code Connect.
Data layerCatalog, tokens, docs index, aliases, patterns. Committed so it reviews like code, hashed so staleness is visible.
Server and surfaceTools over stdio or HTTP, plus a generated AGENTS.md section, llms.txt, skills and editor rules from the same data.

The surface is generated from the same catalog the server reads, so your agent instructions cannot drift from the components you ship. Re-run one command after a release and everything is current again.

From clone to a grounded agent.

The template runs on its own example system the moment you install it, so you can see the behaviour before you configure anything. Then pick how your team consumes the design system.

git clone https://github.com/christophhdesign/open-design-system-mcp-starter.git
cd open-design-system-mcp-starter
npm install
npm run serve

That already works: it serves the fictional acme-elements example. Open the folder in Claude Code and the bundled .mcp.json registers it for you. Ask the agent whether the system has a Card and watch it get corrected.

1

Install your packages here

You do not need a checkout of the design system. The wizard reads the installed package's own declaration files.

npm install -D @your-scope/components @your-scope/foundations @types/react
2

Point the template at them

Detects the barrel, the tokens stylesheet and the READMEs, then writes the config for you.

npx tsx src/cli.ts init --id mysystem \
  --package @your-scope/components \
  --foundations @your-scope/foundations
3

Extract and verify

Reads the packages once and writes the committed data layer, then checks it.

npx tsx src/cli.ts extract
npx tsx src/cli.ts doctor

Read what doctor prints. It tells you how many components came back with a typed props table, how many tokens it found, and whether anything has gone stale. Those counts are how much of your API the server can check for you.

4

Teach the agents your vocabulary

Map the names people reach for to the names you use. It is a few lines of JSON and it prevents more wrong guesses than any other file here.

// data/mysystem/aliases.json
{ "components": [{ "alias": "Switch", "target": "Toggle" }],
  "props":      [{ "alias": "spacing", "target": "gap" }] }
5

Generate the surface into the repos where people build

An AGENTS.md section, llms.txt, a Cursor rule, a Copilot section and two skills, all built from the catalog you just extracted. Re-run it after every release.

npx tsx src/cli.ts generate all --system mysystem --target ../your-app
6

Register it, or share one instance

Per developer over stdio, or one HTTP instance the whole team points at.

claude mcp add mysystem -- npx tsx /abs/path/src/cli.ts serve --config /abs/path/ds.config.json
npx tsx src/cli.ts serve --http --port 3333
claude mcp add --transport http mysystem http://your-host:3333/mcp

Set DS_MCP_TOKEN before exposing the instance beyond localhost, and put TLS in front of it. GET /healthz reports counts, freshness and live sessions.

The parts that are easy to get wrong.

Never fabricate

A miss returns "not in this system", the nearest real names, and the reason. It will not hand back something that merely looks right, which is the failure this whole tool exists to prevent.

Refuse stale ground truth

Every artifact carries a hash of what it was built from. Doctor tells you when a release moved past your data, and the server can refuse to start on it.

Small answers

Each tool has a character budget and logs the size of what it returned, so you can see what grounding your agents costs in context.

Nothing hardcoded

No component, package or framework name lives in the source. If your system does not fit, the fix is a config field or a new adapter, and you can still pull updates.

Your files stay yours

Generated sections sit between markers and hand-written files are never overwritten without an explicit flag.

Light on dependencies

Three at runtime and no build step. The test suite runs offline and never calls a model.

Find out what your agents have been guessing.

The example system answers in a minute. Your own takes about an afternoon, and the first alias you write will tell you a lot about the gap between what your team calls things and what everyone else does.

Get on GitHub