Skip to content
AI

Figma MCP for design work

What the Figma MCP server actually does, the fixed read order, how to slice big frames, why Code Connect is the lever, the design-to-code loop, and an operating model with checklists and a scorecard.

stellae.design

19 min read

TL;DR

  • Figma MCP is not a code generator. It is a machine interface to what has actually been designed: components, variables, layout, and the mapping to production code. The agent still writes the code, in your codebase, under your rules.
  • The read tools have a fixed order: context, then metadata when the context is too big, then a screenshot as the visual baseline. Big frames get sliced by node, never fetched whole.
  • Code Connect is the lever. In Figma's own evaluation it cut tokens by about a third and improved code quality by a full point. Without it the agent knows what a button looks like, not which button to import.
  • Handoff is becoming a loop. Code to canvas and write to canvas send the built thing back into Figma, so the design catches up with reality instead of drifting from it.
  • The quality ceiling is set by the file, not the model. Components, variables, semantic names, Auto Layout and annotations decide what the agent can see.
  • Rate limits, truncation and a couple of open bugs are real. Plan for them, log the tool calls, and never assume the agent received something just because it exists in the file.

First, the name

MCP is the Model Context Protocol, the standard interface through which an agent reaches tools and external context. Figma's server implements it. It is not a plugin format, and it is not "Figma AI". When someone says Figma MCP they mean the server that sits between a Figma file and a client like Claude Code, Cursor, Codex or VS Code, and answers questions about the design in a form a model can use.

That framing matters because the most common mistake is to treat it as a compiler: select a frame, ask for the page, ship it. Figma's own implementation skill says the opposite. The design context the server returns looks like React and Tailwind because that shape reads well for a model, and the agent is expected to translate it into the project's framework, tokens and components. What comes back is evidence about the design. The code is still yours to own.

If MCP itself is new to you, read MCP for designers first: hosts, servers, resources, tools, and why structured context beats a screenshot. If you have not read our Figma to code workflow guide, start there for the human side of handoff. This guide covers what changes when an agent sits in the middle.

What the server actually does

Figma's tools reference is the source of truth and changes often. The shape as of writing:

ToolDirectionWhat it gives you
get_design_contextreadThe structured representation of a selection: layout, styles, component instances, variables. The default entry point.
get_metadatareadA lightweight tree of node ids and names. The map you use when the context is too large.
get_screenshotreadA rendered image of the selection. The visual baseline for validation.
get_variable_defsreadThe variables and styles used in a selection: colour, spacing, type. Useful for drift checks against code tokens.
Code Connect lookupsreadWhich production component implements a Figma component, and how its props map.
generate_figma_designwriteCode to canvas: captures a running interface and turns it into editable Figma layers.
use_figmawriteWrite to canvas: creates and edits native Figma content using your components, variables and Auto Layout. Beta.

Two things to notice. The read tools are cheap and rate-limited. The write tools are new, exempt from most limits, and still changing. Figma's code to canvas and write to canvas pages list which clients support each. At the time of writing, code to canvas is limited to Claude Code and Codex.

There are two ways to run the server. The remote server at mcp.figma.com authenticates with OAuth and has the broadest feature set. The desktop server runs inside the Figma app and is the only place get_variable_defs works. Figma's setup guide covers each client, and several clients now get a Figma plugin that bundles the MCP configuration with an agent skill. Follow those pages rather than a blog post. Setup instructions are the part of this topic that goes stale fastest, and this guide is not going to repeat them.

The read order, and why it is fixed

Figma's own agent skill prescribes a sequence, and it is worth internalising because every failure mode below is a deviation from it.

  1. Context first. Call get_design_context on the selection. This is the structured truth about layout and instances.
  2. Metadata when it is too big. If the context is truncated or enormous, call get_metadata, pick the child nodes that matter, and fetch each one's context separately.
  3. Screenshot as the baseline. Call get_screenshot on the whole selection. Keep it. Everything you build gets compared against it.
  4. Assets after that. Icons and images come last, and only the ones the context references.
  5. Translate, do not paste. Replace generated styling with project tokens. Replace generic elements with existing components. Keep the project's routing, state and data conventions.
  6. Validate against the screenshot. Compare, fix, compare again.

The order exists because step one can fail loudly. Figma's client notes show a real get_design_context response of over 350,000 tokens against a client output limit of 25,000. Raising the limit is a patch. Decomposition is the fix, and it has a second benefit: a slice is small enough to verify on its own.

Big frames get sliced

A production agent should never receive "build this dashboard from the Figma page" as one job. The method that works:

Large frame ├── get_metadata → node map ├── get_screenshot → whole-page baseline, kept ├── identify logical slices → shell, sidebar, content, cards, empty state │ ├── get_design_context → per slice │ ├── implement + test → per slice │ └── compare to baseline → per slice └── integrate → final visual check against the baseline

Slices follow the product's structure, not the layer tree. Navigation shell, one feature section, a modal, the empty and error states. Each slice is a unit of work with its own context call, its own tests and its own comparison. The whole-page screenshot is the only thing that spans them.

This is also how you stay inside the rate limits. Read tools are capped per day and per minute by plan, and Starter plans and View seats get almost nothing. A pipeline that crawls hundreds of nodes on every run will hit the ceiling by lunch. Fetch the slices you are building, cache derived artefacts like token snapshots in the repo, and reserve live calls for information that only the file can give you.

Code Connect is the lever

Without Code Connect, the agent can see that a node is an instance of a component called Button, and it can see the fills and the padding. It still does not know that the correct implementation is @company/ds/Button, which variants exist, or that Figma's state=loading maps to a loading prop. So it rebuilds a button. Then another. Then a third that is almost the same.

Code Connect records the mapping between a Figma component and its production implementation, and the MCP server includes those mappings in the design context. Figma measured the effect in a controlled evaluation across 27 design-to-code tasks: with Code Connect, median token use fell 29.5%, task duration fell 19.6%, and code quality rose a full point on a four-point scale. The driver was coverage. Their well-connected design system had Code Connect snippets in about a fifth of MCP responses; a patchily connected one had them in about a sixteenth, and got proportionally less.

Coinbase ran the same comparison on their own system after bringing coverage up to date, and reported lower token use, faster implementation and, more importantly, agents picking real composite components like a Stepper instead of approximating them from primitives.

Treat these as vendor and customer evidence, not universal benchmarks. But the direction is consistent and the mechanism is obvious: the more the agent can import, the less it has to invent.

Practical order for a team starting from zero:

  • Map the primitives people use on every screen first: Button, Input, Select, Modal, Toast, Tabs, Card, navigation, icons.
  • Weight coverage by usage, not by count. Ninety percent of components mapped means nothing if the missing ten percent is on every page.
  • Prefer Figma's template-based Code Connect files. Recent releases added migration tooling from the older parser-based mappings.
  • Review mappings in pull requests like code. A wrong mapping teaches every future session the wrong import.

Handoff is becoming a loop

The read tools only move information one way. The two write tools change the shape of the work.

Code to canvas takes a running interface, in the browser or a preview, and turns it into editable Figma layers. Figma's help article and workflow lab show the use: an engineer builds a state the design never had, sends it back, and the designer edits the real thing instead of redrawing it from a screenshot. Early passes still need cleanup around Auto Layout and colour mapping, which Figma says openly.

Write to canvas lets the agent create native Figma content directly, using your components, variables and Auto Layout. It is beta, the roadmap points at parity with the Plugin API, and Figma has said it expects the capability to become paid, usage-based functionality at some point. Use it for the things that are tedious by hand: a missing error state, a variant set from a spec, a FigJam flow from a ticket.

Put together, the workflow that Figma now documents runs in a circle:

ticket → FigJam flow → Figma frames → MCP context + Code Connect → implementation ↑ │ └── designer review ← code to canvas / use_figma ← preview build ←──┘

The consequence for the team is that "the design is the source of truth" stops being a rule and becomes a synchronisation problem. The design is authoritative for intent. The code is authoritative for behaviour. The loop exists to keep them from disagreeing for long. A token drift check that compares get_variable_defs output with the code's token file is the smallest useful version of that loop, and it can run in CI.

The file sets the ceiling

Figma's recommendations for files that work well with MCP are the same recommendations a good design-system team makes anyway. They stop being hygiene and become the input format.

  • Components for anything repeated. A group that looks like a card is a group.
  • Variables for colour, spacing, radius and type. A hex value in a fill is a value the agent has to guess a token for.
  • Semantic names. CardContainer, not Group 5. Names travel into the context and become the agent's vocabulary.
  • Auto Layout to express responsive intent. Absolute positions say nothing about what happens at a narrower width.
  • Annotations for behaviour the visuals cannot carry: alignment rules, transitions, what happens on error.
  • Code Connect on the components that matter.

The difference between a clean library and an ad hoc canvas is not a few percent. It is the difference between an agent reusing a Stepper and an agent hand-drawing one out of rectangles. MCP exposes structure. It cannot manufacture structure that is not there.

What is not in the frame

A frame carries a fraction of what a screen has to do. Before an agent touches implementation, the missing parts have to be named, because a model will not infer them and a screenshot will not show them.

  • Loading, empty, error, offline, permission and validation states
  • Long and translated content, text zoom, wrapping and truncation
  • Keyboard order, focus, screen-reader names, reduced motion
  • What actually changes at each breakpoint, not just that a breakpoint exists
  • Business rules, analytics, data contracts, latency

Figma's own workflow example starts by writing the screens, states and transitions down in FigJam before generating anything, and catches a missing failure state on the way back through code to canvas. Do the same. A state inventory at the top of the ticket is the cheapest quality gate in the whole process.

Where DESIGN.md fits

A DESIGN.md file, in the Google Labs format, is a static, portable description of a visual identity: tokens in YAML front matter, rationale in prose. We covered how to write one in DESIGN.md: a design system your agent can read. It is not a Figma MCP feature, and it solves a different problem.

Atlassian ran the comparison that matters. In their test, a login screen built from DESIGN.md alone needed roughly 92% more tokens than the same task with their design-system MCP, took longer, and varied more between runs. The reason is structural: a file has to be loaded whole, an MCP server answers on demand. Their file, even after heavy trimming, was around 80 KB, and agents given only the file were more likely to rebuild a component than import the existing one.

So the division of labour is:

SourceAuthoritative for
Figma libraryWhat has been designed: components, variables, layout
Code repositoryWhat ships: implementation, types, conventions
Code ConnectThe mapping between the two
Agent skills and project rulesHow this team wants the work done
DESIGN.mdPortable visual intent, for contexts where the richer tooling is unavailable
Tests, linters, CIWhat is allowed to merge

DESIGN.md earns its place for art direction, quick prototypes in unfamiliar environments, and interoperability. It is the wrong tool for teaching an agent a 50-component production system. Keep it small, point it at the canonical components, and let MCP carry the detail.

Failure modes to test before you standardise

None of these are reasons to wait. All of them are reasons to run a pilot with logging on.

  • Truncated context. The 350,000-token response above. Symptom: the agent builds the top of the page and hallucinates the rest. Fix: metadata first, slices.
  • Silent omissions. Forum reports of annotations not arriving through the expected calls, and of Code Connect snippets vanishing after a component becomes a variant set. Symptom: a detail exists in the file and never reaches the code. Fix: inspect the tool output when a detail matters, and log every call.
  • Wrong framework mapping. Where both React and Compose mappings exist, the wrong one has been returned. Fix: pin the framework in the project rules and check the first import.
  • Auth and client state. Expired tokens in Cursor are a documented cause of failed calls. Symptom: everything stops working at once. Fix: re-authenticate before you debug anything else.
  • Rate limits. Symptom: a pipeline that worked in the morning fails in the afternoon. Fix: the slicing and caching above, and an explicit budget per task.
  • Third-party servers. The community figma-developer-mcp package had a high-severity command injection in versions before 0.6.3, exploitable through crafted text in a Figma file. This was not Figma's server. The lesson is general: an MCP server has tool privileges on the developer's machine, so it belongs in the same dependency and security review as any other executable.

The operating model

For a team of any size, this is the sequence that has held up. Each step is a prerequisite for the next, which is why teams that skip to step four get disappointing results.

  1. Make the library legible. Normalise names, variants, variables and Auto Layout. Delete or deprecate duplicate components rather than teaching the agent five plausible buttons.
  2. Map the top components with Code Connect. By usage frequency. Ship the first ten, measure, then continue.
  3. Write the project rules. Which packages hold the primitives, which layout abstractions to use, how tokens are referenced, what may never be hardcoded, which validation commands must run. Figma supports custom skills and rules for exactly this. Our How to write a good skill post applies directly.
  4. Pilot on representative work. One component-heavy page, one complex responsive page, one stateful flow, one page with legacy components, one component that does not exist yet. Not a button.
  5. Put the engineering contract in front of the output. Typecheck, lint, unit and component tests, accessibility checks, visual regression, responsive end-to-end, performance budget, preview build, then human review. An agent-authored page passes the same suite a human-authored page would. Storybook's MCP server lets the agent run component and accessibility tests on its own output and fix what fails, which closes the loop on the code side. Playwright's screenshot assertions handle visual regression.
  6. Close the design side of the loop. When implementation reveals a missing state, send it back with code to canvas. When the token file and get_variable_defs disagree, fail the build.
  7. Measure before you automate more. The scorecard below. Expand coverage where the numbers say it pays.

Responsibilities should be explicit, because the loop touches everyone:

ConcernOwner
Intent, states, interaction modelDesigner
File structure, components, responsive intentDesigner, to the design-system team's standard
Semantic tokensDesign-system team, jointly with engineering
Code Connect mappingsDesign-system team, engineering maintains
Project rules and skillsFrontend engineering
MCP installation and security reviewPlatform
Tests and visual acceptanceEngineering builds, designer accepts

Checklists

Before the first agent session

  • Components for every repeated element, no duplicates
  • Variables for colour, spacing, radius, type
  • Semantic layer names on everything the agent will touch
  • Auto Layout on every responsive container
  • Annotations for behaviour the visuals cannot show
  • Code Connect on the ten most-used components
  • Project rules name the primitives, the tokens and the validation commands
  • State inventory written for the ticket

Per implementation

  • Context, then metadata if truncated, then screenshot
  • Sliced by product structure for anything bigger than a component
  • Existing components imported, none rebuilt
  • Tokens referenced, no hardcoded values
  • Compared against the screenshot at 1440 and 390 wide
  • Full test suite green, including accessibility and visual regression
  • Deviations from the design written down, missing states sent back to Figma

Before standardising

  • Pilot ran on the five representative tasks
  • Tool calls logged, truncations and auth failures counted
  • Rate budget per task known and inside plan limits
  • Third-party MCP servers reviewed like any dependency
  • Scorecard baseline recorded

The scorecard

"The screenshots look closer" is an observation, not a result. Measure the things that decide whether the work is faster once review is included.

DimensionMetric
Component reuseShare of eligible UI rendered with approved components
Correct selectionShare of mapped Figma instances resolved to the intended code component
Token adherenceShare of style references using semantic tokens
Visual fidelityMaterial discrepancies against the screenshot per page
Functional correctnessAcceptance-test pass rate
AccessibilityAutomated violations plus manual high-severity findings
Responsive correctnessPass rate across the agreed viewport and content matrix
PerformanceCore Web Vitals deltas against budget
Agent efficiencyTokens, tool calls, elapsed time, cost per accepted change
Human interventionMaterial corrections per implementation
ReworkReview rounds per change
Context failuresTruncations, wrong calls, auth failures per task
DriftOpen Figma-to-code token and component discrepancies, and their age
CoverageCode Connect coverage weighted by component usage

Run paired comparisons, not anecdotes: the same task, model and repository snapshot under codebase only, codebase plus screenshot, plus MCP, plus Code Connect, plus project skill. Reset between runs. Run each condition more than once, because the output is not deterministic. A tool that saves three minutes of generation and costs twenty-five minutes of review is not faster.

The point

Figma MCP is infrastructure for giving an agent evidence about the design and for sending reality back into the file. It rewards the teams that had already made their decisions explicit: components, variables, names, mappings, rules, tests. It does very little for a team hoping the model will make those decisions for them.

The interesting shift is not from design to generated code. It is from design intent that lives in people's heads to design intent that a machine can read at several levels: the canvas through components and variables, the design-to-code relationship through Code Connect, the process through skills, the identity through a small DESIGN.md, and the acceptance criteria through tests. Make those legible, and the agent stops guessing.

Skills to copy

Three skills that encode the practices above, one phase each. Each is a folder with a SKILL.md; drop it into ~/.claude/skills/<name>/ or .claude/skills/<name>/ and invoke it by name, or let the description trigger it. They assume Figma's MCP server is already configured in the client. Read our post on writing skills that change the output before you edit them.

figma-implement

Builds one Figma selection in the project's codebase, in the fixed read order, with slicing and Code Connect.

markdown
---
name: figma-implement
description: >-
  Implement a Figma selection in this codebase through the Figma MCP server.
  Load when given a Figma link or node id and asked to build, port or update
  UI from it. Not for reviewing a finished implementation or checking token
  drift; use figma-drift-check for that.
---

# Figma implement

The MCP response is evidence about the design, not code to paste. This skill
reads in a fixed order, slices anything larger than a component, imports
before it builds, and validates against a screenshot before it reports.

## Before reading

1. Confirm the selection: file, node id, and which frame or component. Ask
   if ambiguous.
2. Write the state inventory the frame will not show: loading, empty, error,
   permission, long content, keyboard, reduced motion. List them in the
   reply before touching code. The user strikes what does not apply.
3. Read the project rules for UI primitives, token file, layout abstractions
   and validation commands. If none exist, ask where the primitives live
   before proceeding.

## Read order

1. `get_design_context` on the selection.
2. If the response is truncated or larger than the client limit, call
   `get_metadata`, choose child nodes by product structure (shell, section,
   modal, state), and call `get_design_context` per node. Never raise the
   client output limit as the first fix.
3. `get_screenshot` on the whole selection. Keep the path; it is the
   baseline for every comparison.
4. Assets only after the context, and only the ones the context references.

## Translate

- Every Code Connect snippet in the response is an import, not a suggestion.
  Use the mapped component and its props. Never rebuild a mapped component.
- Every colour, spacing, radius and type value maps to a project token. A
  raw hex or pixel value in the output is a defect. If no token matches, say
  so and use the nearest, flagged in the report.
- Generic elements in the response (`div` with padding and border) become
  the project's existing primitives where one exists. Search the primitives
  folder before creating anything.
- Keep the project's routing, state and data conventions. The response has
  none; do not invent them.
- Auto Layout in the context describes responsive intent. Implement the
  intent with the project's layout abstractions, not absolute positions.

## Validate

1. Run the validation commands from the project rules: typecheck, lint,
   tests.
2. Render at 1440 and 390 wide, both themes if the project has them.
3. Compare each slice against the screenshot. Fix, compare again. Stop when
   the remaining differences are ones the user should decide on.
4. Confirm every item in the state inventory has an implementation or an
   explicit deferral.

## Report

- Which components were imported via Code Connect, which primitives were
  reused, and anything new that was created, with the reason.
- Values that had no token, as a list.
- Deviations from the screenshot and why.
- States implemented and states deferred.
- Tool calls made, including any truncation or error.

## Never

- Never fetch a whole page in one `get_design_context` call when
  `get_metadata` shows more than a handful of top-level children.
- Never create a component that Code Connect maps to an existing one.
- Never hardcode a value that has a token.
- Never report "matches the design" without a screenshot comparison at two
  widths.
- Never assume a detail reached you because it exists in the file. If an
  annotation or mapping seems missing, say so and ask.

figma-drift-check

Compares the variables a Figma selection uses with the project's token file and reports the differences. Read-only.

markdown
---
name: figma-drift-check
description: >-
  Compare Figma variables to the codebase's design tokens and report drift.
  Load when asked whether design and code tokens still agree, after a
  palette or spacing change, or before a release. Read-only; never edits
  tokens. Not for implementing UI; use figma-implement.
---

# Figma drift check

Design and code drift quietly. A variable gets renamed in Figma, a token
gets tuned in code, and three months later the two disagree on the primary
colour. This skill makes the disagreement visible. It changes nothing.

## Run

1. Ask for the Figma selection to check. A library page or a representative
   screen is enough; the whole file is not necessary.
2. Call `get_variable_defs` on the selection. This tool needs the desktop
   MCP server; if it is unavailable, say so and stop.
3. Locate the code token source: a Tailwind `@theme` block, a tokens JSON,
   CSS custom properties, or a DTCG file. Ask if there is more than one
   candidate.
4. Normalise both sides: names to kebab-case, colours to a single space
   (OKLCH or sRGB hex, not mixed), dimensions to px.
5. Match by name first, then by value for the unmatched remainder.

## Classify

| Class | Meaning |
| --- | --- |
| Match | Same name, same value within tolerance |
| Value drift | Same name, different value |
| Name drift | Same value, different name |
| Figma only | Variable with no code token |
| Code only | Token with no Figma variable |

Tolerance: colours within 1 OKLCH lightness step or 1/255 per sRGB channel;
dimensions exact.

## Report

One table per class, sorted by how many components use the variable if usage
data is available, otherwise alphabetical. Then three counts: matched,
drifted, unmatched. Then one sentence on which side moved, if the git
history or Figma version history makes it obvious.

Value drift on a text or background colour is a contrast risk. Say so and
recommend a contrast audit.

## Never

- Never edit the token file or the Figma variables. Report only.
- Never treat a name-only difference as a value difference.
- Never guess which side is correct. Design is authoritative for intent,
  code for what shipped; the user decides.
- Never run this across an entire large file in one call. Rate limits apply
  to every read.

figma-roundtrip

Sends what was built back into Figma so the design catches up with reality.

markdown
---
name: figma-roundtrip
description: >-
  Send implemented UI back into Figma with code to canvas, or add a missing
  state with write to canvas, so design and code stop diverging. Load after
  an implementation reveals a state, variant or layout the design did not
  have. Not for initial implementation; use figma-implement.
---

# Figma roundtrip

Implementation discovers things the frame did not know: an error state, a
longer label, a layout that had to change at 390 wide. Leaving those only in
code guarantees the next design iteration starts from a stale frame. This
skill puts them back.

## Decide which tool

| Situation | Tool |
| --- | --- |
| Built UI in a browser, designer should edit it | `generate_figma_design` |
| Missing state or variant, buildable from DS components | `use_figma` |
| The change is a value, not a structure | Neither. Report it. |

Check the client supports the tool. Code to canvas is limited to specific
clients; write to canvas is beta. If unsupported, produce a written handoff
note instead and say why.

## Code to canvas

1. Render the state in the browser at the width the design used. One state
   per capture.
2. Call `generate_figma_design` and place the result next to the original
   frame, never over it.
3. Name the new frame `<original name> / <state> (from code, <date>)`.
4. Inspect the result: Auto Layout may be flattened and colours may come
   back as raw values. List what needs cleanup in the report; do not
   silently leave it.

## Write to canvas

1. Confirm the target: which component set, which page.
2. Use only components and variables that already exist in the file. If the
   state needs a colour or size that has no variable, stop and report it
   rather than creating an off-system value.
3. Name the variant by the same convention as its siblings.
4. Take a `get_screenshot` of the result and include it in the report.

## Report

- What was sent back, where it sits in the file, and the frame names.
- Cleanup the designer should expect, per frame.
- Anything that could not be represented and why.
- The state inventory with each item marked designed, built, or sent back.

## Never

- Never overwrite or move an existing frame. New frames beside, always.
- Never write off-system values into the file.
- Never call the roundtrip done without a designer named to review it.

Sources

AIai

Was this article helpful?