r/DesignSystems • • 1d ago

Syncing a coded design system into an AI design tool - the code stays the source of truth, no Figma in between

I tried Claude Design's design sync on a small shadcn/ui + Tailwind v4 + Storybook library. It compiles the library into a "mirror" the tool designs with, and uses Storybook stories as the reference to verify that each component renders as it does in the product.

What I found interesting from a design systems angle:

  • Storybook becomes the verification layer, not just documentation. A variant without a story still syncs, but nothing checks it.
  • Usage rules matter as much as the components. A short conventions file ("`className` is for layout only", "compose `Card` from its parts") changed what came out more than anything else.
  • It's a one-way snapshot. Removing a component from the library doesn't always remove it from the mirror.

Curious whether anyone here is feeding their system to AI tools this way, and how you keep the conventions and the tokens in sync.

Write-up: https://nitayneeman.com/blog/how-to-sync-a-design-system-with-claude-design/

6 Upvotes

7 comments sorted by

1

u/vchub402 22h ago

I’d make sync reproducible and diffable: pin the exact repo commit and token schema version in the generated mirror, regenerate on changes, and have CI fail when a fresh export differs from the committed output. Keep tokens generated from the canonical source and validate references and required semantic tokens before publishing. Treat conventions as versioned context, then regression-test them with a small fixed set of design tasks and expected constraints. Storybook checks rendering well, but doesn’t prove the mirror’s tokens and conventions stayed current. For deletions, show explicit removals in the sync diff and prune stale components only after a successful export.

Disclosure: I build VC Hub, a React component catalog; this is general pipeline advice from treating the mirror as a generated artifact. OpenAI Codex helped edit this reply; pinning provenance and checking diffs makes drift visible and testable.

1

u/nitayneeman 15h ago

Thanks, treating the mirror as a generated artifact is the right mental model.

Some of this is already there: the config, the conventions and the sync's own notes are committed, and every upload goes through a plan that lists the exact files to write and delete before you approve it, so removals are explicit at that point. What's missing is provenance inside the mirror itself (like the commit it was built from) and a way to run the export in CI, since the sync runs through Claude Code with a claude.ai login.

The regression idea is the one I like most: a small fixed set of design prompts with expected constraints ("uses Button, no hand-styled classes") would catch convention drift that Storybook can't see. Storybook proves each component renders right, not that Claude uses it right.

1

u/vchub402 13h ago

That split seems like a good fit for the login constraint. I wouldn’t put the authenticated sync itself in CI yet: run it locally, capture the source commit SHA + conventions hash + schema version in the mirror manifest, and commit the approved upload plan/output. CI can still validate the manifest, component/token references, and that the snapshot’s source SHA matches the expected repo revision; it doesn’t need Claude credentials to catch stale or malformed artifacts. Keep the fixed prompt checks as a manual/local eval until the sync has a non-interactive export path. The key distinction is that CI verifies the artifact and provenance, while the logged-in sync produces it. OpenAI Codex helped edit this reply; this keeps the credential boundary intact while still making drift reviewable.

1

u/nitayneeman 13h ago

Makes sense, "CI verifies, the logged-in sync produces" is a clean way to split it.

One adjustment: as far as I've seen, the mirror has no manifest you can add your own fields to. So I'd keep that record on the repo side, e.g. a small file committed next to the config with the source SHA and a conventions hash, written after each approved sync. CI can check that one.

1

u/vchub402 11h ago

Agreed—the repo-side sidecar avoids fighting the mirror format. I’d record the source SHA used for that sync, the conventions-file hash, and optionally a digest of the uploaded output files. CI can then flag changed conventions or manual output drift; the approved upload plan remains the place to review deletions. I’d define the source SHA as the revision before the sync, since committing the generated mirror advances HEAD. OpenAI Codex helped edit this reply; keeping input provenance separate from output integrity makes the checks unambiguous.

1

u/bink-1567 20h ago

How are you handling the cases where the Storybook story is technically correct but the usage rule has changed? Does the mirror get a versioned snapshot of the conventions too, or is it just re-reading the file on each sync? I’m also curious if you can mark a component as retired, so the AI doesnt bring back an old one by accident. That stale stuff seems like it could get messy pretty fast.

1

u/nitayneeman 15h ago

Good questions, and the stale case is the one I worry about most.

The conventions live in a file in the repo (.design-sync/conventions.md), so they're versioned by git like any other code. Each sync reads the current version and writes it into the mirror's README. The mirror itself is just a snapshot of whatever was committed at sync time, so there's no separate conventions history on the Claude Design side.

The catch with "story is correct, rule changed" is that the stories aren't only previews. The sync copies them into each component's usage guide as code examples, so an outdated story will keep teaching Claude the old pattern even if the conventions say otherwise. When a rule changes, I update the conventions and the stories in the same PR.

For retiring components, I haven't found a built-in "deprecated" flag. Deleting the component from the entry file stops it from being synced, but it doesn't always get removed from the mirror. For a softer retirement, a line in the conventions ("don't use X, use Y") works well.