Soundcheck - Generative Design System

Project: A generative commerce engine that lets ChatGPT, Claude and Gemini compose SiriusXM purchase and subscription screens inside the conversation, without ever writing a price, an entitlement or a payment form.

  • Role: Design Lead, Delivery Lead

  • Methods: Production system, failure-first review, host fidelity study, AI-assisted test harness, live-model evaluation

  • Outcome: Created a system to deflect contact center calls by providing rich account services within browser agents.

Soundcheck Design System

The agent designs the journey. It never closes the sale.

People are starting to shop where they already are: in a chat. For SiriusXM, that means someone asking ChatGPT whether their used car has satellite radio, comparing plans, and subscribing without leaving the conversation.

That created a problem our design system was never built for. A component library assumes a person is choosing the components. Here the one assembling each screen is a model, mid-conversation, with no sense of which values are sacred or which buttons move money.

Asking the model nicely wasn't enough. Even with instructions to defer every figure to the card, a model stated $25.99 a month for a plan that didn't exist. So I set out to design a system whose main user can be trusted with structure, but not with judgment.

Design goals

  1. Let the model compose, never author money. It arranges layout and copy; every price comes from the catalog.

  2. Fail closed. If anything the model composes breaks a rule, the person sees a safe default layout, never a broken or unchecked screen.

  3. One card, every host. The same card renders in ChatGPT, Claude and Gemini, fitting each app without losing SiriusXM's brand.

  4. Rules that can't drift. Every rule is a test, so the design system can't quietly stop being true.

Research process & Iterative Developement

I built the prototype first and treated it as the research instrument: every assumption became something I could break on purpose.

  • Failure-first review. I planted the same wrong price in three places: the model's layout, the model's reply and the backend's own data. The engine rejected the first, flagged the second, and couldn't catch the third. That last result set the project's honest limit.

  • Host fidelity study. I matched each chat app's real interface from screenshots, not memory: Claude's warm ground and reading column, ChatGPT's app card, Gemini's Material dark. A frame that is only chat-shaped turns a prototype into a mockup.

  • AI-assisted test harness. AI generated the cases: a seeded fuzzer that plants one violation per layout, and a live model composing layouts from a contract generated from the vocabulary. The harness runs every flow step × 6 session states × 4 layout sources, 1,344 cells in all, through the real engine. To prove the tests themselves worked, I switched rules off and watched them fail.

Key solutions

Risk decides the tier. The vocabulary has three tiers. Tier 1 is 15 display components the model arranges freely. Tier 2 is 7 controls, each wired to one of 19 fixed intents: no custom code, no arbitrary links. Tier 3 is certified: payment and confirmation. The model can request one, and the engine discards whatever it supplied and inserts the audited component.

Prices are references, not values. The model writes offer.promoPrice, never $25.99. Every screen passes validation and reference checks before it renders; any failure shows a deterministic fallback, which is tested too.

At the money moment, the agent steps out. Commits bypass the model entirely. Each purchase carries a single-use mandate tied to what was on screen, and a downgrade asks for a second identity check first. The permission prompt above it all belongs to the chat app, because a connector that could draw its own consent screen could grant itself consent.

Same card, every host. The card borrows four things from the host: surface, hairline, corner radius and typeface. It never borrows the accent. A purchase button in another brand's colour is a purchase that doesn't look like one, so a test fails the build if the card ever reads a host colour.

Tokens flow from Figma to the chat. Designers edit variables in Figma; Tokens Studio syncs them with GitHub, where every change is a pull request. The build compiles them to CSS and checks branding and 4.5:1 contrast on every host surface, in light and dark, before the MCP server ships the card.


Impact

For customers: answers shaped to their own car, usage and plan, in the conversation they started. Prices they can trust, and a clear, safe screen even when the model gets something wrong.

For the business:

  • A path to chat commerce that billing, legal and brand can sign off on, because the guarantees are tests rather than promises.

  • One engine and one card across three chat apps, instead of three integrations.

  • A design system that works as a live rulebook for AI, with its own browsable docs that run the same checks as production.

The honest limit: provenance, not truth. The engine proves the price on screen is the price the system holds, not that the catalog is right. And it governs the card, not the sentences around it: a scan flags a wrong price in the model's reply, but can't block it.

What's next: agent commerce is shifting. ChatGPT retired in-chat checkout in March 2026 and moved payment back to merchants' own pages. Because payment here is a certified slot, not a screen, it can be filled per host: a wallet where the chat app offers one, SiriusXM's checkout or app everywhere else. The first step is the used-car flow: check compatibility in chat, then hand off to the SiriusXM app to sign up and activate.