Conversational AI interfaces share a fundamental design limitation: the text bubble.
When you ask an AI model a question, it streams back markdown paragraphs. For answering factual questions, explaining a concept, or drafting an email, streaming text works wonderfully.
For complex enterprise workflows, raw text quickly falls apart:
- If an agent analyzes a complex configuration with hundreds of lines, dumping raw JSON or YAML into the chat window forces the user into endless vertical scrolling.
- If an agent generates a multi-year resource cost breakdown, a markdown table is static, hard to read on smaller screens, and cannot be filtered or sorted.
- If an agent needs user confirmation to delete a database or deploy an infrastructure service, asking the user to type "yes" in text is ambiguous and error-prone.
To solve this, we expanded our Model Context Protocol (MCP) architecture beyond simple text tools to support MCP Apps: interactive, sandboxed micro-frontends embedded directly inside the conversational timeline.
Here is the idea, how it performed in production, and what to watch out for.
The Idea: Interactive Tool Micro-Frontends
In the standard Model Context Protocol, when an agent executes a tool, the tool returns plain text, markdown, or a static image.
With MCP Apps, tools can return structured UI metadata pointing to self-contained web components. When the conversational client receives this metadata, it renders an interactive visual component directly inline with the conversation:
- Interactive Data Trees: Instead of dumping hundreds of lines of raw text, the agent renders an interactive tree inspector that allows users to expand nodes, search keys, and copy specific subsections.
- Dynamic Metric Charts & Meters: Rather than describing numbers in text, tools render interactive charts where users can hover over data points, filter metrics, and toggle series.
- Structured Form Wizards: When configuring a complex resource, the agent emits an interactive form widget. The user inputs selections via dropdowns and checkboxes, submitting the completed payload directly back into the agent's context.
- Human-in-the-Loop Approval Cards: For destructive or high-impact actions, the agent displays a clear visual diff card with physical "Approve" and "Reject" buttons, ensuring explicit human confirmation before execution.
How It Worked Well
- Dramatic Usability Improvements: Users could explore dense technical datasets in seconds using interactive filtering and collapsible trees, rather than scrolling through endless walls of text.
- Strict Security Sandboxing: Running third-party UI widgets inside a core enterprise web application presents serious security risks (Cross-Site Scripting, cookie theft). We isolated all widgets inside sandboxed browser frames configured with zero access to the host application's origin, local storage, or session cookies.
- Structured Bi-Directional RPC: Host applications and sandboxed widgets communicate strictly through schema-validated cross-frame messaging. When a user interacts with a widget (such as clicking an approval button or submitting a form), the validated event is sent back to the agent as a clean tool response.
- Pluggable Tool Catalog: Microservice teams could ship their own custom UI widgets alongside their MCP tool definitions without modifying the host chat client's codebase.
What to Watch Out For
- Same-Origin Security Traps: When embedding interactive iframes, never enable both script execution and same-origin access in the sandbox flags. Allowing same-origin access permits the embedded frame to access the parent window's authentication tokens and cookies. Keep widgets in a strictly isolated opaque origin.
- Mobile Responsiveness: Desktop-first charts and wide forms can break layouts on mobile screens or embedded side panels. Ensure all interactive components are designed with fluid width and touch-friendly controls.
- State Desynchronization on Chat Reload: If a user refreshes their browser or reopens an archived conversation thread, what happens to interactive widgets? Ensure your chat client persists both the initial widget properties and the user's completed interactions so historical threads render accurately when reloaded.
- Latency and Bundle Overhead: If every MCP tool downloads megabytes of heavy JavaScript visualization libraries, the chat interface will become sluggish. Provide a shared, lightweight component design system in the sandbox base frame so individual widgets download only minimal business logic.
Top comments (1)
tr.ee/dev-to