DEV Community

Cover image for What a Multimodal AI Workspace Actually Looks Like When It's Built Right
James Anderson
James Anderson

Posted on

What a Multimodal AI Workspace Actually Looks Like When It's Built Right

Here's the shape of a normal hour of AI-assisted work right now.

You're in a chat tool, and it gives you some code. You copy it into your editor. Then you need a quick doc to explain it — different tool, paste the context back in, regenerate. Then someone wants a slide version — a third tool, paste it again. Then a spreadsheet to track the numbers — a fourth. Then a diagram — a fifth. Somewhere in there you also pulled data from an API and a file from Drive.

Notice what you are in that workflow: you're the integration layer. You are the thing manually carrying context between a dozen AI tools that have no idea the others exist. Every handoff is a copy-paste, a re-explanation, a loss of state. The AI is doing the easy parts, and you're doing the tedious glue work between the AI parts.

This is absurd, and we've mostly stopped noticing it because it crept up on us one tool at a time.

The obvious response is "someone should build one workspace that does all of this." And lots of people are claiming to — "all-in-one AI workspace" is on a hundred landing pages. But most of them are a chat box with a few features stapled on. So the interesting question isn't whether a multimodal AI workspace should exist. It's: what are the genuinely hard problems you have to solve for one to actually work — and what does "built right" look like for each?

That's what I want to break down. Not a tour of features; the actual design challenges, and the difference between a real answer and a fake one.

Hard problem 1: routing — knowing which surface the task needs

The naive version of a multimodal workspace is a dashboard with buttons: here's the code editor, here's the doc, here's the slide canvas, here's the spreadsheet. Pick one, then talk to the AI inside it.

That's not a workspace. That's a folder of separate tools wearing a shared skin, and it pushes the hardest cognitive work back onto you: deciding which surface this task needs. The whole promise of "multimodal" is that you shouldn't have to.

The genuinely hard problem is routing — taking a plain request ("turn these notes into a presentation," "clean up this data and chart it," "write and run a script to do X") and figuring out, without a menu, which surface or surfaces the task actually calls for, then summoning them. "Make me a pitch deck" should produce slides, not a markdown wall. "Why is this query slow?" should open a code/analysis context, not a doc.

Built right: the system infers the surface from intent, and gets it right often enough that you stop thinking about surfaces at all. Built wrong: you're back to picking from a menu, which means the "intelligence" is a chat box and everything else is manual.

Hard problem 2: context continuity across surfaces

This is the one that separates a real workspace from a bundle, and it's genuinely difficult.

The entire point of putting code, docs, slides, and sheets in one place is that they share an evolving context. You draft a plan, then build the thing the plan describes, then make slides about what you built, then a spreadsheet to track it — and at no point should you have to re-explain what "the thing" is. The context from step one should still be alive at step four.

Most "all-in-one" tools fail exactly here. Each surface is its own little island with its own conversation, so moving between them means re-establishing context by hand — which is the copy-paste problem again, just relocated inside one app. You've changed the logo on the tabs; you haven't removed the glue work.

Built right: state carries across modes. The doc knows what the code does. The slides know what the doc said. You add a column to the sheet and it understands which numbers you mean, because the context didn't reset when the surface changed. Built wrong: five surfaces, five amnesiac conversations, and you ferrying between them.

Hard problem 3: the right surface, not just a surface

It's not enough to have many surfaces. Each one has to actually be good at its job.

A spreadsheet task needs real spreadsheet affordances — formulas, cells, the ability to recompute — not a static markdown table that calls itself a spreadsheet. A slide deck needs to be a deck you can present and edit, not an image of slides. Code needs to actually run. The failure mode here is ten mediocre surfaces: a toy editor, a fake spreadsheet, slides you can't really use. That's worse than ten specialized tools, because now you've got breadth and no depth.

Built right: each surface is genuinely capable at the thing it's for, so you're not quietly downgrading your tools in exchange for having them in one place. Built wrong: you traded best-in-class tools for the convenience of mediocrity in a single tab.

Hard problem 4: letting it act — safely

A workspace that can only produce text is a document generator. A workspace that can do things — run the code, send the email, change the data, call the API — is far more useful and far more dangerous, and this is where the design gets serious.

The moment the workspace can take actions with real consequences, you're in agent-safety territory, and all the unglamorous guardrails apply: least privilege on what it can reach, human approval for consequential actions (gated by blast radius, not forty rubber-stamp prompts), an audit trail of what it actually did, and ideally a separate check — a reviewer or judge step — before anything irreversible fires. A multimodal workspace that can act without these isn't a productivity feature; it's an incident waiting for a date.

Built right: capability and controllability ship together — the workspace can act, and there are brakes designed in from the start. Built wrong: "look how much it can do!" with no answer for what happens the first time it does the wrong thing.

Hard problem 5: connectors without chaos

The last hard part is the outside world. A workspace that can't reach your real data — your Drive, your repos, your calendar, your database — is a sandbox. So integrations matter. But "100+ connectors" is easy to put on a slide and hard to make coherent.

The failure mode is a junk drawer: a long list of half-working integrations, each with its own auth quirks and inconsistent behavior, most of which nobody uses and several of which are quietly broken. Breadth as a marketing number, not a capability.

Built right: the connectors that exist actually work, behave consistently, and — critically — feed into the shared context and the permission model from problems 2 and 4, rather than being bolted on beside them. Built wrong: a wall of logos and a lot of "connection failed."

What this looks like in practice (with disclosure)

A disclosed note, because it's the honest thing to do: I work on Xenition, which is an attempt to build exactly the kind of workspace I've been describing — one chat that summons the right surface (code, docs, slides, spreadsheet, an image/design canvas), carries context across them, and can take actions through connectors with approval gates and an audit trail around the agent parts.

I'm not going to walk through it as a feature tour, because that's not the point of this piece and you'd rightly tune it out. I'll just say which of the hard problems above were the actually hard ones in practice, since that's the useful part: routing and context continuity (problems 1 and 2) are where almost all the real difficulty lives. Making a system reliably infer the right surface, and keeping one context alive as you move between a doc and code and a sheet, is a genuinely unsolved-feeling problem that you chip away at rather than "finish." The surfaces themselves (problem 3) are a lot of work but a known kind of work. And the acting-safely part (problem 4) is exactly the guardrails conversation — approval gates, audit, a judge agent — which is why I'd argue controllability has to be designed in from day one, not added after the first scary demo.

I mention all this not to pitch it but because the honest version of "here's what built-right looks like" includes "here's where it's still hard," and routing-plus-continuity is the part I'd tell anyone evaluating any tool in this space to poke at hardest. If a multimodal workspace nails the demo but makes you re-explain yourself every time you switch surfaces, it hasn't solved the problem that matters.

The honest counter: when fragmentation actually wins

Now the part that keeps this from being a sales pitch, because it's true and it matters: sometimes ten specialized tools beat one workspace, and you should know when.

If you need the absolute best-in-class tool for one specific thing — the deepest spreadsheet, the most powerful IDE, a design tool a dedicated team has polished for a decade — a general workspace will almost always be shallower at that one thing than the specialist. If you're wary of lock-in, putting everything in one vendor's workspace is a bigger bet than keeping your tools separate and swappable. And if your work is genuinely single-mode — you only ever write code, full stop — the "multi" in multimodal is overhead you don't need.

A multimodal workspace is a bet on integration over specialization: you accept that each surface might be slightly less deep than the best standalone tool, in exchange for the context continuity and the elimination of glue work. That's often a great trade. It is not always a great trade, and anyone telling you it always is, is selling something.

The takeaway

The future of AI tooling probably isn't more tools. We're already drowning — the fragmentation, the copy-pasting, the being-the-integration-layer-yourself. The direction that actually helps is fewer surfaces that adapt to the task, with one context flowing through them.

But "adaptive multimodal workspace" is only as good as how it answers the hard problems: can it route to the right surface without a menu, keep one context alive across modes, make each surface genuinely capable, act with real guardrails, and connect to the outside world coherently? Those five things are the whole game. The branding is noise; those are the questions.

So whether you adopt one of these workspaces, build your own, or decide your specialized stack is better for now — judge it on those five, not on the length of the feature list. "Built right" isn't about how much it can do. It's about whether it removed the glue work without quietly making you the glue in some new way.


Genuinely curious where people land on this: are you consolidating toward one AI workspace, or deliberately keeping a stack of specialized tools — and what's the thing that actually decided it for you? I suspect the answer is mostly "which pain hurt more: the switching, or the shallowness." But I'd like to hear the real ones.

Top comments (0)