DEV Community

mkash25
mkash25

Posted on

Atrium: a tour of the AI investigation platform I run beside Snowflake CoCo

I'm Mayank, an engineer at Snowflake supporting security and connectivity. Atrium is a tool I built for my own support casework. It's internal, so this post has no code and no company specifics: it's an architecture-level tour of how it works and why it's shaped the way it is.


Most AI coding agents are good at investigating. Point one at a problem and it will query data, read logs, and write up an answer faster than I can open the right tabs. What they're not good at is proving their work. In support, a confident answer with a gap in it is worse than a slow answer, because it reaches a customer.

Atrium is my answer to that. It's a lightweight web app that runs right next to my coding agent's terminal session and wraps the whole life of a case: intake, SLA tracking, a structured problem statement, AI insights, an adversarial review of the agent's work by a model from a different family, and a final pass on the customer reply. The agent does the investigating. Atrium frames the work going in and challenges it coming out.

Here's the tour.

Architecture

  CRM mirror ──► ┌─────────────────────────────────────────┐ ◄──► Slack
                 │ Atrium  (Python + Tornado)              │      (validation work)
                 │                                         │
                 │  1. Case workspace   queue, SLA timers  │
                 │  2. AI insights      past cases, causes │
                 │  3. Problem statement                   │
                 │  4. Adversarial review  Mode 1 + Mode 2 │ ◄──► Reviewer model
                 │  5. Draft refinement    tone options    │      (different LLM family)
                 │                                         │
                 └───────┬───────────────────▲─────────────┘
                         │                   │
      versioned feedback │                   │ CLI hooks / events
      packets (as input) ▼                   │
                 ┌─────────────────────────────────────────┐
                 │ Coding agent  (Cortex Code CLI)         │
                 │ investigates the case, drafts the reply │
                 └─────────────────────────────────────────┘

  State: a personal Snowflake database      UI: browser, live over WebSockets
Enter fullscreen mode Exit fullscreen mode

Two things matter in this picture. First, the agent never talks to the reviewer directly: Atrium sits between them, so I decide when a review runs and what goes back. Second, everything flows through the agent's own terminal. Atrium listens to the CLI's hooks and events, and it answers by injecting feedback into the CLI the same way user input arrives.

Why I built it

When an AI agent investigates a support case, its mistakes are rarely loud. They look like a claim with no query behind it, a diagnostic step it quietly skipped, or a jump from symptom to root cause that the evidence doesn't quite support. Each one reads fine on its own. That's the problem.

If I have to re-verify everything the agent did by hand, the time it saved me is spent again. So the real bottleneck in AI-assisted support isn't speed, it's trust: I need to see what the agent did and why before I let its work reach a customer. Atrium is built around making that visible, and around giving the agent a second opinion it can't talk its way past.

Keeping pace with AI investigations

There's a second problem underneath the trust problem: pace. An AI agent can now work through an investigation that used to take a couple of hours to a few days. The engineer who owns the case doesn't get faster at the same rate. They still have to understand the problem, check the agent's findings, and stand behind the answer in front of a customer.

That leaves two bad options. Slow the agent down to human speed, and you lose most of what it offers. Or skim its output and approve it, and you're back to confident answers with gaps in them.

The third option is the one Atrium is built on: give the engineer tooling that keeps pace smartly. Each stage takes a slow part of the engineer's job and makes it fast without making it shallow:

  • Understanding the case becomes reading a structured summary instead of a ticket thread.
  • Checking the work becomes reading a review from an independent model that has already re-run the queries and rebuilt the agent's path.
  • Seeing the whole investigation at once is the next piece, through visualization. More on that at the end.

The agent sets the pace. Atrium's job is to make sure the engineer can keep up with it without cutting corners.

The tour: one case, start to finish

1. Case workspace

Atrium pulls cases from a mirror of our CRM, so the queue, case history, and status live in the same window as the agent session instead of in another browser tab. SLA timers are tracked per case, so I can see what's close to a deadline before I pick the next one up.

2. AI insights

Before I start digging, Atrium surfaces four kinds of insight on the case:

  • Similar past cases, so I'm not solving something that's been solved before.
  • Likely root causes, as a starting set of hypotheses.
  • SLA risk flags, for cases that need attention now.
  • Suggested next steps, as a concrete place to start.

3. Structured problem statement

An LLM turns the raw case into a structured summary with fixed sections:

  • Symptoms and error messages
  • Environment and configuration
  • Timeline and customer impact
  • What's been tried, and current hypotheses
  • Concepts: a plain-language breakdown of the tech stack and jargon involved, so the summary is easy to read even when the case spans unfamiliar territory.

I've tested it in real casework, and it's detailed enough to guide an in-depth investigation without asking the coding agent a single question.

4. Investigation

The investigation itself happens in the coding agent's terminal session, as usual. Atrium doesn't replace the agent or wrap it in its own chat. It listens to the CLI's hooks and events, so it knows what the agent is doing as it happens.

5. Adversarial review

When the agent has a draft, I click to run the review. This is the core of Atrium, so it gets its own section below.

6. Draft refinement

The customer-facing reply gets a final pass on a dedicated refinement screen. The engineer sets tone options, so the reply sounds like the person sending it rather than like a model.

7. Validation in Slack

Some of our validation work happens in Slack. Atrium mirrors that work so it sits next to my casework, and runs a quick sanity check on it.

Why the structured summary matters

Of everything in the tour, the structured problem statement looks the least impressive. It's a summary. But it's the piece that changes how fast an engineer can engage with a case, and I've seen that in three concrete ways.

Faster ramp-up on unfamiliar cases

Support cases don't respect specialties. A connectivity case can turn out to depend on a part of the stack you rarely touch. The Concepts section exists for exactly this: it breaks down the technologies and jargon in the case in plain language, so ramping up means reading one section instead of opening a dozen documentation tabs.

Easier to spot what's missing

A ticket thread hides its gaps. A fixed structure exposes them. When every case is summarized into the same sections (symptoms, environment, timeline and impact, what's been tried), a thin or empty section stands out immediately. If the environment section is vague, that's the next question for the customer, and I know it before the investigation goes down the wrong path.

Confidence on live calls

The hardest moment in support is a real-time engagement: a customer call where you're expected to already understand the problem. Walking in with the symptoms, environment, timeline, and current hypotheses laid out means I can lead the conversation instead of reconstructing the case while the customer waits.

The summary sets the case up for the agent too: as I mentioned in the tour, it's detailed enough to guide an in-depth investigation without asking the coding agent a single question.

The adversarial review, in depth

The review has three design choices that matter more than any prompt: who reviews, what they check, and how the findings get back to the agent.

A reviewer from a different model family

The reviewers run on a different model from a different LLM family than the coding agent. A model asked to judge work in its own style tends to agree with it. This is documented: Panickssery, Bowman, and Feng (NeurIPS 2024) found that LLM evaluators score their own outputs higher than other models' outputs even when human annotators rate them as equal. A reviewer from another family is less likely to share the agent's habits and blind spots, so it's more likely to push back where it should.

The review also only runs when I click to run it. I decide when the agent's work is ready to be challenged, and every review is a deliberate step rather than background noise.

Mode 1: Draft check

Mode 1 tests the answer. It takes the customer-facing draft and cross-verifies the investigation behind it: it cross-references the case material and runs its own SQL queries, independently of the agent, to confirm or counter each point the draft makes. Where the evidence disagrees, or isn't there at all, it raises the point.

The independent queries are the important part. A reviewer that only reads the agent's transcript can check reasoning but not facts. One that goes back to the data can catch a claim the agent never actually verified.

Mode 2: Runtime rebuild

Mode 2 tests the path. It rebuilds the agent's run and checks it against three questions:

  • Was the right flow used? Did the investigation follow the approach this kind of case calls for?
  • Were key skills missed? Mode 2 double-checks both local and remote skills to see whether the agent skipped one it should have applied.
  • Are there gaps or leaps in logic? Places where the agent moved from one conclusion to the next without the evidence in between.

An answer can be correct by luck and a path can look reasonable while the answer is wrong. Running both modes covers both failures.

Feedback as a versioned packet

Findings from both modes go back to the agent as a versioned packet, injected directly into the CLI window the same way user input arrives. The agent doesn't need a special integration to receive a review. It reads the packet like any other instruction and revises its work, and I still review the final reply before it goes anywhere.

Under the hood

Why Tornado

Atrium is a Python app on Tornado, and the choice came down to what the app spends its time doing:

  • Live UI updates. The browser view stays current over WebSockets as the agent works.
  • Async event handling. CLI hooks and events arrive while other work is in flight, and Tornado's async model handles that without blocking.
  • Lightweight and fast to start. It runs beside a working agent session, so it can't be the slow part of my setup.
  • Packet injection. Writing feedback back into the CLI is one more async job alongside everything else.

State

Atrium keeps its state remotely, in a personal Snowflake database, rather than in local files.

From Vesta to Atrium

Atrium is the second generation. The first, Vesta, was a Streamlit app. It proved the ideas, but as features piled up it became heavy and was only ever usable by one person. Rebuilding on Tornado meant giving up Streamlit's convenience in exchange for an app that's fast, light, and designed from the start to sit next to a terminal session instead of being a separate destination.

How I build it

I don't build Atrium by prompting one coding agent line by line. I develop it by delegating tasks to autonomous agents that work through them and pick up new skills as they go. My job is the architecture, the review standards, and deciding what gets built next.

What's next

Visualization: the next layer of the adversarial challenge

Right now the review's findings arrive as text. Text is precise, but it still has to be read, and at the pace AI investigations move, reading is exactly the step that falls behind. Visualization is key to closing that gap, and it's the next part of the adversarial challenge I'm working on. It's at the design stage, built around two views:

  • The investigation path as a graph. Instead of reading about the route the agent took through a case, I'll see it: the steps it took, in order, laid out as a graph.
  • Claim-by-claim verdicts. Each claim in a customer draft, marked as confirmed or countered by the review, so the shape of a draft's evidence is visible at a glance.

The goal is the same as everything else in Atrium: let the engineer keep pace with the agent without giving up any of the scrutiny.

Adversarial checks on validation work

The Slack validation flow currently gets a quick sanity check. Next, I plan to run the same adversarial LLM checks on it for complex casework, so the work that comes through validation gets the two-mode review that cases already get.

Takeaways

  • Let the agent investigate; don't let it grade itself. Put review in a separate model, ideally from a different family.
  • Structure is how engineers keep pace. A fixed-format problem statement speeds up ramp-up, exposes what's missing, and prepares you for live calls.
  • Check the answer and the path separately. Verifying claims against the data and rebuilding how the agent reached them catch different failures.
  • Meet the agent where it already works. Hooks in, feedback injected as input: no custom agent, no new chat window.
  • Keep the human on the trigger. The review runs when I decide the work is ready, and nothing reaches a customer without my sign-off.

If you're building something similar, or you've found other ways to make an agent's work verifiable, I'd like to hear about it in the comments.

Top comments (0)