DEV Community

Cover image for Why I Don't Let My AI Agent Make Meeting Commitments
BALIJA CHAMAKURA SAKETH
BALIJA CHAMAKURA SAKETH

Posted on

Why I Don't Let My AI Agent Make Meeting Commitments

The hardest part of building my meeting agent wasn't extracting commitments from a transcript. It was deciding when an extracted commitment was allowed to become real.
A meeting contains suggestions, tentative promises, unfinished thoughts, and statements that only make sense in context. If an AI turns all of that into authoritative tasks, even a good model can produce a bad system.
So I designed my workflow around a simple rule: the agent can remember, reason, propose, and prepare—but it does not get to approve commitments on my behalf.

The System I Built


The system is a meeting intelligence application that brings together meeting preparation, contextual memory, post-meeting analysis, evidence, commitments, and follow-up workflows.
The backend uses FastAPI services and PostgreSQL. An LLM analyzes meeting notes and transcripts, while Hindsight provides persistent contextual memory. The web application exposes briefings, attendee information, evidence, intelligence, commitments, and follow-up conversations.
The interesting part isn't any individual component. It's the boundary between them.
I didn't want this:
Meeting transcript
↓
LLM
↓
"David committed to fixing the API"
↓
Confirmed commitment

Instead, the workflow is:
Meeting data
↓
Context retrieval
↓
LLM analysis
↓
Proposed commitment
↓
Supporting evidence
↓
Human review
↓
Explicit confirmation
↓
Authoritative state

That extra transition is the core design decision.
Where Hindsight Fits
I use Hindsight on GitHub as the system's contextual memory layer.
Its job is to remember useful information and retrieve relevant historical context when the agent needs it. The repository wraps Hindsight behind an adapter that handles memory-bank creation, retention, recall, health checks, and deletion.
The two operations I care about most are essentially:

await hindsight.retain_memory(
bank_id=bank_id,
content=content,
context=context,
document_id=document_id,
tags=tags,
)
memories = await hindsight.recall_memories(
bank_id=bank_id,
query=query,
tags=tags,
)

This gives the agent continuity.
A later meeting doesn't have to start from zero. Previous decisions, meeting context, attendee information, and other useful history can become available when they're relevant.
The Hindsight documentation describes the underlying memory approach, and Vectorize's explanation of agent memory provides the broader context for why persistent memory matters.
But I keep one distinction very clear:
Hindsight remembers. It doesn't approve.
A recalled memory can influence the context given to the agent, but it doesn't automatically create a new commitment.
That separation matters because old information can be incomplete or stale. Memory should improve reasoning without silently changing the application's authoritative state.

The Commitment Pipeline Has a Hard Stop


After a meeting, the outcome service sends notes and transcript information through the LLM analysis layer.
The model can return structured outcomes, including proposed commitments.
Those commitments are then inserted into PostgreSQL with is_confirmed explicitly set to FALSE.
The important part looks like this:
INSERT INTO public.commitments (
user_id, meeting_id, project_id, owner_name,
description, due_date, status, is_confirmed,
source_excerpt
)
VALUES (
:user_id, :meeting_id, :project_id, :owner_name,
:description, :due_date, 'pending', FALSE,
:source_excerpt
)

That gives me a simple invariant:
Anything produced by automated meeting analysis starts as a proposal.

The model might suggest:
Owner: David
Task: Review API integration
Due: Friday

But the database records:
status = pending
is_confirmed = false

The model's interpretation hasn't become application truth.

Evidence Comes With the Proposal


I also wanted every inferred commitment to remain connected to its source.
The commitment record stores a source_excerpt.
That changes the review process.
Without evidence, I have to ask:
"Do I trust the model?"

With evidence, I can ask:
"Does the source actually support this conclusion?"

Consider:
"I'll review the API integration by Friday."

A model can reasonably turn that into a proposed commitment.
Now consider:
"Could you review the API integration by Friday?"

"I'll see if I have time."

There are still a person, an action, and a deadline in the conversation. A language model may still identify the structure as a possible commitment.
But the evidence tells me that the second conversation is much more ambiguous.
That's why I don't need the model to perfectly solve every ambiguity.
I need it to produce a proposal that I can inspect.

Confirmation Is a Separate Operation


Confirmation isn't a side effect of generating the meeting summary.
It's an explicit operation.
The service exposes a dedicated confirmation path that changes the commitment from unconfirmed to confirmed:
UPDATE public.commitmentsSET is_confirmed = TRUE, updated_at = NOW()WHERE id = :c_id AND user_id = :u_id

So the state transition is straightforward:
LLM inference
↓
Proposed commitment
is_confirmed = FALSE
↓
Human review
↓
Explicit confirmation
↓
is_confirmed = TRUE

There isn't a hidden confidence threshold that automatically confirms a commitment.
There isn't a prompt telling the model to "be careful" and hoping that is enough.
The authorization boundary exists in the application state itself.
That's much easier to reason about and test.

The Same Boundary Applies to Follow-Up Messages

I used the same principle for generated follow-up communication.
The agent can produce a follow-up message after analyzing a meeting, but the application stores that message as a draft.
Conceptually:
LLM
↓
Generated message
↓
Draft
↓
Human review
↓
Send

Generating communication and authorizing communication are different operations.
The agent can handle the tedious part without silently taking the final action.

Making Uncertainty Visible


The backend state is only half the problem.
If the UI presents every generated statement as equally authoritative, users will eventually stop checking it.
The briefing interface therefore distinguishes different epistemic states. Information can be represented as a direct_fact, model_inference, or unverified item.
The system also labels external reconnaissance information as an unverified_assumption that requires confirmation.
That distinction is important.
If I see:
David Miller — VP Engineering

and:
David may object to the proposed architecture.

those statements shouldn't look equally certain.
The first might be directly supported information. The second is an inference.
Uncertainty needs to survive all the way from the reasoning layer to the user interface.

What This Looks Like in Practice

Imagine a meeting where someone says:
"I'll send the revised architecture tomorrow."

The system can extract:
Owner: Priya
Task: Send revised architecture
Due: Tomorrow
Status: pending
Confirmed: false

The relevant transcript excerpt is preserved with the proposal.
I review it and confirm it.
Only then does the commitment become authoritative.
Now imagine:
"Priya could probably send the revised architecture tomorrow."

The model might still propose a commitment.
That's acceptable.
The dangerous behavior would be silently turning that proposal into a confirmed commitment.
The goal isn't to eliminate uncertainty from language models.
The goal is to prevent uncertainty from silently becoming authority.

What I Learned

  1. Extraction accuracy isn't authorization Even a strong extraction model will encounter ambiguous language. Improving the model helps, but removing the approval boundary doesn't. I would rather review an occasional false proposal than discover that an incorrect commitment was silently recorded as fact.
  2. Evidence is more useful than blind confidence A confidence score tells me how strongly the model believes something. A source excerpt lets me inspect why. For commitments, the second is much more useful.
  3. Memory and application state should remain separate Hindsight makes the agent more useful because it can remember relevant history. But recalled context isn't automatically current truth. Memory informs reasoning. The application database owns the actual commitment state.
  4. Put important boundaries in the data model is_confirmed = FALSE is more than a field. It represents an architectural rule: inference does not equal authorization.

The system can enforce that rule independently of how the model behaves.

  1. Human oversight works best as an explicit state transition "Ask the user for confirmation" is a prompt instruction. FALSE → TRUE through an explicit confirmation operation is an application invariant. I trust the second approach more.

The Boundary I Want to Keep

There is a natural pressure when building agent systems to remove every human step.
If the agent can find a commitment, let it confirm the commitment.
If it can write an email, let it send the email.
If it can remember context, let that memory automatically drive decisions.
I don't think those actions are equivalent.
My agent can process the transcript, retrieve relevant history from Hindsight, identify possible commitments, preserve the evidence, explain its reasoning, and prepare the follow-up.
But when an interpretation is about to become an authoritative commitment, I want a person in the loop.
Hindsight gives the agent memory. The LLM provides inference. PostgreSQL holds authoritative state. The human provides the final authority.
That separation is what makes the agent useful without making it the final authority over what people actually agreed to.

Top comments (0)