DEV Community

suissAI
suissAI

Posted on

Proactive by Construction: Why the Next Generation of Agents Must Anticipate, Not Just React

Executive Summary: Current AI assistants are mostly reactive: they wait for a user’s message, run the model, and answer. This works for simple queries but wastes the many idle cycles between interactions. In contrast, a truly proactive agent continuously monitors state, predicts future needs, and acts (or prompts the user) before being asked. Here we argue that instead of bolting on occasional webhooks or notifications, the entire architecture should be designed for proactivity. This “Proactive by Construction” approach minimizes wasted user-initiated messages and shifts heavy work into idle periods.

We compare the reactive, event-driven architecture (e.g. waking on every email) with a state-driven proactive design. We detail how to build the latter using (1) persistent domain state and provenance graphs, (2) explained-information caches, (3) offline-capable UI agents, (4) smart session and sync protocols, (5) a layered security/cost hierarchy (sandbox + cheap classifier + full LLM), and (6) a simulation/risk engine to anticipate critical thresholds. For each, we outline benefits (reduced cost, faster responses, user trust), costs (engineering complexity, storage), security considerations, and implementation notes. We include comparative tables and diagrams (mermaid) of architectures and data flows. All designs are assumed unconstrained (no fixed cloud infra or model limits) and prioritize user utility over technical gimmicks.

“Proactive First” Product Blurb: “Our assistant doesn’t just wait; it stays one step ahead. It monitors your business data, forecasts potential issues, and prepares answers or alerts before you even ask. User prompts become confirmations rather than requests. Proactive First means the system uses idle time on the highest-value tasks so that when you do ask something, the answer is already there.”

Reactive vs. Proactive Architectures

Most deployed agents today follow a chatbot model. They sleep until a user message arrives, then fetch data, call the LLM, and reply. This is simple, but it creates high latency (each turn costs many tokens) and misses the chance to use downtime productively. In contrast, a proactive agent treats idle time as an opportunity to prepare. Instead of user→agent→response, it continuously evaluates the business state, forecasts, and incoming events, and decides when and if to speak up.

Property Reactive/Event-Driven State-Driven/Proactive
Cost High per event (LLM cost on every webhook). Lower long-run cost: expensive LLM calls only when needed, plus lightweight tracking. Enables batching work in idle moments.
Latency High at query time (full LLM run). Faster responses when asked (answer often precomputed).
User Effort High (must trigger each query). Low: fewer manual queries, more automatic summaries/alerts.
Privacy Many messages/queries over network. Local processing and caching means less raw data sent out.
Complexity Simpler design, but fragile at scale (cost/security issues). Higher system complexity (state management, sync, simulations) but much richer behavior.
Proactivity None – only reacts. Full: anticipates tasks, can initiate actions or ask clarifying questions.

In short, reactive agents expend expensive LLM calls on all inbound events (95% of which are noise). State-driven agents, by contrast, filter and summarize locally, predict what’s needed, and only then prompt the LLM or user. This division of labor (cheap local vs. occasional expensive reasoning) yields huge savings and faster turnarounds.

System Overview (Architecture)

Below is a high-level architecture of our proactive agent system. It contrasts with a simple “agent-as-backend” model. Here, the UI-local agent and background agents handle much logic client-side and sync to the backend only deltas (events or requests). The backend remains the authority on domain state and data.

flowchart LR
  subgraph Frontend
    WA[WhatsApp/Mobile Chat] 
    Web[Web GUI / Desktop App]
    WA -- Message→ AgentUI(UI-Local Agent)
    Web -- Action→ AgentUI
  end
  subgraph Mobile/Shop
    LocalDevice((Local Device))
    AgentUI -. offline .-> LocalDB((Local Store))
  end
  subgraph Backend
    BackendSystem((Core Backend\nDomain State & DB))
    ProvenanceDB((Explained Data & Provenance))
  end
  UserSentinel(UserSentinelAgent) -- heartbeats → AgentUI
  AgentUI -- delta/sync → BackendSystem
  BackendSystem --> ProvenanceDB
  AgentUI <--> LocalDB

The UI-local agent (in browser or app) maintains session state, local cache, and offline interaction. It communicates with (and is supervised by) a UserSentinelAgent via periodic SessionACKs (heartbeats). Only when something meaningful occurs (or periodically) does it sync updates to the BackendSystem, which persists authoritative state and explanations. The backend can update the UI only by responding to these syncs (front-end initiated sync model), not by continuously pushing.

State-Driven Proactivity and Session Management

A core design is to separate authentication/presence from interaction preference. Each user session has:

  • Identity & Authentication: The user logs in once (web or WhatsApp).
  • Session (Interactive Mode): This expires after inactivity (e.g., 1 hour) unless renewed. The UI agent (or WhatsApp bot) will send a prompt: “Your session is inactive. Continue?”. If user says “yes”, the UserManagerAgent renews it; if “no” or timeout, session ends.
  • Channel Presence: The user can remain authenticated but choose to go silent. For example, a store manager might remain logged in (on desktop or via WhatsApp) but disable proactive messages. In that case, the system stays “passive” – it will not message unsolicited – but continues receiving and caching data.
  • Authorization: Even if the session is active, before any action (e.g. sending an email or modifying data), a state/policy check ensures the action is valid (permissions, budget, risk). This extra layer, suggested by practitioners, prevents rash actions.

Concretely, the UI agent sends a SessionACK (a simple heartbeat including last activity) to UserSentinelAgent on the backend every minute. The sentinel checks how long since last real user reply. If near timeout, it triggers the “continue session?” prompt. If timeouts fully, it closes the session. This keeps both sides in sync: the frontend initiates syncs (sending events/deltas to backend), rather than backend continuously polling clients. The UI agent can continue operating offline (answering questions from local cache) and simply queues a delta to send later.

Benefits: Reliable session control, user stays in control of inactivity. Separating idle presence from authentication avoids spamming the user (they can go “quiet but logged in”). Costs: Slight added complexity in managing heartbeat logic and session flags. Security: This avoids stale credentials or orphan actions – every critical action is re-checked on the backend. Implementation: Store a last-ACK timestamp per session in backend. On each frontend interaction reset it. Sentinel runs a timer to expire sessions if no ACK.

Explained-Information and Provenance Graphs

Central to proactivity is the notion of self-explaining state. Every important derived Information (e.g. “Forecast for next month: $18,450”) is stored along with its provenance: exactly what raw values and operations produced it. We structure this as a graph of value identities. For example:

{
  "information_context": "sales.forecast",
  "information_id": "2026-10-04",
  "value": { "forecast": 18450 },
  "provenance": {
    "values": {
      "777": { "value": 142.50, "source": "sales.123", "field": "total", "version": "2026-09-30T12:00:00Z" },
      "1432": { "value": 89.90,  "source": "sales.124", "field": "total", "version": "2026-09-30T12:05:00Z" },
      "111": { "value": 37,     "source": "stock.981", "field": "available", "version": "2026-10-01T08:00:00Z" },
      "4":   { "value": "pharmacy", "source": "customer_segment.42", "field": "type", "version": "2026-08-15T00:00:00Z" }
    },
    "operations": [
      { "operation": "aggregate", "inputs": [777, 1432] },
      { "operation": "join",      "inputs": [777, 111] },
      { "operation": "forecast",  "inputs": [9001, 111, 4], "model": "ARIMA-v2" }
    ]
  },
  "explanation": {
    "pt-BR": "Explicação detalhada em português...",
    "audio": "(link to pre-rendered audio)",
    "short": "Resumo breve...",
    "detailed": "Explicação completa detalhada..."
  }
}
Enter fullscreen mode Exit fullscreen mode

In this scheme, each raw or intermediate value has a unique ID (e.g. 777, 1432). From the ID we know exactly which source record and field it came from (and even the timestamp/version of that record). The provenance graph traces how each value fed into operations to produce others (see Mermaid example below). This lets the agent reconstruct, deterministically, how any information was computed, without re-running the raw data queries.

flowchart LR
  subgraph Provenance
    A["sales.123.total = 142.50 (v2026-09-30)"] 
    B["sales.124.total = 89.90 (v2026-09-30)"]
    C["stock.981.available = 37 (v2026-10-01)"]
    D["customer_segment.42.type = \"pharmacy\" (v2026-08)"]
    E["aggregate(142.50, 89.90) = 232.40 (id 9001)"]
    F["join(232.40, 37) = 269.40 (id 9002)"]
    G["forecast(269.40, 'pharmacy') = 18450"]
  end
  A --> E
  B --> E
  E --> F
  C --> F
  F --> G
  D --> G

Benefits: The system always knows where each piece of information came from. When the user asks “Why is the forecast $18,450?” we don’t blind-call an LLM on the raw database. Instead we feed the above structured provenance into a context prompt to generate a natural-language explanation. This ensures explanations match the true data lineage. We also generate different forms (“short” vs “detailed”, audio, localized languages) up front and store them.

Costs: Storing the provenance graph increases storage (especially if values/versions are stored). Querying it on the fly needs careful indexing. But these are standard issues with event-sourced or knowledge-graph databases. We offset cost by reducing repeated LLM calls and by compressing provenance in JSONB blobs (see below).

Security: The provenance data are derived from business data, so they carry no secret beyond what the user is authorized to see. The main risk is that an LLM might hallucinate explanations beyond the given facts; by bounding the LLM context to the provenance and using a deterministic “explain” prompt, we prevent falsification of the chain. (One can include provenance hashes to verify no tampering of stored graphs.)

Implementation: Use a hybrid relational/JSONB store. Key tables hold stable entity state and base events, while a JSONB column holds each Information’s provenance and explanations. Include a provenance_hash (e.g. SHA256 of the sorted provenance entries) to detect if the underlying data changed; if so, invalidate the cached explanation. Queries by information_context and information_id retrieve value and provenance quickly (index these fields).

Explained-Information Cache

Once an explanation is generated for a given information (e.g. a forecast or anomaly alert), we cache it. That is, for each (information_id, version, locale, detail_level) we store the text (and audio) explanation. Next time anyone asks the same question, or a semantically equivalent one, we can replay the cached response. In practice:

  • Provenance Hash: Whenever underlying data change, the provenance graph hash changes. Only if hash is identical and context unchanged do we reuse the old explanation. Otherwise we prompt the LLM anew and update the cache. This avoids stale explanations and unnecessary LLM calls.
  • Semantically-indexed: In parallel, we keep a local “reverse index” mapping key concepts (like product names or metrics) to their explanations, to serve local intent queries (see next section).
  • Locale/Format Variants: We store both short and long versions, text and audio, so the UI can pick the appropriate format without rerunning the LLM.

Benefits: Dramatically lowers LLM usage. Once explained, answers are instant. It also ensures consistency in replies. Costs: Extra storage (but text is cheap). Slight complexity in cache invalidation logic. Security: No new issue – the cache just reduces recomputation. Notes: This is akin to “explainable materialized views” in a database, a technique known from data lineage research.

UI-Local Agents and Offline-First Operation

We strongly advocate an offline-first UI agent design. The UI (browser app or shop terminal) stores all viewed data locally, so the user’s context and requested reports are instantly available even without network. For example, when a shop employee requests a sales report, the UI fetches it once and keeps it. If they ask follow-up questions (“What contributed to the trend?”), the agent can answer from its local cache of values and explanations. This is facilitated by:

  • Local Data Store: Use browser IndexedDB or mobile SQLite as the “second replica” of relevant state. This store holds the last sync of information and provenance for that user’s scope. (E.g. a JSONB-like blob for each ExplainedInformation.)
  • Optimistic UI: User actions (e.g. “ask agent”) are applied immediately to the UI state and queued to the backend. This optimistic update means the interface is always responsive.
  • Sync-on-Demand: When online, the UI agent pushes its unsynced events or deltas to the backend (not vice versa). The backend returns any new updates. If offline, the UI continues to operate on cached data. This matches the offline-first principle: “the network is allowed to fail, but the application should still be useful”.

Thus each client (phone/web) becomes its own mini-agent: it can answer many questions from its cache (including anything explained previously), prompt the user to re-authenticate if needed, and even simulate risk scenarios locally. We can even ship small “sleeping” agents on devices that wake on schedule and process data.

Benefits: The assistant feels immediate (no lag waiting on network for each query). It also scales effortlessly: users' requests don’t hit the central backend unless needed. Costs: More complexity in client code and data sync logic. Must handle conflict resolution (resolve if same data updated in parallel clients – use last-write-wins or merge strategies). Security: Local data must be encrypted or sandboxed, as it may include sensitive info. Implementation: Use proven PWA/local-first frameworks (service workers, IndexedDB, CRDT libraries). Key references: “Offline-first frontends make the app part of a distributed system”.

SessionACK/UserSentinel Pattern

To bridge front-end autonomy with centralized control, we introduce a session heartbeat pattern. The UI agent and WhatsApp bot periodically send a minimal SessionACK message to the UserSentinelAgent on the backend. This message might include the user ID, channel (WhatsApp/Web), last action ID, and a timestamp. The sentinel uses this to:

  • Track if the user is still active.
  • Trigger reauthentication prompts (if nearing inactivity timeout).
  • Optionally carry minimal UX metrics (e.g. “user dismissed 3 suggestions”).

On the backend, UserSentinelAgent simply updates the session’s last-active timestamp. It does not query the frontend; it passively waits for ACKs. If an ACK is overdue, it starts session termination. Conversely, the front-end never blocks waiting for the backend on every event – it locally updates state and asynchronously sends a short “what happened” notice later (just the essential event IDs or state changes).

Benefits: Ultra-lightweight server. Avoids polling. Maintains security (no user action happens without backend confirmation). Costs: Requires the backend to trust the sequence of ACKs (they should be signed or validated). Security: We guard against replay attacks by including a monotonically increasing counter or nonce in each ACK. If an attacker tries to spoof an ACK, it will have an invalid sequence number or be rejected by the sentinel.

Sandboxed Edge + Decision Models + LLM Loop

Inspired by [uriwa’s Reddit proposal], we use a three-tier pipeline for external events (emails, webhooks, sensor streams):

  1. Sandboxed Edge Code (pre-filter): Incoming events (e.g. email text, chat mentions) are first processed by a tightly sandboxed script (no loops, no arbitrary code execution). This extracts key fields and calls a decision model (a small inference engine). Because it’s sandboxed, even if an attacker injects malicious content, it cannot execute on our server or leak secrets.

  2. Decision Model (System 1): This is a compact classifier (e.g. a small transformer or tuned model) that answers yes/no or scores "Should I wake the agent?". For example, “Does this email need our attention?”. Unlike a generative model, it emits a bounded label, and runs in milliseconds at a fraction of the cost (∼1/100th). This filters out ~95% of noise for near-zero cost.

  3. Full Agent Loop (System 2): Only if the decision model says “yes” do we notify the agent. Critically, rather than blasting the user with a cold message, we inject a structured notification into the existing conversation thread. The agent then wakes with full context (persona, history, tools) and can formulate a proposal. E.g. “I drafted a reply…” The user confirms, and the agent acts.

This layered approach means: cheap code handles routine parsing, a tiny AI filters events, and the big LLM only runs for real tasks. It solves the “cost trap” and “security trap” noted on Reddit (where arbitrary webhooks would otherwise spam or expose the system).

Benefits: Orders-of-magnitude cost reduction. Better security: sandboxing prevents prompt injection from executing sensitive actions. Users only get contacted for true positive events. Costs: Building and maintaining the decision model; ensuring it stays high-precision (missed events = lost opportunity). We mitigate that by continuous learning (feedback from user: “no action needed” vs “good alert”). Implementation: Use a policy engine (like OpenAI’s new classification APIs) or a small distilled model. Run it in a WASM or custom sandbox. The decision model’s queries are few-token prompts (“Is this email urgent?”) to minimize token use.

Simulation/Scenario Engine (Risk Bands and Mitigation)

To be truly proactive, the agent not only reports current stats but also simulates future scenarios using cached data. For example, given past sales trends and current inventory, the agent can probe “What if sales slow by 10%?” or “What if supply is delayed 5 days?” and see how KPIs would change. In practice, we:

  • Use the stored historical data (and any domain model, e.g. sales forecasting) to project different futures (optimistic, normal, pessimistic). This might happen nightly or whenever new data arrives.
  • Tag ranges of outcomes as safe, warning, or dangerous. For example, if inventory falls below threshold X, label that a “risk zone”.
  • When current data approaches a risk band, trigger an alert: “Inventory is nearing low threshold (27 units). If trend continues, we predict a stockout in 3 days.”
  • Even before that, suggest mitigations: the agent knows from simulations what would avert the risk (e.g. “If you expedite order by 2 days, stockouts avoidable; or you could push a sale on slow products to maintain revenue”).

This is essentially building a what-if engine on cached data. Because it uses historical and current numbers already in the database, the agent can run these probes offline (or cheaply on a separate process) and surface the results as pre-generated reports or alerts. The user’s questions become more exploratory (“show me the high and low scenarios”) rather than constructive.

Benefits: Dramatically shifts the assistant from passive to preventive. Users get warned in advance of problems and informed of solutions. Costs: Increased computation for simulations (likely done during low-load periods). More complex logic to define risk bands and actions. Security: Fully on-user’s data, no extra risk. Notes: This aligns with research in proactive health/finance systems that run “Monte Carlo scenario analysis” on the side. Here the outputs (tables or charts) can be cached in the ExplainedInformation store with provenance.

Local Intent Matching and Decision Primitives

In the UI, an agent-like component listens to the user’s free-form queries (in chat or voice) and handles them as follows:

  • Regex + Semantic Matcher: A fast local engine first checks if the query matches a known intent template (e.g. “why was X down?”, “show me [metric] for [period]”, “what triggered [alert]?”). These can use simple regex or keyword-based matching at first pass (O(1) cost).
  • Information Registry: If an intent is recognized (say, “Why this forecast changed?”), the system immediately retrieves the relevant ExplainedInformation and its provenance, and returns the answer (possibly filling in variables). The LLM is not needed because this is a structured lookup.
  • Proactive Decision Primitives: If no template matches, a slightly larger model or LLM may classify “Should the agent ask the user for clarification?”, “Should the agent act on this data?”, or “When is the best time to prompt?” These can be framed as binary or multi-choice questions (akin to Tang et al’s framework of deciding to stay silent, ask, assist, or act). For instance:
    • ShouldAct? (Is now the correct moment to intervene on this information?)
    • BestNextAction? (Given current tasks, what should the agent do or say now?)
    • BestTimeToAct? (Is this immediate, or better scheduled for later/overnight?). Using such “System 1” predicates lets the agent throttle itself and avoid unnecessary chatter.

Only if the intent remains ambiguous do we defer to the full System-2 loop. This multi-layered local reasoning (regex → semantic classifier → LLM) mirrors the Reddit pattern: cheap checks first, costly calls last.

Benefits: Very low-latency responses for common queries. Drastically fewer LLM calls. Costs: Building and curating intent templates; training light classifiers. Security: Local matching never leaks data. Implementation: Store a mapping of intent keywords → information IDs in the frontend. Use regex to catch obvious questions. Keep an on-device mini LLM (or API) for semantic matching if needed.

Data Storage: JSONB + Relational

Practically, we store explained information and provenance in a relational DB with JSONB fields. For example:

  • Table information: (information_id PK, context, timestamp, value JSONB, provenance JSONB).
  • Table explanations: (info_id FK, version, locale, audience, explanation JSONB, provenance_hash).
  • Indices on (context, timestamp) and on provenance_hash (since many queries will check for changed data).

This hybrid model retains the query efficiency of SQL (e.g. find all info for given date range) while storing flexible, nested provenance and explanation data in JSONB (which PostgreSQL can index on keys).

Benefits: Flexible schema (we can evolve provenance shape without migrating whole tables). Powerful JSONB indexing allows quick lookup by info ID or by provenance attributes if needed. Costs: More complexity in schema design. JSONB queries can be slower than pure SQL, but most reads are by key. Security: Standard DB security applies; ensure JSON content is validated on insert.

Tables of Comparison

1. Reactive vs. State-Driven (Proactive)

Dimension Reactive (Webhook) State-Driven (Our Design)
Cost High API & LLM cost (invoke on every event). Low – cheap filtering + batched processing. LLM calls only for true alerts.
Latency Slow per request (full prompt, tool loop). Fast answers for precomputed info. Interactive queries often answered locally.
Privacy Many messages to cloud, more exposure. Mostly local. Only deltas and flagged events sent to backend.
Complexity Simple design, but scales poorly (cost, security issues). More components (cache, probes, sync), but efficient and robust.
Proactivity None (user must ask for updates). Full (anticipates needs, alerts proactively).
Trust & UX Users often spam queries if they don’t see updates. Users trust agent to speak up only when needed.

2. Component Responsibilities & Data Flow

Component Responsibility Data Flow
WhatsApp/Mobile User interaction via chat/voice. Sends user inputs to UI agent; receives text/audio reply.
Web UI / Desktop Visual interface, optional mic/voice for shops. Sends clicks/voice to UI agent; displays charts, prompts.
UI-Local Agent Client-side agent harness. Manages session, cache, local Q&A, and UI updates. Reads/writes LocalDB; sends SessionACK & deltas to BackendSystem; receives updated state or explanations.
LocalDB On-device storage of user’s data (cache of info & provenance). Queried by UI agent for fast local answers; syncs with Backend.
UserSentinelAgent Backend presence monitor. Handles timeouts and session control. Receives SessionACK pings; triggers session renewal prompts or logout.
BackendSystem Authoritative domain logic and data. Applies events, updates state. Receives deltas/requests from UI; emits updates and notifications back.
ExplainedData Store Database of all explained information, provenance, cached explanations. Queried by backend and UI for any needed info. Updated when new info created.
Sandbox Layer (Optional edge service) Runs untrusted code or parsing safely. Processes external webhooks (emails etc.) before passing to agent.
Decision Model (Edge) Fast classifier filtering events. Takes parsed event info; returns “wake” boolean.
Agent Loop Full LLM reasoning when needed (chat loop, tools). Operates on user-turn or triggered notifications.
Simulation Engine Offline scheduler for what-if analysis. Reads historical data; writes predicted metrics/alerts to Backend.

Key Citations and Sources

  • Reactive vs. real assistant: “Most people build chatbots: user sends message, model answers, then sleeps. A real assistant does not sit idle – it watches inboxes and pings you with drafted replies when needed.”.
  • Cost/Security pitfalls of webhooks: 95% of emails are noise, and spinning an LLM on each wastes huge money. Also, untrusted webhooks can inject dangerous commands. The three-tier solution (sandbox + classifier + LLM) addresses both.
  • Efficacy of proactivity: Proactive agents can use idle time to reduce dialogue turns by ~15% and error rates by ~28%. Well-crafted proactive behaviors can convert idle cycles into greater user utility; by contrast, poorly timed action hurts trust.
  • Human-study insight: Users prefer a “sleep-time” assistant that does tasks offline even if imperfect. A single mis-timed action can sharply erode trust, so timing and relevance matter as much as correctness.
  • Proactive decision theory: Agents should choose “silent vs ask vs assist vs act” based on value, delay cost, and risk. This formalizes the ShouldAct?/BestNextAction?/BestTimeToAct? primitives.
  • Offline-first design: “Offline-first architecture starts from a different assumption: the network is allowed to fail, but the app should still be useful.” The frontend becomes part of a distributed system. Data is kept local, UI optimistically updates, and network is just a sync layer.
  • Sandbox and filters: “Sandboxed edge code” with no host privileges prevents prompt injections. A lightweight decision model can be orders of magnitude cheaper than an LLM.
  • Explained information: We construct each info item as “value + provenance + explanation”. This aligns with ideas in data provenance and explainable AI: every fact carries its origin so explanations are faithful.
  • Tables & diagrams: The mermaid figures above illustrate (a) the component architecture and (b) a sample provenance graph. We also provide a JSON schema snippet for ExplainedInformation (the content of the JSON example).

In summary, Proactive by Construction demands that we stop treating proactivity as an optional “plugin” and instead embed it in every layer: from state storage to UI to decision policies. By doing so, the agent learns to use downtime, minimize manual queries, and explain its knowledge, ultimately delivering a smarter, faster, and more trustworthy assistant.

Top comments (0)