DEV Community

Cover image for Why We Built Foresight: Stopping AI from Committing Commercial Suicide
Tanmaya sree Chirra
Tanmaya sree Chirra

Posted on

Why We Built Foresight: Stopping AI from Committing Commercial Suicide

If you give an LLM access to a sales inbox, it will eventually agree to an impossible contract.

Last week, we tested what happens when a prospect sends an aggressive negotiation email: "Can you give us 40% off and get us live into production in two weeks?" A standard prompt-chained LLM happily replied that it could make that work if the customer signed by Friday. In the real world, sending that email would have triggered a contractual disaster: our engineering team requires at least 21 days for secure VPC peering, our finance team caps rep discounts at 15%, and our security lead had already barred all production data ingestion until an overdue compliance audit was reviewed.

Stateless LLMs fail at enterprise sales because they suffer from acute context amnesia. They treat each incoming message as an isolated conversational turn, completely blind to the commitments, constraints, and politics established over months of prior calls.

To fix this, we built Foresight, an enterprise deal intelligence engine that uses Hindsight to maintain an active, self-updating memory bank of the entire customer relationship. Instead of simply generating agreeable text, Foresight intercepts incoming demands, checks them against historical commitments and company policies, detects multi-variable collisions, and dynamically updates its strategy when real-world facts change.

Here is how we designed the system, the architectural trade-offs we encountered, and why agent memory must act as a computational governor rather than a simple passive search index.


The System Architecture: How Foresight Fits Together

Enterprise sales deals are not linear chats; they are multi-month distributed state machines. A single deal involves half a dozen stakeholders with competing agendas:

  • The Champion wants roadmap features delivered yesterday.
  • The Technical Evaluator cares about latency, architecture, and deployment constraints.
  • The Security Gatekeeper (CISO) cares about compliance, data residency, and audit certifications.
  • The Economic Buyer (CFO) enforces budget ceilings and hates multi-year lock-in.
  • Procurement plays hardball on discounts and payment terms.

To model this reality without building a bloated microservice architecture, we designed Foresight around three distinct layers:

┌─────────────────────────────────────────────────────────────┐
│ 1. Historical Interaction Ingestion (Hindsight Memory Bank)  │
│    Ingests call transcripts, emails, and commitments into   │
│    a dedicated deal memory bank via retain().               │
└──────────────────────────────┬──────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ 2. The Collision Engine                                      │
│    Cross-checks incoming demands against:                   │
│    • Active commitments & fulfillment status in memory       │
│    • Static company baseline policies (SLAs, margins)       │
│    Categorizes conflicts into Security, Timeline, and Price.│
└──────────────────────────────┬──────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ 3. The Dynamic Reflection & Strategy Loop                   │
│    When real-world state changes occur (e.g. SOC-2 sent),   │
│    memory mutates live, re-evaluating advice from           │
│    defensive blocker to offensive closing plan.             │
└─────────────────────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

Rather than feeding an entire 8-week call history into a prompt window—which burns tokens, degrades reasoning attention, and costs a fortune—we use Vectorize agent memory to maintain persistent structured entities: stakeholders, explicit objections, competitor mentions, and an auditable commitment ledger.


Core Technical Story: Moving from Naive RAG to Collision Detection

Most developers building sales tooling reach immediately for naive Retrieval-Augmented Generation (RAG): embed previous call summaries in a vector database, perform cosine similarity search on the incoming prompt, and feed the top-k chunks into the context window.

During our early testing, naive RAG broke down completely. Consider this incoming request from the customer's procurement lead:

"Can you give us 40% off and get us live into production in two weeks?"

If you run a standard vector search against the past 8 weeks of call notes, your embedding model will retrieve chunks containing words like "discount", "pricing", and "timeline".

What does it miss? Call #2 from 31 days prior.

In Call #2, the customer's CISO (Nadia Chen) stated:

"No production data can touch your platform without our security team reviewing your audited SOC-2 Type II report."

Semantically, "SOC-2 Type II audit" has almost zero cosine similarity to "40% discount". A traditional RAG pipeline ignores it. Yet operationally, that single security mandate is a hard fatal blocker: agreeing to a two-week deployment is contractually impossible because the security team hasn't even begun reviewing the compliance report.

To solve this, we abandoned simple semantic retrieval and built a deterministic Multi-Constraint Collision Engine.

Instead of asking the LLM "How should I answer this email?", the system evaluates the request against three discrete constraint checks:

  1. 🔴 Security Gate: Are there unfulfilled compliance commitments or unresolved access blockers?
  2. 🟠 Timeline Feasibility: Does the requested go-live date violate engineering lead times (4–6 weeks standard) or crash into documented customer code freezes?
  3. 🟡 Commercial Alignment: Does the requested discount violate delegation-of-authority limits, ignore established budget caps, or fail to account for competitor bids?

Code-Backed Implementation: How We Integrated Hindsight

Let's look at how this is implemented in our codebase.

1. Ingesting Deal History with retain()

In memory.py, we initialize the deal bank and index every conversation, stakeholder sensitivity, and commitment into Hindsight:

class DealMemoryBank:
    def __init__(self, deal_path: str = ACME_DEAL_PATH, kb_path: str = COMPANY_KB_PATH):
        self.deal_data = self._load_json(deal_path)
        self.company_kb = self._load_json(kb_path)
        self.commitments = self.deal_data.get("commitments", [])
        self.interactions = self.deal_data.get("interactions", [])

        # Initialize Hindsight client
        self.hindsight_client = None
        self._init_hindsight()

    def retain(self, content: str, category: str, metadata: Optional[Dict[str, Any]] = None):
        """Retains an observation, commitment, or interaction into deal memory."""
        if self.hindsight_client:
            self.hindsight_client.retain(
                bank_id="acme-deal",
                text=content,
                context=category
            )
Enter fullscreen mode Exit fullscreen mode

Each interaction is tagged with its chronological context. Crucially, promises made by sales reps are extracted into structured Commitments with an explicit initial state: "Overdue" or "Pending", accompanied by the audit note: "No fulfilment recorded".

2. Multi-Constraint Collision Evaluation

In collision_detector.py, the engine intercepts the customer's request and cross-references active commitments and baseline company rules:

def check_collisions(self, customer_request: str, deal_context: Dict[str, Any]) -> Dict[str, Any]:
    commitments = deal_context.get("commitments", [])
    kb_deploy = self.company_kb.get("deployment_rules", {})
    kb_pricing = self.company_kb.get("pricing_and_discount_rules", {})

    # Check 1: Security Gate (SOC-2 Type II Report)
    soc2_comm = next((c for c in commitments if c.get("id") == "COMM-02"), None)
    is_soc2_completed = soc2_comm and soc2_comm.get("status") == "Completed"

    collisions = []

    if not is_soc2_completed:
        collisions.append({
            "severity": "RED",
            "badge": "🔴 Security Blocker",
            "title": "Unfulfilled SOC-2 Mandate for Production Access",
            "description": "Nadia Chen (CISO) explicitly mandated that NO production data can touch Veridian without reviewed SOC-2 Type II audit report.",
            "evidence": [
                "Call #2 (31 days ago): Nadia stated 'No production data can touch your platform without our security team reviewing your audited SOC-2 Type II report.'",
                "Commitment COMM-02: 'SOC-2 Type II Audit Report Delivery' is currently OVERDUE (31 days elapsed)."
            ],
            "impact": "CRITICAL: Promising deployment before SOC-2 sign-off triggers an immediate security veto."
        })
    else:
        collisions.append({
            "severity": "GREEN",
            "badge": "✅ Security Cleared",
            "title": "SOC-2 Type II Report Delivered",
            "description": "Security gate unlocked for staging and production onboarding."
        })

    # Check 2: Timeline Feasibility
    if any(k in customer_request.lower() for k in ["two weeks", "2 weeks", "14 days"]):
        collisions.append({
            "severity": "ORANGE",
            "badge": "🟠 Timeline Conflict",
            "title": "Unrealistic 2-Week Deployment Request vs. 4-6 Week Baseline",
            "description": "Veridian standard onboarding takes 4 to 6 weeks. No 2-week promise was ever recorded.",
            "evidence": [
                "Call #3: SA confirmed standard deployment is 4 to 6 weeks due to VPC peering.",
                "Call #9: Marcus confirmed ACME has an annual production freeze starting October 15."
            ],
            "impact": "HIGH: Agreeing creates catastrophic delivery failure and SLA breach risk."
        })

    return {"collisions": collisions}
Enter fullscreen mode Exit fullscreen mode

Notice the evidence array. Every single collision badge is tied directly to a specific historical call, speaker, and timestamp. In enterprise environments, human operators do not trust an AI that asserts "This is risky". They trust an AI that shows the exact receipt from Call #2.

3. Dynamic Memory Reflection

The most important capability of Hindsight agent memory is that memory is mutable. It updates when real-world actions occur.

When the sales rep finally emails the compliance package to the prospect's security team, they click "Mark SOC-2 Sent to Nadia" in the UI. Here is what happens under the hood:

def update_commitment_status(self, deal_context: Dict[str, Any], commitment_id: str, new_status: str, notes: str) -> bool:
    commitments = deal_context.get("commitments", [])
    for comm in commitments:
        if comm.get("id") == commitment_id:
            comm["status"] = new_status
            comm["status_notes"] = notes
            comm["last_updated"] = datetime.now().isoformat()

            # Retain this critical state change into Hindsight memory
            reflection_text = (
                f"COMMITMENT STATE CHANGE: '{comm['title']}' for recipient '{comm['recipient']}' "
                f"changed to [{new_status}]. Note: {notes}"
            )
            self.hindsight_client.retain(
                bank_id="acme-deal",
                text=reflection_text,
                context="commitment_update"
            )
            return True
    return False
Enter fullscreen mode Exit fullscreen mode

When this state mutation is saved, the Collision Engine immediately recalculates. The 🔴 Security Blocker dissolves into a ✅ Security Cleared notification, and the downstream response generation pivots automatically.


Results and Behavior: The Live Interaction

To verify the system, we ran our benchmark enterprise deal: ACME Corp (an 8-week history comprising 9 calls, 5 stakeholders, a \$72k quote, and an active competing bid from NimbusAI at \$48k).

Interaction 1: Before Memory Reflection (Defensive Triage)


When the aggressive request ("40% off + 2-week deploy") arrives:

  • Foresight output:
    • 🔴 Security Blocker: Missing SOC-2 Report (Overdue by 31 days).
    • 🟠 Timeline Conflict: 2 weeks requested vs. 4–6 week standard.
    • 🟡 Commercial Alignment: 40% discount (\$43,200) violates margin policy and is unnecessarily low given Linda's \$50k budget ceiling.
  • Strategic Advice Generated:
    • "⛔ DO NOT agree to deployment dates or price concessions yet. Your immediate priority is delivering the SOC-2 report to Nadia Chen."

Interaction 2: After Memory Reflection (Offensive Negotiation)


The rep clicks "Mark SOC-2 Sent to Nadia". Hindsight records the reflection. The collision check re-runs automatically:

  • Foresight output:
    • ✅ Security Cleared: SOC-2 delivered; security gate unlocked.
    • 🟠 Timeline Conflict: Adjusted to offer a phased 3-week express onboarding with dedicated solution architects.
    • 🟡 Commercial Alignment: Counter-offers at \$49,000 ARR.
  • Strategic Advice Generated:
    • "🎯 PROCEED WITH COMMERCIAL NEGOTIATION. Target price: \$49,000 (respects Linda's \$50k budget ceiling and beats NimbusAI's \$48k offer) with an agreement to sign before the October 15 code freeze."

The draft customer reply updates instantly, switching from a defensive delay to a confident commercial closing pitch.


Lessons Learned

Building Foresight surfaced several non-obvious engineering realities about memory-augmented agents:

1. Negative Grounding Prevents Hallucinated Certainty

When tracking commitments, never let your agent declare: "The report was never sent." If an interaction happened outside the recorded system, that absolute claim destroys user trust. Instead, our ledger outputs: "No fulfilment recorded". That subtle wording change reflects epistemic humility: the agent only claims knowledge of recorded events.

2. Decouple Static Policy from Dynamic Memory

Early on, we tried storing company policies (like minimum deployment timelines and discount delegation matrices) inside the same memory bank as call transcripts. The LLM regularly confused company rules with customer statements. Separating static company ground truth (company_kb.json) from dynamic deal memories (acme_deal.json) eliminated cross-contamination completely.

3. Multi-Variable Constraints Beat Pure Probabilistic Generation

You should not rely on an LLM's next-token prediction to decide whether a discount is legally permissible. Use deterministic code or structured Pydantic schemas to validate hard boundaries (budgets, compliance rules, timeline baselines), and let the LLM handle the natural language synthesis within those validated bounds.


Conclusion

Stateless chatbots treat conversations as ephemeral noise. But in high-stakes enterprise workflows, conversation history is a web of promises, liabilities, and leverage.

By integrating Hindsight into our deal intelligence architecture, we gave our agent the ability to remember what was promised, flag dangerous collisions, and dynamically adapt its strategy when real-world milestones are reached. That is the difference between an AI that makes reckless promises and one you can actually trust with your business.


Explore the full implementation on GitHub.

Top comments (0)