<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gautam Raju Saripalli</title>
    <description>The latest articles on DEV Community by Gautam Raju Saripalli (@gautam_rajusaripalli_e0f).</description>
    <link>https://dev.to/gautam_rajusaripalli_e0f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150815%2F59ea9f0a-94f1-470b-b335-530de4fba571.jpg</url>
      <title>DEV Community: Gautam Raju Saripalli</title>
      <link>https://dev.to/gautam_rajusaripalli_e0f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gautam_rajusaripalli_e0f"/>
    <language>en</language>
    <item>
      <title>Why Procurement Memory Needs Three Data Sources, Not One</title>
      <dc:creator>Gautam Raju Saripalli</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:23:53 +0000</pubDate>
      <link>https://dev.to/gautam_rajusaripalli_e0f/why-procurement-memory-needs-three-data-sources-not-one-3dek</link>
      <guid>https://dev.to/gautam_rajusaripalli_e0f/why-procurement-memory-needs-three-data-sources-not-one-3dek</guid>
      <description>&lt;p&gt;I worked on the data layer for VendorPulse. The problem that hit me early was simple: a single procurement dataset is useless for agent memory.&lt;/p&gt;

&lt;p&gt;We started with DataCo (180k shipments: late delivery risk, real vs. scheduled days). It’s great for logistics, but it has no rejection rates, no excuses, and no vendor context. It tells you a delivery was late—not why it matters.&lt;/p&gt;

&lt;p&gt;We added the Vendor Performance dataset (2.3M purchases: vendor invoices, freight costs). This gave us price and invoice dates, but lacked seasonality.&lt;/p&gt;

&lt;p&gt;Both datasets are accurate, but neither captures what a procurement manager actually remembers:&lt;br&gt;
"Last July, SteelCore was 22 days late during the monsoon with a 12% rejection rate and claimed 'truck breakdown'. They used the exact same excuse in August."&lt;/p&gt;

&lt;p&gt;That sentence never exists inside an ERP. To fix this, we had to build a third source: synthetic procurement context.&lt;/p&gt;

&lt;p&gt;The Three-Layer Data Model&lt;br&gt;
Here is the three-layer data architecture I implemented:&lt;/p&gt;

&lt;p&gt;Procurement KPI → Vendor / PO History (Structured Facts)&lt;/p&gt;

&lt;p&gt;Data: PO_ID, vendor name, quantity, price, delivery date.&lt;/p&gt;

&lt;p&gt;Storage: SQLite (deterministic query layer).&lt;/p&gt;

&lt;p&gt;DataCo → Logistics Context (Macro Patterns)&lt;/p&gt;

&lt;p&gt;Data: Aggregate logistics patterns, a 54.8% late delivery baseline, First Class 0% on-time performance, monsoon delay vectors.&lt;/p&gt;

&lt;p&gt;Purpose: Provides realistic seasonality.&lt;/p&gt;

&lt;p&gt;Synthetic Context → Demo / Historical Narrative (Qualitative Memory)&lt;/p&gt;

&lt;p&gt;Data: Excuses, rejection reasons, true cost after penalty, human operational notes.&lt;/p&gt;

&lt;p&gt;Purpose: Enables meaningful semantic recall in Hindsight.&lt;/p&gt;

&lt;p&gt;The Data Pipeline&lt;br&gt;
I built a pipeline to merge factual numbers with narrative context to create episodic experiences for our memory store:&lt;/p&gt;

&lt;h1&gt;
  
  
  data_processing_pipeline.py - merging real + synthetic
&lt;/h1&gt;

&lt;p&gt;import pandas as pd&lt;br&gt;
from faker import Faker&lt;/p&gt;

&lt;p&gt;dataco = pd.read_csv("DataCoSupplyChainDataset.csv")&lt;br&gt;
vendor = pd.read_csv("vendor_invoice.csv")&lt;/p&gt;

&lt;h1&gt;
  
  
  Create an episodic experience that Hindsight will store
&lt;/h1&gt;

&lt;p&gt;def build_experience(row):&lt;br&gt;
    return {&lt;br&gt;
        "vendor": row["Vendor"],&lt;br&gt;
        "material": row["Category"],&lt;br&gt;
        "month": row["Order_Month"],&lt;br&gt;
        "delay": row["Days_Real"] - row["Days_Scheduled"],&lt;br&gt;
        "rejection": row["Defect_Rate"],&lt;br&gt;
        "excuse": generate_excuse(row["Vendor"], row["Month"]), # e.g., truck breakdown, customs&lt;br&gt;
        "true_cost": row["Price"] + (row["Delay"] * penalty_per_day)&lt;br&gt;
    }&lt;br&gt;
Why Hindsight for this?&lt;br&gt;
Because real-world procurement data is messy by design. Phrases like "Truck breakdown" and "Vehicle failure in heavy rain" must match semantically. A traditional key-value store or relational DB can't bridge that gap; Hindsight does.&lt;/p&gt;

&lt;p&gt;Resources: Hindsight GitHub | Hindsight Documentation&lt;/p&gt;

&lt;p&gt;Retention Quality Controls Recall Quality&lt;br&gt;
If you retain low-context tuples like:&lt;br&gt;
JSON&lt;br&gt;
{"vendor": "X", "late": 22}&lt;br&gt;
...you get poor recall.&lt;/p&gt;

&lt;p&gt;However, if you retain rich, natural narrative strings like:&lt;/p&gt;

&lt;p&gt;Plaintext&lt;br&gt;
"Structural Steel from SteelCore in July monsoon: +22 days late, 12% rejected, excuse truck breakdown, flagged by Apex team"&lt;br&gt;
...you unlock excellent semantic matches for future queries.&lt;/p&gt;

&lt;p&gt;Implementation:&lt;br&gt;
Python&lt;/p&gt;

&lt;h1&gt;
  
  
  retain.py - what I actually store in memory
&lt;/h1&gt;

&lt;p&gt;hindsight.retain(&lt;br&gt;
    text=f"Vendor {vendor} delivered {material} in {month}: "&lt;br&gt;
         f"{delay} days late, {rejection}% rejected. "&lt;br&gt;
         f"Reason: {excuse}. Outcome: {outcome}. Lesson: {lesson}",&lt;br&gt;
    metadata={"vendor": vendor, "material": material, "month": month}&lt;br&gt;
)&lt;br&gt;
Key Takeaway&lt;br&gt;
The best memory system will fail if your source data lacks a narrative.&lt;/p&gt;

&lt;p&gt;Real datasets give you credible numbers.&lt;/p&gt;

&lt;p&gt;Synthetic context gives you recallable narrative.&lt;/p&gt;

&lt;p&gt;You need both. DataCo and Vendor KPIs provide baseline credibility, while synthetic excuses provide semantic recallability.&lt;/p&gt;

&lt;p&gt;The Result&lt;br&gt;
When we query Hindsight for "Structural Steel October high priority", it doesn't just return average historical delay—it surfaces the July and August failures linked to repeated excuses.&lt;/p&gt;

&lt;p&gt;The baseline evaluation sees a 4.1 rating. Agent memory sees the pattern.&lt;/p&gt;

&lt;p&gt;Note: The human procurement manager remains the ultimate decision-maker—we never auto-purchase. We simply ensure the next evaluation remembers last July.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>architecture</category>
      <category>data</category>
    </item>
  </channel>
</rss>
