DEV Community

Karthik Nandagiri
Karthik Nandagiri

Posted on

Building ParcelGuard AI: Architecture and Evidence-Based Delivery Recovery

A delivery attempt fails because the recipient is unavailable. The operations team now has to decide what to try next: confirm another delivery time, coordinate a handoff, or choose a different recovery path. A useful system should make relevant prior outcomes visible without confusing a remembered story with verified evidence, and it should leave the decision and any real-world action with a person.

ParcelGuard AI explores that boundary. It is a memory-driven delivery-recovery prototype built around PostgreSQL operational records, a Spring Boot API, a React dashboard, and optional Hindsight recall. It is not a dispatch platform: it does not contact customers, dispatch couriers, or execute delivery actions.

A small, explicit architecture

The frontend is a React application built with Vite. Its in-app navigation switches between overview, incidents, experience history, approvals, and Hindsight comparison views. A shared API utility sends JSON requests to the Spring Boot service at http://localhost:8080. The backend is organized around Spring MVC controllers, services, JPA entities, and repositories. PostgreSQL stores incidents, experiences, and approval records; Hindsight is an external memory service.

flowchart LR
    Ops[Operations user] --> UI[React + Vite dashboard]
    UI -->|JSON over HTTP| API[Spring Boot REST API]
    API --> IC[DeliveryIncidentController]
    API --> EC[DeliveryExperienceController]
    API --> AC[RecoveryApprovalController]
    IC --> RS[RecoveryRecommendationService]
    RS -->|Query and validate evidence| PG[(PostgreSQL)]
    RS -->|Optional recall| HMS[HindsightMemoryService]
    EC -->|Save and verify outcomes| PG
    EC -->|Retain verified memory| HMS
    HMS <--> HC[Hindsight Cloud]
    AC --> AS[RecoveryApprovalService]
    AS --> PG

The main request path starts with GET /api/incidents/{id}/recommendation. DeliveryIncidentController.getRecommendation delegates to RecoveryRecommendationService.recommend. The service loads the incident, queries experiences for the same failure reason, applies eligibility checks, optionally recalls memories, and returns a RecoveryRecommendationResponse containing the incident ID, failure reason, action, supporting experience IDs, explanatory message, and memory insights.

PostgreSQL is the evidence boundary

An experience is not recommendation evidence merely because it exists. POST /api/experiences/{incidentId} stores a recovery action and outcome as unverified. A separate human verification request marks the record verified. During recommendation generation, the repository orders matching verified experiences by creation time and ID, newest first; the service then keeps only successful outcomes, non-audit actions, and nonblank recovery actions.

The relevant filtering is implemented in RecoveryRecommendationService.recommend:

List<DeliveryExperience> successfulExperiences = experiences.stream()
        .filter(DeliveryExperience::isVerified)
        .filter(e -> e.getOutcome() != null
                && "DELIVERED_SUCCESSFULLY"
                .equalsIgnoreCase(e.getOutcome().trim()))
        .filter(e -> !isAuditExperience(e))
        .filter(e -> e.getRecoveryAction() != null
                && !e.getRecoveryAction().isBlank())
        .toList();
Enter fullscreen mode Exit fullscreen mode

If the eligible list is empty, the API returns no recommendation and no supporting IDs, even if Hindsight returns text that sounds persuasive. That behavior is important: an unverified recollection cannot establish that a recovery action worked.

Hindsight ranks references, not prose

When parcelguard.hindsight.enabled is true, RecoveryRecommendationService builds a query from the incident’s failure reason and description and calls HindsightMemoryService.recall. On experience verification, DeliveryExperienceController.verifyExperience attempts to retain a memory containing the PostgreSQL experience ID, incident ID, failure reason, action, outcome, and verification status. If remote synchronization fails, the PostgreSQL row remains verified; memorySynced is only set after a successful retain call.

Hindsight can affect ranking only through an explicit Experience ID: <number> reference that matches a candidate already accepted by the PostgreSQL filters. The selected action and supporting ID are then read from that PostgreSQL entity. Otherwise, the service falls back to the newest eligible experience:

Optional<DeliveryExperience> hindsightSelection =
        selectRecalledEligibleExperience(
                successfulExperiences, memoryResults);
DeliveryExperience selected = hindsightSelection
        .orElse(successfulExperiences.get(0));
Enter fullscreen mode Exit fullscreen mode

Recall is optional and wrapped so runtime failures do not remove PostgreSQL recommendations. A returned memory is context, not proof of improved delivery success. The Hindsight Comparison screen makes this distinction visible; switching modes requires changing the backend property and restarting Spring Boot. The frontend does not toggle backend configuration.

For background on the external component, see Hindsight on GitHub, the Hindsight documentation, and Vectorize’s overview of agent memory.

A clearly simulated example

The service tests use mocked candidates rather than real deliveries: a newer eligible record with the action “Newest recovery action” and an older eligible record with “Recalled recovery action.” A mocked Hindsight response explicitly references the older record’s ID. The test asserts that the response contains the older PostgreSQL action and ID. Separate tests assert that disabled Hindsight, unmatched IDs, ineligible records, or recall failure retain the newest-eligible fallback. These labels are test fixtures, not evidence that a parcel was delivered.

For a manual demonstration, use a separate development database and a dedicated Hindsight bank. Add two clearly marked simulated historical experiences with the same failure reason, different actions, and synthetic successful outcomes; verify both through the UI so their memories carry the generated PostgreSQL IDs. Then use a simulated target incident with that failure reason. With Hindsight disabled, the expected baseline is the newest eligible experience. With it enabled, the action changes only if Hindsight recalls the older eligible ID. If it does not, fallback is correct. Avoid synthetic records in operational data and never create an approval just to show ranking.

The dashboard and human decision

The overview shows API-derived incident and approval counts, recent incidents, experience totals, and recommendation highlights. Incident details pair the suggested action with supporting experience IDs and their source records, plus clearly labeled supplementary Hindsight insights. The experience view distinguishes verified from unverified records and successful, failed, and rescheduled outcomes. The comparison page captures Before/After responses for a selected incident and keeps them associated with the incident and mode in browser sessionStorage.

Approval is a separate workflow. RecoveryApprovalService.createApproval calls the recommendation service and refuses to create a request without both a recommendation and supporting experience IDs. The resulting status is PENDING; reviewApproval accepts only APPROVED or REJECTED for a still-pending record. There is no downstream delivery executor.

if (recommendation.recommendation() == null
        || recommendation.supportingExperienceIds().isEmpty()) {
    throw new ResponseStatusException(
            HttpStatus.CONFLICT,
            "No evidence-backed recommendation is available");
}
Enter fullscreen mode Exit fullscreen mode

Operations overview with live incident and recovery summary data.

Incident detail with its recommendation, verified supporting experience, and supplementary Hindsight context.

Live Before/After captures for one incident; Hindsight adds memory context while the recommendation action remains evidence-backed.

Engineering lessons

Separate retrieval from authority. Hindsight is useful for surfacing a past record, but only PostgreSQL decides whether that record is verified, successful, and eligible. An ID match provides a controlled bridge between recall and evidence.

Make failure behavior predictable. Hindsight timeouts, malformed responses, and unmatched references fall back to the newest eligible PostgreSQL record. This preserves recommendation availability without quietly relaxing the evidence bar.

Verification and synchronization are different states. PostgreSQL is saved first. A remote retain failure is logged and leaves the record available as verified evidence, while synchronization can be retried. This avoids making an external dependency the owner of operational truth.

Keep approval explicit. A recommendation is not an instruction to dispatch. Requiring supporting evidence to create a pending approval keeps the human decision visible in the data model and UI.

Testing and limits

Automated backend coverage uses JUnit, Mockito, Spring integration tests, and an in-memory H2 database. The recommendation service tests exercise matched older IDs, disabled recall, unmatched and ineligible IDs, and recall failures; API integration tests cover verification and approval flows. The last recorded .\mvnw.cmd test run completed 17 tests with no failures. These results are not live end-to-end validation: integration tests mock Hindsight, use H2 rather than PostgreSQL, and do not prove provider recall behavior for a configured bank. No measured operational improvement is claimed.

ParcelGuard AI remains a prototype. It has no authentication or role-based access controls, its JPA configuration uses ddl-auto=update, and Hindsight recall quality depends on an external service. Future work could include migration-managed schemas, access control, explicit memory provenance and retention controls, and broader tests against isolated PostgreSQL and Hindsight environments—without making memory itself the authority for delivery outcomes.

Top comments (0)