DEV Community

Cover image for GoWithProof: I Built My Friend an AI That Argues Against Its Own Recommendations
Dayanand garg
Dayanand garg

Posted on

GoWithProof: I Built My Friend an AI That Argues Against Its Own Recommendations

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🀝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Choosing a place to go should be easy.

I built GoWithProof for a close friend I regularly make plans with. Whenever we decide where to eat, drink, or spend an evening around Delhi/Gurgaon, one of us inevitably ends up doing the same research loop: Google Maps, Reddit, restaurant apps, recent reviews, and then an AI search on top of all of it. We rarely trust a single source enough to make the decision.
That recurring problem was the reason for this project. I wanted to build something specifically for the way we actually make plans togetherβ€”not another generic β€œbest restaurants near me” app.

  • Search Google Maps
  • Check ratings
  • Open Reddit
  • Search for recent discussions
  • Look at restaurant platforms
  • Read the newest reviews
  • Read the worst reviews
  • Ask an AI
  • Discover that every source says something different
  • Still not know where to go

The problem is not that recommendations are unavailable.

The problem is that there are too many recommendations, from sources with different incentives, different levels of freshness, and different blind spots.

So I built GoWithProof.

Recommendations with receipts.

GoWithProof is an AI-assisted decision engine for choosing places.

Instead of asking an LLM:

"What's the best pub in Gurgaon?"

and trusting whatever appears next, GoWithProof does something closer to how I would research the decision myself.

It:

  1. understands what you actually care about,
  2. discovers plausible places,
  3. researches recent reviews and live web evidence,
  4. deliberately searches for complaints and counter-evidence,
  5. extracts grounded claims,
  6. measures evidence quality and disagreement,
  7. ranks candidates deterministically,
  8. and shows both why you should consider a place and why you might not want to go there.

The feature I care about most is called Devil's Advocate.

Most recommendation systems spend their effort finding reasons why their recommendation is good.

GoWithProof also asks:

What could make this recommendation wrong?

If a place looks great for craft beer but recent visitors repeatedly complain about service, noise, pricing, crowds, or reservations, I want to know that before we leave home.

It doesn't require a perfect prompt

A user can start with something vague:

"Find me a pub in Gurgaon."

GoWithProof progressively asks for the information that will materially change the decision.

For example:

"Friday around 8 PM."

"Around β‚Ή1500 per person."

"Craft beer."

It builds a structured decision context across the conversation instead of forcing users to express every constraint in one giant prompt.

During development, I also learned that specific preferences need to stay specific.

For example:

craft beer != drinks
live music != music
Enter fullscreen mode Exit fullscreen mode

If someone asks for craft beer, a generic cocktail bar should not automatically count as a strong match.

Preserving that meaning throughout clarification, discovery, research, and ranking became one of the most important parts of the system.


Demo

πŸ‘‰ Live Demo: https://go-with-proof.onrender.com

You can try something like:

Dinner in CP tomorrow at 7 PM, 4 people, around β‚Ή1000–1200 each, great food, vegetarian options, and quiet enough to talk.

Or something simpler:

Find me a pub in Gurgaon with good craft beer.

The application will clarify missing details when necessary, research candidates, and return evidence-backed recommendations.


Code

The complete project is open source:

πŸ‘‰ Repository: https://gitlab.com/gargdaya/go-with-proof

The project is a small monorepo:

go-with-proof/
β”œβ”€β”€ frontend/
β”‚   └── Next.js + React + TypeScript
β”‚
β”œβ”€β”€ backend/
β”‚   └── Bun + Elysia + TypeScript
β”‚
β”œβ”€β”€ .gitlab-ci.yml
└── render.yaml
Enter fullscreen mode Exit fullscreen mode

GitLab CI runs the automated checks and production build, while the frontend and backend are deployed separately on Render.


How I Built It

One of the main design decisions was to avoid building a system where an LLM receives a prompt and simply returns three restaurant names.

I wanted to use the model where language models are strong, while keeping the actual decision process understandable and testable.

The architecture looks roughly like this:

User
  β”‚
  β–Ό
Conversational clarification
  β”‚
  β–Ό
Gemma
intent + constraints
  β”‚
  β–Ό
DecisionContext
  β”‚
  β–Ό
Candidate discovery
  β”‚
  β–Ό
SerpApi
  β”‚
  β”œβ”€β”€ Google Maps
  β”œβ”€β”€ Reviews
  └── Web/community search
          β”‚
          β–Ό
Counter-evidence searches
          β”‚
          β–Ό
Gemma evidence extraction
          β”‚
          β–Ό
Grounding + Zod validation
          β”‚
          β–Ό
Deduplication + normalization
          β”‚
          β–Ό
Deterministic TypeScript ranking
          β”‚
          β–Ό
Recommendations
  β”‚
  β”œβ”€β”€ Why it fits
  β”œβ”€β”€ Confidence
  β”œβ”€β”€ Concerns
  └── Devil's Advocate
Enter fullscreen mode Exit fullscreen mode

There are really three different jobs in that pipeline:

Understand

What does the user actually want?

Research

What evidence exists about the candidates?

Decide

How well does each candidate fit the current situation, and how much evidence do we actually have?

Keeping those responsibilities separate made the system much easier to reason about.


Gemma: Language Understanding, Not an Oracle

Gemma is at the core of GoWithProof, but it does not get to arbitrarily choose the final answer.

I use two Gemma models for different workloads.

Gemma 3 4B

The smaller model handles conversational clarification.

This operation happens frequently and is latency-sensitive.

Its job is to understand statements such as:

"tomorrow around seven"
"somewhere we can talk"
"around β‚Ή1200 each"
"craft beer"
"vegetarian options"
Enter fullscreen mode Exit fullscreen mode

and turn them into structured context.

Gemma 3 27B

The larger model is used where understanding messy natural language becomes more important: extracting structured evidence from reviews and search results.

Its job is to turn unstructured text into evidence such as:

{
  "category": "drinks",
  "sentiment": "positive",
  "claim": "The beer quality was praised.",
  "supportingQuote": "Superb quality beer....",
  "firsthand": true,
  "confidence": 0.8
}
Enter fullscreen mode Exit fullscreen mode

But Gemma does not calculate the final recommendation score.

That is deliberate.

The model interprets language.

The application makes the decision.

Things such as:

  • freshness weighting,
  • preference matching,
  • negative ratios,
  • source diversity,
  • practical constraints,
  • evidence quantity,
  • disagreement,
  • and confidence

are handled deterministically in TypeScript.

That gives me something I can test and reason about.


SerpApi: Connecting the System to Live Evidence

A recommendation engine isn't very useful if its knowledge stops at model training time.

GoWithProof uses SerpApi as its live research layer.

It provides information used for:

  • place discovery,
  • Google Maps results,
  • business metadata,
  • ratings and review counts,
  • operating information,
  • recent reviews,
  • low-rated reviews,
  • Google Search results,
  • community discovery,
  • and deliberate counter-evidence searches.

The key idea is that SerpApi provides evidence, not the final answer.

GoWithProof does not assume:

"Google Maps ranked this first, therefore it must be the best."

Instead, those results become inputs into a larger decision process.

That distinction matters.


I Search for Reasons NOT to Recommend a Place

This is probably the most important part of the project.

For a candidate, GoWithProof doesn't only run normal searches.

It deliberately generates searches such as:

"<place>" bad experience
"<place>" overrated
"<place>" crowded
"<place>" expensive
"<place>" noisy
Enter fullscreen mode Exit fullscreen mode

Those searches can also depend on what the user actually cares about.

If the user says:

"I want somewhere quiet enough to have a conversation."

then evidence about noise becomes especially important.

If they want craft beer, evidence about the beer matters more than generic evidence about cocktails.

This became the Devil's Advocate part of GoWithProof.

Instead of only asking:

"How can I justify recommending this?"

the system also asks:

"What evidence would make me change my mind?"


Every AI Claim Needs a Receipt

One of my favorite bugs from the weekend changed the architecture.

During a live test, a review said something positive about vegetarian food.

Gemma produced a claim saying both:

the vegetarian food was good and the music was good

The source excerpt did not actually support the music claim.

It was a small hallucination.

But for GoWithProof, that is a serious problem.

The entire product is based on the idea that recommendations should be inspectable.

So I changed the evidence model.

Every extracted claim now contains a supportingQuote.

Conceptually:

type Evidence = {
  sourceType:
    | "google_maps_review"
    | "reddit"
    | "official"
    | "web";

  sourceUrl: string;
  sourceDate?: string;

  category:
    | "food"
    | "service"
    | "noise"
    | "price"
    | "crowd"
    | "ambience"
    | "drinks"
    | "wait_time"
    | "cleanliness"
    | "other";

  sentiment:
    | "positive"
    | "negative"
    | "mixed"
    | "neutral";

  claim: string;
  supportingQuote: string;
  firsthand: boolean;
  confidence: number;
};
Enter fullscreen mode Exit fullscreen mode

The backend verifies that the supporting quote actually exists in the source material before accepting the evidence.

If an extracted claim cannot be grounded in the retrieved text, it should not survive the pipeline.

That was an important lesson:

LLMs can be excellent evidence interpreters, but evidence still needs verification.


Negative Evidence Is First-Class Data

I also wanted to avoid the classic recommendation pattern:

4.5 stars β€” highly recommended!

A place can have thousands of excellent reviews while still having a recurring problem that matters specifically to you.

So GoWithProof tracks negative evidence explicitly.

For every candidate it can calculate things such as:

total evidence
positive evidence
negative evidence
negative ratio
recent negative evidence
source diversity
Enter fullscreen mode Exit fullscreen mode

The ratio matters more than simply counting complaints.

For example:

5 negative signals out of 8
Enter fullscreen mode Exit fullscreen mode

is very different from:

5 negative signals out of 80
Enter fullscreen mode Exit fullscreen mode

even though the raw negative count is identical.

Freshness matters too.

A complaint from last month should generally matter more than a complaint from four years ago.


Match and Confidence Are Not the Same Thing

This became another important lesson.

Imagine a place appears to perfectly satisfy what the user asked for.

It is:

  • the right kind of venue,
  • in the right area,
  • open at the requested time,
  • and associated with the requested preference.

That can make it look like an excellent fit.

But what if the system only managed to collect one useful piece of evidence?

Then these two statements are very different:

This place looks like a good fit.

and:

We have strong evidence that this place is a good fit.

GoWithProof therefore separates recommendation fit from evidence confidence.

A candidate can effectively be:

Strong potential fit Β· Low confidence

instead of turning sparse evidence into false certainty.

When the evidence is too weak, the application explicitly labels it:

Weak evidence

I think recommendation systems should be allowed to admit uncertainty.


Disagreement Is Useful Information

Sources do not always agree.

Instead of averaging disagreement away, GoWithProof can track it by category.

Imagine the research produced:

Food
8 positive
1 negative

Service
6 positive
5 negative

Noise
1 positive
7 negative
Enter fullscreen mode Exit fullscreen mode

Those categories tell very different stories.

Food looks consistently strong.

Service is disputed.

Noise looks like a legitimate concern.

A single average score hides all of that.

GoWithProof preserves disagreements so the user can decide whether they matter.


Failure Is Part of Live Web Research

Real-time research is messy.

During development I encountered:

  • empty search results,
  • malformed model output,
  • network failures,
  • provider timeouts,
  • missing dates,
  • duplicate URLs,
  • evidence that could not safely be attributed,
  • and community searches with no useful result.

I didn't want one unavailable source to make the whole recommendation fail.

Each research channel therefore has its own coverage state.

Conceptually:

{
  "recentReviews": "ok",
  "lowRatedReviews": "ok",
  "community": "empty",
  "counter": "ok"
}
Enter fullscreen mode Exit fullscreen mode

The system can still return useful partial research.

And importantly, it can tell the user when evidence is incomplete instead of silently pretending everything succeeded.


The First Research Pipeline Took 85 Seconds

One of my first successful live research runs took around 85 seconds.

It technically worked.

But nobody wants to stare at a loading spinner for a minute and a half while deciding where to eat.

The first implementation was performing too much research across too many candidates.

So I changed the orchestration.

Instead of deeply researching everything immediately, the pipeline became closer to:

Broad discovery
      β”‚
      β–Ό
Candidate filtering
      β”‚
      β–Ό
Smaller research set
      β”‚
      β–Ό
Parallel source collection
      β”‚
      β–Ό
Evidence extraction
Enter fullscreen mode Exit fullscreen mode

Independent searches run concurrently with bounded concurrency.

Candidate research is capped.

Partial results can survive the research time budget.

One later live run dropped from roughly 85 seconds to around 15 seconds.

That was a useful engineering lesson:

Better prompting cannot compensate for inefficient orchestration.


Progressive Clarification Was Harder Than Expected

The clarification system also went through several iterations.

Initially, every answer was sent back through Gemma.

That caused a funny but important failure.

The UI asked:

"What matters most?"

The user clicked:

"Drinks"

The backend failed to preserve that answer correctly.

So it asked:

"What matters most?"

again.

And again.

The fix was simple once it was obvious:

Do not use AI to interpret information your own UI already knows.

Quick-choice answers are now merged deterministically.

Gemma is reserved for genuinely free-form language such as:

"good beers and somewhere we can actually talk"
Enter fullscreen mode Exit fullscreen mode

That reduced unnecessary inference and removed an entire class of clarification loops.


Specific Intent Should Stay Specific

Another problem appeared when:

craft beer β†’ drinks
live music β†’ music
Enter fullscreen mode Exit fullscreen mode

Those transformations are technically reasonable abstractions, but they destroy information that matters to the decision.

A person asking for craft beer does not simply want "drinks".

A person asking for live music does not necessarily want a DJ.

So GoWithProof preserves specific traits alongside broad categories.

For example:

{
  "priorities": ["drinks"],
  "traits": ["craft_beer"]
}
Enter fullscreen mode Exit fullscreen mode

That information can then influence discovery and ranking.

It also prevents unrelated evidence from being treated as more relevant than it really is.

For example, a complaint about cocktails should not automatically mean the craft beer is bad.


Testing

The backend has tests around the areas I considered most dangerous:

  • context accumulation,
  • explicit corrections,
  • timezone-aware dates,
  • quick-choice merging,
  • clarification loops,
  • model validation,
  • evidence grounding,
  • deduplication,
  • unknown source dates,
  • partial provider failure,
  • freshness weighting,
  • negative ratios,
  • source diversity,
  • disagreement detection,
  • ranking,
  • confidence,
  • and weak-evidence handling.

By the time I froze the backend for submission:

86 automated tests were passing

along with successful type checks and production builds.

For a weekend project, writing those tests saved much more time than they cost.


Tech Stack

Frontend

  • Next.js
  • React
  • TypeScript
  • Tailwind CSS
  • shadcn/ui
  • Lucide

Backend

  • Bun
  • TypeScript
  • Elysia
  • Zod

AI

  • Gemma 3 4B
  • Gemma 3 27B
  • Hugging Face Inference

Live Research

  • SerpApi
  • Google Maps results
  • Google Maps reviews
  • Google Search and community discovery

Deployment

  • Render Web Service for the frontend
  • Render Web Service for the backend
  • GitLab CI/CD

Both production services are described in render.yaml.

GitLab CI performs verification before deployment:

git push
   β”‚
   β–Ό
GitLab CI
   β”‚
   β”œβ”€β”€ install
   β”œβ”€β”€ tests
   β”œβ”€β”€ typecheck
   └── production build
          β”‚
          β–Ό
        Render
       /      \
 frontend    backend
Enter fullscreen mode Exit fullscreen mode

There is no PM2, Docker, Redis, queue, vector database, or giant agent framework involved.

For this project I deliberately preferred a small system whose decisions I could still understand.


Why Does Open Innovation Matter?

GoWithProof is fundamentally about questioning opaque authority.

The application doesn't say:

"Trust this recommendation because AI said so."

It says:

"Here is the evidence I found. Here is what supports the recommendation. Here is what contradicts it. You decide."

I wanted the AI layer to follow the same philosophy.

Using an open-weight model like Gemma means the language-understanding layer isn't inseparably tied to a single proprietary recommendation system.

The architecture keeps clear boundaries:

Gemma
"What does the user mean?"
"What does this evidence say?"

SerpApi
"What is happening on the live web?"

Application code
"How fresh is this?"
"How relevant is it?"
"How much independent evidence exists?"
"Where do sources disagree?"
"How strongly should we recommend it?"
Enter fullscreen mode Exit fullscreen mode

That separation gives me control over the actual decision logic.

I can change model sizes.

I can change inference providers.

I can inspect and iterate on prompts.

I can test deterministic scoring independently.

And because Gemma is open-weight, the model itself can eventually run in environments where sending every user interaction to a closed hosted model is not desirable.

A completely closed recommendation API would give me an answer.

Open components let me build a decision process.

That difference is the entire point of GoWithProof.

Don't ask for authority. Ask for evidence.


What I Learned

I started the weekend thinking I was building a recommendation engine.

By the end, I realized I was really building a small decision engine.

The interesting questions were not:

"How do I ask an LLM for the top three restaurants?"

They were:

"How do I know whether this evidence is actually about this restaurant?"

"What happens when recent reviews disagree with the overall rating?"

"Should five negative reviews matter if there are eighty positive ones?"

"How do I distinguish a place that looks right from a place I have strong evidence for?"

"What happens when one research source fails?"

"How do I stop the model from turning 'craft beer' into generic 'drinks'?"

"How do I prevent an extracted claim from saying more than its source actually said?"

Those questions made the project much more interesting than the original idea.

The biggest lessons were:

  • Discovery and verification are different problems.
  • Specific user intent should remain specific.
  • Negative evidence deserves equal treatment.
  • Freshness matters.
  • Source disagreement should be exposed, not averaged away.
  • A model-generated claim needs provenance.
  • Match quality and evidence confidence are different concepts.
  • Partial failure is normal in live web research.
  • Deterministic code is often better than another LLM call.
  • Sometimes "I don't have enough evidence" is the best recommendation.

What I'd Build Next

There is a lot more I would like to explore with GoWithProof.

Future versions could include:

  • post-visit feedback,
  • gradual preference learning,
  • remembering what kinds of places a user actually enjoys,
  • comparing places friends have already suggested,
  • stronger cross-platform candidate discovery,
  • deeper Reddit/community research,
  • on-demand verification of a specific place,
  • recommendation history,
  • streaming research progress,
  • trip and itinerary decisions,
  • and eventually decisions beyond restaurants entirely.

But those are extensions.

The core idea is already:

Discover
   ↓
Verify
   ↓
Challenge
   ↓
Decide
Enter fullscreen mode Exit fullscreen mode

Prize Categories

Best Use of Gemma

Gemma is a core part of GoWithProof rather than an optional chatbot layer.

Gemma handles natural-language clarification and turns messy live research into structured evidence.

I use a smaller Gemma model for latency-sensitive clarification and a larger model for deeper evidence interpretation.

The final ranking remains deterministic, making the boundary between model interpretation and application decision-making explicit.

Best Use of SerpApi

SerpApi is GoWithProof's connection to the live world.

It powers place discovery, Maps information, review retrieval, organic search, community discovery, and the deliberate counter-evidence searches that make Devil's Advocate possible.

Without live search, the product would simply be another model recommending places from stale training knowledge.

SerpApi turns it into a research system.

Best Use of Render

The complete production application runs on Render.

Both the Next.js frontend and Bun/Elysia backend are deployed as independent Render Web Services.

The backend running on Render orchestrates live SerpApi research, Gemma inference, validation, evidence processing, and ranking.

Deployment configuration is committed as infrastructure-as-code through render.yaml, while GitLab CI runs tests, type checks, and production builds before deployment.


I built GoWithProof because I wanted a recommendation system I could disagree with.

Not one that hides uncertainty behind a confident answer.

Not one that assumes the highest-rated result must be right.

And not one that asks me to trust an AI without showing me where its answer came from.

Sometimes the most useful thing an AI can say isn't:

"Go here."

It's:

"This looks promising. Here's the evidence. Here's what worries me. Now you decide."

That's GoWithProof.

Recommendations with receipts.

Top comments (0)