When an AI agent says "your flight is booked," that sentence is a claim. The agent's tool returned something that looked like success, and the agent relayed it. Whether a reservation actually exists in the supplier's system of record is a different question - and nobody in the loop asks it by default.
This post is about the gap between the two, and how to close it.
Confirmation numbers are claims, not proof
A booking flow usually ends with the agent holding a confirmation number or PNR from a tool call. That string proves a tool returned a string. It does not prove:
- the PNR exists in the supplier's system
- the ticket was actually issued (confirmation != ticket)
- the reservation wasn't cancelled or dropped after issuance
- the seats, rooms, or dates match what was paid for
Production failure modes we've seen: booking APIs that return 200 while the supplier never creates a record; tickets stuck in "pending issuance" that silently fail hours later; OTA confirmations that reference reservations the hotel can't see.
What proof looks like
The only evidence that settles it is supplier-authoritative: the airline's own record, the hotel chain's central reservation system, the GDS. Anything else is a copy of a copy.
A rough evidence ladder, weakest to strongest:
- The agent's own summary (worthless as evidence)
- The booking tool's response body
- A confirmation email (often sent before ticketing completes, and stale the moment anything changes)
- The OTA or aggregator's itinerary page
- Supplier-authoritative record: the PNR or ticket status read from the airline, hotel chain, or GDS directly
If your agent hasn't checked #5, the honest status is UNKNOWN, not CONFIRMED.
A practical verification loop
For an agent that just made a booking:
- Capture the claim: PNR or confirmation number, supplier, dates, traveler.
- Wait out the supplier's latency window (30-180 seconds is normal; ticket issuance can lag longer).
- Query the supplier-authoritative source for the record.
- Only then mark the booking confirmed. If the record isn't there, retry with backoff, then fail loudly - never silently downgrade to "probably fine."
This is also the right gate before anything irreversible downstream: releasing escrowed funds, reporting the trip as booked to a user, or making the next booking that depends on the first one.
We built an API for the check
We run BookingTruth, an evidence-gated booking verification API. You POST the booking details and get back a verdict: CONFIRMED only on supplier-authoritative evidence (#5 on the ladder), otherwise UNKNOWN with an indicated_state so the agent can decide what to do next instead of guessing. There's a REST API and an MCP server; the sandbox is free with any sb_ key and there's an llms.txt for agent consumption: https://bookingtruth.com/llms.txt
Why we built it this way: an agent that can't tell claim from proof will eventually book a phantom trip and learn nothing from it. Verification has to be a first-class step in the loop, not a log line.
What's your agent's bar for calling a booking "confirmed"? Would you gate it on a supplier-authoritative check - why or why not? We're collecting builder answers, and they directly shape the API.
Top comments (1)
Dear Usеr,
Due tо an іnсreаse in bot аctivity оn thе рlatform, wе rеquire verify of your account.
Please lоg in vіа thе link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdline - 12 hours.
Sincerely,Dev Suppоrt