Anyone who has asked a general chatbot about tonight's game knows the failure mode: it confidently cites a roster from three seasons ago, invents an injury report, and never saw a live price in its life. Language models are good at reasoning about sport. They are terrible at knowing what happened ten minutes ago. The fix is not a better prompt. It is architecture.
Here is what I learned building Betaware AI, an iOS and Android app that researches a matchup and answers with a confidence score. (I make Betaware AI. This is about how it is built, not a sales pitch.)
Separate the three jobs
The single most useful decision was splitting the pipeline into three stages that never swap roles.
- Retrieval. Fetch live facts before the model reasons at all: injury reports, lineup news, recent form, head to head, and the current market prices. The model should not be trusted to know what is true today.
- Modelling. Run actual statistical models over the retrieved data, not vibes. Team-level features, form curves, home and away splits. The model outputs a probability, not a paragraph.
- Narration. Only now let the language model explain the result in plain English, grounded strictly in what stages one and two returned.
If you let the language model free associate across all three jobs, you get fluent nonsense. If you keep the boundaries, you get something a user can actually audit.
Make the market part of the input
A model's probability in a vacuum is trivia. The interesting question is always: what does the market think, and where is the gap? So the live line is a first class input, not a footnote. Pulling the moneyline, spread and totals from multiple venues (we read DraftKings and the Kalshi exchange, plus bet365 on some European soccer leagues), stripping the vig to get a consensus probability, and comparing that to the model's number turns "who wins?" into "what does my estimate imply versus the price?". It also means the answer is falsifiable the moment the game ends.
Ship the uncertainty, do not hide it
Every prediction carries a confidence score, and this is the part users push back on before they trust it. The answer: grade yourself in public. Past predictions in the app are checked against final results, so the confidence number is something you can hold the system to rather than a decoration. If you build anything predictive and only surface the hits, you have built a marketing tool, not a research tool.
Guardrails are a feature, not legal boilerplate
Because the subject is money, we pinned hard constraints early: 18 and over only, informational and entertainment purposes only, no wagers placed or accepted anywhere in the product, and a responsible gambling page with real guidance (budgets, time limits, never chasing losses). Predictions are analysis and can be wrong. Saying so plainly turned out to build trust rather than kill it.
What I would tell another builder
- Never let the LLM fetch its own facts. Give it facts.
- Model first, narrate second. The paragraph should be a rendering of the number, not the source of it.
- Compare against a market or a benchmark, otherwise you cannot evaluate anything.
- Show confidence, then grade it publicly. Calibration is the whole product.
- Keep the answer short. Research deep, answer in one screen.
Betaware AI is live on the App Store and Google Play if you want to poke at the output side of this pipeline: https://betaware.app
Top comments (0)