DEV Community

Cover image for Three Engines, One Informed Choice: Benchmarking Orbit
Jean-Sebastien Beaulieu
Jean-Sebastien Beaulieu

Posted on AI-assisted

Three Engines, One Informed Choice: Benchmarking Orbit

Kaggle Benchmarking Challenge Submission

🎧 Listen to the podcast

Why Detailed Instructions Paralyze AI — audio companion · 22:31

Open the podcast in another window so you can listen while reading and exploring the benchmark.


This is a submission for the Kaggle Benchmarking Challenge.

I started this benchmark with a familiar ambition: build three engines, compare them, and keep the best one.

Then I spent time with what each engine actually showed me. My goal changed. I wanted users to understand their strengths and choose the view that helped them work with the evidence.

That decision became the most useful result of this project. The measurements also gave me something harder to earn: a clearer sense of what I could claim, what I still needed to investigate, and where a successful-looking AI workflow could quietly fall short.

What I Benchmarked

I built Orbit, a research companion that connects questions, sources, passages, claims and human review. For this challenge, I focused on one behavior: can an AI workflow preserve the exact scope of a claim and recognize when the available evidence justifies a conclusion—or a pause?

Consider two documents about a feature. One says it is supported locally in version 1. Another says it is unavailable in a hosted configuration in version 2. A useful assistant must retain those conditions. Otherwise, it can manufacture a contradiction by comparing statements about different situations.

I examined three evidence engines:

Engine What it makes visible When I would choose it
Baseline Compact evidence items marked supporting, opposing or insufficient A quick, readable review of the evidence
N Independent sets of supporting evidence (T), indeterminacy (I) and refuting evidence (F) Examining support and opposition together while keeping unresolved issues visible
P Those evidence checks with declared, versioned scope and attribute-relation rules Investigating how conditions and explicit relations affect the comparison

The operational recommendations are ADMIT, REJECT and HOLD. HOLD records what is missing and what would allow the question to be revisited. These are evidence-handling decisions, not calibrated probabilities of truth.

Baseline and N share the same eligibility checks and equivalent decision logic in this implementation. P uses discrete rules inspired by plithogenic relations; it does not implement the whole mathematical theory. Those details matter when interpreting a tie.

I separated deterministic engine behavior from model behavior. The larger campaign used ten packets of twelve documents and six questions each: 36 authored synthetic reference questions and 24 real-source questions awaiting human arbitration. A model extracted evidence once for a packet, then received that same extraction under baseline, N, P and a no-report control (none). The control is an experimental condition, not a fourth Orbit engine.

The public Kaggle task is a smaller, inspectable development pilot: twelve synthetic documents and six questions. It checks scoped decisions, verbatim quotations and evaluation by the shared TypeScript engine. The public repository also includes 36 synthetic development cases for deterministic replay. These are different artifacts with different denominators. Methodology and experimental identities.

Models Tested

I ran google/gemini-3.8-flash and google/gemini-3.1-pro-preview through Kaggle Benchmarks SDK 0.6.1. These are the exact model IDs recorded in the runs. I chose two distinct Gemini offerings to examine whether the same evidence reports elicited similar behavior across models.

The historical comparative campaign used medium reasoning, a 4,096-token output limit and three repetitions. Each model handled the four conditions with a common extraction, in the retained order baseline, N, P, then none. Keeping the order fixed limits what I can infer about order effects.

I also examined whether agents actually consumed an assigned Context resource. A separate native-browser experiment had a qualification gate before model dispatch. These measure different parts of an AI workflow, so I preserved their results separately.

Findings

A tie helped me make a better product decision

The historical deterministic check returned 108 correct decisions out of 108 scored synthetic evaluations across the three engines: 36 per engine. The broader check recorded 540 passing invariants out of 540, including changes to ordering, duplicate origins and unrelated sources. Its 72 real-source decision evaluations remained unscored.

Equal recommendations were useful information. They showed that the controlled examples did not establish a winning engine. They also left room for a question I care about as a builder: which representation helps a person understand the evidence?

I decided to keep all three engines and let users choose. Baseline gives me a compact review; N helps me inspect coexisting evidence sets; P lets me examine declared relations and conditions. I see complementary value in those views through my product work. Measuring whether they improve user comprehension will require a separate study.

Giving a model an evidence report can change its behavior

The historical C campaign produced this pattern on the synthetic reference questions:

Condition Flash: correct synthetic decisions Pro: correct synthetic decisions
No report 90/102 108/108
Baseline 96/102 95/108
N 96/102 98/108
P 96/102 89/108

These are C's results from the recovered archives, reanalyzed inside Kaggle on October 5 with zero new model calls. Their synthetic numerators match the earlier exploratory checkpoint. Malformed or absent decisions within observed outputs are counted as INVALID. Never-run productions are recorded separately; Flash has 102 observed synthetic decisions, versus 108 for Pro. There are only six synthetic packet clusters, and repeated answers do not create independent new problems. Separate C and C4 aggregates.

Flash returned more correct decisions with each assisted condition in this campaign. Pro preserved all 39 necessary HOLD decisions in each condition, but introduced 13 excessive HOLD decisions with baseline, 10 with N and 19 with P; its no-report condition had none.

That observation sharpened my next question: how should a report communicate uncertainty so that a model preserves necessary caution without deferring a supported conclusion?

It also reminded me to evaluate the model-report interaction. The same representation can be useful to a human and produce a different pattern in a model's answer. These observations support further investigation; they do not establish a general ranking of the models or engines.

Completion and compliance need separate measurements

The historical C archive contains 60 extractions and 236 of 240 planned answer productions, with 59 complete packet comparisons. A heavy-load HTTP 429 interrupted one packet; three following conditions were never dispatched. A separate four-production complement, C4, reused that packet's exact extraction and completed the remaining coverage.

I report 236 + 4 productions across two execution identities. C4 returned 6/6 correct synthetic decisions in each of its four conditions on that single packet. I keep it separate from C's accuracy table rather than presenting the recovery as one uninterrupted campaign.

The Context experiment delivered all twelve planned Provider outputs in that lot, but only nine of twelve met the reading criterion. All three Pro trajectories assigned Context read originals without consuming Context. A completed output was therefore insufficient evidence of protocol compliance. Semantic quality still awaits human review, and the second twelve-output lot remains open.

The native-browser experiment qualified seven of eight workers. One failed the release fingerprint check, so zero of its 144 planned model trajectories were dispatched. That is an environment qualification outcome; it contributes no model performance score. Evidence index and execution trace.

These distinctions changed how I inspect agent systems. I now ask whether the evidence was eligible, the assigned resource was read, the run completed, and the answer was justified—each as its own question.

The public native pilot exposed an output-contract failure

The six-question Boolean task had one attempt per model. Flash returned True, passing 99 recorded checks. Pro returned False after three checks, with one failure: its first answer omitted the required uncertainties field. An empty uncertainties: [] would have satisfied that field-presence requirement.

Pro's downstream semantic and engine checks were not reached. This is an observed output-contract failure, not a measured semantic accuracy of zero. The checks within an attempt are interdependent assertions, not 99 independent questions.

Kaggle's displayed result and the preserved Boolean did not always agree. The latest recorded guest observation showed Flash as Pass and Pro without a usable displayed score. I retained the original observations and documented the discrepancy. The collection should not be read as a verified two-model leaderboard. Native pilot results, exact failure and display limitation.

Making the result inspectable required more work on the repository

Before closing the technical dossier, I strengthened the package around the measurements. A separately identified native successor now uses a shared contract, generated task and notebook, strict JSON and Unicode bounds, and atomic diagnostic storage that preserves the original outcome without triggering another model attempt.

I also added a checker that validates all 486 detailed replay checks, recomputes their outcomes and rejects missing or duplicate identities and inconsistent summaries. The package validator checks the hashes of 46 frozen historical files, keeping the original observations intact. CI runs the software checks and synthetic replay without model credentials.

The October 8 isolated local reproduction passed 486/486 checks: 108 decisions, 324 metamorphic checks and 54 relation checks across the three engines. It used the existing 36 synthetic development cases and made zero model calls. Its scope differs from the broader historical check's 540 invariants. The final software suite passed 70 tests; generation parity and read-only package validation also passed.

Those checks make the dossier easier to inspect and reproduce. They do not add model measurements. The strengthened native successor remains prepared and locally tested, with no execution or publication on Kaggle. Reproduction receipt, technical closeout and remaining limits.

What I would measure next

I would complete human arbitration of the 24 real-source questions, test new held-out packets with balanced condition order, and finish the separate Context and native-browser investigations. I would also measure whether people can explain and resolve a scoped conflict more accurately using baseline, N or P.

That last experiment would directly examine the value behind my decision to retain user choice. The current results give me a concrete next question, with the strengths and limits of each view explained clearly.

My Benchmark

Orbit: scoped evidence and user choice — Kaggle Benchmark

The collection contains the synthetic development pilot and the native Boolean task. Its hosted model results belong to those task identities; the larger historical campaign is documented through separate receipts and aggregates.

Public GitHub repository: code, methodology, evidence and execution trace

The README explains the experimental units, exact model IDs, source fingerprints, reproduction steps and remaining limitations. The package contains frozen engine source, authored synthetic cases, sanitized execution receipts and the analysis-only preparation notebook. Real-source snapshots, private model outputs and credentials are excluded from the public export.

You can also explore Orbit and inspect its original implementation.

Authorship and collaboration

I used Codex as a research partner for code mapping, source comparison, evidence organization, consistency checks, technical review and editorial preparation. I formulated the intent, defined the scope, interpreted the results and made the public decisions. Judgment, responsibility, authorship and final approval remain mine.

This was my first benchmark, and it became a substantial learning experience. I am proud of the work and of the clearer questions it helped me formulate, independently of the competition outcome.

I began by trying to choose an engine. I came away wanting to give users an informed choice—and a record they could inspect for themselves.

Top comments (0)