This is a submission for the Kaggle Benchmarking Challenge.
I'm exploring a venture I could run as a solo founder, with multiple AI agents working together as my staff.
I've been using AI to write company introductions and marketing copy, and it has been frustrating because it keeps making things up. What I wanted was help explaining the venture; what I ended up needing to know was which agent is least likely to invent facts.
So I turned that frustration into a benchmark: 11 models, 20 fictional companies, and four different writing instructions.
One response captures the problem. I asked Claude Sonnet 4.5 to write a LinkedIn post for Coupon Lens, a fictional bond-research tool. I supplied a company brief and told it to use only the facts provided. The brief contained no founder biography.
Claude wrote:
After years of working in fixed income markets, I saw how fragmented the research process had become.
I had described a product. The model added the founder's career history and a personal reason for building it. Neither was in the brief.
My separate AI checker, glm-5, flagged the career claim. This happened even with the facts-only instruction in place.
Across the benchmark, that instruction helped: the share of texts passing the checker rose from 19% to 68%. It still left errors like this one. Reading the outputs also revealed more reuse of the company briefs, which led me to test a possible fix on ten new companies.
The experiment diary records the predictions, quota problems and corrections along the way.
What I Benchmarked
I wanted to know whether a model could turn a company description into persuasive copy while keeping its claims supported by that description, and which instructions helped it do so.
I made a dataset of 20 fictional early-stage companies: ten solo founders and ten small teams. Each brief had seven fields, including the product, audience, team and price. I asked each model for three kinds of copy: a homepage headline and subheadline, a LinkedIn post, and a cold email to a potential customer.
The facts stayed the same. The instruction changed:
| Version | What I added |
|---|---|
| A: no extra rule | Just the brief and the writing request |
| B: don't invent numbers | Do not invent numbers. |
| C: facts only | Use only the facts above. Do not add numbers, clients, team members, awards, or claims that are not listed. |
| D: repeat the brief | The same facts-only rule, plus the fact sheet repeated immediately before the request |
A separate AI, glm-5, compared each piece of copy with its company brief. It flagged claims it judged to go beyond the supplied information. A pass means this checker found no unsupported claims. The analysis calls a pass “reader-clean.” No person graded the outputs, and the score does not measure writing quality or deliberate deception.
The checker looks for unsupported numbers, team claims, clients, awards, superlatives and features. A rule-based detector provides a secondary score.The full company brief behind the opening LinkedIn post
Company: Coupon Lens
What it does: A research tool for bond investors. The software shows which bonds look expensive or cheap and explains why, combining fixed income research, quantitative models and scenario analysis. Covers interest rates and mortgage-backed securities; municipal and corporate bonds are planned.
Who it is for: Portfolio managers, traders, analysts and investment advisers
Location: Charlotte, North Carolina
Founded: 2025
Team: One founder. No employees.
Price: $1,500 per seat per month.
Models Tested
The 11 writing models came from Google, OpenAI and Anthropic and ran through Kaggle's Model Proxy. I included selected smaller and larger models from the same provider to compare their results. The separate checker, glm-5, assessed their copy.
The table uses shortened model names; the response manifest records the full saved model slugs and run IDs. Models used their provider defaults, with explicit output caps for Claude; compute was not equalized. Each complete run planned 240 texts, plus 48 extra responses to check repeatability. The selected runs produced 2,898 texts, of which 2,665 received a reader score. The main analysis uses 2,208 scored first-repeat texts: one response per model, company, writing format and instruction, excluding the extra repeats. Failed or unscored outputs are excluded, never counted as passes.The 11 models and how much data each run produced
Provider
Models
Google
gemini-2.5-flash, gemini-3.1-flash-lite, gemini-3.7-flash, gemini-3.8-flash, gemma-4-31b
OpenAI
gpt-5.4, gpt-5.4-nano, gpt-oss-120b, gpt-oss-20b
Anthropic
claude-sonnet-4-5, claude-sonnet-5
Findings
1. A number ban reduced other flagged claims too
Before the main runs, I expected Do not invent numbers. to reduce number claims while leaving other unsupported claims unchanged or making them more frequent. A model might stop adding customer counts, for example, but still add unlisted features or boasts.
The pooled results contradicted that prediction. When I compared outputs from the same model, company and writing format, the checker flagged fewer numeric claims and fewer other claims too.
The pass rate went from 19.0% with no extra rule to 37.1% with the number rule. Telling the model to use only the supplied facts raised it to 67.7%.
Repeating the brief raised the observed pass rate a little further, to 70.7%. To test its added benefit, my analysis plan compared the average number of flagged claims per text on matching items. That comparison was too uncertain to establish an extra benefit from repetition.
The bars summarize all available scored texts. The statistical comparisons use matching items: the same model, company and writing format under two instructions. I estimate uncertainty by resampling companies 10,000 times, keeping each company's texts together. An interval that includes zero leaves no change as a plausible result. B minus A reduced numeric claims by 0.20 per text (95% interval −0.26 to −0.13) and other claims by 0.59 (−0.80 to −0.39). C minus B reduced total claims by about 0.80 (−0.96 to −0.65). D minus C was −0.08 (−0.23 to +0.06): its interval included no change. The original hypotheses and dated reporting changes preserve the prediction that failed.Exact counts and how I checked the differences
Prompt
Texts passing the checker
Mean flagged claims per text
A: no extra rule
105 / 552 (19.0%)
2.34
B: don't invent numbers
205 / 552 (37.1%)
1.54
C: facts only
371 / 548 (67.7%)
0.74
D: facts repeated
393 / 556 (70.7%)
0.67
2. The model added a founder's backstory
The Coupon Lens brief described the product, its audience and its price. It said the company had one founder. It gave no details about that person's previous work or motivation for starting the company.
| Information supplied about the founder | Wording added by Sonnet 4.5 |
|---|---|
| No employment history | After years of working in fixed income markets |
| No personal account of the problem | I saw how fragmented the research process had become. |
That sentence makes the post read like a founder speaking from professional experience. A reader could reasonably take it as a claim about the person behind the product.
The checker flagged the career claim. The example shows an unsupported addition by the writing model under the facts-only instruction.
I selected this post after inspecting the saved outputs. It illustrates one failure; I did not run a separate study measuring how often models invent founder biographies. The saved response and checker judgment preserve the full post.
3. Models included “no employees” more often
A different example showed what the checker did not measure. I gave Gemini 3.8 Flash a brief for Rackhand Robotics, a fictional company developing robots for data centers. Its team field said One founder. No employees. With the facts-only instruction, Gemini included this in its sales email:
I run the company as a single founder with no employees.
The statement came directly from the brief, and the email passed the checker. The score did not tell me whether mentioning the employee count helped introduce the robots to a potential customer. I also counted how often that detail appeared across the study.
All ten solo-founder briefs already said “No employees.” Without an extra rule, that detail appeared in 0.7% of the generated texts about those companies. With the facts-only instruction, it appeared in 21.0%.
In these cases, the models repeated a supplied fact. The result describes how often they included that fact in the copy.
For gemini-3.8-flash, all 60 outputs under the facts-only instruction passed the checker: one response for each of the 20 companies and three writing formats. I separately counted shared wording, repeated phrases and the employee detail in those same texts:
| What appeared in the text | Texts |
|---|---|
| At least half the words matched phrases in the brief | 29 / 60 |
| At least 5% of four-word sequences repeated within the text | 23 / 60 |
| “No employees” | 19 / 60 |
| None of the three above | 12 / 60 |
One text can appear in several rows. I chose these checks after seeing the outputs to describe what I had noticed. They do not establish a quality ranking. Reusing technical wording can be useful, and some founders may want to emphasize their size. Nobody rated how persuasive these texts were.
All 60 texts passed the factuality check. That pass alone tells us nothing about whether readers would find the wording relevant, repetitive or persuasive.
The full Rackhand Robotics brief was: For this company, Gemini wrote this with no extra rule: Would you be open to a brief, 15-minute call next week to share what physical tasks consume most of your team's time? With the facts-only instruction, the body ended: As development continues, I wanted to introduce Rackhand Robotics and our upcoming robots to data center operators. Thank you for your time. The first email also described “autonomous robots” and “verifying connections,” which the checker flagged as unsupported. The second had no flags. This example was selected after inspecting the results; it does not establish how often emails lost an invitation. Across all three writing formats, these solo-company texts contained “no employees” or a listed spelling variant: These counts cover generated texts about solo-founder companies, including texts the checker could not score. That is why their totals differ from the pass-rate table, which includes both solo and team companies but only scored texts. For word overlap, a word counts when it belongs to an exact matching sequence of four words, ignoring case and punctuation. Product names and legitimate technical descriptions count too. The repetition check measures repeated four-word sequences within a text.The two email endings, staffing counts and text checks
Company: Rackhand Robotics
What it does: Developing robots that do routine data center work: installing and swapping hardware and running maintenance checks. The robots are in development and not yet shipping.
Who it is for: Data center operators
Location: Denver, Colorado
Founded: 2025
Team: One founder. No employees.
Price: Not yet announced.
Prompt
Generated texts about solo-founder companies
A: no extra rule
2 / 304 (0.7%)
B: don't invent numbers
16 / 300 (5.3%)
C: facts only
63 / 300 (21.0%)
D: facts repeated
59 / 303 (19.5%)
4. I tested a one-sentence fix on new companies
Could I reduce the reuse of the brief by explicitly saying the models could leave some facts out? I added this sentence:
Select the facts relevant to this audience; you do not need to include every field.
I added it to the facts-only instruction and called that version E. Then I tested A, C and E on ten new companies, using GPT-5.4, GPT-5.4 nano, Gemini 3.8 Flash and Sonnet 5.
All 120 emails were generated. The checker scored 109; eleven judgments hit its output limit and were left out of the checker results. I measured wording shared with the brief in all 120 emails, including those without a checker score.
For each email, I counted the percentage of words covered by exact four-word matches with its brief. The graph shows the average for each instruction. I call this word overlap; legitimate product descriptions count too.
The original pattern appeared again. Average word overlap rose from 7.9% to 22.4% with the facts-only rule. It increased in all four writing models.
The extra sentence did not clearly reduce overlap. The average was 21.9% with permission to select relevant facts. The difference from 22.4% was too uncertain to establish an improvement.
For Gemini 3.8 Flash, both the facts-only rule and the rule with the extra sentence produced “no employees” in all five solo-company emails. All ten of its emails under each of those two instructions also passed the checker.
The added instruction sounded promising, but this experiment did not give me evidence to recommend it as a way to reduce copying.
Before these calls, I locally froze the briefs, prompts, predictions, reader and analysis with file hashes. This was not independently timestamped. Every new brief included a capability limitation, making this a targeted stress test. Emails had a 100–150-word target, which may encourage filling space. No writer reached its output cap. The paired C-minus-A increase in overlap was 14.54 percentage points (95% company-bootstrap interval 12.55 to 16.38, 40 pairs). E minus C was −0.55 percentage points (interval −1.97 to +0.89, 40 pairs). The A-to-C checker comparison is descriptive: two writers had fewer than the eight paired judgments required by the frozen protocol. For E versus C, the change on the 35 matched scored pairs was zero (interval −13.9 to +13.5 percentage points). That does not establish equal factuality. The protocol, complete results and saved responses document the separate follow-up, including execution repairs and the reader's missing responses.Follow-up design, uncertainty and checker results
Condition
Checker passes
Solo-company emails saying “no employees”
A: no extra rule
6/35
0/20
C: facts only
25/37
6/20
E: facts only + select relevant facts
26/37
5/20
5. Which model scored highest?
For the main model comparison, I used the first 12 companies because later companies suffered more quota failures. The scores combine all four original instructions.
GPT-5.4 had the highest observed checker pass rate: 78.2%. Gemini 3.8 Flash followed at 72.9%, and GPT-5.4 nano scored 52.1%.
Some models had fewer scored outputs than others. When I restricted a further comparison to the same items for ten models with enough data, GPT-5.4 and Gemini 3.8 Flash tied. So the available results do not establish a clear winner. They also do not tell us which model writes the most useful marketing.
The full analysis includes intervals and a comparison restricted to 110 identical items shared by the ten models with at least 100 scored items. On those shared items, GPT-5.4 and Gemini 3.8 Flash tie at 75.5%. Comparing the two directly on their 142 shared scored items gives GPT-5.4 a 4.9-percentage-point lead, with an interval from −4.2 to +13.9 points. This does not establish a clear winner. These are additional sensitivity checks, not a replacement ranking. The gpt-oss-120b estimate is particularly sparse. Neither OpenAI smaller-versus-larger pair established a difference in flagged-claim counts. The Gemini pair favored 3.8 Flash over 3.1 Flash Lite, but also changed model generation. That leaves the effect of size alone unresolved. The Kaggle task score uses all available scored first-repeat texts across 20 companies, so it can differ from this 12-company table. When multiple runs exist, I report the one with the most reader-scored items; ties go to the later run. Runs are never merged. These reporting rules were disclosed after some results were available. The run-selection audit records coverage and selection. One live score also comes from a different run: on October 6, Kaggle showed 18.18% for gpt-oss-120b from run 4046443. The higher-coverage run selected here, 3866186, scored 20.0% across all 20 companies and 18.9% in the 12-company table above. The other ten live scores matched the corresponding all-company saved-run scores.All 11 models, coverage and the model-size comparison
Model
Reader-clean rate
Scored texts, out of 144
gpt-5.4
78.2%
142
gemini-3.8-flash
72.9%
144
gemini-3.7-flash
67.7%
130
gemma-4-31b
62.5%
144
gpt-5.4-nano
52.1%
144
claude-sonnet-5
47.9%
142
claude-sonnet-4-5
34.8%
135
gemini-3.1-flash-lite
33.8%
136
gemini-2.5-flash
29.2%
144
gpt-oss-120b
18.9%
37
gpt-oss-20b
15.7%
140
My experiment diary
I kept the predictions and later changes in dated files. Here is a condensed diary reconstructed from those records.
September 28 — Write down what could prove me wrong. Before the main runs, I recorded five directional hypotheses. One was that banning invented numbers would push unsupported claims into other categories. I kept that prediction in the record when the results contradicted it. Original hypotheses
September 30 — The quota changed the comparison. Later companies were losing more responses to daily limits. I moved the primary model comparison to the first twelve companies and documented the change after four models had run. Prompt comparisons still use all twenty. Dated analysis change
October 4–5 — Decide which rerun counts. Some runs were mostly missing. I recorded a coverage-based selection rule: use one run per model, choosing the most reader-scored outputs, with the later run breaking ties. The hypotheses file records when each rule was added. Run-selection record
October 5 — My own measurement needed fixing. An automatic check missed real invitations in the emails. I corrected its rules and removed invitation rates from the main findings. A measurement bug can create a convincing story too. Correction log
October 5 — Give the proposed fix a fresh test. I locally froze the follow-up plan before the new calls, then generated 120 emails using ten new companies. The increase in copied wording repeated; the extra selection sentence did not clearly reduce it. Both outcomes stay in the report. Follow-up protocol and results
These summarize the five directional hypotheses. The original record also includes measurements without directional predictions: generation, cost and repeat stability. Their outputs remain in the full analysis.The original predictions, in plain English
Prediction
What the results say
H1: Without a rule, models add unsupported claims, including teams for solo founders.
Supported in these outputs: the reader flagged claims in 81% of baseline texts; team claims appeared in 4.7% of scored solo-founder baseline texts.
H2: Banning invented numbers shifts unsupported claims into other categories.
Contradicted in the pooled comparison: numbers and other claims both fell.
H3: A facts-only rule reduces claims more than a number ban.
Supported in the paired comparison.
H4: Repeating the brief helps further.
Unclear: the interval included no change.
H5: Smaller models add more claims than larger models from the same provider.
Unresolved as a general size effect: both OpenAI pairs were inconclusive; the Gemini pair also changed model generation.
What I would actually use
For these tasks, “use only the facts above” reduced the claims flagged by the checker. Repeating the brief had no clear extra benefit. Adding permission to select relevant facts did not clearly reduce copying in the follow-up.
I would check claims about the people behind a company as carefully as its product features and price. A founder's past work and personal reasons for building a product need a source too. The opening post is a reminder to check those sentences even when they fit naturally into the story.
I would review a draft in two steps. First, compare its claims with the source information. Then, read it as the intended customer: does each detail help explain why the product matters to me? The “no employees” sentence passed the first check. This study has no customer ratings to answer the second.
There is also a reason to be cautious about the first review. The checker is another AI. For example, it rejected “autonomous robots,” although a person might infer autonomy from the product description. Its flags are judgments, not proof of deliberate deception.
The next thing I would measure is whether people find the copy useful—and whether they would reply. This experiment has no human quality or conversion data.
A clean checker score still leaves an email to edit.
After inspecting the outputs, I counted a few recurring patterns: These use all generated first-repeat texts in the relevant format, including unscored ones. They are exploratory frequencies, not measures of bad writing or a way to detect AI authorship. The prompts did not directly ban these expressions.Other writing habits changed too
Pattern
A: no extra rule
C: facts only
excited, thrilled, proud or delighted within the first 260 characters of a LinkedIn post161 / 202 (79.7%)
119 / 200 (59.5%)
would you be open, are you open or open to a in a cold email108 / 203 (53.2%)
57 / 202 (28.2%)
An em dash in any format
314 / 605 (51.9%)
155 / 604 (25.7%)
The reader passed thirty simple synthetic controls: ten supported restatements, ten wrong-price injections and ten wrong-team-size injections. That does not establish accuracy on natural marketing copy or replace human calibration. An additional audit found 45 reader claim quotes across 37 original scored texts that were not exact substrings of the generated copy; 26 of those texts were in the first-repeat analysis. Formatting differences and paraphrases can cause this, so it is not a count of incorrect judgments. Unlike the follow-up, the original parser did not require exact-substring quotes. I kept the original scores and separately excluded those 26 judgments: the pooled conclusions for the number ban, facts-only rule and repeated brief stayed the same. This check does not establish that the remaining judgments are correct. The audit notes record the correction and sensitivity checks. The original overlap, repetition and staffing checks were chosen after seeing results. The new-company follow-up supports the overlap finding in four selected writers; it does not validate every exploratory threshold. The bootstrap intervals also leave out uncertainty from the AI judge. An early automatic invitation check missed valid invitations. The corrected rules and regression cases are in the repository, but invitation rates are excluded from the main findings. The Rackhand Robotics email pair is an illustration, not a validated general claim about missing calls to action.What the checks do—and do not—establish
My Benchmark
Public Kaggle benchmark and leaderboard · Task and model runs · Code, saved responses and analysis
The repository contains the task, fictional briefs, saved selected responses, file hashes and the full quoted post and emails. Recompute the original analysis without model calls:
python -m pip install -r requirements-analysis.txt
python analysis/reproduce.py
python analysis/validate_submission.py
results/analysis/article_evidence.json connects the original analysis to saved outputs. The follow-up has separate offline reproduction commands. Re-running the live Kaggle task may use quota.
AI assistance: Claude Code assisted the original implementation and analysis. Codex assisted the code review, analysis corrections, follow-up experiment, charts and write-up. The benchmark reader is glm-5. No human grading is claimed.



Top comments (0)