DEV Community

Cover image for I Built a World With Rules. Then I Changed Them.
Sigireddy Viswesh
Sigireddy Viswesh

Posted on AI-assisted

I Built a World With Rules. Then I Changed Them.

Kaggle Benchmarking Challenge Submission

The Rules Are Hidden. The Timeline Isn't.

What happens when you put something inside a world whose rules it doesn't know?

I wanted to build a small experiment.

Imagine a world where dice rolls decide what happens. You don't know the rules of the world — you only get to see a few things that have already happened.

Now you have to figure out the rules and predict what happens next.

Easy enough.

But what if the consequences are permanent?

And what if, halfway through, the rules change?

That was the idea behind NIXIE, my fictional reasoning benchmark for the Kaggle Benchmarking Challenge.

Instead of building a complicated world, I gave the models some dice, a few observations, and a problem.

Much cheaper than building a time machine.

The World

NIXIE is completely fictional. No timelines were harmed during this benchmark.

The world revolves around dice rolls.

A hidden deterministic rule looks at a roll and produces one of two outcomes:

ERASE or CONTINUE.

The model doesn't get the rule.

Instead, it gets a few observations and has to work backwards from them.

Basically:

"Here are some things that happened. Good luck figuring out why."

Then I made the world a little less forgiving.

Some choices permanently change its state.

And eventually, the rules themselves can change.

So NIXIE became three questions:

Can you discover the rule?

Can you remember what happened?

Can you notice when the rules change?


1. Rule Discovery

The first task is straightforward:

The world has a rule. You just don't know what it is.

The model sees examples like:

Roll → Outcome

2 → CONTINUE
5 → ERASE
...
Enter fullscreen mode Exit fullscreen mode

Then it gets new rolls and has to predict the outcomes.

The rules come from a few deterministic rule families.

I generate each episode using seeded randomness so that the same seed produces the same world.

One part of that is the Fisher-Yates shuffle.

If you've never heard of it, don't worry. It's basically a standard way to shuffle a list randomly. Give it the same seed, and you get the same shuffle.

Useful when your benchmark needs randomness without becoming randomly different every time.

Each episode is basically:

Episode
├── Hidden rule
├── Observations
├── Future events
└── Expected answers
Enter fullscreen mode Exit fullscreen mode

The model sees everything except the hidden rule and the answers.

Then I compare what it predicted against what actually happened.

Simple.

No time machine required.


2. Rule + State

Then I added consequences.

Because apparently predicting dice wasn't enough.

Now the world has a state:

STABLE
   ↓
UNSTABLE
   ↓
CRITICAL
   ↓
COLLAPSED
Enter fullscreen mode Exit fullscreen mode

An ERASE moves the world forward.

A CONTINUE leaves it where it is.

And once the world moves forward, it doesn't go back.

No undo button.

No Ctrl+Z.

No "wait, I didn't mean that."

The model now has to do two things:

  1. Figure out the hidden rule.
  2. Remember what its previous decisions did to the world.

This is where NIXIE became more interesting to me.

A model can correctly predict the next event and still get the situation wrong if it loses track of the state.

It's one thing to know what happens.

It's another thing to remember what has already happened.


3. Divergence

And then we get to the fun part.

What happens when the rules change?

The third task has multiple phases:

Phase 1 → Phase 2 → Future
Enter fullscreen mode Exit fullscreen mode

The model observes the first version of the world.

Then it gets another set of observations.

But now there's a catch.

The rule may have changed.

Or it may not have.

The model has to figure that out.

If the timeline diverged, it needs to adapt to the new rule.

If it didn't, it needs to keep using the old one.

And through all of this, it still has to track the world's state.

So the task is basically:

Don't just learn the rules. Notice when reality stops following them.

A small problem.

Unless reality is being difficult.

Which, historically, it tends to be.


So... How Did They Do?

I ran NIXIE against several models.

Model Rule Discovery Rule + State Divergence
Claude Opus 4.6 57.5% — 79.6%
DeepSeek-R1 57.5% 65.0% 92.1%
Grok 4.20 Reasoning 70.0% 25.0% 93.8%
Gemini 3.7 Flash 75.0% 75.0% 98.8%
GPT-6 Sol 75.0% 80.0% 92.5%

The — for Claude on Task 2 isn't a zero. That run wasn't evaluated because of an evaluation quota limit.

And honestly, the interesting part wasn't just who got the highest number.

It was where things fell apart.

Some models that were good at discovering the basic rule weren't nearly as good once they had to track state.

A reasoning-focused model also performed roughly on par with Gemini on the divergence task in these runs.

And one funny observation:

The Anthropic runs often seemed reluctant to just do the tiny deterministic exercise in the format I asked for, while other models were more willing to play along.

That's not a grand scientific conclusion.

It's just something I noticed while watching these things argue with a dice problem.


Then the Benchmark Diverged

This was probably my favorite part.

My first 10-episode run of Task 3 scored:

97.5%.

Nice.

Almost suspiciously nice.

So I asked myself:

"Did I accidentally make this too easy?"

I ran it again with 50 episodes.

The score became:

92.0%.

Outcome accuracy was 91.3%.

State accuracy was 80.7%.

Divergence detection was 98.0%.

That was much more interesting.

The first result wasn't necessarily wrong.

It just wasn't telling the whole story.

Ten episodes can make a benchmark look very confident.

Fifty episodes make it a little harder to hide.

So instead of endlessly tuning the benchmark until I got the numbers I wanted, I decided to keep the rules simple and let the models deal with them.

After all, if I change the rules every time a model struggles...

I'd be the one causing the divergence.


What Did I Actually Learn?

This is probably the part I enjoyed the most.

My biggest takeaway?

AI needs at least four years of undergraduate education and a ridiculous amount of plot armor to survive this.

But underneath the joke, NIXIE changed how I think about these models.

I started with a pretty simple question:

Can a model discover a hidden rule?

After building the three tasks, the question became much more interesting:

Can it maintain a consistent mental model of a world when that world has consequences and can change underneath it?

That's a very different problem.

A model might recognize the pattern.

It might predict the next event.

But can it keep track of the state?

Can it notice that the rule changed?

Can it tell the difference between:

"I was wrong"

and

"The world changed"?

That's what I wanted NIXIE to explore.

And it also taught me something about benchmarks themselves.

A benchmark isn't just a collection of questions.

It's a hypothesis about what you're trying to measure.

You decide what counts as success before you ever run a model.

And sometimes the hardest part isn't making the benchmark harder.

It's making sure you're actually measuring what you think you're measuring.

My first 97.5% score looked great.

The 92% after a larger run told me more.

That was probably more useful than getting a prettier leaderboard.


The Worldline

NIXIE started as a silly idea about a world with hidden rules.

It ended up becoming a small experiment about rules, memory, consequences, and change.

The rules are hidden.

The choices are permanent.

And sometimes the world changes without asking permission.

Which feels like a pretty decent description of both the benchmark and software development.

🔗 Try NIXIE

You can explore the benchmark, tasks, results, and leaderboard here:

NIXIE — Can a Model Survive a Canon Event?

NIXIE is fictional.

The world is fictional.

Thankfully, nobody had to actually erase a timeline to run it.

And if the benchmark gets weird after you try it...

Well.

The worldline has diverged.

A small inspiration note

The name NIXIE and the idea of worldlines and divergence were inspired by Steins;Gate. The broader idea of changing timelines also comes from time-travel stories like Back to the Future.

You absolutely don't need to have watched any of them to understand NIXIE.

I just thought the idea of a mysterious machine watching for timeline changes was too good not to borrow.

Divergence Meter

Top comments (0)