I’ve been building .me around a fairly simple idea:
Knowledge is not just data. It is data plus the relationships that make changes meaningful.
Most of my previous tests of this idea used synthetic workloads. That’s useful when testing things like fan-out, dependency depth, recomputation, and early cutoff.
But synthetic tests eventually raise an uncomfortable question:
What happens with somebody else’s data?
So I took the GTFS Madrid workload used in research on incremental knowledge graph construction and adapted it to run on .me.
What started as a benchmark ended up being more interesting than I expected.
We effectively built a small, evolving knowledge graph of Madrid’s transportation data inside .me.
And then we changed the world underneath it.
⸻
First: what did .me actually “learn”?
“Learn” is probably the wrong technical word.
There was no machine learning involved. No model was trained.
Instead, .me acquired a structured representation of a domain.
GTFS contains transportation data describing things such as:
- routes
- trips
- services
- calendars
- stop times
The adapter loaded those entities into the .me namespace.
Conceptually:
gtfs
├── services
│ ├── service_1
│ ├── service_2
│ └── ...
│
├── routes
│ ├── route_1
│ └── ...
│
├── trips
│ ├── trip_1
│ ├── trip_2
│ └── ...
│
└── stopTimes
├── ...
└── ...
But simply storing GTFS records in a tree wouldn’t be particularly interesting.
The important part was connecting them.
⸻
From data to dependencies
A transportation system contains relationships.
A trip belongs to a route.
A trip uses a service.
A service calendar determines when that trip can operate.
Stop times belong to trips.
Conceptually:
service
│
▼
trip
│
▼
stop_time
The adapter expressed relevant relationships as explicit .me dependencies.
A tiny version of the same idea looks like this:
me.A(2)
me.B(3)
me"="
C is no longer merely another stored value.
It is knowledge derived from A and B.
If we change:
me.A(5)
.me knows that C depends on A.
We can inspect that relationship:
me.explain("C")
and get information corresponding to:
sourcePath: A
dependsOn: [A, B]
k: 1
recomputed: [C]
The GTFS adapter takes this basic mechanism and applies it across a much larger domain.
⸻
Building the Madrid knowledge universe
I ran the benchmark at three scales.
At scale 1, the loaded GTFS snapshot contained 2,364 stop times.
At scale 10:
23,640 stop times.
At scale 100:
236,400 stop times.
The scale-100 base snapshot included approximately:
routes 1,300
trips 13,000
stop_times 236,400
calendar 500
calendar_dates 7,000
Alongside that source knowledge, the adapter created indexes and derived relationships used by the dependency graph.
So at this point .me wasn’t looking at a toy A → B → C example anymore.
It had a structured transportation universe containing hundreds of thousands of GTFS records and relationships.
⸻
Then we changed Madrid
This is where the experiment becomes interesting.
The GTFS benchmark doesn’t only provide a static dataset.
It provides evolving versions of that dataset.
I compared the base version with seed0 and converted the differences into neutral mutations:
CREATE
UPDATE
DELETE
Then those mutations were applied to .me one by one.
Across the three scales, the complete sweeps contained:
Scale 1 151 mutations
Scale 10 1,304 mutations
Scale 100 13,401 mutations
At scale 100:
13,401 / 13,401 mutations were measurable.
No mutations were skipped.
Now we could ask a much more interesting question than “how fast can you read a node?”
We could ask:
When one fact in this knowledge universe changes, how much other knowledge does that change actually affect?
⸻
n versus k
This distinction is central to .me.
Let:
n = total knowledge loaded into the system
and:
k = knowledge affected by a particular mutation
These are not necessarily the same variable.
Imagine a graph containing a million pieces of knowledge.
Changing one value doesn’t necessarily affect a million things.
Maybe it affects three.
Maybe twenty.
Maybe 100,000.
The topology determines that.
So instead of assuming that a mutation should somehow involve the entire knowledge universe, .me keeps explicit dependencies and follows the affected region.
Conceptually:
knowledge universe
n
┌─────────────────┐
│ │
│ Δ │
│ │ │
│ ▼ │
│ [ k ] │
│ │
└─────────────────┘
The experiment was asking what happens to k when we make n much larger.
⸻
The surprising part
Here are the results.
Scale 1
stop_times: 2,364
mutations: 151
k p50: 1
k p95: 18.5
k max: 27
recompute p50: 0.004 ms
recompute p95: 0.073 ms
Scale 10
stop_times: 23,640
mutations: 1,304
k p50: 1
k p95: 18
k max: 27
recompute p50: 0.003 ms
recompute p95: 0.086 ms
Scale 100
stop_times: 236,400
mutations: 13,401
k p50: 1
k p95: 18
k max: 27
recompute p50: 0.003 ms
recompute p95: 0.131 ms
Put differently:
n
2,364
↓
23,640
↓
236,400
k p95
18.5
↓
18
↓
18
k max
27
↓
27
↓
27
The dataset became approximately 100× larger.
The affected-set distribution essentially didn’t.
That is the result I find most interesting.
⸻
DELETE was even more stable
DELETE mutations gave us a particularly clean pattern.
Across scale 1, scale 10, and scale 100:
DELETE k p50 = 18
DELETE k max = 19
Exactly the same.
So while the surrounding knowledge universe became roughly two orders of magnitude larger:
Scale 1 18 / 19
Scale 10 18 / 19
Scale 100 18 / 19
the dependency neighborhood touched by those mutations stayed local.
⸻
This does NOT mean everything is O(k)
This is where I want to be careful.
It would be tempting to look at these numbers and say:
“.me mutations are O(k).”
We haven’t demonstrated that.
In fact, the benchmark exposed a bottleneck that makes that claim inappropriate.
There are at least two different things happening when we mutate this knowledge universe:
mutation
│
├── structural application
│
└── dependency propagation
Dependency recomputation remained very small.
Structural mutation handling did not always behave that way.
At scale 100, DELETE apply operations were often measured in seconds. One observed operation took roughly 22 seconds to apply structurally while its dependency recomputation was around 0.37 ms.
Recomputation for these mutations was typically below 1 ms, although larger outliers were observed.
That gives us a much more useful picture:
mutation
│
├── structural apply/delete
│ │
│ └── currently appears sensitive to corpus size
│
└── dependency propagation
│
└── remained localized to k
That’s not something I want to hide.
It’s something I want to profile.
The benchmark showed both a property that appears to work and another part of the implementation that needs work.
That’s exactly what I want from a benchmark.
⸻
Loading also grows with n
The same distinction appears when initially constructing the knowledge universe.
Measured loading times were approximately:
Scale 1 0.33 s
Scale 10 2.9 s
Scale 100 32.5 s
So it would also be wrong to say:
“.me performance is independent of n.”
Clearly it isn’t.
A better model of what we’re observing is:
.me
┌─────────┴─────────┐
│ │
build universe propagate change
│ │
▼ ▼
n k
Building a bigger world costs more.
The interesting observation is what happens after that world exists and one thing inside it changes.
⸻
CREATE is a special case
CREATE required another important design decision.
A dependency can point to something that already exists.
But what happens when an entirely new member enters a set?
There is no previous dependency edge from that nonexistent member to invalidate.
For this benchmark, CREATE operations are therefore represented using explicit membership/index slots.
Conceptually:
member absent
│
▼
slot = 0
CREATE
│
▼
slot = 1
│
▼
aggregate changes
│
▼
dependency wave
The CREATE mutations in this adapter produced:
k = 1
But that’s important to interpret correctly.
It does not mean GTFS CREATE naturally has k=1.
It means this adapter models set membership explicitly, and the membership transition produces a one-node dependency wave.
That distinction matters if somebody wants to reproduce or challenge the result.
⸻
The graph can recompute without changing
Another thing showed up in the calendar mutations.
Sometimes a dependency was affected and therefore recomputed, but its resulting value remained the same.
For example:
recomputed: 1
changed: 0
Other mutations produced:
recomputed: 27
changed: 27
This gives .me another useful distinction:
affected
≠
changed
A node may need to be inspected because one of its inputs changed.
If its effective value remains unchanged, propagation can stop there.
That’s early cutoff.
It was only a small portion of this particular workload — roughly 0.5% of measured nodes — but seeing it emerge from external data rather than a synthetic test was useful.
⸻
Did .me “learn” Madrid?
Not in the machine-learning sense.
No neural network was trained.
No statistical model inferred the GTFS structure for us.
We explicitly mapped the domain.
But something important did happen:
raw GTFS
│
▼
structured knowledge
│
▼
relationships
│
▼
dependencies
│
▼
derived knowledge
│
▼
changing source facts
│
▼
affected knowledge
.me acquired a structured, executable representation of part of the Madrid transportation domain.
And because the relationships were represented as dependencies, changes to the source data had computational meaning.
That’s much closer to what I mean when I call .me a knowledge graph.
The interesting object isn’t just:
trip = {...}
It’s:
this trip uses this service
this service has this calendar
these values derive this state
this stop time belongs to this trip
this mutation can affect these derived facts
The graph contains not only values, but relationships about how those values participate in other knowledge.
⸻
How this compares to the original research
The GTFS Madrid workload comes from work on incremental knowledge graph construction.
It’s important not to turn this into an invalid benchmark comparison.
.me is not pretending to be an RML/RDF construction engine.
And .me recomputation timings should not be placed next to IncRML construction timings and presented as though both systems performed the same operation.
They don’t.
Instead, we reused the evolving source workload to examine a different mechanism.
Conceptually:
GTFS mutation
│
┌──────────┴──────────┐
│ │
incremental KG .me
construction │
│ dependency graph
│ │
▼ ▼
construction metrics affected set k
Same changing source world.
Different question.
The question for .me was:
Does the amount of dependency propagation follow the size of the entire knowledge universe, or the portion actually affected by the change?
Across scale 1 → 10 → 100, the evidence from this workload points toward the latter.
⸻
What did we actually demonstrate?
The strongest statement I think the experiment currently supports is:
On this adaptation of the GTFS Madrid mutation workload, increasing the loaded dataset roughly 100× did not increase the affected dependency-set distribution for the measured mutations.
That’s deliberately narrower than:
“.me is O(k).”
We haven’t proven the latter.
What we observed was:
n × ~100
k p50 1 → 1 → 1
k p95 18.5 → 18 → 18
k max 27 → 27 → 27
So at least in this experiment:
the size of the knowledge universe and the size of a mutation’s dependency wave behaved like different variables.
That’s a much more interesting result to me than simply saying the benchmark was fast.
⸻
Now I want to break it
The next experiment shouldn’t just make n bigger.
It should deliberately manipulate n and k independently.
For example:
fixed n
increasing k
then:
increasing n
fixed k
and eventually:
increasing n
increasing k
If recomputation follows k as we independently vary those dimensions, we’ll have much stronger evidence about the behavior of the dependency engine.
The structural DELETE bottleneck also needs profiling.
And GTFS should not be the last external workload.
I’d like to attack different properties of .me with different independent datasets and benchmarks: dependency propagation, cascading updates, provenance, topology, revocation, audiences, and structural privacy.
The objective isn’t to find workloads that make .me look good.
It’s to find the workload that breaks the model.
Because if the model survives, then we have learned something much more useful than another benchmark number.
⸻
The .me kernel:
https://github.com/neurons-me/.me
GTFS Madrid benchmark:
Top comments (0)