Three nodes, one database, the client sends UPSERT. Two things must hold at once: admit a write only when the cluster is synchronized, and never let a network cut leave different versions of the same data in the journals.
In a Raft-style stack nodes elect a leader and only the leader writes; each leadership period has a term; when the leader is gone, they elect again. Grid admits writes differently: there is no election. Nodes exchange a scalar phase (coupled oscillators after Kuramoto) and separately confirm each operation with a checksum (digest) of its contents. That protocol is ORCHID.
Recovery after failure means bringing phases back together and collecting digest agreement — not running another election.
Context: what Grid means in this article
Grid is a distributed SQL database: hot data in memory, a mutation journal (OpLog) and sealed snapshots on disk, and ORCHID between nodes. A single node with disk and no peers is a normal mode; replication is opt-in.
| Piece | Role |
|---|---|
| SQL, catalog | DDL/DML, schema |
| Memory | Hot working set |
| Disk (OpLog + sealed GridMap) | Survive restart, keep commits |
| Cluster (ORCHID + replication) | Who may write and what is durable |
| Client |
grid:// (and tooling JDBC on the same wire) |
The rest of this article is write admission only. SQL dialect, sealed layout, and client APIs belong elsewhere; start at quick-start.
Two questions — two mechanisms
Keep them apart: they fail and are observed differently. Phase does not replace the checksum.
| Phase synchronization | Checksum agreement | |
|---|---|---|
| Question | May the cluster propose a write right now? | Is this the same operation the others see? |
| Measure | Order parameter R ∈ [0, 1] over phases θ |
Matching digest |
| Scope | Live nodes in the local site | Nodes from the configuration (+ remote voters when enabled) |
| Failure |
R below threshold → no proposal |
No majority on one checksum |
| Writer | Among synced nodes, minimal nodeId
|
Same node sends the proposal |
Solo write (N = 1) is allowed only if the peer list was empty from the start. If peers were configured and then the network failed, the survivor does not become a solo writer — otherwise a minority would keep committing in isolation.
Defaults: order-threshold: 0.85, tick-ms: 10, natural-freq-hz: 1.0, coupling: 15, digest-quorum: MAJORITY.
How R is computed
Order parameter over phases:
R = | (1/N) * Σ_k exp(i · θ_k) |
Closer phases push R toward 1. In code — mean cos/sin and hypotenuse:
private double computeR() {
double sx = Math.cos(phase);
double sy = Math.sin(phase);
int n = 1;
for (PeerView peer : peers.values()) {
if (!peer.seen || !countsForPhase(peer.id)) {
continue;
}
sx += Math.cos(peer.phase);
sy += Math.sin(peer.phase);
n++;
}
return Math.hypot(sx / n, sy / n);
}
Every tick-ms a node updates its phase and exchanges it with local-site peers. Free-running angular speed is natural-freq-hz; pull strength is coupling. Convergence needs the free advance per tick to stay well below one:
2 · π · natural-freq-hz · tick-ms / 1000 ≪ 1
Defaults (1.0 Hz, 10 ms) satisfy that. At natural-freq-hz = 50 with the same tick, phase jumps too far per step: R falls below threshold on a healthy cluster. Lowering order-threshold “to stop the errors” hides desync instead of fixing frequency / tick / coupling.
Multi-DC once. R is local-site only; remote sites do not enter local R. SYNC_VOTERS_ACROSS_DC adds digest confirmations from another site (with a WAN timeout) — it is not Kuramoto-over-WAN and not a substitute for local phase lock.
How a write proceeds
While R ≥ order-threshold and the node is the writer (minimal nodeId among synced nodes):
- Digest = hash of previous
opSeqand operation bytes. - Proposal is broadcast; peers ACK the same digest or NACK.
- Majority is taken from the configuration (not “whoever answered this second”).
forgetPeerdoes not shrink majority size. - The record is appended to the OpLog and confirmed on disk before the commit broadcast and before map visibility.
- Only then is the row visible to readers as committed.
admit(op):
if configuration has peers but none are reachable → refuse
if R below threshold → refuse
if this node is not the writer → refuse / redirect to writer
digest = hash(prevOpSeq, op bytes)
broadcast proposal
wait for majority on the same digest
append journal, confirmPersisted
broadcast commit
apply into the map
Linearizable reads in the Jepsen sense are from the writer and only from already committed map state. Client connection samples live in quick-start, not here — they break the protocol narrative.
When nodes lose connectivity
| Configuration | Situation | Who writes |
|---|---|---|
| N = 3 | One node unreachable | The other two (majority) |
| N = 3 | Split 2 / 1 | Only the side of two |
| N = 3 | Two nodes unreachable | The lone survivor does not write |
| N = 2 | Link down | Nobody writes |
| any | Link restored | Phases re-lock; lagging node catches up the journal; no election |
| any | Restart | Reads stored tip, replays OpLog tail |
There is no “promote myself because I see nobody.” Client role change uses ServerMeta / PROMOTE_NOTIFY and rediscoverWriter(), not host hunting in the URL.
From the model: commit only with majority; one journal slot does not get two different payloads.
Evidence
TLA+ / TLC. Single-site and remote-voter specs run in CI: fixed voter set, one phase-ranked writer, no two values on one opSeq, minority does not commit under partition.
Jepsen. Register and append workloads under network and process faults (one site and multi-site). That is consistency evidence, not a TPS claim: clients must not observe divergence, and the journal must not fork against ORCHID invariants.
Lab host numbers
Two nodes (writer + replica), fsync: true, load over grid:// (JMeter; sync JDBC tooling is not the load path). Host: Intel Core 7 240H (10C/16T), 64 GiB RAM, NVMe.
Living anchors and ≈95% regression floors on the band lower bound:
| Profile | Mix | Avg / band | ≈95% floor |
|---|---|---|---|
| WRITE_ONLY | 100% UPSERT | 4676 ops/s | ≈4442 |
| READ_ONLY | EQ 90% / JOIN 10% | 52261…59430 ops/s | ≈52261 |
| Capacity QG | EQ 50 / UPSERT 40 / JOIN 10 | 8772…11519 ops/s | ≈8333 |
Numbers are tied to this host and fsync: true. WRITE ceiling includes ORCHID wait (orchidWaitP50/P99). Method: docs/en/performance/capacity-slo.md.
Closing
ORCHID admits a write when phases are close enough (R ≥ threshold) and a configuration majority confirmed one digest. No elected leader, no term. The trade-off is explicit: without majority the write is refused; a two-node split writes nowhere — better a client error than two journals.
Source and releases: github.com/GenCloud/grid-sql.
On this topic: orchid-consensus.md, bio-inspired.md, docs/spec/orchid/, benchmarks/jepsen/, capacity-slo.md.
Try it. Spin up one node with the quick-start (JDK 25, Maven, grid://) and run your own UPSERT — no sidecache, no leader election in the write-admission path. Questions and bugs: Issues in the same repository.
Top comments (0)