DEV Community

Cover image for Consensus without leader election: ORCHID in Grid (Kuramoto phase + quorum)
Antharas
Antharas

Posted on Originally published at habr.com

Consensus without leader election: ORCHID in Grid (Kuramoto phase + quorum)

Three nodes, one database, the client sends UPSERT. Two things must hold at once: admit a write only when the cluster is synchronized, and never let a network cut leave different versions of the same data in the journals.

In a Raft-style stack nodes elect a leader and only the leader writes; each leadership period has a term; when the leader is gone, they elect again. Grid admits writes differently: there is no election. Nodes exchange a scalar phase (coupled oscillators after Kuramoto) and separately confirm each operation with a checksum (digest) of its contents. That protocol is ORCHID.

Recovery after failure means bringing phases back together and collecting digest agreement — not running another election.


Context: what Grid means in this article

Grid is a distributed SQL database: hot data in memory, a mutation journal (OpLog) and sealed snapshots on disk, and ORCHID between nodes. A single node with disk and no peers is a normal mode; replication is opt-in.

Piece Role
SQL, catalog DDL/DML, schema
Memory Hot working set
Disk (OpLog + sealed GridMap) Survive restart, keep commits
Cluster (ORCHID + replication) Who may write and what is durable
Client grid:// (and tooling JDBC on the same wire)

The rest of this article is write admission only. SQL dialect, sealed layout, and client APIs belong elsewhere; start at quick-start.


Two questions — two mechanisms

Keep them apart: they fail and are observed differently. Phase does not replace the checksum.

Phase synchronization Checksum agreement
Question May the cluster propose a write right now? Is this the same operation the others see?
Measure Order parameter R ∈ [0, 1] over phases θ Matching digest
Scope Live nodes in the local site Nodes from the configuration (+ remote voters when enabled)
Failure R below threshold → no proposal No majority on one checksum
Writer Among synced nodes, minimal nodeId Same node sends the proposal

Solo write (N = 1) is allowed only if the peer list was empty from the start. If peers were configured and then the network failed, the survivor does not become a solo writer — otherwise a minority would keep committing in isolation.

Defaults: order-threshold: 0.85, tick-ms: 10, natural-freq-hz: 1.0, coupling: 15, digest-quorum: MAJORITY.


How R is computed

Order parameter over phases:

R = | (1/N) * Σ_k exp(i · θ_k) |
Enter fullscreen mode Exit fullscreen mode

Closer phases push R toward 1. In code — mean cos/sin and hypotenuse:

private double computeR() {
    double sx = Math.cos(phase);
    double sy = Math.sin(phase);
    int n = 1;
    for (PeerView peer : peers.values()) {
        if (!peer.seen || !countsForPhase(peer.id)) {
            continue;
        }
        sx += Math.cos(peer.phase);
        sy += Math.sin(peer.phase);
        n++;
    }
    return Math.hypot(sx / n, sy / n);
}
Enter fullscreen mode Exit fullscreen mode

Every tick-ms a node updates its phase and exchanges it with local-site peers. Free-running angular speed is natural-freq-hz; pull strength is coupling. Convergence needs the free advance per tick to stay well below one:

2 · π · natural-freq-hz · tick-ms / 1000  ≪  1
Enter fullscreen mode Exit fullscreen mode

Defaults (1.0 Hz, 10 ms) satisfy that. At natural-freq-hz = 50 with the same tick, phase jumps too far per step: R falls below threshold on a healthy cluster. Lowering order-threshold “to stop the errors” hides desync instead of fixing frequency / tick / coupling.

Multi-DC once. R is local-site only; remote sites do not enter local R. SYNC_VOTERS_ACROSS_DC adds digest confirmations from another site (with a WAN timeout) — it is not Kuramoto-over-WAN and not a substitute for local phase lock.


How a write proceeds

While R ≥ order-threshold and the node is the writer (minimal nodeId among synced nodes):

  1. Digest = hash of previous opSeq and operation bytes.
  2. Proposal is broadcast; peers ACK the same digest or NACK.
  3. Majority is taken from the configuration (not “whoever answered this second”). forgetPeer does not shrink majority size.
  4. The record is appended to the OpLog and confirmed on disk before the commit broadcast and before map visibility.
  5. Only then is the row visible to readers as committed.
admit(op):
  if configuration has peers but none are reachable → refuse
  if R below threshold → refuse
  if this node is not the writer → refuse / redirect to writer
  digest = hash(prevOpSeq, op bytes)
  broadcast proposal
  wait for majority on the same digest
  append journal, confirmPersisted
  broadcast commit
  apply into the map
Enter fullscreen mode Exit fullscreen mode

Linearizable reads in the Jepsen sense are from the writer and only from already committed map state. Client connection samples live in quick-start, not here — they break the protocol narrative.


When nodes lose connectivity

Configuration Situation Who writes
N = 3 One node unreachable The other two (majority)
N = 3 Split 2 / 1 Only the side of two
N = 3 Two nodes unreachable The lone survivor does not write
N = 2 Link down Nobody writes
any Link restored Phases re-lock; lagging node catches up the journal; no election
any Restart Reads stored tip, replays OpLog tail

There is no “promote myself because I see nobody.” Client role change uses ServerMeta / PROMOTE_NOTIFY and rediscoverWriter(), not host hunting in the URL.

From the model: commit only with majority; one journal slot does not get two different payloads.


Evidence

TLA+ / TLC. Single-site and remote-voter specs run in CI: fixed voter set, one phase-ranked writer, no two values on one opSeq, minority does not commit under partition.

Jepsen. Register and append workloads under network and process faults (one site and multi-site). That is consistency evidence, not a TPS claim: clients must not observe divergence, and the journal must not fork against ORCHID invariants.


Lab host numbers

Two nodes (writer + replica), fsync: true, load over grid:// (JMeter; sync JDBC tooling is not the load path). Host: Intel Core 7 240H (10C/16T), 64 GiB RAM, NVMe.

Living anchors and ≈95% regression floors on the band lower bound:

Profile Mix Avg / band ≈95% floor
WRITE_ONLY 100% UPSERT 4676 ops/s ≈4442
READ_ONLY EQ 90% / JOIN 10% 52261…59430 ops/s ≈52261
Capacity QG EQ 50 / UPSERT 40 / JOIN 10 8772…11519 ops/s ≈8333

Numbers are tied to this host and fsync: true. WRITE ceiling includes ORCHID wait (orchidWaitP50/P99). Method: docs/en/performance/capacity-slo.md.


Closing

ORCHID admits a write when phases are close enough (R ≥ threshold) and a configuration majority confirmed one digest. No elected leader, no term. The trade-off is explicit: without majority the write is refused; a two-node split writes nowhere — better a client error than two journals.

Source and releases: github.com/GenCloud/grid-sql.
On this topic: orchid-consensus.md, bio-inspired.md, docs/spec/orchid/, benchmarks/jepsen/, capacity-slo.md.

Try it. Spin up one node with the quick-start (JDK 25, Maven, grid://) and run your own UPSERT — no sidecache, no leader election in the write-admission path. Questions and bugs: Issues in the same repository.

Top comments (0)