DEV Community

David D
David D

Posted on

Building a Real-Time Collaborative Code Editor with Yjs and CRDTs

Pair programming over a screen share works right up until the other person needs to type. A multiplayer code editor — Google Docs, but for code — is one of those projects that looks simple and turns genuinely hard the moment you get into it. This post walks through the architecture I'd use, the decisions that matter, and the parts most tutorials skip.

The naive version, and why it breaks

The obvious design: every keystroke becomes an operation like { pos: 42, insert: "a" }, you broadcast it, and everyone applies operations in arrival order. It works with one editor and one connection.

It falls apart the moment two people type at the same time:

  • Two inserts at the same offset have no defined order.
  • A delete racing an insert corrupts every offset after it.
  • A dropped packet or a reconnect leaves the replicas diverged, permanently.

You can patch this with a server that serialises and rewrites every operation before broadcast (OT — operational transform). It's the approach Google Docs used for years, and it is hard: the transform logic depends on the full operation history and a central authority that is always online.

CRDTs: converge without asking permission

A CRDT (Conflict-free Replicated Data Type) is a structure where replicas can be updated independently and still converge to the same state — no coordination, no locking, no central transform. For text, Yjs implements a sequence CRDT (YATA) with a tiny wire format and excellent performance.

That property is what makes offline editing and flaky connections survivable: a client applies local edits instantly and reconciles whenever it can reach the network.

The architecture

Monaco Editor
   ↕  y-monaco (MonacoBinding)
Y.Doc  (Y.Text, Y.Map)
   ↕  y-websocket provider
Sync server  ──  Awareness (cursors, names, colors)
Enter fullscreen mode Exit fullscreen mode
  • Monaco renders and edits text. It knows nothing about multiplayer.
  • y-monaco binds a Y.Text to the Monaco model, so local edits become CRDT operations and remote updates become Monaco edits.
  • y-websocket carries the CRDT updates plus an awareness channel (who's in the room, where their cursor is).
  • The sync server relays messages between clients in a room and can persist updates.

The core, in code

npm i yjs y-monaco y-websocket monaco-editor
Enter fullscreen mode Exit fullscreen mode
import * as monaco from "monaco-editor";
import * as Y from "yjs";
import { MonacoBinding } from "y-monaco";
import { WebsocketProvider } from "y-websocket";

const doc = new Y.Doc();
const type = doc.getText("monaco");

const editor = monaco.editor.create(document.getElementById("editor"), {
  value: "",
  language: "javascript",
});

const provider = new WebsocketProvider(
  "wss://your-server",
  "room-42", // one room == one document
  doc
);

new MonacoBinding(
  type,
  editor.getModel(),
  new Set([editor]),
  provider.awareness
);
Enter fullscreen mode Exit fullscreen mode

That's the whole collaboration core. Two clients in room-42 now share one document, and the awareness object is already wired up.

Presence and live cursors

Awareness is a separate, ephemeral channel — it never enters the document history. Each client publishes a small state object and you render the others:

provider.awareness.setLocalStateField("user", {
  name: "Aditi",
  color: "#22c55e",
});

provider.awareness.on("change", () => {
  const states = provider.awareness.getStates();
  // for each client id in `states`, draw a cursor + selection
});
Enter fullscreen mode Exit fullscreen mode

Style each remote cursor with that client's colour. Because awareness is separate from the CRDT, a disconnected user just disappears — it never pollutes the document.

Persistence

The document lives in memory by default; restart the server and it's gone. Two layers make it durable:

  1. Update log — append every update to storage (Postgres, LevelDB).
  2. Snapshots — periodically compact the log into a full document state.

On room load, replay the latest snapshot plus any later updates. y-leveldb and y-postgresql give you this out of the box.

Running untrusted code safely

The moment there's a Run button, you're executing someone else's code. Never eval it in your own page or process. Isolate it:

  • a throwaway container (or gVisor / Firecracker microVM),
  • no network, read-only filesystem, CPU/memory/time limits,
  • stream stdout/stderr back over a socket.

This is the one place where a small mistake becomes a breach, so treat the sandbox as a hard boundary, not a feature.

The operational bits people forget

  • One room per document, routed by document id.
  • Sticky sessions or pub/sub. With multiple server instances, fan updates out with Redis pub/sub (or the y-redis provider).
  • Reconnection. The provider resyncs automatically — test it by killing Wi-Fi mid-edit.
  • Undo. Scope the Y.UndoManager to the local origin so Ctrl-Z never undoes a teammate's typing.
  • Only sync what should sync. Share the text/CRDT state; keep local UI state out of it.

What to test before you call it done

  • Two clients typing at the same offset.
  • Disconnect one, edit on both, reconnect.
  • Five-plus clients in one room.
  • Paste a large file while someone edits.
  • Kill the server and restart it — does the room survive?

If those pass, you have something real.


If you'd rather start from a working implementation than assemble it piece by piece, the complete real-time collaborative code editor with source code and a full project report is here.

And if you're still deciding what to build, there's a breakdown of final year project ideas with architecture notes for each.

Top comments (0)