DEV Community

Selçuk Karayel
Selçuk Karayel

Posted on

One process, one account — the architecture behind rebuilding a 28-year-old IRC network in Go

``In the first post I showed what the platform does. This one is about how it's built: how the code is laid out, what Go made easy, which problems turned out to be hard, and how each one was solved.

The constraint that drove everything: a person must be the same person no matter how they connect. A raw IRC client, a browser app and a website all share one identity, one inbox, one social graph and one set of permissions, and none of them is a second-class client.

1. One process instead of a fleet

The classic IRC setup is an ircd plus a separate services daemon, linked over a server-to-server protocol, with a website off to the side reading a copy of the data. Every piece has its own view of the world, and most of the bugs live in keeping those views in sync.

I went the other way. One Go binary contains the IRC server, the services (accounts, channels, social graph, admin) as ordinary Go packages, the HTTP side (website, JSON APIs, admin panel, browser gateway) and a Lua runtime for bots.

Services aren't linked to the server; they are the server. There's no burst on reconnect, no desync, and no protocol between them. A service handler reads the same in-memory user and channel structures the server uses and writes to the same database.

The cost is a single failure domain. Deploys are built around it: a short-lived bridge process opens the same port with SO_REUSEPORT while the new binary starts, so the HTTP side never goes dark.

2. How the code is laid out

About 300k lines of Go (excluding 265 test files), split by responsibility rather than by protocol:

Package Lines Role
core 23k IRC server: connections, channels, modes, SASL, IRCv3, WebRTC signalling, port sharing
services 25k NickServ, ChanServ, the social service, admin service, Lua engine
veri 4k the data layer: the only place allowed to touch shared tables
content 75k site domain logic: feed, groups, games, content, notifications
web/handlers, web/render 53k HTTP handlers, templates, admin API
oyun 9k game engine: rule packages, drawing packages, bridges
wasm/chat, wasm/tema, wasm/sa 100k the three browser programs (TinyGo)

Two rules keep this manageable. First, dependencies point inwards: handlers call content, content calls veri, and nothing below reaches up. Second, one data layer owns the shared records: accounts, settings and messages are written through veri, so the IRC side and the web side can't disagree about what a record looks like.

The external dependency list is short on purpose: pion/webrtc, gopher-lua, modernc.org/sqlite, bluemonday, minify, webpush-go, a GeoIP reader, and the x/ packages.

3. What Go made easy

A goroutine per connection, a channel per client. Each IRC client gets a reader goroutine and a bounded outbound channel. Broadcasting to a channel is just a non-blocking send into every member's queue. The select-with-default pattern gives you slow-consumer protection almost for free (see below).

IRC and HTTP in the same process with no ceremony. A raw TCP/TLS listener and net/http live side by side and share memory directly. The website can ask the IRC server "is this user online right now?" with a function call, not an RPC.

One static binary. modernc.org/sqlite is a pure-Go SQLite, so there's no cgo, no shared libraries and no build matrix. Deploying the server means copying one file.

The same language on the client. TinyGo compiles Go to WebAssembly, so the browser programs share types and helpers with the server instead of re-describing them in another language.

Profiling a live system. pprof on a running server turned every performance argument into a measurement. Both of the biggest wins below came from a single profile.

Low-level control when needed. net.ListenConfig.Control gives access to the raw socket, which is all it takes to set SO_REUSEPORT for the deploy bridge, with separate build-tagged files for Linux and BSD.

4. The hard parts, and how they were solved

SQLite under concurrent load

SQLite is a great single-file database, but it has one writer. The first version shared one connection (SetMaxOpenConns(1)) for everything. A stress test with pprof showed that with 32 concurrent requests, half of all waiting time was spent in database/sql waiting for that one connection. Reads were queueing behind other reads.

The fix separates the two paths:

  • One writer connection for all writes and transactions.
  • A separate read pool in WAL mode, where readers don't block each other or the writer. The pool is opened with query_only(ON), so an accidental write through it fails loudly instead of quietly succeeding.
  • A small page cache per connection plus a shared mmap. Sixteen connections with a 32 MB cache each had pushed memory towards 915 MB while the same cold pages were read over and over. A small per-connection cache with shared memory-mapped I/O fixed both.

A global lock in the scripting engine

Bots run in a single Lua state, and a Lua state isn't goroutine-safe, so it sits behind one lock. The original design called the script hook synchronously from the sender's own connection loop. A bot command that made an outbound HTTP call held the lock for seconds, and in that time every user on every channel was waiting behind it. Chat lag with no obvious cause.

The fix moves scripting off the hot path: every message goes into a bounded, ordered queue drained by a single worker, which keeps message order for the bots. If the queue is full, the message simply isn't offered to the scripts (and is counted). A user's message is never delayed by a bot. Per-account rate limits on bot commands stop anyone from flooding the queue on purpose.

Slow consumers

One client on a bad mobile connection must never slow down a busy channel. Each client's outbound queue is a buffered channel whose size can be set per connection class. Sending into it is non-blocking; if the buffer is full, that client is disconnected with SendQ exceeded, and everyone else carries on. The rule is simple: the server never waits for a client.

Constraints of Go in the browser

TinyGo made the browser side possible, but it asks for discipline:

  • Binary size. The webchat is 5.8 MB (1.9 MB gzipped). Reflection-heavy packages are avoided; JSON is parsed by the browser's own JSON.parse and read through syscall/js instead of encoding/json.
  • 32-bit int. TinyGo's wasm target uses a 32-bit int, so hashes are kept in explicitly sized unsigned types and modulo arithmetic happens on the unsigned side.
  • One thread. Goroutines are cooperative in the browser. Long work has to yield, and everything that talks to the network is written as callbacks from the event loop.

In return we get a strict Content-Security-Policy: no inline scripts, no on* attributes, no eval. The only JavaScript on the page is the WASM loader, and every UI event goes through one delegation layer that dispatches on data-* attributes.

5. Identity across three protocols

Client Credential
Raw IRC client SASL (PLAIN / EXTERNAL) or NickServ
Website session cookie
Browser app short-lived token

All three end in one resolver that returns an account. Features never branch on the transport; they only ask who is this. Account names also go through one normalisation function everywhere a name becomes a map key or a database key, because case folding differs between browsers, Go's standard library and SQL.

6. Messages: one store, delivery decided by presence

There is one message store, and every client writes to it. Where a message lands depends on the recipient's connection state at that moment, not on the sender's client:

Situation Result
Recipient connected Delivered live as a protocol message; the stored copy is marked read at once
Sender on the web, recipient connected Stored, then pushed into the recipient's live session through a localhost-only internal bridge
Recipient offline Stored unread, and surfaced in the inbox on their next visit, from any client

Deciding by recipient presence rather than sender surface removed a whole family of "the message went to the wrong place" problems. Offline history is kept only between mutual friends, as a privacy rule.

7. Social features on top of protocol primitives

Follows, friendships, groups and notifications are a service inside the IRC server, and they reuse IRC primitives wherever possible. A group is a channel: roles are channel permissions, and announcements and events are delivered into the channel as messages. The website adds a page on top; it isn't the source of truth.

So the social layer inherits everything the chat layer already has: permissions, moderation, bans and flood limits.

8. A pluggable transport, with the server kept honest

The browser app reaches the IRC server through a gateway. The transport between browser and gateway is pluggable and picked from the admin panel: HTTP long-polling (what runs in production), a WebRTC data channel, or WebSocket. Whatever the transport, the gateway connects with WEBIRC, so the server sees the real client IP and applies the same bans and limits as for any native client.

9. Real-time media as separate systems

  1. 1:1 calls: WebRTC peer to peer, with signalling relayed by the server and short-lived TURN credentials.
  2. Moderated voice rooms per channel: media is peer to peer, but room state (speaking, raised hands, mutes) lives on the server, so moderation is authoritative.
  3. One-to-many streaming over WHIP/WHEP: screen, game or camera straight from the browser, or from OBS.

Keeping these as separate systems, each with its own signalling, state and failure modes, means a problem always points at one small system instead of one big "media" module.

10. Games: the server is the only authority

The table games share one engine. Each game is a rule package (legal moves, validation, outcome), a drawing package that knows nothing about the game, and a thin bridge. Seats, tables, rankings and moderation are shared.

The server never sends the game state to the browser. It computes a separate view for each seat: in a card game you get your own hand, your opponents' hands arrive only as a count, and spectators see none. The client draws the view and sends back a move chosen from the legal moves the server supplied.

11. Numbers

Single VPS, 6 vCPUs and 12 GB of RAM:

  • Binary: 35 MB, statically linked, with no cgo
  • Webchat WASM: 5.8 MB (1.9 MB gzipped); site WASM: 1.5 MB (0.5 MB gzipped)
  • About 300k lines of Go and 265 test files; about 13k lines of sandboxed Lua for bots and games
  • About 550 MB RSS
  • One SQLite database, with one writer connection and a WAL read pool

What's next

  • Load. A much larger network may move onto this platform, so the gateway is being load-tested for several thousand concurrent browser clients. That will be the next post, with numbers.
  • Always-on sessions. A bouncer built into the server, so a user stays in their channels and keeps their history with every client closed.

If you've run long-polling at a few thousand clients in Go, or pushed SQLite harder than this, I'd like to hear what broke first.

Try it without an account: https://www.yudum.net/webchat/
``

Top comments (2)

Collapse
 
kyisaiah47 profile image
kyisaiah47 •

How are you checkpointing WAL when a read stays open for a long time?

Collapse
 
yudumnet profile image
Selçuk Karayel •

Good question. We mostly rely on SQLite's automatic checkpointing (the default wal_autocheckpoint of 1000 pages, which runs a PASSIVE checkpoint on commit) and on keeping read transactions short. Feed queries read their rows into memory and close the cursor before doing any follow-up lookups, so no reader holds a snapshot while other queries run. Pooled read connections also have a one-hour max lifetime.
A PASSIVE checkpoint can't get past a reader that's still open, but because our reads are short, it catches up on a later commit. What PASSIVE doesn't do is shrink the -wal file: it grew to its high-water mark and stayed there (about 25 MB for us). As of today's release, a scheduled wal_checkpoint(TRUNCATE) runs at a quiet hour, and journal_size_limit is set on the writer connection, so the file gets reset regularly. If a long reader blocks the truncate, it simply retries the next night.