The parts that actually eat weeks are invisible: draw call budgets, GLB asset pipelines, skeletal animation compatibility across avatar sources, framework upgrades, and realtime networking that survives a bad connection. Nobody writes tutorials about these, and they don't show up in demos — but they decide whether your world runs on a phone or dies on load.
I spent the last couple of years building a browser-based, self-hosted 3D virtual world (~320k lines, vanilla JS + Three.js + Node.js, PC/mobile/tablet/XR). This post is a pit-by-pit teardown of the unsexy parts, with the real numbers from each fix. Everything below is measured from reproducible acceptance scripts in the repo — not estimated.
Pit 1: Draw calls quietly kill your frame rate
My first world scene hit 3,064 draw calls. On desktop it stuttered; on mobile it was a slideshow.
What actually helped, in order of impact:
- Merge static geometry and cut redundant material instances
- LOD for grouped assets — 118 asset groups got low-poly variants capped at ≤100 faces, cutting triangle counts by 71%–87%
- Parse model files inside a Worker instead of the main thread After the full pass: 3,064 → 837 draw calls (−73%). One more non-obvious one: changing the number of lights in a scene used to trigger a full shader recompile that froze the frame for ~9 seconds. Precompiling variants and keeping light counts stable brought that to 0. Pit 2: GLB assets are heavier than they look The models that artists hand you are not browser-ready. My asset pipeline now does what I call "GLB surgery":
- Texture downscaling and recompression: −68% to −92% texture size
- Mesh simplification to ≤100 faces per low-LOD variant
- Stripping unused nodes, morph targets, and embedded material bloat The rule I landed on: nothing enters the world folder without going through the pipeline. One 40MB "final" model can cost you more loading time than a hundred code changes. Pit 3: Four avatar sources, four skeleton conventions Mixamo, Ready Player Me, VRoid, and Root Motion clips all name bones differently and map retargeting differently. Getting one to work is a weekend; getting all four to coexist in one world is a project. The fix was a skeleton compatibility layer with per-source bone mapping, validated by a regression suite: 67/67 test cases pass across the four sources. If you're starting from scratch, build this harness before wiring your first avatar — retrofitting it after content depends on inconsistent skeletons is painful. Pit 4: Three.js r128 → r185 Staying three majors behind felt safe — until browser policy changes forced the move. The upgrade brought a color management pipeline change (every material's color shifted until corrected), and I had removed third-party CDN dependencies entirely, which meant writing 87 shim symbols for APIs that moved or disappeared. Lesson: pin your own copies of everything, and budget for upgrades as ongoing maintenance, not a one-time event. The web platform changes under you whether you move or not. Pit 5: Realtime networking must survive reality Multiplayer demos are recorded on good Wi-Fi. Real users are on elevators and subways.
- Reconnection with state recovery: our weak-network self-healing suite passes 9/9 scenarios (drop mid-move, drop mid-sync, resume after sleep, etc.)
- Voice chat needs a cap: we use a slot system (~10 concurrent speakers ≈ 1.3 Mbps) instead of unlimited mesh audio Pit 6: Letting AI agents walk into the world This is the part I think is genuinely differentiated. An AI agent connects with a domain + a key, and joins your world as a humanoid character — with coordinates, visible to real users, able to walk, talk, and guide visitors. The server only sends structured JSON; the visitor's browser renders everything. The counterintuitive part is the cost curve: because the server never renders anything, a visible AI is cheaper than an invisible one. Measured: ~1 KB/s per agent; 100 concurrent agents ≈ 0.079 CPU cores. Idle agents auto-disconnect after 5 minutes. Integration is three steps: publish a .well-known/virtual-world-agent.json descriptor → exchange a key for a 15-minute token → connect to /ws/agent. A zero-dependency Node client ships in examples/agent-client/. What it cannot do (and I'd rather say it up front):
- The AI cannot see — it gets a structured radar (observe, 200m range), never pixels
- No speech recognition or synthesis built in
- No hosted knowledge base — bring your own LLM and memory
- No terrain collision handling for agents; they can't touch or move world assets On the roadmap (planned, not shipped): an agent store, MCP server packaging for worlds, agent-to-agent (A2A) protocol, and cross-world federated agent networks. The pattern behind all six pits None of these problems show up in week one, and all of them block real users. That's the actual cost of "building a 3D world from scratch" — the renderer is maybe a tenth of the work. So I packaged the rest as a base layer you can deploy yourself: Layer What's in it Rendering & assets Three.js r185 / WebGL2, GLB pipeline, LOD, skeleton compatibility, 3D Gaussian Splatting support Performance & multi-device Draw call budget, shader precompile, Worker parsing, PC / phone / tablet / XR Realtime & federation Multiplayer sync, self-healing reconnect, slot-based voice, cross-world federation AI integration The agent system above, admin UI for keys and agents You write the layer on top: your scenes, your logic, your users. Licensing, stated plainly: the source is open — run it locally for personal use, free. Connecting it to the network, using it commercially, or federating worlds requires a paid license. Updates and support follow a subscription. (See the repo's LICENSE for the exact terms.) Honest current shortcomings
- First load is still heavy for large worlds; an asset audit found 79.3 MB of orphaned template assets (cleanup designed, not yet deployed) that should cut initial cross-world load from tens of seconds to ~1–2s
- Mobile performance is good but not yet profiled on low-end Android
- The frame loop still uses a fixed per-frame delta; low-FPS devices get worse motion feel than they should Source & repositories
- GitHub: https://github.com/miduo100/3d-virtual-world
Top comments (2)
The draw call reduction from 3,064 to 837 is impressive, but the weak-network recovery and avatar compatibility work are just as important. These are the engineering details that separate a cool demo from something people can actually use.
Appreciate it — and yes, the boring parts ended up mattering more. The skeleton compatibility layer alone took longer than the draw call work, because the four avatar sources disagree on bone naming and retargeting in ways you only find at runtime. The regression harness (67 cases) is now my most valuable test asset.