We started with Three.js and WebGL. Once we had models rendering in the browser, we worked through the usual optimizations: loading in chunks, reducing detail at a distance, batching draws, and skipping occluded geometry.
As that work progressed, changes began to affect one another. Drawing less required more checks. Keeping the display stable meant dealing with visibility decisions that changed from frame to frame. A fix in one part meant another round of checks elsewhere.
Eventually, we decided to rebuild our BIM rendering core with WebGPU. Looking back, several attempts that didn't work out as expected help explain why.
If it's marked invisible, why are we still drawing it?
One investigation in the old viewer stands out. The diagnostics showed that many elements had been marked invisible, but the amount of drawing hadn't gone down.
We traced the problem to a mismatch in granularity. The visibility logic could decide that an individual element didn't need to be drawn. The data being used to display it, however, didn't always let us remove that element independently. A more precise decision wasn't necessarily something the renderer could act on.
We tried adding checks for some of those elements. It saved no drawing work and caused more changes in their visibility state. We removed that change.
This was a limitation of how our old implementation organized data and applied its decisions. Blaming WebGL alone would have missed the problems in our own implementation. But adding more conditions to the visibility logic wasn't going to fix the layer underneath it.
Investigations like this led us to break down the work in each frame again. The CPU had to traverse the scene, update state, and prepare drawing commands. The main thread also had to handle mouse input and the application UI. Drawing only one room didn't mean the work needed to decide what to draw was cheap.
We can group that work into a few broad categories. Some of it overlaps, so the timings can't simply be added together:
After that, we started asking one more question when reviewing an optimization: did the work we claimed to have saved actually disappear from the steps that followed?
WebGPU gave us more room to divide the work
WebGL already let us use the GPU for drawing. What made WebGPU worth considering was the flexibility of its compute and drawing interfaces.
Compute shaders can process batches of work that lend themselves to parallel execution. Storage buffers can hold the data those tasks need. Indirect draws can read their drawing parameters directly from a buffer.
In a typical GPU-driven design, some filtering and draw preparation can happen on the GPU. That creates opportunities to reduce CPU work per object and avoid certain round trips between the CPU and GPU. Chrome's introduction to WebGPU discusses these capabilities as well.
Moving work has a cost, though. Data has to be prepared and transferred, and tasks may have to wait for one another. Small, scattered tasks may not be worth moving. Large batches of independent, similar decisions are more promising candidates.
We also considered replacing only the renderer while keeping our existing data preparation. We set that option aside during evaluation. If the data organization and repeated processing stayed the same, a new graphics API would inherit those costs.
The rebuild therefore included how we prepared data before it reached the renderer. We needed to reconsider what belonged on the GPU, what should stay on the CPU, and what work could be avoided earlier.
Shared geometry still needs to be drawn
Imagine a scene with a thousand identical chairs. They share the same geometry and material, but sit in different positions.
Storing a complete copy of the geometry for every chair would duplicate a lot of data. Instancing lets the chairs share geometry while retaining their own positions. Three.js's InstancedMesh supports this kind of use.
But all the chairs visible on screen still need to be drawn. A smaller file doesn't remove the pixels they cover or the occlusion between them.
Batching involves a similar tradeoff. Combining many objects into one draw batch can reduce submission work. If that batch spans several rooms, viewing one room may also cause geometry from the others to be submitted.
BIM adds another requirement: individual elements must remain usable. Two identical valves still need to be selected, hidden, and inspected separately. Large batches can reduce draw calls while making those operations harder to support.
Culling has a cost too
Inside a building, many elements are hidden behind walls. Not drawing them sounds like a straightforward optimization.
First, the engine has to establish that they really are hidden.
The view frustum, spatial hierarchies, and occlusion information can all help narrow down what needs processing. Each approach has different costs and works best in different situations. The most detailed check isn't necessary for every object.
Consider a CPU that issues an occlusion query and immediately waits for the GPU to return the result. It may save some drawing, but it also gives up an opportunity for CPU and GPU work to overlap. A long enough wait can make the whole frame slower. NVIDIA's discussion of occlusion culling covers this problem.
Correctness matters just as much. When the camera turns a corner, newly exposed equipment needs to appear promptly. A coarse decision mustn't hide pipes behind glass. When visibility is uncertain, drawing a little extra is usually preferable to leaving out an element.
Move the camera above the building and the tradeoff changes again. Much more of the scene is visible, so there is less drawing to eliminate, while the checks still cost something. An approach that works well indoors may offer little benefit from above.
Fewer draw calls, more GPU time
A draw call is a request to draw geometry. Counting draw calls is useful when looking at submission work, but that count doesn't describe the full rendering cost.
We ran into this again shortly after moving to WebGPU.
Once the minimal viewer could display a model, we tried merging more draw batches. The reasoning was straightforward: fewer calls should be faster. The call count did fall. GPU time increased, though, and the model took longer to appear.
We didn't adopt that experiment for the main rendering path, but kept it as a reference. We hadn't established exactly why it was slower. What we could tell was that further merging hadn't helped that version.
A single draw call can still contain a great deal of geometry. Overlapping transparent objects, shadows, and post-processing can also keep the GPU busy.
Resolution provides a simple example. Double both the width and height of an image, and the pixel count quadruples. Rendering time won't necessarily quadruple, but the GPU has more pixels to process.
WebGPU render bundles let us reuse previously recorded drawing commands, avoiding repeated encoding. Indirect draws let the GPU obtain drawing parameters from a buffer. Both are useful, and neither addresses every cost.
To validate a change like this, we still need to repeat the same interaction: does dragging feel smoother, does the UI pause, and has the image quality changed?
How much space does a 4K texture take?
The number of bytes downloaded often differs substantially from the memory a resource occupies at runtime, including GPU memory.
Take a hypothetical 4096 × 4096 texture. With uncompressed RGBA8 at 4 bytes per pixel, the pixel data for that resolution level alone occupies:
4096 × 4096 × 4 = 67,108,864 bytes, or 64 MiB.
That excludes other resolution levels and additional resources. GPU texture compression changes the footprint again. The download size can't be used directly as an estimate of runtime texture memory.
These costs are worth keeping separate:
A model loading successfully once doesn't settle the question. We also need to check repeated model switches. Long sessions, canceled loads, and closed views all put resource management to the test. A graphics API upgrade doesn't handle that work for the application.
A black screen in another browser
After integrating the new core into the viewer page, we encountered a black screen. The page worked in our development environment, but failed to render in an embedded browser.
We eventually found that our resource usage exceeded what that environment allowed. Our development machine exposed more generous limits, so it hadn't revealed the problem. After adjusting resource usage and testing again, rendering worked in that environment.
“Supports WebGPU” isn't the whole story. The browser, operating system, GPU, and driver all affect the capabilities available. Once something works on the development machine, the target environments still need to be tested.
Framework maintenance belongs in this calculation too. Three.js already provides WebGPURenderer, so choosing WebGPU still leaves room to use a framework. Taking direct responsibility for more of the rendering also means taking on its debugging, compatibility, and resource management.
We kept both viewers available during the transition. We compared the same models from the same camera positions and gradually brought the interactions across. Only after the new viewer became the default did we archive the old one.
We kept the batching experiment that had made rendering slower, too. If we revisit that direction, we can first look at what we changed and why we didn't carry it forward.
Rebuilding with WebGPU gave us more control over where computation happens and how drawing is organized. Whether a particular change helps still comes down to opening that model and moving the camera back to the spot where it used to stutter.


Top comments (0)