DEV Community

Cover image for Build a Colorful Wave of Cubes on the GPU with bgfx
puffball1567
puffball1567

Posted on

Build a Colorful Wave of Cubes on the GPU with bgfx

GPU Implementation with bgfx is a series about building a variety of GPU effects and techniques with bgfx. We'll start with small, working demos and explore animated shapes, drawing many objects efficiently, image effects, and GPU computation. Each article will connect what you see on screen to the code that makes it happen, explaining how the CPU and GPU share the work along the way.

For this first installment, let's make something move.

The scene below contains a 25 × 25 grid of cubes. Waves travel through the grid, lifting the cubes and shifting their colors. The CPU places the cubes on a flat grid; a vertex shader calculates their vertical movement on the GPU.

625 colored cubes rising and falling in a wave, rendered by the accompanying bgfx application

This is a capture of the working demo, not a mockup. Press Space to pause, Up/Down to change the amplitude, and R to reset it.

We'll build this effect step by step using bgfx's public API. You need some C++ familiarity, but no previous shader experience. The full demo source and build instructions accompany the article; the snippets below come from that implementation, with surrounding setup omitted where stated.

Representing 3D shapes with polygons

The scene shows cubes, but our program does not simply ask the GPU to “draw a cube here.” We describe the cube's surface as smaller faces and supply their positions and connections. Let's look at how that representation works.

A polygon is a shape with straight sides. Triangles and quadrilaterals are both polygons; in this demo, we divide every surface into triangles. Three points that do not lie on a straight line define one flat face. With four or more points, the points might not all lie in the same plane. This makes triangles a convenient building block for 3D surfaces, and they are the units used by our drawing setup.

To describe one square face of a cube, split it along a diagonal. Here is that face viewed straight on. The numbers 0 through 3 identify its corner points:

3 -------- 2
|        / |
|      /   |
|    /     |
|  /       |
0 -------- 1

Triangle A: 0 → 1 → 2
Triangle B: 0 → 2 → 3
Enter fullscreen mode Exit fullscreen mode

Together, the triangles cover the whole square without a gap. The diagonal shows where we split the surface; it does not need to appear as a visible line. A cube has six square faces, so 6 faces × 2 triangles = 12 triangles describe its entire surface. We do not need to fill its interior with smaller cubes. Drawing the surface and using depth to determine which face is in front gives it a solid appearance.

Each corner point is a vertex. In 3D, its position is described by three numbers: (x, y, z). Their reference point is the origin, at (0, 0, 0).

Let's establish the axis directions used in this demo. Imagine looking horizontally toward the origin from the negative-Z side. From that reference viewpoint, the directions are:

Axis Positive direction (increasing values) Negative direction (decreasing values)
X Right Left
Y Up Down
Z Away from you, beyond the origin Toward you

For example, (1, 0, 0) is one unit from the origin along positive X, (0, 1, 0) is one unit up, and (0, 0, 1) is one unit along positive Z. X and Z describe positions along the ground plane; Y describes height. These are this demo's conventions, not rules shared by every 3D application.

The actual demo camera looks at this space from above and at an angle, so positive X does not necessarily point directly right on screen. “Right, up, and away” in the table describe our reference viewpoint. Moving the camera does not change the axes in the scene. The cube's vertex coordinates will use its own center as the origin, while positions used to arrange the cubes will use the center of the entire grid. We do not rotate the cubes in this demo, so the axis directions in these two coordinate systems stay aligned.

Positions alone do not tell us which points form a face, so we also supply connections: “use 0, 1, and 2 for one triangle; use 0, 2, and 3 for another.” A shape represented by vertices and their face connections is called a mesh. We will prepare one cube mesh.

Curved surfaces, such as spheres and characters, can also be approximated with many triangles. Our cube's faces are flat, so just two triangles per face describe their shape. This small mesh will be our starting point for drawing and animation on the GPU.

What runs on the CPU, and what runs on the GPU?

Our program has three pieces:

CPU / C++
  Upload one cube mesh during initialization
  Each frame: supply time, amplitude, and a position for each cube
       |
       v
GPU / vertex shader
  Calculate the wave height at that cube's center
  Move each corner upward or downward by that height
  Transform the result into clip space
       |
       v
Rasterization → GPU / fragment shader
  Color the covered fragments and shade the cube faces
       |
       v
Depth testing and color output → window
Enter fullscreen mode Exit fullscreen mode

A vertex shader calculates where each mesh vertex goes for drawing. The GPU assembles triangles from the resulting vertices and the supplied connections. Rasterization determines which screen pixels those triangles cover: it connects the faces described by coordinates to the grid of pixels that makes up the image. A fragment shader then calculates color for the covered samples. Depth testing also determines visibility, so a hidden face's color does not simply appear on screen.

The shaders in this project use bgfx's GLSL-like shader language and are compiled with its shaderc tool. bgfx shader tooling

We will reuse one mesh, but submit a separate draw for each cube. This makes the connection between a cube's position and its shader inputs easy to follow. It does not make 625 cubes a single draw call. Instancing is a useful next step once this version works.

1. Project directory structure

We will organize the program in a directory named demo-wave-cubes, separating the C++ code from the shaders that run on the GPU:

demo-wave-cubes/
├── main.cpp              window, mesh, input, drawing loop
├── demo_support.h        vertex types and shader-loading helpers
├── build.sh              shader and C++ build commands
├── make_overlay.py       generate the transparent title/help image
└── shaders/
    ├── vs_wave.sc        vertex shader for wave displacement
    ├── fs_wave.sc        color, lighting, edge accents
    ├── varying.def.sc    vertex inputs and shader-stage connections
    ├── vs_overlay.sc    vertex shader for positioning the title image
    └── fs_overlay.sc    fragment shader for drawing the title image
Enter fullscreen mode Exit fullscreen mode

main.cpp will prepare the geometry and draw inputs; the shaders in shaders/ will calculate the cubes' movement and color. The title and keyboard help will use a transparent image generated by make_overlay.py.

Building the project will generate the executable, compiled shaders, and overlay image in build/. We will cover the build and run commands after the implementation. First, let's prepare the geometry for one cube.

2. Describe one cube

First, define what one vertex contains. This is the actual type in demo_support.h, included by main.cpp:

struct Vertex { float x, y, z, nx, ny, nz; };
Enter fullscreen mode Exit fullscreen mode

x, y, and z are the local position. nx, ny, and nz describe a normal: a direction pointing out of the face, used later for lighting. The top face has the normal (0, 1, 0); the face pointing along positive X has (1, 0, 0).

The demo creates four vertices for each of six faces: 24 vertices in total. A geometric cube has only eight corners, but one corner belongs to three faces with different normals. Keeping separate vertex records gives each face a flat appearance.

We encode the connections from the diagram as indices: numbers that refer to entries in the vertex array, starting at zero. In our drawing setup, each group of three indices specifies one triangle. For the first face:

Four face vertices: 0, 1, 2, 3
First triangle:    0, 1, 2
Second triangle:   0, 2, 3

6 faces × 2 triangles × 3 indices = 36 indices
Enter fullscreen mode Exit fullscreen mode

The first face's index array is {0, 1, 2, 0, 2, 3}. Both triangles refer to vertices 0 and 2, so they reuse those vertex records within the face. The next face has four separate vertices and uses {4, 5, 6, 4, 6, 7}. Repeating this for all six faces gives the demo's 36 indices. 24 counts vertex records, 12 counts triangles, and 36 counts references to vertices; these numbers describe different things.

main.cpp fills std::vector<Vertex> vertices and std::vector<uint16_t> indices with that data. Then it describes the vertex layout and creates the buffers:

bgfx::VertexLayout layout;
layout.begin()
    .add(bgfx::Attrib::Position, 3, bgfx::AttribType::Float)
    .add(bgfx::Attrib::Normal, 3, bgfx::AttribType::Float)
    .end();

// Copy the CPU arrays into memory managed by bgfx.
auto vb = bgfx::createVertexBuffer(
    bgfx::copy(vertices.data(), uint32_t(vertices.size() * sizeof(Vertex))),
    layout);
auto ib = bgfx::createIndexBuffer(
    bgfx::copy(indices.data(), uint32_t(indices.size() * sizeof(uint16_t))));
Enter fullscreen mode Exit fullscreen mode

Vertex stores six floats in this order: x, y, z, nx, ny, nz. The vertex buffer receives their bytes, without the C++ member names. We therefore need to tell bgfx which part represents a position and which part represents a normal. layout describes how to read that data.

The first .add(bgfx::Attrib::Position, 3, bgfx::AttribType::Float) says, “read the first three floats as a position.” The next .add(bgfx::Attrib::Normal, 3, bgfx::AttribType::Float) says, “read the next three floats as a normal.” Here, 3 counts the components of one position or direction, not the number of vertices. begin() starts the layout definition, and end() finalizes it.

One vertex:    [ x, y, z   | nx, ny, nz ]
Layout meaning:[ position | normal     ]
Shader input:  [ a_position| a_normal   ]
Enter fullscreen mode Exit fullscreen mode

For example, a vertex on the top face might contain {0.46f, 0.46f, 0.46f, 0.0f, 1.0f, 0.0f}. The first three values become the position (0.46, 0.46, 0.46), and the last three become the upward normal (0, 1, 0). Our shader receives them through inputs named a_position and a_normal. This also connects to the POSITION and NORMAL definitions in varying.def.sc, which we will examine later.

In this build, a float occupies four bytes: 12 bytes for the position and 12 for the normal, making 24 bytes per vertex. The layout also records this distance from one vertex to the next, called the stride. With multiple vertices in the buffer, the same reading pattern repeats every 24 bytes.

That is why createVertexBuffer receives both the data and layout: the data supplies the numbers, and the layout describes their arrangement and meaning. The layout does not rearrange the numbers, so the order and types in the .add() calls must match the actual Vertex records.

vb is a bgfx::VertexBufferHandle; ib is a bgfx::IndexBufferHandle. They identify resources managed by bgfx. They are not pointers into the vectors. bgfx::copy copies the supplied bytes, so their later use does not depend on keeping the vectors' backing storage alive. Buffer and memory API

These buffers are created once, before the frame loop. Every cube uses them.

3. Give each cube a position

The previous section prepared one cube centered at (0, 0, 0). We will reuse that shape to make a grid of 25 by 25 cubes: 625 in total. Recall our reference directions: X is left/right, Z is depth, and Y is height.

Start with grid indices

Number the positions from 0 to 24 along both X and Z. The code calls these integer grid indices x and z. They are not yet coordinates measuring distance in the scene.

For example, (x, z) = (0, 0) selects the first position along both directions. (12, 12) selects the middle. Counting from zero makes index 12 the thirteenth position, the middle of 25 positions.

A top view mapping the 25-by-25 grid indices to cube centers, with a second diagram translating a cube by 1.12 along X

Each dot in the upper diagram is a cube center. Viewed from above, X runs across the page and Z runs up the page; Y is perpendicular to it. This is a different viewpoint from the demo's angled camera.

Turn each index into a center position

Using the indices directly as coordinates would place all cubes on the positive side of the origin. Subtract 12 to center the grid around zero:

Index:            0    1   ...   12   ...   23   24
Subtract 12:    -12  -11   ...    0   ...   11   12
Enter fullscreen mode Exit fullscreen mode

Then multiply by the center spacing, 1.12. Each cube extends from -0.46 to +0.46 on each axis, making its side length 0.92. Spacing the centers by 1.12 leaves a gap of 1.12 - 0.92 = 0.20 between adjacent faces before displacement. This spacing is our design choice, not a bgfx constant.

Here is the calculation along X. The same calculation applies to Z.

Grid index x Center around zero: x - 12 Center X: (x - 12) * 1.12
0 -12 -13.44
11 -1 -1.12
12 0 0.00
13 1 1.12
24 12 13.44

Grid indices (13, 12) therefore give the cube center (1.12, 0, 0). Indices (12, 13) give (0, 0, 1.12). For now, every center's Y coordinate stays zero.

Move the whole cube to that position

Moving every vertex by the same amount is called translation. To place the center at (1.12, 0, 0), add 1.12 to every vertex's X coordinate:

Vertex within cube       Translation       Vertex in scene
(0.46, 0.46, 0.46)   +  (1.12, 0, 0)  =  (1.58, 0.46, 0.46)
Enter fullscreen mode Exit fullscreen mode

The lower diagram shows the center and the rightmost vertex moving by the same amount. Applying the same translation to every vertex preserves the cube's shape and size.

We describe this operation to the shader with a model matrix. A matrix is a rectangular arrangement of numbers; in 3D graphics, matrices represent operations such as translation, rotation, and scaling. Here we only translate. Coordinates relative to the cube's own center are called local coordinates; coordinates in the scene containing the entire grid are world coordinates. Our model matrix describes the conversion from local to world coordinates.

Inside the per-cube drawing loop, the implementation uses:

// x and z are grid indices, each ranging from 0 to 24.
float model[16];
bx::mtxTranslate(model, (x - 12) * 1.12f, 0.0f, (z - 12) * 1.12f);
bgfx::setTransform(model);
Enter fullscreen mode Exit fullscreen mode

model is an array of 16 floats holding a 4 × 4 matrix. bx::mtxTranslate writes the matrix into its first argument, model. The next three arguments specify the X, Y, and Z translation. We do not need to calculate the individual matrix entries ourselves.

bgfx::setTransform(model) sets the matrix for the draw we will submit next. This call alone does not draw a cube. We will pair it with bgfx::submit later, and the vertex shader will use the matrix to calculate each vertex's position. The original vertex buffer stays unchanged.

Repeating this setup and drawing the shared cube inside the nested x and z loops fills the grid. So far, the CPU has only arranged the cubes on a flat plane. Next, we will pass time to the GPU and add the vertical wave motion.

4. Pass time into the shader

A uniform supplies a value that stays the same across a draw. We need a clock and an amplitude, so one four-component float value is enough.

Create its handle once, beside the buffers:

auto uWave = bgfx::createUniform("u_wave", bgfx::UniformType::Vec4);
Enter fullscreen mode Exit fullscreen mode

Then fill its values during each frame:

// time is elapsed seconds; amplitude starts at 1.35.
const float wave[4] = {time, amplitude, 0.0f, 0.0f};

// Set this before each cube's submit; the final two components are unused.
bgfx::setUniform(uWave, wave);
Enter fullscreen mode Exit fullscreen mode

The corresponding shader declaration is:

uniform vec4 u_wave;
Enter fullscreen mode Exit fullscreen mode

The name u_wave connects the C++ uniform with the shader declaration. Its .x component carries time and .y carries amplitude. Pausing freezes the C++ clock; the shader continues rendering with the same input values. Uniform API

5. Calculate the wave at the cube's center

Here is the movement calculation from vs_wave.sc:

// A local origin transformed by the model matrix is the cube's center.
vec3 center = mul(u_model[0], vec4(0.0, 0.0, 0.0, 1.0)).xyz;
float radius = length(center.xz);
float height = u_wave.y * (
    sin(radius * 0.65 - u_wave.x * 1.8)
    + 0.35 * cos(center.x * 0.5 + center.z * 0.35 + u_wave.x));
Enter fullscreen mode Exit fullscreen mode

u_model[0] is bgfx's predefined model matrix, populated through setTransform. .xz selects the two horizontal coordinates, and length measures their distance from the grid's center. Predefined uniforms

Read the first wave as three controls:

  • radius * 0.65 changes the phase as we move outward, producing rings.
  • -time * 1.8 shifts those rings over time.
  • amplitude controls the vertical distance traveled.

The smaller cosine wave adds variation across X and Z, so the result is not perfectly circular. These coefficients are artistic choices in this demo, not physical constants.

For a concrete check, put the center at (0, 0, 0) and time at zero. The sine term is zero and the cosine term is one, so the height is 1.35 × 0.35 = 0.4725. Setting amplitude to zero makes the height zero everywhere.

Why use the center, rather than the current vertex position? All corners of a cube need the same displacement. Sampling the wave separately at each corner would deform the cube. Here, every vertex invocation for a given cube independently calculates the same height, keeping it rigid.

6. Move each vertex and project it

Next, the shader transforms the actual corner and adds that height:

vec3 world = mul(u_model[0], vec4(a_position, 1.0)).xyz;
world.y += height;
gl_Position = mul(u_viewProj, vec4(world, 1.0));
Enter fullscreen mode Exit fullscreen mode

a_position comes from the vertex buffer. world is the position after placing that cube in the grid. u_viewProj combines the camera view and projection; gl_Position is the clip-space output consumed by the graphics pipeline.

The CPU creates the camera matrices once because this demo has a fixed camera:

float view[16], projection[16];
bx::mtxLookAt(view, {27.0f, 25.0f, -32.0f}, {0.0f, 0.0f, 0.0f});
bx::mtxOrtho(projection, -25.0f, 25.0f, -11.82f, 11.82f,
    0.1f, 120.0f, 0.0f, bgfx::getCaps()->homogeneousDepth);
bgfx::setViewTransform(0, view, projection);
Enter fullscreen mode Exit fullscreen mode

An orthographic projection keeps the grid's distant cubes from shrinking with perspective, which suits this graphic, diagram-like composition. View 0 is the scene's drawing group. The title overlay uses a separate view.

7. Give the wave color and depth

The vertex shader passes three values onward: the world position, the face normal, and the original local position. varying.def.sc declares the connection between the shader stages:

vec3 a_position : POSITION;
vec3 a_normal : NORMAL;
vec3 v_normal : TEXCOORD1;
vec3 v_world : TEXCOORD2;
vec3 v_local : TEXCOORD3;
Enter fullscreen mode Exit fullscreen mode

This excerpt shows the wave inputs and outputs; the complete file also contains the overlay's UV coordinates. Here TEXCOORD1, for example, identifies an interpolated data channel. It does not require a texture. The vertex shader assigns v_normal = a_normal, v_world = world, and v_local = a_position; the fragment shader receives interpolated values across each triangle.

In fs_wave.sc, three offset cosine curves produce the red, green, and blue components:

float phase = v_world.x * 0.022 + v_world.z * 0.018 + v_world.y * 0.055;
vec3 color = 0.55 + 0.45 * cos(6.2831853 *
    (vec3(phase, phase, phase) + vec3(0.02, 0.35, 0.65)));
Enter fullscreen mode Exit fullscreen mode

The phase depends on height as well as horizontal position. That makes a cube change color as it rises and falls, without uploading a new per-cube color from C++.

For lighting, compare the surface direction with one fixed light direction:

float light = 0.32 + 0.68 * max(dot(normalize(v_normal),
    normalize(vec3(-0.45, 0.85, -0.55))), 0.0);
Enter fullscreen mode Exit fullscreen mode

The dot product grows as the two normalized directions align. max(..., 0.0) prevents negative lighting, while 0.32 keeps the unlit faces visible. This is a simple directional-light effect, with no cast shadows. Since our model matrices only translate, the normals remain valid as supplied. Adding rotation or nonuniform scaling would require transforming the normals too.

The full shader also adds a narrow bright accent near each face edge. It measures how close the local position is to two cube boundaries. This makes the cubes easier to separate visually; it does not add bevel geometry, bloom, or a physical material simulation.

8. Submit the same mesh with different inputs

The program handle is created before the frame loop by our own loadProgram helper in demo_support.h. It reads vs_wave.bin and fs_wave.bin, creates a shader from each with bgfx::createShader, and combines them using bgfx::createProgram(v, f, true). The full helper checks failures; the final true transfers responsibility for releasing those shader references to the program.

With the resources ready, the inner draw loop is:

// model and wave are the values constructed in the preceding steps.
bgfx::setTransform(model);
bgfx::setVertexBuffer(0, vb);
bgfx::setIndexBuffer(ib);
bgfx::setUniform(uWave, wave);
bgfx::setState(BGFX_STATE_WRITE_RGB | BGFX_STATE_WRITE_A |
               BGFX_STATE_WRITE_Z | BGFX_STATE_DEPTH_TEST_LESS);
bgfx::submit(0, program);
Enter fullscreen mode Exit fullscreen mode

The first 0 in setVertexBuffer selects vertex stream zero; the 0 in submit selects view zero. They name different things. Depth writing and testing keep a farther cube from painting over a nearer one. This introductory mesh disables face culling; adding consistent winding and culling is a separate improvement.

After all cube draws and the overlay draw, call:

bgfx::frame(); // Advance the frame after recording this frame's drawing work.
Enter fullscreen mode Exit fullscreen mode

The next frame supplies a new time. We never rebuild or upload the cube mesh to animate the wave. The CPU still builds matrices and records separate draws; moving an equation into a shader does not remove submission overhead. Drawing API

9. Build and run the source

With the implementation assembled, we can build the source and see it in motion.

The accompanying build targets Linux, SDL2/X11, and bgfx's OpenGL backend. XWayland also works in the environment used here. It uses a bgfx API 159 SDK with matching bx/bimg libraries and shaderc, plus a C++20 compiler. Other backends need matching shader binaries and platform setup; this particular build does not select them automatically.

Prepare the SDK and other dependencies using the README's build instructions, then run these commands from demo-wave-cubes:

# The directory must contain bgfx/, bx/, and bimg/.
export BGFX_SDK=/path/to/bgfx-sdk
bash build.sh
./build/wave-cubes
Enter fullscreen mode Exit fullscreen mode

build.sh compiles the shaders and generates the overlay image and C++ executable. Running ./build/wave-cubes opens the animated field of 625 cubes. Press Space to pause, Up/Down to change the amplitude, R to reset, and Esc to exit.

Check the effect instead of trusting the picture

The supplied validation script renders screenshots through bgfx, then compares their pixels. It checks that a zero-amplitude field stays unchanged when time changes, and that a nonzero-amplitude field does change. The capture path reads this application's framebuffer, not the desktop.

The demo was built and exercised on Linux with SDL 2.0.20 and a bgfx API 159 SDK. bgfx reported OpenGL 4.3; the display's driver reported an accelerated AMD RENOIR device with Mesa 23.2.1. The six-second capture contains 180 frames sampled at a fixed 1/30-second simulation step. That is an animation export setting, not a measured performance claim.

Try changing just one term:

  • Set amplitude to zero to recover the flat grid.
  • Remove the cosine term to see only circular waves.
  • Increase 0.65 to fit more rings into the same space.
  • Remove the Y term from the color phase to keep the color field stationary.

Rebuild after changing a shader. Its .sc source is compiled into a .bin asset; editing the source alone does not replace the binary loaded by the application.

Where to take it next

We now have a small effect with a clear division of work: C++ supplies shared geometry and draw inputs, the vertex shader supplies movement, and the fragment shader supplies appearance.

The next practical question is how to draw this grid with fewer CPU submissions. A follow-up can replace the per-cube draw loop with instance data while preserving the same wave equation. After that, render targets and compute shaders open up effects that need intermediate images or persistent simulation state.

Clay Board Style System

Clay Board Style System is my CSS-inspired foundation for native GUI toolkits. It includes an optional GPU host and a bgfx adapter through bgfxim, connecting custom GPU content with native layout and input. If you want to take shader experiments into an application UI, see its GPU host documentation.

Top comments (0)