Writing a voxel engine from the ground up teaches you computer graphics in the most raw, visceral way possible. When you place a block in a 3D world, you quickly realize that naive cube rendering (12 triangles per block) will choke any modern GPU the moment you render a modest landscape of 16x16 chunks.

In this retrospective, I break down the core architectural decisions behind Voxel Engine 2 and how we achieved smooth 60+ FPS chunk streaming with dynamic terrain modification.

1. The Naive Problem: Geometry Explosion

Consider a standard 16x16x256 chunk. That’s 65,536 potential voxels.

  • If rendered naively: 65,536 * 6 faces * 2 triangles = 786,432 triangles per chunk.
  • With a standard render distance of 16 chunks around the player (33x33 = 1089 chunks), you would be asking the GPU to push 850 million triangles per frame.

Clearly, optimization must start before a single byte reaches the Vertex Buffer Object (VBO).

graph LR
    Noise["3D Simplex Noise"] --> VoxelGrid["3D Array of Voxel Data"]
    VoxelGrid --> FaceCulling["Occlusion Culling<br/>(Drop hidden interior faces)"]
    FaceCulling --> GreedyMesh["Greedy Meshing<br/>(Merge contiguous coplanar quads)"]
    GreedyMesh --> GPU["OpenGL VAO / VBO Upload"]

2. Occlusion & Greedy Meshing

The first step is face culling: only generate quad faces between a solid block and an air/transparent block. This immediately eliminates ~90% of interior faces.

The second, most impactful optimization is Greedy Meshing (first popularized by Robert O’Leary):

  1. Slice the chunk into 2D cross-sections along the X, Y, and Z axes.
  2. Form a 2D mask of face types for each slice.
  3. Iterate through the mask, finding the maximum rectangular width and height of identical, coplanar adjacent faces.
  4. Merge them into a single, scaled quad polygon.
  5. Zero out the visited slots in the mask and repeat until the slice is cleared.

Results:

  • Geometry reduction: From ~80,000 vertices per chunk down to 1,200 – 3,500 vertices.
  • Draw call batching: Chunks are uploaded as single indexed VBOs (glDrawElements).
// Greedy meshing slice generation preview
for (int d = 0; d < 3; ++d) {
    int i = (d + 1) % 3;
    int j = (d + 2) % 3;
    
    // Iterate over slice depths and evaluate neighbor occlusion
    for (x[d] = -1; x[d] < CHUNK_SIZE; ++x[d]) {
        computeMask(mask, x, d, i, j);
        generateGreedyQuads(vertices, mask, x[d], d, i, j);
    }
}

3. Asynchronous Chunk Streaming

Terrain generation should never block the main render loop. We implemented a worker thread pool:

  • Worker Threads: Compute 3D Perlin/Simplex noise, determine biomes, and generate the raw voxel arrays.
  • Mesh Dispatcher: Runs greedy meshing asynchronously in memory.
  • Main Thread: Receives the raw vertex buffers and handles the quick OpenGL context upload (glBufferData).

4. Key Takeaways

  1. Memory Locality is King: Flattening 3D arrays (chunk[x + y*WIDTH + z*WIDTH*DEPTH]) dramatically reduced cache misses during meshing.
  2. Raycasting Block Picking: Using a fast voxel traversal algorithm (Amanatides & Woo) allows instant, sub-millisecond block placement and destruction.

The full source code and engine experiments are available in the voxel-engine2 repository.