Writing a voxel engine from the ground up teaches you computer graphics in the most raw, visceral way possible. When you place a block in a 3D world, you quickly realize that naive cube rendering (12 triangles per block) will choke any modern GPU the moment you render a modest landscape of 16x16 chunks.
In this retrospective, I break down the core architectural decisions behind Voxel Engine 2 and how we achieved smooth 60+ FPS chunk streaming with dynamic terrain modification.
1. The Naive Problem: Geometry Explosion
Consider a standard 16x16x256 chunk. That’s 65,536 potential voxels.
- If rendered naively:
65,536 * 6 faces * 2 triangles = 786,432 triangles per chunk. - With a standard render distance of 16 chunks around the player (
33x33 = 1089 chunks), you would be asking the GPU to push 850 million triangles per frame.
Clearly, optimization must start before a single byte reaches the Vertex Buffer Object (VBO).
graph LR
Noise["3D Simplex Noise"] --> VoxelGrid["3D Array of Voxel Data"]
VoxelGrid --> FaceCulling["Occlusion Culling<br/>(Drop hidden interior faces)"]
FaceCulling --> GreedyMesh["Greedy Meshing<br/>(Merge contiguous coplanar quads)"]
GreedyMesh --> GPU["OpenGL VAO / VBO Upload"]
2. Occlusion & Greedy Meshing
The first step is face culling: only generate quad faces between a solid block and an air/transparent block. This immediately eliminates ~90% of interior faces.
The second, most impactful optimization is Greedy Meshing (first popularized by Robert O’Leary):
- Slice the chunk into 2D cross-sections along the X, Y, and Z axes.
- Form a 2D mask of face types for each slice.
- Iterate through the mask, finding the maximum rectangular width and height of identical, coplanar adjacent faces.
- Merge them into a single, scaled quad polygon.
- Zero out the visited slots in the mask and repeat until the slice is cleared.
Results:
- Geometry reduction: From ~80,000 vertices per chunk down to 1,200 – 3,500 vertices.
- Draw call batching: Chunks are uploaded as single indexed VBOs (
glDrawElements).
// Greedy meshing slice generation preview
for (int d = 0; d < 3; ++d) {
int i = (d + 1) % 3;
int j = (d + 2) % 3;
// Iterate over slice depths and evaluate neighbor occlusion
for (x[d] = -1; x[d] < CHUNK_SIZE; ++x[d]) {
computeMask(mask, x, d, i, j);
generateGreedyQuads(vertices, mask, x[d], d, i, j);
}
}
3. Asynchronous Chunk Streaming
Terrain generation should never block the main render loop. We implemented a worker thread pool:
- Worker Threads: Compute 3D Perlin/Simplex noise, determine biomes, and generate the raw voxel arrays.
- Mesh Dispatcher: Runs greedy meshing asynchronously in memory.
- Main Thread: Receives the raw vertex buffers and handles the quick OpenGL context upload (
glBufferData).
4. Key Takeaways
- Memory Locality is King: Flattening 3D arrays (
chunk[x + y*WIDTH + z*WIDTH*DEPTH]) dramatically reduced cache misses during meshing. - Raycasting Block Picking: Using a fast voxel traversal algorithm (Amanatides & Woo) allows instant, sub-millisecond block placement and destruction.
The full source code and engine experiments are available in the voxel-engine2 repository.