The Fork in the Road

A frame is produced when the CPU finishes constructing the draw calls and command buffers, and the GPU finishes executing them. Whichever one finishes last determines when the frame ships. If the CPU is the slow half, the GPU sits idle waiting for work. If the GPU is the slow half, the CPU races ahead, submits work, and then stalls waiting for the queue to drain. Either way you lose frames, but the fix for one is useless against the other.

The test is simple in principle: reduce GPU load by lowering resolution or turning off expensive effects and watch what happens to frame time. If frame time drops proportionally, the GPU was the bottleneck. If frame time barely moves, the CPU is. A compatibility layer adds wrinkles to both sides of that test.

What to measure and when

FROM THIS ENTRY
GPU load testdrop resolution or effects; if frame time falls proportionally, GPU is the bottleneck
CPU load testif frame time is unchanged after reducing GPU load, CPU is the bottleneck
warm cacheshader-compilation spikes disappear once the cache is populated; always measure after warmup
frame-time graph vs. average FPSspikes are invisible in averages; a per-frame graph is required

On the CPU side, every translated API call costs something. A Direct3D call that would be a near-zero overhead on its native platform becomes a state-diffing, format-converting, command-rewriting exercise before it ever touches the command buffer. That overhead sits squarely in the CPU frame budget. A game that was borderline CPU-bound on Windows can tip over the edge on a translation layer even before the GPU has been asked to do anything different.

On the GPU side, shader compilation stutter is the dominant layer-specific cost. The first time a shader variant is needed, it must be compiled to the target ISA, and that compilation can stall the pipeline for multiple frames — spikes that show up clearly in a frame-time graph but vanish in an average frame rate number. Once the cache is warm those spikes go away, which is exactly why you measure on a warm run.

A mechanical keyboard pushed aside, one cable across it
FIG. 2Input starts here and ends at a photon; every queue in between is part of the latency budget.PHOTO: FOX ^.ᆽ.^= ∫ / PEXELS

The second GPU-side issue is that translated shaders are not always as efficient as hand-written ones for the target architecture. A HLSL shader translated to SPIR-V and then compiled by a Mesa or vendor driver may generate more instructions or hold registers differently than a native shader would. The resulting GPU time per draw call is higher, and the headroom before the GPU becomes the bottleneck is narrower.

Getting the answer right matters because the solutions diverge completely. A CPU-bound workload needs fewer, cheaper draw calls — batching, culling, reduced translation overhead. A GPU-bound workload needs lower resolution, reduced shader complexity, or a better shader cache hit rate. Applying the GPU fix to a CPU-bound problem buys nothing. The diagnosis comes first.

A graphics card backplate and connectors in macro
FIG. 3The card the shaders are finally compiled for — whichever hardware the program believed in.
FILED UNDER: FRAME TIMESHORT ENTRY