The Fork in the Road
A frame is produced when the CPU finishes constructing the draw calls and command buffers, and the GPU finishes executing them. Whichever one finishes last determines when the frame ships. If the CPU is the slow half, the GPU sits idle waiting for work. If the GPU is the slow half, the CPU races ahead, submits work, and then stalls waiting for the queue to drain. Either way you lose frames, but the fix for one is useless against the other.
The test is simple in principle: reduce GPU load by lowering resolution or turning off expensive effects and watch what happens to frame time. If frame time drops proportionally, the GPU was the bottleneck. If frame time barely moves, the CPU is. A compatibility layer adds wrinkles to both sides of that test.
What to measure and when
FROM THIS ENTRY| GPU load test | drop resolution or effects; if frame time falls proportionally, GPU is the bottleneck |
|---|---|
| CPU load test | if frame time is unchanged after reducing GPU load, CPU is the bottleneck |
| warm cache | shader-compilation spikes disappear once the cache is populated; always measure after warmup |
| frame-time graph vs. average FPS | spikes are invisible in averages; a per-frame graph is required |
On the CPU side, every translated API call costs something. A Direct3D call that would be a near-zero overhead on its native platform becomes a state-diffing, format-converting, command-rewriting exercise before it ever touches the command buffer. That overhead sits squarely in the CPU frame budget. A game that was borderline CPU-bound on Windows can tip over the edge on a translation layer even before the GPU has been asked to do anything different.
On the GPU side, shader compilation stutter is the dominant layer-specific cost. The first time a shader variant is needed, it must be compiled to the target ISA, and that compilation can stall the pipeline for multiple frames — spikes that show up clearly in a frame-time graph but vanish in an average frame rate number. Once the cache is warm those spikes go away, which is exactly why you measure on a warm run.

The second GPU-side issue is that translated shaders are not always as efficient as hand-written ones for the target architecture. A HLSL shader translated to SPIR-V and then compiled by a Mesa or vendor driver may generate more instructions or hold registers differently than a native shader would. The resulting GPU time per draw call is higher, and the headroom before the GPU becomes the bottleneck is narrower.
Getting the answer right matters because the solutions diverge completely. A CPU-bound workload needs fewer, cheaper draw calls — batching, culling, reduced translation overhead. A GPU-bound workload needs lower resolution, reduced shader complexity, or a better shader cache hit rate. Applying the GPU fix to a CPU-bound problem buys nothing. The diagnosis comes first.

