Why the first pass always costs
A shader is a small program that runs on the GPU. It describes how geometry is transformed, how surfaces are lit, how pixels are coloured. The source — or an intermediate representation such as SPIR-V — travels with the game, but that intermediate form is not what the GPU executes. The GPU needs machine code specific to its own architecture, and that translation step, pipeline compilation, happens on the host CPU.
In an ideal world it would happen before the scene starts. In practice, most engines defer it: compile when the shader is first needed, on whatever thread is available, and hope the stall is short enough to go unnoticed. It is rarely short enough. The first time a particular material, shadow technique, or post-process combination is encountered, the driver stops and compiles. That pause is shader compilation stutter — visible as a freeze of anywhere from a few milliseconds to several seconds, typically at the moment the player enters a new area, triggers an explosion, or encounters an enemy type for the first time.
Under a compatibility layer the problem is compounded. The game submits shaders in a Direct3D or OpenGL format; the layer must translate them before passing anything to the host graphics API. On Linux with Vulkan underneath, that means HLSL bytecode becomes SPIR-V, which then becomes the GPU vendor's own binary. Every translation adds latency, and the compile happens in the same frame the shader is first needed. The hitch is larger, because there is more work.
How the stall unfolds
FROM THIS ENTRY- Shader arrivesthe program submits a shader it has not used this session
- Translationcompatibility layer converts it to the host API's format (e.g. HLSL → SPIR-V)
- CompilationCPU compiles GPU-vendor machine code from the translated source
- GPU starvationGPU sits idle waiting for the compiled pipeline
- Frame spikeframe time jumps; may be visible as a freeze of milliseconds to seconds
- Cache writeresult saved to disk; next session this path is skipped
What the cache actually does — and does not do
The canonical fix is a pipeline cache. After the first compilation, the result — GPU-specific binary code — is written to disk. Next launch, if the binary is present and valid, the driver loads it directly rather than recompiling. The stutter disappears.
The qualifications matter. The cache entry is valid only when the exact combination of shader source, driver version, GPU, and driver-side flags matches what produced it. A driver update can invalidate every entry. So can a game update that ships new shaders. A user who moves the game to a different machine loses the entire cache. On the first run after any of those events, every stutter returns.
The shader cache has its own invalidation logic, and the rules are stricter than most users realise.
This is why pre-compilation pipelines exist: rather than waiting for a scene to demand a shader, the engine — or the compatibility layer, or a background process — tries to compile everything before gameplay begins. The loading screen that appears to hang is sometimes doing exactly this. The trade-off is that pre-compilation requires knowing, in advance, which shaders in which states will be needed. Engines that generate shader variants dynamically — blending material parameters at runtime — cannot fully pre-compile. There will always be some variant that only the player's specific choices produce.

Where time actually goes during the hitch
Shader compilation is CPU work, not GPU work. During a stall the GPU is typically starved — it has finished its queue and is waiting for the CPU to hand it something to execute. Frame time spikes. The GPU utilisation graph drops. To a performance counter, this looks like the GPU going idle, which is misleading: the machine is not underloaded, it is blocked.
Because the stall is CPU-bound, thread count matters more than GPU tier. A compilation that takes two hundred milliseconds on two cores may take thirty on eight, if the driver parallelises the work. Compatibility layers that dispatch compilation work to background threads can reduce foreground hitches dramatically, though they introduce a different problem: if the background compile is not finished when the draw call arrives, the layer must either stall anyway or substitute a placeholder pipeline and render incorrectly until the real one is ready.
The stutter is not a mystery. It is the inevitable cost of compiling GPU programs just-in-time, delayed only by whether that compilation happened in this session or a previous one. Where the milliseconds go ultimately leads here: a frame that should take sixteen milliseconds takes two hundred because a translation nobody scheduled arrived at the worst possible moment.
What breaks the cache
FROM THIS ENTRY| Driver update | vendor binary is no longer valid for new driver |
|---|---|
| Game update | new or changed shaders shipped with a patch |
| Hardware change | cached binary does not match new GPU architecture |
| First run on a new machine | no cache exists yet |

Shimwork is an independent publication. We are not a software vendor, distributor, or support service.
