The first pass lies
Every compatibility layer does work the first time it sees something. Shaders compile. The registry file gets parsed and cached in memory. Disk paths are resolved and their case-corrected versions remembered. The loader maps libraries it has never mapped before. All of that cost lands on the first few seconds of a run — and if your benchmark starts timing from launch, you are measuring setup, not steady state.
This is not a flaw in the layer. It is how every caching system behaves. The question is whether your reported number reflects the workload you care about or the workload that only happens once.
What to pull out
FROM THIS ENTRY| Cold run | the first execution pass; caches empty, one-time setup costs included; real but unrepresentative as a steady-state figure |
|---|---|
| Warm run | subsequent passes after caches are populated; reflects persistent per-call overhead rather than setup cost |
| The rule | run once to warm, discard, then measure; state which pass was reported |
A cold run tells you something real: it is what a first-time launch feels like, and shader compilation stutter is exactly this effect expressed as a hitch inside a scene rather than a slow start before one. But a cold run reported as a performance figure, without that label, is misleading. The layer will never be that slow again in the same session, and usually not in the next session either, once caches are warm.
The warm run is what you get after the layer has done its one-time work. Shaders are resident. The path resolver has its answers. Memory is mapped. From here, the translation overhead you are seeing is the overhead that persists — the per-call cost of rewriting an API call into another form, the state tracking that runs on every draw, the scheduler seeing threads it was not tuned for. That is the number worth reporting, because it is the number that does not go away.

The practical rule is simple: run once to warm the cache, discard that result, then run the measurement. State which pass you are reporting. A repeatable run with a warm cache and a stated configuration is the minimum condition for a number that means anything. Without it, you are timing the layer's first impression of your workload, which is the worst impression it will ever make.

