The three things a number needs
A benchmark from a compatibility layer is not like a benchmark from native hardware. The layer is a moving piece: shader compilation happens, caches warm, translation paths diverge depending on the exact API sequence a title chooses. Run the same scene twice in a row and the second number will usually be faster — not because the hardware improved, but because the first run did work the second run did not have to. Report the first number and you are measuring the cost of being new. Report the second without saying so and you are hiding it.
Repeatability is the foundation. That means the same scene, the same camera path, the same frame window, and — critically — the same starting state. If a built-in benchmark loop is available, use it every time; it removes the variable of human input and gives identical geometry to the GPU on every pass. If there is no loop, a fixed replay or a scripted input sequence is the next best option. Anything that varies between runs is noise that cannot be distinguished from signal.
The warm cache question is where most published numbers go wrong. Shader compilation stutter produces a spike on first encounter that disappears once the compiled result is stored. A cold run captures that spike; a warm run does not. Both are real measurements of real moments — they are just measurements of different things. The honest approach is to state which one you are reporting, and why. A number from a cold run tells you about first-launch experience. A number from the warm run tells you about sustained gameplay after the layer has settled. Neither is wrong; unlabelled, either is useless.
The three requirements
FROM THIS ENTRY| Repeatable run | fixed scene, fixed camera path, identical starting state every time |
|---|---|
| Warm cache | shader compilation and translation caches settled before the reported pass |
| Stated configuration | OS, kernel, driver version, layer version, API target, active options |
Configuration is the other half of the sentence
A frame time or a frame rate is only half a statement. The other half is everything about the system that produced it: the host operating system and its kernel version, the graphics driver and its exact release, the compatibility layer version, the graphics API the layer is translating into, and whether any layer-level options — asynchronous shader compilation, DXVK state caches, threaded rendering paths — were enabled or disabled. Omit any of those and the number cannot be reproduced by anyone else, which means it cannot be verified, which means it is an anecdote dressed as data.
Driver version matters more than most people assume. A single driver release can shift frame time on a specific title by a meaningful margin in either direction, because the driver itself contains optimisation paths that are title-aware. A result from three driver versions ago may be perfectly accurate for that driver and completely misleading for the current one. Version-pin everything.
Averages, meanwhile, conceal the frame that cost four times what the others did. A compatibility layer that translates well on average but hitches badly on a specific API call will produce a fine average and a broken experience. Report the one-percent low alongside the mean, because smoothness is not average performance — it is the worst performance that happens often enough to feel like a pattern. A mean without a distribution is not dishonest exactly, but it is incomplete in a way that flatters.

What an honest result looks like
State the configuration in full. Run at least once cold to capture first-encounter cost, then run warm and label the warm result separately. Use a fixed scene or replay. Record frame time, not only frame rate, and include the one-percent low. If you changed any layer option from its default, name the option and say what you changed it to.
That is it. It is not a large burden. It is the minimum information a second person needs to disagree with you — and a number nobody can disagree with is not a measurement; it is a claim.
Key distinctions
FROM THIS ENTRY| Cold run vs warm run | different measurements of different moments; both valid, both must be labelled |
|---|---|
| Mean vs one-percent low | average hides the hitches; the low captures what smoothness actually feels like |
| Frame time vs frame rate | rate is derived; time shows the outliers a rate average erases |

Shimwork is an independent publication about how compatibility layers work. It is not a software vendor, distributor, or support service.
