The three things a number needs

A benchmark from a compatibility layer is not like a benchmark from native hardware. The layer is a moving piece: shader compilation happens, caches warm, translation paths diverge depending on the exact API sequence a title chooses. Run the same scene twice in a row and the second number will usually be faster — not because the hardware improved, but because the first run did work the second run did not have to. Report the first number and you are measuring the cost of being new. Report the second without saying so and you are hiding it.

Repeatability is the foundation. That means the same scene, the same camera path, the same frame window, and — critically — the same starting state. If a built-in benchmark loop is available, use it every time; it removes the variable of human input and gives identical geometry to the GPU on every pass. If there is no loop, a fixed replay or a scripted input sequence is the next best option. Anything that varies between runs is noise that cannot be distinguished from signal.

The warm cache question is where most published numbers go wrong. Shader compilation stutter produces a spike on first encounter that disappears once the compiled result is stored. A cold run captures that spike; a warm run does not. Both are real measurements of real moments — they are just measurements of different things. The honest approach is to state which one you are reporting, and why. A number from a cold run tells you about first-launch experience. A number from the warm run tells you about sustained gameplay after the layer has settled. Neither is wrong; unlabelled, either is useless.

The three requirements

FROM THIS ENTRY
Repeatable runfixed scene, fixed camera path, identical starting state every time
Warm cacheshader compilation and translation caches settled before the reported pass
Stated configurationOS, kernel, driver version, layer version, API target, active options

Configuration is the other half of the sentence

A frame time or a frame rate is only half a statement. The other half is everything about the system that produced it: the host operating system and its kernel version, the graphics driver and its exact release, the compatibility layer version, the graphics API the layer is translating into, and whether any layer-level options — asynchronous shader compilation, DXVK state caches, threaded rendering paths — were enabled or disabled. Omit any of those and the number cannot be reproduced by anyone else, which means it cannot be verified, which means it is an anecdote dressed as data.

Driver version matters more than most people assume. A single driver release can shift frame time on a specific title by a meaningful margin in either direction, because the driver itself contains optimisation paths that are title-aware. A result from three driver versions ago may be perfectly accurate for that driver and completely misleading for the current one. Version-pin everything.

Averages, meanwhile, conceal the frame that cost four times what the others did. A compatibility layer that translates well on average but hitches badly on a specific API call will produce a fine average and a broken experience. Report the one-percent low alongside the mean, because smoothness is not average performance — it is the worst performance that happens often enough to feel like a pattern. A mean without a distribution is not dishonest exactly, but it is incomplete in a way that flatters.

A thermal probe taped to a heatsink
FIG. 2A probe taped to the heatsink: ten minutes in, the same machine is a different machine.

What an honest result looks like

State the configuration in full. Run at least once cold to capture first-encounter cost, then run warm and label the warm result separately. Use a fixed scene or replay. Record frame time, not only frame rate, and include the one-percent low. If you changed any layer option from its default, name the option and say what you changed it to.

That is it. It is not a large burden. It is the minimum information a second person needs to disagree with you — and a number nobody can disagree with is not a measurement; it is a claim.

Key distinctions

FROM THIS ENTRY
Cold run vs warm rundifferent measurements of different moments; both valid, both must be labelled
Mean vs one-percent lowaverage hides the hitches; the low captures what smoothness actually feels like
Frame time vs frame raterate is derived; time shows the outliers a rate average erases
A bench power supply and meter beside an open machine
FIG. 3Supply and meter beside the machine: absolute numbers first, percentages only once the baseline is stated.

Shimwork is an independent publication about how compatibility layers work. It is not a software vendor, distributor, or support service.

FILED UNDER: BENCHMEDIUM ENTRY