The machine changes as it runs
Processors do not run at their rated clock speed by default — they run above it, for as long as the thermal headroom permits. Intel calls this Turbo Boost, AMD calls it Precision Boost, and the details differ, but the principle is identical: extract extra performance from transient thermal budget, then step back when the silicon gets hot. That step-back is not a crash or a failure. It is designed behaviour, and it happens quietly, usually somewhere between three and twelve minutes into a sustained load.
The consequence for benchmarking is awkward. A two-minute run catches the machine at its fastest. A twenty-minute run catches it at its thermally settled speed. Neither is wrong, but they are measuring different things, and if you do not know which state the machine was in during your measurement window, the number is genuinely ambiguous.
What changes and when
FROM THIS ENTRY| Boost state | processor runs above rated clock while thermal headroom allows; duration typically three to twelve minutes of sustained load |
|---|---|
| Thermal steady state | the lower, stable clock speed reached once the silicon reaches equilibrium |
| Boost step-back | the designed, automatic reduction in clock speed as temperature rises; silent, not an error |
| Translation overhead | extra CPU work the compatibility layer adds, which itself consumes thermal budget |
Cooling compound ages, fans accumulate dust, chassis airflow paths change when a cable is rerouted or a vent is partly blocked. All of these shift the thermal ceiling over months. A benchmark result from the same machine six months apart is not necessarily a regression in software — it may be a confession from the hardware.
The compatibility layer adds its own heat. Translation overhead is CPU work that a native run never does, and that work lands on the same cores, raising their temperature and tightening the boost budget available for everything else. The effect is not enormous, but it is real, and it means that where the milliseconds go in a translated workload depends partly on how long the workload has been running. A spike measured at minute one is a different datum from the same spike at minute fifteen.

Thermal drift also explains a recurring frustration with reproducibility: two runs of the same benchmark on the same hardware, same software, same config, separated by ten minutes, produce meaningfully different frame times. The machine is the variable. Controlling for it means warming the machine to steady state before the measurement window opens — not before the benchmark launches, but before the numbers are recorded. The warm-up is not ceremony; it is calibration.
A number taken from a cold run is a best case. Whether that is useful depends entirely on what you are trying to learn.
Why runs disagree
FROM THIS ENTRY| Cold run vs. warm run | same machine, same software, different thermal state; results are not comparable |
|---|---|
| Hardware confounds | aged thermal compound, dust, airflow obstructions all lower the thermal ceiling over time |
| Measurement window | recording begins after steady state, not after benchmark launch; warm-up is calibration, not ceremony |

