The distance between a key press and a photon, and the queue in between

A key press is not an event — it is the start of a pipeline. The switch closes, the USB controller polls it (or the PS/2 interrupt fires), the OS notifies the application, the application reads it, the logic updates, a frame is rendered, the frame is queued for the display, and the display draws it. Each of those stages has a budget. Miss any of them and a human feels it, even if they cannot name what they feel.

The pipeline, broken out

FROM THIS ENTRY
  1. USB poll intervalthe hardware floor before software sees anything
  2. OS input queueserviced in the window-message thread; a stalled frame delays it
  3. Layer re-dispatchthe synthetic event added by the compatibility layer; small but real
  4. Frame queue depthwhere translated graphics paths commonly introduce 16–33 ms of hidden lag
  5. Pre-rendered frame limitthe knob that trades CPU-GPU parallelism for lower latency

Every handoff costs time

The USB polling interval is already a floor: full-speed HID devices poll at 1 ms by default, but many peripherals ship configured at 8 ms or even 125 ms intervals, and that delay is invisible in any frame-time graph. Above that sits the OS input queue, which on most desktop systems is serviced in the same thread loop that processes window messages — meaning a stalled frame can delay input processing even when the CPU is otherwise idle.

Inside a compatibility layer the handoff multiplies. A translated application does not receive a raw kernel event; it receives a synthetic event fired by the layer after the layer has received, interpreted, and re-dispatched the host OS notification. That extra hop is usually measured in tens of microseconds, not milliseconds, so it is rarely the dominant term. What the layer can introduce that genuinely hurts is frame queue depth.

When the graphics API translation path — one API into another — builds a submission queue that runs one or two frames deep for throughput reasons, the display shows frame N while the application is already computing frame N+2. The image on screen is technically correct. The relationship between it and the last input event is not. A player moves the mouse; the cursor catches up one or two frames later. At 60 Hz that is 16–33 ms of hidden lag sitting inside the pipeline, attributable not to hardware, not to the network, but to queue depth.

Pre-rendered frame limits, where they are exposed, constrain this. Turning the limit to one frame recovers that latency at the cost of CPU-GPU parallelism: the CPU idles briefly waiting for the GPU to finish before it can begin the next submission. Where the milliseconds go is the broader context; input latency is simply one of the lines in that ledger, and often the line that hides in averages while the one per cent low catches everything else.

The correct way to measure it is end-to-end: a hardware latch on the input event, a photodiode on the display, a high-speed camera between them. Everything else is a model.

A monitor showing a flat test pattern in a dim room
FIG. 2A flat test pattern holds every variable still except the one being measured.
A mechanical keyboard pushed aside, one cable across it
FIG. 3Input starts here and ends at a photon; every queue in between is part of the latency budget.PHOTO: FOX ^.ᆽ.^= ∫ / PEXELS
FILED UNDER: FRAME TIMESHORT ENTRY