The two problems are not the same problem

Emulation is impersonation. When a program written for one processor architecture runs on a different one, every instruction it issues must be intercepted, decoded and re-expressed as instructions the host CPU understands. QEMU doing this in full-system mode, or a classic console emulator replicating a MIPS or PowerPC core on an x86 machine, is paying a real price for every single operation the guest code performs. Even with dynamic recompilation — where hot paths are translated into native code blocks and cached — the overhead is structural. The emulator must maintain a model of the guest machine: its registers, its memory map, its interrupt lines. That model costs cycles to keep consistent, and the gap between guest and host widens with every architectural quirk that has no direct analogue.

A translation layer is a different problem entirely, and the confusion between the two costs people clarity every time they try to reason about performance. When a Windows application runs under a compatibility layer on Linux, the application's machine code is not being interpreted. The processor is running it directly. The x86-64 binary that was compiled for Windows is still x86-64, and a modern desktop Linux system is also x86-64. There is nothing to translate at the instruction level. What differs is everything above the instructions: the system calls, the runtime libraries, the graphics API, the filesystem conventions, the windowing model. The layer sits at those boundaries and rewrites the requests, not the code.

This is why the cost model is so different. Emulation pays per instruction. Translation pays per boundary crossing. If a tight inner loop runs for a million iterations without touching any operating-system service or API, the layer does no work during those iterations. The application is simply running. The cost appears at the edges — the moment the program asks the kernel for something, or hands a draw call to the graphics stack, or tries to open a file. Those crossings are real work, and they are not free, but they are far fewer than the instructions between them.

How the costs stack up

FROM THIS ENTRY
  1. Emulation costpaid per CPU instruction the guest program executes
  2. Translation costpaid per boundary crossing (syscall, API call, file operation)
  3. Instruction-level executionruns natively in a translation layer; no per-instruction overhead
  4. Shader translationtwo compilation passes: HLSL → SPIR-V (layer's cost), SPIR-V → GPU machine code (driver's cost, same as native)
  5. Threadingscheduler-personality mismatch is a cost the layer cannot fully absorb

What the layer actually rewrites

The syscall boundary is the sharpest example. A Windows application calls into ntdll.dll, which eventually issues an interrupt or syscall instruction that the kernel is expected to handle. On Linux, the kernel knows nothing of Windows system call numbers — the numeric identifiers for file operations, thread creation, timer queries and everything else are simply different, and many concepts do not map one-to-one. The layer catches these before they reach the kernel, translates the call number and argument layout, issues the equivalent Linux syscall, and then translates the result back into the shape the caller expects. The program never learns that anything unusual happened.

Graphics is where the translation workload is heaviest, and where the cost model becomes most visible in practice. A Windows game issues Direct3D calls. The layer — in practice, a library of shader translation and API mapping components — catches every draw call, every state change, every resource binding, and rewrites it as Vulkan commands. This is not interpretation of instructions; it is transformation of a structured protocol. Direct3D and Vulkan are both designed to drive the same class of hardware, so the mapping exists, but it is not trivial. State that the Direct3D runtime tracks implicitly must be made explicit for Vulkan. Descriptor sets must be assembled. Shaders written in HLSL must be compiled to SPIR-V. Each of these is a defined, finite task, but together they happen thousands of times per frame.

The shader compilation step is where the distinction from emulation becomes almost philosophical. An emulator running a game from a different architecture would have to translate the shader code as part of translating everything else. A translation layer receives the shader source or bytecode from a program that compiled it for real hardware — the same hardware the host machine has — and must re-express it for a different graphics API, then hand it to the real driver. The driver's compiler then produces the final machine code for the actual GPU sitting in the machine. There are two compilation steps, not one, and the second one is identical to what a native application would trigger. The first step is the overhead the translation layer owns; it is real, and it is why shader compilation stutter appears in translated games just as it does in native ones, sometimes more severely.

A bare processor die and socket pins, macro
FIG. 2The contract in physical form: pins are the ABI — position, order and width, with no names attached.PHOTO: PIXABAY / PEXELS

Where the cost actually lands

Because translation pays at boundary crossings rather than per instruction, its performance profile is unusual. CPU-bound workloads that spend most of their time in computation and rarely surface to the operating system can run at near-native speed — the layer is nearly invisible. What suffers is anything that crosses boundaries at high frequency: file I/O in tight loops, frequent small allocations routed through the Windows heap manager, heavy use of synchronisation primitives that have no exact Linux equivalent, or graphics workloads that hammer the API with small draw calls rather than batching them efficiently.

Threading assumptions are another source of hidden cost. Windows and Linux schedulers make different guarantees about how threads are woken, how priority is expressed, and how spinlocks behave. A program that was tuned for one set of assumptions runs under the other and may spend more time in the scheduler than its authors expected. The layer cannot fully correct for this; it can translate the API call, but it cannot change the scheduler's personality.

None of this is emulation overhead. A program that is slow under translation is slow for one of three reasons: the boundary crossings themselves are expensive in absolute terms; the mapping between two APIs is lossy and requires workarounds; or the host system's behaviour at those boundaries differs enough from the expected behaviour that the program's internal assumptions break down. Each of these is diagnosable and often improvable. The instruction-level cost of emulation, by contrast, is a floor set by the architecture gap — you pay it everywhere, regardless of what the program is doing.

The practical consequence is that a well-translated program can sit within a few percent of native performance on workloads that matter, while an emulated one is always fighting the overhead of representation. A compatibility layer succeeds when the program it runs has no idea it is being translated — when the boundary crossings are fast enough and faithful enough that the program's model of the world matches the responses it receives. That match, not the absence of rewriting, is what the engineering aims for.

The two regimes

FROM THIS ENTRY
CPU-bound inner loops with few API callsHigh-frequency boundary crossings (small draw calls, tight I/O loops, frequent allocations)
near-native under translationthe real performance surface of the layer
Ribbon and power cables crossing an open case
FIG. 3Everything crosses somewhere. The layer's work happens at seams like these, not in the code between them.

Cedega is an independent publication about how compatibility layers work. It is not a software vendor, distributor, or support service, and does not distribute, licence or recommend any software product.

FILED UNDER: THE BOUNDARYLONG READ