Lockstep: Bit-Exact, Verifiable Transformer Inference at Parity Across Devices, for Dense and Mixture-of-Experts Models
Perica Glavas · Maksym Yakovenko · Logan Allen
Production transformer serving has no single correctness relation at the bit level: equivalent kernels, reduction orders, batch compositions, compilers and devices all legitimately return different floating-point strings, so an auditor must either reproduce the provider's environment or accept a tolerance. Lockstep makes the served inference function itself canonical. A versioned contract fixes every implementation-selected floating-point choice, and correctness becomes bit equality to a public function that an optimized serving engine and a portable CPU reference evaluate alike. Implemented as a vLLM plugin, five dense model families serve at or above stock throughput and Qwen3-30B-A3B at 1.32-1.37x, with identical output digests across Hopper, Ada and Blackwell GPUs and the reference reproduced bit for bit on x86 and ARM. The contract core is mechanized in Lean 4.