Roadmap conversations tend to run on qubit counts. Ours included: Libra is designed for 256 error-corrected logical qubits built from more than 10,000 physical qubits at a 10-6 logical error rate, arriving on Amazon Braket in 2028, with a Next Gen Gigaquop-class system planned for 2028/2029.
Those numbers are the entry ticket, not the machine. Between a peer-reviewed architecture and a computer that returns an answer sits a layer that rarely makes the slide: compilers, schedulers, calibration specs, and simulators faithful enough that you can trust what they tell you before the hardware exists. Three talks at the 2026 QuEra Quantum Alliance meeting came at that layer from different angles and converged on the same conclusion. Qubit count is not the thing standing between us and a useful machine.
This post draws on those three talks, given at the QuEra Quantum Alliance 2026 Annual Meeting (QQA 2026) by Yuval Baum of Q-CTRL, who spoke on heterogeneous fault-tolerant architectures, and by Stefan Krastanov and Phillip Weinberg, who covered QuEra's compiler and simulation stack.
Most of Your Qubits Are Doing Nothing
Yuval Baum of Q-CTRL opened with a claim worth sitting with: software design is trailing hardware progress. We design codes and asymptotics from the top down, he argued, then hope the hardware team makes it run. What is missing is the bottom-up work of building software that can actually execute these workloads.
His example was the Gidney implementation of Shor's algorithm. It requires roughly 1014 cycles across about 1,400 qubits, and more than 99.5% of those qubit-cycles are idle. The qubits are not computing. They are waiting for other qubits to finish, while the machine burns microsecond cycles and distance-25 syndrome extraction to hold them in place, pushing a staggering volume of data at the decoder for the privilege. If the dominant cost of your algorithm is waiting, a faster gate is not the fix.
Baum also cautioned against one-dimensional resource estimates. Pushing Shor's physical-qubit requirement down toward ten thousand looks like progress until you plot the second axis and find a runtime of 43,000 days, roughly 120 years. Trade-offs are fine. Hiding them is not.
Q-CTRL's proposed answer borrows from classical computing: rather than one uniform sea of identical qubits, a heterogeneous architecture of small, fast QPU modules (each with its own magic state factory, each O(1) in size no matter how large the algorithm grows) connected to a hierarchy of memory, with the weight of scaling carried by memory rather than by compute. Idle logical qubits get moved to slower, cheaper storage instead of being maintained at full speed in the processor.
Making that concrete requires a compiler that schedules an algorithm onto a user-defined architecture all the way down to machine instructions, modeling interconnect rates, error rates, and the timing mismatch between modules running at different clock speeds. That is Q-CHESS, their Quantum Compiler for Heterogeneous Execution, Scheduling, and Synthesis. The result Baum showed was orders-of-magnitude reduction in algorithmic logical error for the same algorithm, unchanged, simply scheduled onto a better architecture. The error budget moves as you iterate: idling stops dominating, state transfer across interconnects takes over, and with further tuning you arrive at the regime where you are limited by the T gates you genuinely need.
Which is why he closed on synthesis. A real algorithm needs billions, possibly trillions, of unitaries synthesized into Clifford+T. At a minute each, classical compilation takes days before the quantum computer starts. Q-CTRL plans to release a synthesizer targeting near-optimal T counts in well under a millisecond.
Layers Exist so Teams Can Move Independently
Stefan Krastanov, who leads the QEC layer of our compiler stack, addressed a question we get from collaborators who want access further down: why all the abstraction?
The usual assumption is that layering protects intellectual property. His answer was that the separation is primarily engineering. He wants a stable interface below him just as much as an external user does. The QEC experts and the atomic-physics and hardware engineers are working on genuinely different problems, and each needs to iterate fast without waiting on the other.
The pipeline lowers application code into abstract QEC gadgets, chosen from a set of architecture options; then into primitive QEC gadgets, fine-tuned for how a given code is implemented on the trapping hardware and aware of placement in the atom grid; then into what he calls the logical placement dialect, where you can subdivide zones, assign responsibilities, and size buffers without working in absolute coordinates. Colleagues describe that layer as something like virtual memory for the machine. Below it, lowering targets the embedded controllers: real pulses, real atom moves, real calibration.
A concrete example of what lowering means here: a transversal CNOT between two code blocks stays a single simple gadget all the way down, because the microarchitecture makes it cheap. An in-block T gate on one logical qubit expands into allocating a resource block from an asynchronously replenished buffer, performing an automorphism (a permutation of logical qubits, which at the atomic level is just movement), then a teleported T that itself lowers further into gates, measurements, and conditional operations.
All of this sits on Kirin, our open-source compiler infrastructure (Kernel Intermediate Representation Infrastructure), which is general-purpose and not specific to quantum computing. Each lowering step is parameterized by a machine-readable description of the architecture, so a new architecture does not mean rewriting the compiler passes. That parameterization is also what makes resource estimation possible before Libra is built, through a mix of coarse operation counting and estimates informed by physical simulation of the atom moves themselves.
A Digital Twin, Not a Circuit Simulator
Phillip Weinberg picked up where the compiler stops. Simulating a fault-tolerant machine, he argued, is no longer a matter of running a circuit and reading the output.
Three things change at this scale. Atom movement, long treated as effectively free, becomes something QEC has to account for. Some moves depend on measurement outcomes, so they cannot all be scheduled ahead of time and the processor has to decide in real time. And the machine becomes asynchronous: atom reloading, decoding, and magic state production each run on their own timelines. Understanding the quantum circuit is not enough. You need to understand what the CPU, the FPGAs, the optics, the atoms, and the decoder are doing to each other.
PPVM, the Pauli Propagation Virtual Machine, is our first public piece of that. It combines stabilizer simulation via generalized tableaus for near-Clifford circuits at scale, a model of the interacting quantum and classical runtime, and noise sampling that covers the failure modes specific to atoms: single-atom loss during gate operations, correlated loss when qubits are entangled, Rydberg gates executed after an atom has already been lost, and loss on reset.
The most interesting design decision he described was about constraint. Rather than allowing arbitrary atom moves, the hardware is specified as a defined set of pre-calibrated lanes. Compose a dense enough graph of those calibrated moves and you still get all-to-all connectivity, whether you are moving individual atoms or entire code blocks, without asking the hardware team to keep an unbounded space of arbitrary trajectories in calibration. His phrasing was that it keeps his hardware friends sane. The neutral-atom QEC testbed we have installed in Japan already works this way, with lane assignment done by hand today and automation as the goal.
He also demoed Bloqade Studio QEC, a drag-and-drop interface for assembling small QEC experiments (think LabVIEW for error correction) with PPVM underneath doing the simulation and our decoder integration building the detector error model as the program runs. Build a memory experiment, step through it, sweep the noise, and watch the logical error rate after decoding. As he put it, the logical error rate is the most accurate resource estimate you can possibly get.
Co-Design Runs in Both Directions
Put the three talks together and a single loop appears. The compiler asks the hardware team for a set of moves. The hardware team answers with what can actually be calibrated and held stable inside a finite field of view. The digital twin answers whether that floor plan, and the schedule the compiler produced for it, hits the target for a real application. Then the loop runs again.
Baum's heterogeneous architectures and our fixed lanes are the same instinct approached from opposite ends: constrain the machine enough that software can reason about it, and build software good enough to evaluate an architecture before anyone builds it. Both are ways of answering a question that used to be unanswerable, namely whether a design choice made today will still look correct at 10,000 physical qubits.
None of this is glamorous. It is also the difference between a fault-tolerant computer and a fault-tolerant computer that anyone can use. Kirin, PPVM, and Tsim are on GitHub today, and we co-organized the first Workshop on Compilation, Emulation and Verification of Neutral Atom Computing (CEVNAC 2026) alongside IEEE Quantum Week in Toronto, because this is not a stack any one company should be building alone.




.webp)
.png)

