होम/ब्लॉग/Google Taught Willow to Calibrate Itself While It Runs Error Correction
HardwareError CorrectionIndustry

Google Taught Willow to Calibrate Itself While It Runs Error Correction

Google Quantum AI and DeepMind published a Nature paper showing a reinforcement learning agent recalibrates the Willow processor using the same error-detection data error correction already produces, setting a new logical error rate record for the surface code.

FreeQuantumComputing
·· 8 min read

Google Quantum AI and Google DeepMind published a paper in Nature on July 10, 2026 showing that a reinforcement learning agent keeps a quantum processor calibrated using data the processor was already producing, rather than pausing to run a separate tune-up routine, according to The Quantum Insider and the paper itself. The work ran on Google's Willow superconducting processor and was led by Volodymyr Sivak and Alexis Morvan, with 299 authors credited in total.

The problem this solves

Every quantum processor drifts. Control parameters, the microwave pulse amplitudes, frequencies, and coupling strengths that drive each qubit, slowly shift due to temperature changes, material aging, and electronic noise. The standard fix is periodic recalibration: pause the machine, run a dedicated tune-up sequence, then resume. That works, but it treats calibration and computation as separate activities competing for the same hardware time, and it only catches drift at the moments you choose to check for it.

Quantum error correction already generates a stream of error-detection events as a normal part of running any error-corrected circuit. Google's insight was that this data is not only useful for correcting the logical qubit's state. It also tells you, continuously, how well the underlying hardware is behaving, which is exactly the signal a calibration system needs.

What the reinforcement learning agent does

The system trains a reinforcement learning agent to read the error-detection events produced during normal error-correction cycles and use them to continuously adjust more than 1,000 control parameters, without stopping the computation to do it. Google tested the approach on distance-5 and distance-7 surface codes and a distance-5 color code on the 105-qubit Willow chip, and ran simulations extending to a distance-15 surface code involving roughly 40,000 parameters, to check whether the approach holds up at a scale beyond what current hardware supports.

The results at the tested scale: a 20% reduction in logical error rate beyond what conventional calibration achieves, even after exhaustive manual tuning. Under artificially injected hardware drift, meant to simulate the kind of degradation a real system experiences over hours of operation, the reinforcement learning approach cut the logical error rate by 24% and made performance 2.4 times more stable than static calibration. Adding decoder-parameter adaptation on top pushed that to a 31% error reduction and 3.5 times more stability. On the distance-7 surface code specifically, the system reached a logical error rate of 7.72 × 10⁻⁴ per cycle, which the authors report as a new record for that code family.

Why this matters more than the specific numbers

This is the first demonstration of reinforcement learning controlling error correction at the scale of a full error-corrected processor. Earlier experimental work applied similar reinforcement learning techniques to isolated gates or to bosonic codes, smaller, more contained systems. Running it across a full surface-code processor with over 1,000 live control parameters is a meaningfully harder engineering problem, and the fact that it improved on hand-tuned calibration rather than merely matching it is the part worth sitting with.

Drift is a problem that gets worse, not better, as machines scale up. A processor with a few dozen qubits is recalibrated by hand often enough to stay ahead of drift. A processor with thousands of qubits, the scale IBM's roadmap and others are explicitly building toward, cannot rely on manual tune-up passes keeping pace with how many parameters need adjusting. A system that recalibrates itself continuously, using data the machine produces anyway, is a more plausible path to keeping a much larger processor stable than scaling up the current manual approach.

What is still unproven

This result was demonstrated on Google's own Willow hardware and published in a peer-reviewed venue, which puts it on firmer ground than a vendor press release. It does not yet tell you whether the same reinforcement learning approach transfers cleanly to a different qubit modality, a different code family beyond the ones tested, or a processor an order of magnitude larger than Willow's 105 qubits. The distance-15 simulation results are exactly that: simulations, not a demonstration on real hardware at that scale.

What is established is narrower and still significant: on real hardware, at the scale tested, learning from the error-correction data a processor already generates beats static calibration, and does so by a wide enough margin to set a new record for the surface code. Our error correction explainer and logical qubits and fault tolerance piece cover the broader context for why keeping a large processor calibrated is as hard a problem as building the qubits in the first place.

संबंधित पोस्ट