Steven GellerQuantum Computing, End to End

Book contents

Current section

Part VI. Reliability and Fault Tolerance

  1. Decoherence and Error Channels
  2. Error Mitigation vs Error Correction
  3. Stabilizers and Syndrome Measurement
  4. Repetition, Bit-Flip, Phase-Flip, and Shor Codes
  5. Surface Codes and Threshold Intuition
  6. LDPC, Bosonic, Cat, GKP, and Topological Approaches
  7. Decoders and Real-Time Classical Control
  8. Logical Qubits and Reliable Operations
  9. Why Useful Quantum Computers Are Systems Engineering Projects

Part VI. Reliability and Fault Tolerance · Chapter 52

Logical Qubits and Reliable Operations

Grouping a thousand noisy qubits into a block gives you a bigger noisy object, not a logical qubit. What earns the name is error suppression you can measure. This chapter is the checklist that tells the difference between a logical-qubit milestone and a relabeling exercise.

Lab
In this chapter 17 sections

Demonstrate repeated logical preparation, idle memory, measurement, and required gates under the same protocol, then show logical failure decreases as protection grows while reporting physical qubits, distance, cycles, logical operation time, decoder conditions, and uncertainty.

A one-time encoded state, postselected error-detection run, or logical memory result does not establish a universal logical gate set. Logical and physical error rates are comparable only when the operation, duration, decoding, state preparation, measurement, and confidence method are aligned.

Logical capability evidence ladderEvidence advances from encoded preparation through memory, measurement, gates, and workload sequences while overhead rises.preparememorymeasuregatesequenceprotection levelevery rung retainsbaseline • decodercycles • uncertainty
Figure 52.1. Logical-operation evidence card and suppression fixture: The evidence card distinguishes preparation, memory, gates, measurement, sequences, and their physical baseline.

A logical qubit needs an operation contract

List preparation, idle, measurement, Clifford, entangling, and non-Clifford capabilities as separate claims.

A logical qubit spreads one qubit of information across many physical qubits so that no single failure destroys it. But the value was never the grouping. The value is that operations on the logical qubit fail less often than operations on the physical qubits underneath — by enough, for long enough, to matter to a workload.

Logical quantum information is encoded across physical degrees of freedom and judged by operation-specific logical failure behavior. [nielsen-chuang] [gottesman-stabilizer]

A logical qubit is an experiment with a denominator

“Logical” identifies an encoding and a decoding rule; it does not by itself establish protection. The minimum record names the code, physical elements, stabilizers or checks, number and cadence of rounds, decoder, accepted and discarded shots, and the physical comparator. A postselected experiment can demonstrate that an encoded subspace and syndrome information exist, but its discarded runs still count against a claim about reliable computation. Publish both conditional error among accepted shots and unconditional probability of delivering a usable result.

The comparator must perform the same task for the same duration. A logical memory held for (R) correction rounds should be compared with the best specified unencoded memory over the corresponding wall-clock interval, not with a one-gate physical error. If several physical qubits or repeated measurements form the comparator, describe them. Break-even means that the encoded result beats that declared baseline on the declared metric; it is not a universal threshold crossed forever.

Reporting rule: Publish numerators and denominators for every rejected shot, leakage event, detected fault, and decoded failure. Give exact confidence construction and the number of independent calibration periods. Without those fields, a reader cannot distinguish protection from selection or estimate whether the observed advantage will survive a longer operation sequence.

Choose the physical baseline carefully

Define the best unencoded operation with matched duration and protocol rather than an arbitrary physical-qubit number.

So the practical question for any announcement is short: does this logical qubit reduce risk for a computation anyone actually wants to run? Everything else — the encoding, the code distance, the demonstration — is evidence toward or away from that answer.

Fault-tolerant capability requires reliable logical operations and not merely encoded storage or detection. [fault-tolerant-roads]

Distance scaling is stronger evidence than one favorable point

For a code family, below-threshold behavior should improve as distance or another protection parameter increases, after holding the noise process and operation definition sufficiently stable. Plot logical failure with confidence intervals against distance for several physical conditions. A single larger code that beats a smaller code is encouraging, but fabrication changes, calibration selection, decoder tuning, and different shot counts can explain part of the improvement. Record those changes rather than fitting one clean exponential through incomparable points [surface-codes-2012].

Use exposure units that match the claim. Memory can be reported per round and per unit time; an operation should be reported per logical operation, including the correction rounds and ancilla work surrounding it. When a sequence contains NN operations with small, roughly independent failure probabilities pip_i, ipi\sum_i p_i is a useful first-order budget. It is not a substitute for testing correlations, coherent accumulation, leakage, or fault propagation through a gate.

Suppression across protection levels

Build a synthetic distance-series plot with uncertainty and explain what trend constitutes error suppression.

The factor is not a footnote; it decides when the machine becomes buildable. Second, usefulness. A circuit survives only if its operations collectively stay inside the error budget:

Surface-code evidence should relate logical performance to distance, cycles, decoder, and physical error assumptions. [surface-codes-2012]

Memory, measurement, and gates occupy different evidence rows

A logical memory experiment asks whether encoded information survives. A logical measurement asks whether an encoded observable is read with the claimed reliability. A one-logical-qubit gate asks whether an operation preserves the code space and achieves an error below its comparator. A two-logical-qubit gate adds fault propagation and usually more demanding geometry or ancilla preparation. A non-Clifford operation introduces injection, distillation, code switching, or a platform-specific alternative. Do not fill one row by citing another.

Process-level characterization should include more than truth-table success. Coherent over-rotation can look acceptable on basis states and accumulate badly in algorithms. Leakage can evade a Pauli error model. State-preparation and measurement errors can contaminate an apparent gate result. Use randomized or tomographic tools only with their assumptions stated, and pair aggregate metrics with targeted diagnostics. The conclusion should say exactly which operation and error notion were tested.

Overhead and time accompany the rate

Report physical count, cycles, wall time, decoder, postselection, and yield next to logical error.

A logical-qubit milestone matters exactly when it improves that product for operations a target workload needs — deeper circuits, more two-qubit gates, lower failure probability. A milestone that improves neither is science, which has its own value, but it is not yet utility.

Physical overhead and runtime must be carried alongside logical error when assessing workload relevance. [gidney-ekera-2021] [full-stack-review]

Repeated correction changes the claim

One syndrome round can detect a constructed fault; fault tolerance requires repeated operation in which measurement errors are themselves handled. The controller must timestamp syndromes, the decoder must link events across rounds, and resets must return ancillas without corrupting data. Demonstrate enough rounds to expose temporal correlations, heating, drift, decoder backlog, and leakage. If the run is shortened to remain inside a favorable calibration window, that boundary belongs in the result.

Fault-tolerant preparation and gates require that a small number of physical faults not spread into an uncorrectable logical fault. Verify this structurally in the circuit and empirically where possible. An encoded operation with lower observed error but no fault-containment argument may be a useful logical operation without yet being a fault-tolerant one. The distinction is technical, not rhetorical [fault-tolerant-roads].

Memory, gates, and sequences are separate rungs

Prevent idle-lifetime evidence from silently validating active logical operations or long algorithms.

Suppose a target workload needs on the order of fifty logical qubits, many reliable two-qubit logical gates, and a total failure probability below a stated threshold. An announcement arrives: a team has demonstrated a logical qubit. The evaluation note asks:

Comparative benchmark claims require matched tasks and explicit uncertainty rather than isolated best metrics. [benchmarking-2025]

Read a milestone table without borrowing confidence

Give each row five fields: task, accepted denominator, comparator, uncertainty, and evidence class. Add blanks for operations not demonstrated. A memory result does not populate two-qubit gates; a detected-error result does not populate corrected operation; a simulator result does not populate hardware. If an organization changes code, decoder, or physical platform between milestones, preserve separate rows rather than drawing a continuous trend line.

Suppose a synthetic distance series reports physical error 0.8%, logical memory failure per round of 1.2%, 0.55%, and 0.31% at distances three, five, and seven. The downward trend supports suppression with distance under that fixture, but the distance-three point does not beat the physical comparator. Call the evidence “scaling without demonstrated break-even at every distance,” not simply “a logical qubit.” If the same packet reports a logical two-qubit gate at 2%, leave the reliable-gate conclusion open until its matched physical implementation, interval, and confidence bound are supplied.

A workload consumes an error budget

Translate operation-specific logical rates into a named sequence while exposing independence approximations.

If the note cannot answer most of these, the milestone may be real and still be irrelevant to the workload. Both halves of that sentence matter.

Logical quantum information is encoded across physical degrees of freedom and judged by operation-specific logical failure behavior. [nielsen-chuang] [gottesman-stabilizer]

Convert evidence into a workload reliability claim

An application needs a sequence-level success probability, not a milestone label. Allocate failure among preparation, memory intervals, gates by class, measurements, and classical control. Use upper confidence bounds when provisioning, not only point estimates. Then test whether correlations or shared failure modes make the allocation optimistic. Resource estimates such as large-scale factoring studies are useful precisely because they expose how logical operation counts and target failure drive code and factory requirements [gidney-ekera-2021].

The result of this chapter's exercise should therefore be a sparse evidence matrix and an honest sequence budget. A blank cell is safer than a transferred number. The next experiment is the one that fills the binding cell: perhaps repeated fault-tolerant measurement, a logical entangling gate, or an operation at a larger distance. That is how “logical qubit” becomes a testable engineering statement rather than a ceremonial milestone.

An operation-level acceptance checklist

For every claimed logical operation, ask six questions. What ideal channel is the target? Which encoded input ensemble was tested? Which correction rounds, ancillas, and decoder decisions are included? What physical or logical baseline performs the same task over the same duration? Which uncertainty interval applies to the difference? Which faults can violate the model? The record must also state postselection and leakage treatment. A process estimate conditioned on remaining in the code space cannot be presented as unconditional delivered reliability.

Sequence tests are a useful cross-check. Construct identity-equivalent sequences of increasing length and interleave the target operation, then compare observed decay with a model derived from the independently measured operation budget. Disagreement can reveal coherent accumulation, context dependence, or correlated correction failures. Do not force agreement by fitting every mechanism into one effective depolarizing parameter; preserve the residual and investigate it.

A milestone passes the chapter's standard when a reader can reproduce the evidence-row classification from the published fields. If the lower confidence bound still beats the declared comparator, the break-even conclusion is supported for that task. If larger distance improves the fitted rate while absolute break-even remains unresolved, report suppression only. If memory is protected but the gate row is blank, say “logical memory demonstrated” and make the gate the next proof target.

Interleave operations before claiming a processor

A memory curve does not certify a logical gate. Insert preparation, idle, gate, measurement, and reset into sequences that resemble the intended workload, and retain a separate failure denominator for every operation class. Randomized sequences can expose accumulation, while targeted circuits test propagation through the specific logical primitive. Compare protection levels using identical decoding, inclusion, and timing rules. If the larger code changes the schedule or creates more deadline misses, the overhead belongs in the logical-operation result rather than in a footnote.

Claim-to-source ledger

Logical quantum information is encoded across physical degrees of freedom and judged by operation-specific logical failure behavior. [nielsen-chuang] [gottesman-stabilizer]

Fault-tolerant capability requires reliable logical operations and not merely encoded storage or detection. [fault-tolerant-roads]

Surface-code evidence should relate logical performance to distance, cycles, decoder, and physical error assumptions. [surface-codes-2012]

Physical overhead and runtime must be carried alongside logical error when assessing workload relevance. [gidney-ekera-2021] [full-stack-review]

Comparative benchmark claims require matched tasks and explicit uncertainty rather than isolated best metrics. [benchmarking-2025]

Logical-operation evidence card and suppression fixture

Format: Machine-readable experiment schema plus generated synthetic distance-series plot and workload error-budget table; reuse `resources.py` only for labeled sensitivity context.

Artifact acceptance contract
inputoutputreject when
assumptions, units, source/date, workloadraw and derived values, uncertainty, commandunits or comparison scope are missing
synthetic fixture labeled syntheticdeterministic record and PASS lineattributed to real hardware
named baselinesame task and denominatormetric or evidence class differs
REQUIRED = {"operation", "baseline", "duration_s", "distances", "cycles", "uncertainty", "postselection"}
def classify(card, rates):
    missing = REQUIRED - card.keys()
    if missing or tuple(card["distances"]) != tuple(sorted(rates)):
        raise ValueError(sorted(missing) or "distance mismatch")
    values = [rates[d] for d in card["distances"]]
    trend = "suppression" if all(a > b for a, b in zip(values, values[1:])) else "worsening" if all(a < b for a, b in zip(values, values[1:])) else "flat"
    return trend, "memory-evidence" if card["operation"] == "memory" else "operation-evidence"
card = {"operation":"memory", "baseline":"physical-idle", "duration_s":.001, "distances":[3,5,7], "cycles":1000, "uncertainty":"binomial-95", "postselection":False}
baseline = classify(card, {3:2e-3, 5:7e-4, 7:2e-4})
boundary = classify(card, {3:7e-4, 5:7e-4, 7:7e-4})
counterfactual = classify(card, {3:2e-4, 5:7e-4, 7:2e-3})
assert (baseline[0], boundary[0], counterfactual[0]) == ("suppression", "flat", "worsening")
assert baseline[1] == boundary[1] == counterfactual[1] == "memory-evidence"
print(f"PASS: 52 logical evidence trends={baseline[0]}/{boundary[0]}/{counterfactual[0]} claim={baseline[1]}")

Verification: Schema requires operation/baseline/duration/distance/cycles/uncertainty/postselection; fixture tests distinguish suppression, flat, and worsening trends and never reports a logical capability from memory data alone.

Commissioned exercise

Prompt: Classify three synthetic distance-series datasets for logical idle memory and one logical gate, then estimate a 10,000-operation sequence failure under a declared approximation.

Deliverable: Evidence cards, uncertainty plots, classification, resource columns, and workload budget with assumptions.

Pass condition: Suppression is claimed only where confidence-supported trends improve; gate and memory evidence remain separate; overhead and duration appear beside rates.

Verifiable solution

Format: Reference classifications, plots, and error-budget calculation with an independence caveat.

Verification: Tests recompute intervals and trend labels and reject cards missing operation or comparator metadata.

The synthetic distance series falls from 2e-3 at d=3 to 7e-4 at d=5 and 2e-4 at d=7. The operation budget assigns 2e-4 to memory, 3e-4 to measurement, and 5e-4 to two-qubit operations, summing to the declared 1e-3 workload target without treating memory evidence as gate evidence.

Companion work

Artifacts for this chapter

These entries resolve to checked-in local source. Commands are reproduced exactly from the chapter manifest, and source-embedded fixtures are exported as direct downloads.

  1. Reproduce or test

    python3 tools/validate_briefs.py --briefs data/editorial_briefs_36_63.json --from 36 --through 63 --check-rewritten-sources --execute-artifacts

Provenance

Sources and review

  1. Michael A. Nielsen and Isaac L. Chuang. Quantum Computation and Quantum Information. Cambridge University Press. 2010textbook
  2. Daniel Gottesman. Stabilizer codes and quantum error correction. California Institute of Technology / arXiv. 1997doctoral thesis
  3. Earl T. Campbell, Barbara M. Terhal, and Christophe Vuillot. Roads towards fault-tolerant universal quantum computation. Nature. 2017peer-reviewed review
  4. Austin G. Fowler et al.. Surface codes: Towards practical large-scale quantum computation. Physical Review A. 2012peer-reviewed review
  5. Craig Gidney and Martin Ekerå. How to factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits. Quantum. 2021primary peer-reviewed resource estimate
  6. Lieven M. K. Vandersypen et al.. A look at the full stack. Nature Reviews Physics. 2021peer-reviewed perspective
  7. Timothy Proctor et al.. Benchmarking quantum computers. Nature Reviews Physics. 2025peer-reviewed perspective

The load-bearing claims in the chapter are mapped inline to this registered source set. A citation supports only the bounded claim beside it.

Cite this chapter