LCU¶
Build the circuit PREP, SELECT, PREP† whose all-zero control block is sum_j c_j U_j / alpha, with alpha = sum_j |c_j|, from complex coefficients c_j and dense unitaries U_j (Childs and Wiebe, arXiv:1202.5822v1, Lemma 2, Fig. 1 and Theorem 3). Import the functions from nwqlib.subroutines.lcu. prepare_lcu_data checks the coefficients and unitaries, and build_lcu_prepare, build_lcu_select and build_lcu_circuit build PREP, SELECT or their composition from that data.
import numpy as np
from qiskit.quantum_info import Operator
from nwqlib.subroutines.lcu import build_lcu_circuit
I = np.eye(2)
Z = np.diag([1.0, -1.0])
lcu = build_lcu_circuit([0.75, -0.25], [I, Z])
alpha = lcu.data.coefficient_l1_norm
block = Operator(lcu.circuit).data[::2, ::2] # control qubit is bit 0
print(alpha, lcu.preparation_l2_error)
print(np.round(alpha * block.real, 12))
1.0 0.0
[[0.5 0. ]
[0. 1. ]]
0.75 I - 0.25 Z is diag(0.5, 1), and direct PREP has zero preparation error.
What the circuit encodes¶
- PREP prepares
sum_j sqrt(|c_j| / alpha) |j>on the control register, withalpha = sum_j |c_j|. SELECT applies(c_j / |c_j|) U_jon addressj, so the coefficient phases sit in SELECT and PREP encodes only nonnegative magnitudes (Childs and Wiebe, Sec. II, after Eq. (4)). This is the single-register standard form of Low and Chuang, arXiv:1610.06546v3, Lemma 5 and Eq. (10), and the caseP_L = P_Rof the subnormalization of a linear combination of block encodings in Gilyén, Su, Low and Wiebe, arXiv:1806.01838v1, Lemma 52, with each phase-adjustedU_ja(1, 0, 0)block encoding of itself (Definition 44). - The control register precedes the system register, so it occupies the low-order bits of Qiskit's little-endian index, and the block entry for system basis index
ssits at indexs * 2**num_control_qubits. - The control register has
Paddresses, the smallest power of two at least the numberNof terms. TheP - Npadding addresses have zero PREP amplitude and act as the identity, without stored matrices or controlled gates. - A nonzero global scale of the coefficients changes
alphaand keeps the normalized coefficient direction. The defaultcoefficient_atol=0therefore rejects only zero total weight. A supplied positive threshold applies to that total weight and never removes individual coefficients.
Build the circuit¶
prepare_lcu_data ¶
prepare_lcu_data(coefficients: Any, unitaries: Any, *, coefficient_atol: float = 0.0, max_bytes: int = DEFAULT_MAX_BYTES, max_work: int = DEFAULT_MAX_LCU_WORK) -> LCUData
Check LCU coefficients and unitaries, and move each coefficient's phase into its unitary.
The result holds the PREP amplitudes sqrt(|c_j| / alpha) with
alpha = sum_j |c_j| and the phase-adjusted unitaries
(c_j / |c_j|) U_j that SELECT applies, ready for
build_lcu_prepare
and build_lcu_select.
Coefficient count, finite values and nonzero total weight are checked
before unitary conversion.
Parameters:
-
coefficients(array_like) –One-dimensional complex coefficients
c_j. -
unitaries(Sequence[array_like]) –One square unitary matrix
U_jper coefficient, of power-of-two dimension D. -
coefficient_atol(float, default:0.0) –Default
0.0. Threshold on the total coefficient 1-norm. The default rejects only a zero norm, so a representable global rescaling is still accepted. This threshold does not prune individual terms. Every nonzero coefficient contributes its phase. -
max_bytes(int, default:DEFAULT_MAX_BYTES) –Default 10 GB (decimal,
10_000_000_000bytes). Byte limit for the coefficient and PREP arrays, the supplied matrices and the later dense SELECT, including the synthesized branch circuits it keeps. -
max_work(int, default:DEFAULT_MAX_LCU_WORK) –Default
1_000_000_000. Work limit for these checks and the later dense SELECT, including the synthesis of each controlled branch.
Returns:
-
data(LCUData) –Coefficients, PREP amplitudes, phase-adjusted unitaries and register sizes.
data.coefficient_l1_normisalpha.
Raises:
-
ValueError–If inputs have inconsistent shapes or all coefficients are zero, or if the PREP or later dense SELECT would exceed
max_bytesormax_work.
Examples:
For 0.5 X - 0.5i Z, alpha = 1, both PREP amplitudes are
sqrt(0.5) = 0.7071..., and SELECT applies X and -i Z:
>>> import numpy as np
>>> from nwqlib.subroutines.lcu import prepare_lcu_data
>>> X = np.array([[0, 1], [1, 0]])
>>> Z = np.diag([1, -1])
>>> data = prepare_lcu_data([0.5, -0.5j], [X, Z])
>>> print(data.coefficient_l1_norm, data.num_control_qubits)
1.0 1
>>> print(np.round(np.real(data.prep_amplitudes), 4))
[0.7071 0.7071]
>>> print(np.allclose(data.phase_adjusted_unitaries[1], -1j * Z))
True
build_lcu_circuit ¶
build_lcu_circuit(coefficients: Any, unitaries: Any, *, preparation_backend: LCUPreparationBackend = 'direct', mps_max_bond_dim: int | None = None, mps_threshold: float = 1e-14, mps_num_layers: int = 2, max_bytes: int = DEFAULT_MAX_BYTES, max_work: int = DEFAULT_MAX_LCU_WORK) -> LCUCircuit
Build the PREP-SELECT-PREP^dagger circuit whose all-zero control block is sum_j c_j U_j / alpha.
The function runs prepare_lcu_data, build_lcu_prepare and
build_lcu_select and composes PREP, SELECT and the inverse of PREP,
with alpha = sum_j |c_j|.
Parameters:
-
coefficients(array_like) –Complex LCU coefficients
c_j. -
unitaries(Sequence[array_like]) –One D-square unitary matrix
U_jper coefficient, with D a power of two. -
preparation_backend(str, default:'direct') –Default
"direct", exact PREP."mps_circuit"uses MPS disentangling, as forbuild_lcu_prepare. -
mps_max_bond_dim(int | None, default:None) –Default
None(no cap). Rank cap for MPS PREP. -
mps_threshold(float, default:1e-14) –Default
1e-14. Singular-value threshold for MPS PREP. Each singular value at or below it is dropped from the TT-SVD of the unit-norm coefficient state, keeping at least one per bond. -
mps_num_layers(int, default:2) –Default
2. Number of MPS disentangling layers. -
max_bytes(int, default:DEFAULT_MAX_BYTES) –Default 10 GB (decimal,
10_000_000_000bytes). Byte limit for the supplied matrices, PREP and the dense SELECT with its kept synthesized branch circuits. -
max_work(int, default:DEFAULT_MAX_LCU_WORK) –Default
1_000_000_000. Work limit for the input checks, the SELECT synthesis and the TT-SVD of MPS PREP.
Returns:
-
lcu(LCUCircuit) –The circuit in
lcu.circuit, targeting the normalized all-zero blocksum_j c_j U_j / sum_j |c_j|, andalphainlcu.data.coefficient_l1_norm. Direct PREP realizes the coefficient amplitudes exactly, and SELECT realizes each controlled branch to binary64 rounding when the branch is unitary to rounding. MPS PREP approximates the coefficient amplitudes, so the realized block can differ. Its error remains unevaluated in the preparation metadata, andpreparation_l2_errorisNone.
Examples:
The control register is the low-order bits, so the block is read at
stride 2. alpha times the block equals 0.5 X - 0.5i Z:
>>> import numpy as np
>>> from qiskit.quantum_info import Operator
>>> from nwqlib.subroutines.lcu import build_lcu_circuit
>>> X = np.array([[0, 1], [1, 0]])
>>> Z = np.diag([1, -1])
>>> lcu = build_lcu_circuit([0.5, -0.5j], [X, Z])
>>> block = Operator(lcu.circuit).data[::2, ::2]
>>> alpha = lcu.data.coefficient_l1_norm
>>> print(np.allclose(alpha * block, 0.5 * X - 0.5j * Z))
True
build_lcu_prepare ¶
build_lcu_prepare(data: LCUData, *, preparation_backend: LCUPreparationBackend = 'direct', mps_max_bond_dim: int | None = None, mps_threshold: float = 1e-14, mps_num_layers: int = 2, max_bytes: int = DEFAULT_MAX_BYTES, max_work: int = DEFAULT_MAX_LCU_WORK) -> QuantumCircuit
Build the PREP circuit that loads the LCU coefficient state on the control register.
Parameters:
-
data(LCUData) –Output of
prepare_lcu_dataorprepare_lcu_gate_data. -
preparation_backend(str, default:'direct') –Default
"direct", the exact magnitude tree ofbuild_qiskit_state_preparation."mps_circuit"uses the MPS disentangling circuit ofbuild_mps_circuit_state_preparation. -
mps_max_bond_dim(int | None, default:None) –Default
None(no cap). Rank cap for MPS PREP. -
mps_threshold(float, default:1e-14) –Default
1e-14. Singular-value threshold for MPS PREP. Each singular value at or below it is dropped from the TT-SVD of the unit-norm coefficient state, keeping at least one per bond. The constants registry gives its reason. -
mps_num_layers(int, default:2) –Default
2. Number of MPS disentangling layers. -
max_bytes(int, default:DEFAULT_MAX_BYTES) –Default 10 GB (decimal,
10_000_000_000bytes). Limit on the known coefficient and PREP arrays. It does not bound SDK memory. -
max_work(int, default:DEFAULT_MAX_LCU_WORK) –Default
1_000_000_000. Limit on the PREP work and the TT-SVD of MPS PREP.
Returns:
-
circuit(QuantumCircuit) –Circuit targeting
sum_j sqrt(|c_j| / alpha) |j>from|0...0>. Direct PREP realizes this ideal relation. Finite-layer MPS PREP approximates it with unevaluated circuit error. TT-SVD discarded weight alone does not bound that circuit error.
build_lcu_select ¶
build_lcu_select(data: LCUData, *, max_bytes: int = DEFAULT_MAX_BYTES, max_work: int = DEFAULT_MAX_LCU_WORK) -> QuantumCircuit
Build SELECT, which applies the dense unitary of branch j when the control register holds j.
This entry serves small dense validation tables of unrelated multi-qubit unitaries. Pauli and product-formula callers use their gate-level multiplexor compilers and do not go through this dense construction.
The whole controlled matrix of each branch, with its address as the
control state, is synthesized exactly (see
Controlled dense unitaries).
SELECT therefore equals the ideal relation to binary64 rounding when
every branch is unitary to rounding. A branch that passes Qiskit's
constructor check (numpy.allclose of U^dagger U with the identity,
atol 1e-8 and rtol 1e-5) with a larger defect is realized as its unitary
polar factor, which differs from it by the order of that defect. For
U = expm(-i t G) with random Hermitian G, t from 1e-10 to 10, one to
three system qubits and one to three controls, the largest entry error
was 4e-14, also after lowering to U and CX. A branch on
m = log2(P D) qubits takes at most (25/96) 4**m - 2**m + 4/3 CX
for m >= 3, the largest count that Qiskit's own synthesis of the same
controlled matrices reached in tests.
Parameters:
-
data(LCUData) –Output of
prepare_lcu_data, which holds the dense unitaries. -
max_bytes(int, default:DEFAULT_MAX_BYTES) –Default 10 GB (decimal,
10_000_000_000bytes). Byte limit for the supplied matrices, the arrays of one branch synthesis and the synthesized circuits the branches keep, checked before the first synthesis. -
max_work(int, default:DEFAULT_MAX_LCU_WORK) –Default
1_000_000_000. Work limit for the unitarity checks and the synthesis of each branch, checked before the first synthesis.
Returns:
-
circuit(QuantumCircuit) –Circuit applying branch
jwhen the control register storesj.
Raises:
-
ValueError–If the dense unitary tuple is empty or does not cover the supplied term count. Extra padding branches act as identity.
prepare_lcu_gate_data ¶
prepare_lcu_gate_data(coefficients: Any, *, system_dimension: int, coefficient_atol: float = 0.0, max_bytes: int = DEFAULT_MAX_BYTES, max_work: int = DEFAULT_MAX_LCU_WORK) -> LCUData
Check LCU coefficients and compute the PREP data for a SELECT built from your own gates.
The coefficient handling is that of
prepare_lcu_data,
for callers whose SELECT branches are inserted as controlled gates
rather than dense UnitaryGate matrices. phase_adjusted_unitaries is
empty, and each branch gate must carry its own coefficient phase. The
PREP register still encodes only nonnegative coefficient magnitudes.
The returned data cannot be passed to build_lcu_select.
Parameters:
-
coefficients(array_like) –One-dimensional complex coefficients.
-
system_dimension(int) –Dimension D of the system register that the branch gates act on, a power of two.
-
coefficient_atol(float, default:0.0) –Default
0.0. Threshold on the total coefficient 1-norm, as forprepare_lcu_data. -
max_bytes(int, default:DEFAULT_MAX_BYTES) –Default 10 GB (decimal,
10_000_000_000bytes). Limit on the coefficient and PREP arrays. -
max_work(int, default:DEFAULT_MAX_LCU_WORK) –Default
1_000_000_000. Limit on the PREP work.
Returns:
-
data(LCUData) –Coefficients, PREP amplitudes and register sizes, with an empty
phase_adjusted_unitaries.
Results¶
LCUData ¶
LCUData(*, original_coefficients: tuple[complex, ...], positive_coefficients: tuple[float, ...], phase_adjusted_unitaries: tuple[ndarray, ...], coefficient_l1_norm: float, prep_amplitudes: tuple[complex, ...], term_count: int, padded_term_count: int, num_control_qubits: int, num_system_qubits: int, system_dimension: int, register_order: tuple[str, str] = ('lcu_control', 'system'))
Checked LCU coefficients, PREP amplitudes and phase-adjusted unitaries.
prepare_lcu_data and
prepare_lcu_gate_data
return it. alpha is coefficient_l1_norm. The fields below are
read-only.
Attributes:
-
original_coefficients(tuple[complex, ...]) –The N supplied complex coefficients
c_j, one per term, before their phases are absorbed. -
positive_coefficients(tuple[float, ...]) –The magnitudes
|c_j|that PREP encodes, zero-padded to P entries. -
phase_adjusted_unitaries(tuple[ndarray, ...]) –The N D-square complex128 matrices
(c_j / |c_j|) U_jthat SELECT applies, withU_junchanged forc_j = 0. Only the N supplied matrices are stored, and theP - Npadding addresses act as the identity. Empty for gate-level SELECT data (prepare_lcu_gate_data), whose branch gates carry their own coefficient phases. -
coefficient_l1_norm(float) –alpha = sum_j |c_j|, the subnormalization of the encoded block, in the units of the coefficients. -
prep_amplitudes(tuple[complex, ...]) –The P complex amplitudes
sqrt(|c_j| / alpha)that PREP prepares on the control register, zero on padded addresses. -
term_count(int) –Number N of supplied terms before padding.
-
padded_term_count(int) –Number P of control-register basis states, the smallest power of two at least N.
-
num_control_qubits(int) –Number
a = log2 Pof control qubits. -
num_system_qubits(int) –Number
log2 Dof system qubits. -
system_dimension(int) –Dimension D of each unitary
U_j. -
register_order(tuple[str, str]) –Register names in circuit order, control register first.
LCUCircuit ¶
LCUCircuit(*, circuit: QuantumCircuit, data: LCUData, preparation_backend: str = 'direct', preparation_metadata: Mapping[str, Any] = dict(), preparation_l2_error: float | None = 0.0)
A PREP-SELECT-PREP^dagger circuit with its coefficient data and PREP error.
build_lcu_circuit
returns it. The circuit is circuit, and alpha is
data.coefficient_l1_norm. The fields below are read-only.
Attributes:
-
circuit(QuantumCircuit) –The built circuit PREP, SELECT, PREP^dagger on
a + log2 Dqubits, control register first. With direct PREP its all-zero control block issum_j c_j U_j / alpha. MPS PREP approximates that block. -
data(LCUData) –The
LCUDatathe circuit was built from: coefficients, phase-adjusted unitaries, dimensions and register order. -
preparation_backend(str) –PREP construction used in
circuit,"direct"or"mps_circuit". -
preparation_metadata(Mapping[str, Any]) –JSON-like record of that same PREP construction, such as its method name, error model and, for
"mps_circuit", the TT-SVD summary. -
preparation_l2_error(float | None) –2-norm distance between the state PREP prepares and the target coefficient state, phases included, in the ideal-gate model, as reported by the PREP construction. 0 for direct PREP, which makes no algorithmic approximation, and for a single term, which needs no PREP circuit. None for MPS PREP of two or more terms, whose circuit error this construction does not evaluate. It has the same phase convention as coherent PREP, inverse PREP and controlled use.
Accuracy and limits¶
- A supplied zero-weight unitary that is not the identity stays a real SELECT branch and goes through the normal unitarity check.
- Every function takes
max_bytes(default 10,000,000,000) andmax_work(default 1,000,000,000). Before any matrix conversion or synthesis they check the coefficient and PREP tables, the supplied matrices and, for a dense SELECT with a control register, the exact synthesis of each controlled branch, forNterms,Paddresses and system dimensionD. A branch is synthesized from itsPD-square controlled matrix. WithM = PDandm = log2(M), each branch counts11 M³ + (m² + 5m + 256) M²work units,256 M² + 65536bytes of working arrays, which one branch at a time uses, and176 M² + 16384bytes for the circuit it keeps. - The default
max_workis a guard against runaway planning work, and a dense SELECT above it needs an explicitly raisedmax_work. TheDEFAULT_MAX_LCU_WORKrow of Engineering constants gives the derivation and the measured times and memory behind these terms. Work counts scalar operations, not time. - Pass the limits to each stage you call.
build_lcu_circuitforwards the same values to its stages. The workspace of the tensor library used by layered MPS PREP is not counted.
Source map¶
Code paths are relative to nwqlib.subroutines. Equation and section numbers refer to the listed arXiv versions.
| Scientific step | Source | Location | Code |
|---|---|---|---|
| Two-term PREP-SELECT-PREP† circuit and its extension to general combinations | Childs and Wiebe, arXiv:1202.5822v1 | Lemma 2, Fig. 1 and Theorem 3 | lcu.core.build_lcu_circuit |
| Complex weights absorbed as phases of the unitaries | Childs and Wiebe, arXiv:1202.5822v1 | Sec. II, after Eq. (4) | lcu.core.prepare_lcu_data |
Single-register PREP state and alpha as the coefficient 1-norm |
Low and Chuang, arXiv:1610.06546v3 | Lemma 5 and Eq. (10) | lcu.data._coefficient_bookkeeping |
| Subnormalization of a combination of block encodings | Gilyén et al., arXiv:1806.01838v1 | Lemma 52 | Composition rule in the conventions |
Branch j fires on control value j |
Qiskit little-endian ctrl_state convention |
lcu.core._controlled_branch_gate |
|
| Exact synthesis of each controlled branch, demultiplexed at the top because the controlled matrix is block diagonal in each control qubit | Shende, Bullock and Markov, arXiv:quant-ph/0406176v5, with the block-ZXZ steps of Krol and Al-Ars, arXiv:2403.13692v2 | Theorem 12. Controlled dense unitaries lists the other locators | qiskit_compat.controlled, _dense_synthesis.controlled_unitary_circuit |
| Input and construction limit checks | NWQLib, registered as DEFAULT_MAX_LCU_WORK in Engineering constants. The synthesis terms follow the recursion of the exact synthesis |
_dense_synthesis.controlled_synthesis_size |
lcu.core._admit_lcu |