Exact dense synthesis¶
This page is for contributors and for readers who check NWQLib's CX counts and work laws. It describes how NWQLib builds a circuit that equals a dense unitary matrix to rounding, and its controlled form, where Qiskit's own synthesis is not exact. It also covers how the dense_control_route option of LCHS, QLS and build_control_diagonal_generator_encoding chooses between the two controlled routes, how each construction counts the synthesis work against its limits, and how an uncontrolled multiplexer is inverted. The user guide for blocks is Compose blocks.
Controlled dense unitaries¶
Native control in lowering (building the Qiskit circuit from a Program), the controlled constructions of LCHS, QLS and QSP, the dense LCU SELECT and the QPE powers all call subroutines/qiskit_compat.controlled. Qiskit 2.5.2 controls a composite gate by unrolling its definition and controlling each gate, and it unrolls a dense UnitaryGate on two or more qubits through its own synthesis, which is not exact to rounding. That synthesis replaces the Weyl coordinates of a two-qubit unitary by those of a special class whenever the average gate fidelity of the replacement is at least 1 - 1e-9. A two-qubit unitary close to the identity thus loses its entangling part, with entry errors measured up to 4.2e-5 (dependency issues). Qiskit's quantum Shannon decomposition also omits multiplexed rotations of at most 1e-10 rad.
controlled therefore first replaces each such unitary in the definition, at any depth, by the circuit of subroutines/_dense_synthesis.dense_unitary_circuit, which equals it to rounding. No operator is formed to check the result. A two-qubit unitary uses its KAK decomposition in the magic basis and at most three CX. A larger one uses the block-ZXZ quantum Shannon decomposition down to two-qubit blocks, with the CX savings of Krol and Al-Ars, arXiv:2403.13692v2, and optimization A.2 of arXiv:quant-ph/0406176v5. A generic n-qubit unitary then takes (22/48) 4**n - (3/2) 2**n + 5/3 CX, the same count as Qiskit's qs_decomposition. A two-qubit block with a small second Weyl coordinate can keep one more CX, because A.2 cannot remove its last coordinate exactly.
Only a Weyl coordinate, a one-qubit gate or a block entry within 2**-46 of zero, of the identity or of an exact special value is replaced by that value, and each replacement changes the operator by an amount of that order (ROUNDING_WINDOW, engineering constants). tests/test_dense_synthesis.py checks the entry error for one to five qubits, the CX counts and the special classes, and tests/test_scientist_lchs.py compares controlled LCHS branches with SciPy exponentials.
A UnitaryGate that is controlled directly, as in the dense LCU SELECT, the dense QPE powers and the coherent QPE powers, is synthesized from its whole controlled matrix by subroutines/_dense_synthesis.controlled_unitary_circuit. Qiskit's own UnitaryGate.control synthesizes the same matrix by quantum Shannon decomposition and keeps the result when numpy.allclose with atol=1e-7 and rtol=1e-5 accepts it, and otherwise falls back to an Isometry circuit. For one-qubit X, Y and Z rotations by 1e-8 to 1e-2 with one to three controls, the kept results erred by up to 7.9e-6 per entry, and for RX(1e-7) with three controls the fallback raised ValueError (Qiskit 2.5.2, dependency issues). For a random 512 × 512 unitary with one control, dependency issues compares the time and gate count of this synthesis with those of Qiskit's qs_decomposition and UnitaryGate.control.
The controlled matrix is block diagonal with respect to each control qubit, so the exact synthesis demultiplexes it at the top (Shende, Bullock and Markov, arXiv:quant-ph/0406176v5, Theorem 12) and synthesizes the two halves without A.2. On m >= 3 qubits this takes at most (25/96) 4**m - 2**m + 4/3 CX, the largest count that Qiskit 2.5.2 reached for the same controlled matrices in tests, and at most (11/16) 4**m - 2**(m+1) instructions. A base matrix that is not unitary to rounding is realized as its unitary polar factor, which differs from it by the order of its unitarity defect.
Controlling each gate of the synthesized base unitary instead takes about four times as many CX with one control, the case of every QPE power. In tests on up to seven qubits it took fewer only with four or more controls, for example 1476 instead of 4140 CX for a two-qubit branch with five controls. For U = expm(-i t G) with t from 1e-10 to 10, one to three system qubits and one to three controls, the entry error stayed below 4.1e-14, also after lowering to U and CX. The software versions, the platform and the generators G of these tests were not recorded, so this value describes those tests and is not an error bound.
subroutines/_dense_synthesis.controlled_synthesis_size derives the work and bytes that the LCU and QPE limit checks and the coherent QPE builder count for this synthesis. inverse() of the resulting gate follows Qiskit's ControlledGate.inverse and synthesizes the adjoint through UnitaryGate.control again, while subroutines/qiskit_compat.inverse_realized_gate, which lowering uses for adjoint blocks, reverses the exact definition.
An uncontrolled UnitaryGate stays a matrix in the logical circuit, and Aer applies that matrix directly. Before compiling, the following targets replace each dense UnitaryGate outside control-flow blocks and annotated operations by the same exact synthesis (qiskit_compat.exact_dense_unitaries): the Nexus preparation, NWQ-Sim, IonQ with its QIS gateset, IBM at optimization levels 0 and 1, and Aer with a noise model whose gate basis omits the unitary instruction. Optimization levels 2 and 3 of NWQ-Sim, IonQ and IBM resynthesize two-qubit blocks with Qiskit's own synthesis, so they receive the logical circuit as given, and remote compilers such as Nexus apply their own passes. The MPS state preparation (subroutines/state_preparation/mps_circuit.py) synthesizes its two-qubit gates with dense_unitary_circuit too.
Dense control route¶
In LCHS, QLS and build_control_diagonal_generator_encoding, the dense_control_route option selects one of two routes by which a construction adds controls to a dense unitary on two or more qubits (subroutines/_dense_synthesis.select_dense_control_route).
| Value | Route |
|---|---|
"gatewise" |
Synthesize the unitary with dense_unitary_circuit and let Qiskit control each synthesized gate |
"whole_matrix" |
Synthesize the controlled matrix with controlled_unitary_circuit, as for a directly controlled UnitaryGate |
"auto" (default) |
The whole-matrix route for one control and the gate-wise route for two or more (AUTO_WHOLE_MATRIX_MAX_CONTROLS, registered with its derivation in engineering constants) |
The option is part of the Method and of the Plan's identifier, and planning, the limit checks, the resource laws and lowering read the same route. For Haar-random unitaries on 2 to 5 qubits (Qiskit 2.5.2) the whole-matrix route took 3.2 to 4.0 times fewer CX than the gate-wise route with one control, 3.6 to 3.8 times fewer with two and 1.9 to 2.1 times fewer with three. Its classical work, about 88 M3 units with one control against about 28 M3 for the synthesis of the M-square unitary alone, grows about eight times and its bytes four times per added control. Both routes give the controlled matrix to rounding (tests/test_dense_control_route.py).
The route applies to three constructions that hold the dense matrix: an LCHS dense_exact branch on its address bits, the controlled QLS query of a planned dense dilation and a dense_dilation child that the combine qubit of a two-child joint generator controls. On the whole-matrix route an LCHS constant-source branch is one unitary, its evolution times the matrix of its input preparation, and a joint-generator branch is controlled part by part, the child through its controlled dilation and the diagonal rotation gate-wise, with the branch's global phase as a phase gate on the control (qiskit_compat.controlled).
A one-qubit unitary keeps Qiskit's exact definition on both routes. A supplied circuit, such as a native encoding, a supplied preparation or the ADAPT reference, is controlled gate-wise on every route, because its matrix is not known and it may hold operations without one. The parity control of a QSP pass controls the whole pass and is gate-wise on every route.
Limit checks of the exact synthesis¶
subroutines/_dense_synthesis.dense_synthesis_size gives the work and bytes of one dense_unitary_circuit on m qubits, 113 M**3/4 + (m**2 + 16 m + 512) M**2 units, 320 M**2 + 65536 working bytes and 352 M**2 + 16384 kept bytes with M = 2**m, derived in its docstring and registered in engineering constants. controlled_synthesis_size gives those of a directly controlled matrix. subroutines/qiskit_compat.dense_synthesis_widths lists the syntheses that a rewrite would make, without making them. A Haar-random unitary on 8 qubits took 0.57 s and one on 9 qubits 2.6 s on one core (Apple-silicon Mac, Qiskit 2.5.2, one thread, dense_synthesis_size docstring).
Gate-wise control law¶
On the gate-wise route, when a construction controls a composite gate that holds dense unitaries, qiskit_compat.controlled replaces each by its exact synthesis, and Qiskit then controls every synthesized gate, once for every place where the synthesis occurs in the gate. This control step has its own law, the gate-wise control law of gatewise_control_counts and gatewise_control_size.
For one synthesis on m qubits with k controls, gatewise_control_counts bounds three counts: the gates that Qiskit unrolls (the gate count of dense_synthesis_gate_census plus the global phase), the instructions that it emits, and the heavy instructions among them, which hold angles or are Python objects. Qiskit 2.5.2 turns a U gate into 1, 21, 86, 193 and 829 instructions with 1, 2, 3, 4 and 8 controls and an RZ into 1, 8, 36, 64 and 276, and a CX or the global phase into one and an H into seven for every k (CONTROLLED_U_INSTRUCTIONS and CONTROLLED_RZ_INSTRUCTIONS store these counts for 1 to 64 controls).
gatewise_control_size counts 2048 units of planning work for each unrolled gate and 16 for each emitted instruction. Its kept-byte law allows 96 bytes for each instruction and 1024 more for each heavy one. qiskit_compat.dense_control_counts adds these counts over the occurrences of the dense unitaries in a gate.
For Haar-random syntheses on 4 to 7 qubits with one to eight controls (Qiskit 2.5.2, Apple-silicon Mac, one thread), the law counted 1.0 to 1.19 times the emitted instructions, and Qiskit's control call took 4.4 to 35 ns per unit of planning work against 1.6 to 15 ns per unit for the synthesis itself. Both laws feed the same limits, which count units of planning work, not time. Together, the kept-byte laws of the synthesis and of this step allowed at least 1.7 times the memory that the kept controlled gates held, and the working and kept bytes of this step allowed at least 2.1 times the peak increase of resident memory during the call, wherever that increase exceeded 0.1 MB. For a random 512 × 512 unitary, the size of the dense dilation of a 256-dimensional QLS system, the synthesis gave 326,146 gates, 119,383 of them CX, in 2.6 s, Qiskit's control took 3.8 s with one control and 27 s with two, and the controlled circuit held 391,679 and 2,914,498 instructions (Qiskit 2.5.2, Apple-silicon Mac, one thread), where the law counts 424,445 and 3,400,957.
Where each construction checks the synthesis¶
The table lists the constructions that synthesize dense unitaries and how each checks the cost of those syntheses against its limits.
| Construction | Syntheses | Checked against |
|---|---|---|
| QLS query of a dense dilation with valued controls, in the Hermitian dilation of a non-Hermitian inverse and in both shortcuts | The forward and the adjoint dilation on log2(p) + 1 qubits for padded dimension p, each once per Run, and Qiskit's control of each with one control, or two in shortcut_dilation. On the whole-matrix route, which "auto" takes for one control, the forward and the adjoint controlled dilation instead |
QLS.max_work and max_bytes at planning |
Supplied QLS encoding (encoding.native) in a query with valued controls, and a supplied RHS circuit in a shortcut |
The dense unitaries of the supplied circuit and Qiskit's control of them, gate-wise, in each controlled specialization, two for the encoding, two for the RHS of shortcut_native_svp and four for that of shortcut_dilation |
QLS.max_work and max_bytes at planning |
LCHS dense_exact SELECT with address bits, with or without a constant source |
One on the q system qubits per physical branch, none for q = 1, and Qiskit's control of each branch with the address bits. On the whole-matrix route one controlled branch on q + a qubits per physical branch on q >= 2 qubits, after the matrix of a constant-source branch's input preparation is formed | LCHS.max_select_work and max_bytes at planning |
| LCHS compiled QSP SELECT with a dense-dilation child | One per dense child beside another child, which the combine qubit controls once and the parity qubit again at every query of a pass, or, on the whole-matrix route, one controlled dilation per such child, which the parity qubit controls at every query. A lone child has its forward and adjoint synthesized by the control of each pass, which controls it at every query | LCHS.max_select_work and max_bytes at planning |
build_qsp_evolution_encoding and build_control_diagonal_generator_encoding |
Those in the passes or branches that they control, and Qiskit's control of every occurrence, except the second control of a gate that an earlier control unrolled, which LCHS planning counts | Their own max_work and max_bytes |
Dense QPE powers, build_coherent_qpe_circuit and the dense LCU SELECT |
One controlled matrix per power or branch | QPE max_work and max_bytes at planning, the builder's own limits and _admit_lcu |
| ADAPT reference preparation from a supplied circuit | The dense unitaries of the circuit and Qiskit's control of them, once per Run for the Hadamard-test and joint-state queries | ADAPT.max_products and max_bytes at planning |
Explicit ADAPT verification of a quantum Plan with a supplied reference circuit |
The dense unitaries of the uncontrolled reference, once | AdaptVerificationOptions.max_products and max_bytes, before the synthesis |
Controlled transform (transform_block) of a dense-dilation encoding, a supplied native encoding or a supplied preparation, such as a FixedGCIM basis state |
The dense unitaries of the base and Qiskit's control of them, once per base and Run, shared by its controlled forward and adjoint | The Run's ExecutionLimits.max_synthesis_work, or the max_synthesis_work of lower_qiskit, before the synthesis. The transformed record reserves the same work in its construction_work |
| LCHS representative sampling | One per active dense-dilation child of the QSP representative, with Qiskit's control of it. A dense exact SELECT is reported by its structural upper-bound record and synthesizes nothing | sample_resources(max_build_work=...) |
| Basis lowering by NWQ-Sim, Nexus, IonQ with the QIS gateset, IBM at levels 0 and 1, and Aer with a noise model whose basis omits the unitary instruction | Each distinct uncontrolled dense matrix once per Run object, through the Run's synthesis cache | The Run's ExecutionLimits.max_synthesis_work, a total over its preparations, for the matrices missing from the cache, before the first synthesis of each preparation |
| MPS preparation on n qubits | At most layers * (n - 1) two-qubit syntheses |
The builder's max_svd_work and max_bytes, before the first synthesis, and in LCHS also the layered MPS upper bound, which exceeds their work for n >= 2 |
QLS charges the forward and the adjoint synthesis of a controlled dense query and their control to max_work and max_bytes at planning, which by default admits at most 64 padded coordinates. LCHS charges its dense_exact branches and dense-dilation QSP children to max_select_work and max_bytes, whose default admits at most 4 branches on 6 system qubits, 8 on 5, 32 on 4 and 128 on 3 among power-of-two branch counts for the test problems of _dense_select_problem in tests/test_lchs_resource_structural_law.py.
Synthesis cache of a Run¶
Basis lowering synthesizes through Run._exact_dense_unitaries, which keeps the circuit of every synthesis until the Run is closed, keyed by the shape, dtype and bytes of its matrix. A later preparation that holds the same matrix, even in new gate objects, receives a copy of that circuit, which equals a new synthesis because the synthesis is deterministic. A reopened Run starts with an empty cache.
The Plan of a QLS inverse of an 8-dimensional Hermitian dense system with a random dense observable has 35 readout settings, and its 35 preparations on NWQ-Sim make two syntheses, one for each of its two distinct dense unitaries.
The Run reserves the syntheses that basis lowering or a controlled transform makes on the charge of the preparation that makes them (the work that preparation counts against the Run's limits), in one journal commit for each batch before its first synthesis starts. A dense unitary becomes known only when the Plan is lowered, which follows the preparation charge, so a refusal of its synthesis comes after that charge. The refused preparation stays charged against max_total_circuits and leaves its experiment unprepared. After run.extend_limits(max_synthesis_work=...) raises the limit, the next call of the static workflow or of the Method's controller prepares the experiment again.
Qiskit also controls the other gates of a controlled composite, such as the projector phases and multiplexors of a QSP pass or the input preparation of a controlled LCHS source branch, and a backend later lowers the controlled circuit to its gate basis. These costs lie outside the laws above, and Limitations and open work keeps them open.
Adjoint multiplexers¶
An uncontrolled, complete multiplexer can be adjointed by conjugate-transposing each stored target matrix and keeping the same address order. Aer applies that table directly when its target supports the native multiplexer instruction (Aer 0.17.2 QubitVector::apply_multiplexer, source). The table route copies each conjugate-transposed 2×2 entry, allocating 64 * 2**k new array bytes for k controls, and forms no dense operator. The adjoint identity holds for the stored matrices. Equality to an inverse also requires unitary table blocks, so a native forward/adjoint pair has the table's Gram defect in addition to execution error.
An outer control expands the multiplexer into a gate circuit. QSP evolution and controlled QLS queries therefore construct the inverse by reversing the realized forward definition before copying repeated child queries. This reuses the forward UCG synthesis. Gate-wise control still processes each query occurrence. Uncontrolled native consumers, including Lanczos readout and direct block adjoints, keep their adjoint tables.
inverse_realized_gate(..., native_ucg=False) selects the reversed-definition route when a caller knows the inverse will be controlled or decomposed. With native_ucg=True, native execution remains conditional on the later transformations and backend target. A backend that decomposes an adjoint table constructs a separate numerical synthesis whose error must be assessed in that representation.
Source map¶
The CX counts and work laws above rest on these sources. The selection relations of blocks are mapped in Compose blocks.
| Relation | Source | Location | Code |
|---|---|---|---|
| KAK decomposition of a two-qubit unitary in the magic basis, with the real and imaginary parts of the symmetric unitary sharing a real eigenbasis | Shende, Markov and Bullock, arXiv:quant-ph/0308033v3 | Proposition IV.3 and its proof | subroutines._dense_synthesis._weyl, subroutines._dense_synthesis._local_factors |
| Three-CX circuit of exp(i(a XX + b YY + c ZZ)), with every rotation angle negated for Qiskit's rotation convention | Vatan and Williams, arXiv:quant-ph/0308006v3 | Sec. V and Fig. 6 | subroutines._dense_synthesis._two_qubit_plans |
| Two-CX circuit after a diagonal that makes the trace of gamma real | Shende, Markov and Bullock, arXiv:quant-ph/0308033v3 | Proposition V.2 | subroutines._dense_synthesis._two_qubit_up_to_diagonal |
| Block-ZXZ decomposition and demultiplexing | Krol and Al-Ars, arXiv:2403.13692v2 | Secs. 4.1 and 4.2, Eqs. (5)-(10) | subroutines._dense_synthesis._block_zxz, subroutines._dense_synthesis._demultiplex |
Demultiplexing a block-diagonal matrix, U0 ⊕ U1 = (I ⊗ V)(D ⊕ D†)(I ⊗ W) with U0 U1† = V D² V†, the top step for a controlled unitary |
Shende, Bullock and Markov, arXiv:quant-ph/0406176v5 | Theorem 12 | subroutines._dense_synthesis._demultiplex, subroutines._dense_synthesis.controlled_unitary_circuit |
CX count (25/96) 4**m - 2**m + 4/3 and instruction count (11/16) 4**m - 2**(m+1) of a controlled unitary on m >= 3 qubits |
NWQLib derivation from the recursion | _dense_synthesis module docstring (CX) and controlled_synthesis_size docstring (instructions) |
subroutines._dense_synthesis.controlled_synthesis_size |
| Closing CX of the outer multiplexors merged into the central block as CZ gates, similar to optimization A.1 of Shende, Bullock and Markov, arXiv:quant-ph/0406176v5 | Krol and Al-Ars, arXiv:2403.13692v2 | Sec. 5.2, Eq. (11) | subroutines._dense_synthesis._split, subroutines._dense_synthesis._append_gates |
| Diagonal moved between consecutive two-qubit blocks (optimization A.2) | Shende, Bullock and Markov, arXiv:quant-ph/0406176v5 | Appendix A | subroutines._dense_synthesis._apply_a2 |
CX count (22/48) 4**n - (3/2) 2**n + 5/3 of a generic n-qubit unitary |
Krol and Al-Ars, arXiv:2403.13692v2 | Abstract and Sec. 5.3 | subroutines._dense_synthesis.dense_unitary_circuit |
Work 113 M**3/4 + (m**2 + 16 m + 512) M**2, working bytes 320 M**2 + 65536 and kept bytes 352 M**2 + 16384 of one exact synthesis on m qubits |
NWQLib derivation from the recursion | dense_synthesis_size docstring |
subroutines._dense_synthesis.dense_synthesis_size |