Engineering constants¶
This page records the engineering constants of NWQLib: guards, tolerances, safety factors, schedule defaults and work and byte budgets. Each entry gives the value, the site that defines it, the reason for the value and the condition for revisiting it. The definition site carries a matching comment (contribution policy). Constants taken from papers, such as (23/48) 4^n and the Jacobi-Anger coefficients, are documented with their source equations in docstrings. Sites are paths from the repository root.
The entries distinguish derived formulas from untuned defaults. A finite tolerance can define which inputs a numerical check accepts, or a chosen approximation, without certifying the physical result. No default listed here authorizes an additional dense property check.
Defaults that change your results¶
These defaults decide the size, accuracy or stopping point of a computation. The linked section gives each value and its reason.
| Defaults | What they decide | Section |
|---|---|---|
ExecutionLimits: max_simulation_qubits, simulator_memory_mb, max_total_circuits, max_total_shots, max_direct_amplitudes, max_synthesis_work |
The width and memory of a local simulation, and the circuits, shots, direct state preparations and dense-synthesis work of one Run | Explicit workflow and reference controls |
DEFAULT_MAX_BYTES |
The byte limit of known arrays, stored Results, inspection and transport | Shared byte default |
ExpectationMethod.max_classical_products, ExpectationMethod.sampling_shots |
Classical work of an expectation value, and the shots per measured group for a requested sampling error | Expectation |
Lanczos krylov_dimension and overlap cutoff |
Size of the Krylov subspace and which projected directions the solve keeps | Chebyshev Lanczos defaults |
LCHS approximation_tolerance, kernel exponent, source-time quadrature and Low–Somma shift |
Error allocated to the kernel integral and its quadrature, and the work of the source term | LCHS defaults and error allocation |
QPE schedules of QCELS, SPE, RFE and RWPE, automatic time step and minimum_overlap |
Number of times, draws or steps, the evolution time and the verification overlap | QPE defaults |
| QLS target margin, condition-number floor and degree limits | The polynomial that approximates the inverse and its phase solve | QLS 1/x fit and phase pipeline, QLS construction and readout limits |
ADAPT-GCiM theta, t_user, max_iterations and stopping tolerances |
Selection rounds, the working basis and when ADAPT stops | ADAPT-GCiM defaults and tolerances |
| QHD Method defaults, box refinement and augmented-Lagrangian defaults | Grid, steps, time, schedule, initial state, refinement levels and the constrained outer loop | QHD Method and refinement defaults, QHD augmented-Lagrangian defaults |
| Example notebook sizes and settings | Problem sizes, qubit counts and run times of the examples | Example instance sizes |
Terms used on this page¶
| Term | Meaning |
|---|---|
| admission, admit | The check, before work starts, that accepts or rejects an input, a planned operation or its data against a limit or a domain |
| charge | The work or bytes that an admission counts against a limit |
| work unit | A counted scalar or element operation of a size law, used by limits such as max_work. It is a planning quantity, not measured CPU time |
| law | A formula that gives a gate count, work count or byte count from the problem sizes |
| allowance, envelope | An amount added to a law for objects the law does not count (allowance), or an upper bound (envelope) |
| census | A count of gates, Pauli terms, term pairs or stored objects that a law or an admission uses |
| preparation record | The record (PreparedArtifact) that a backend writes about the circuit or host kernel it prepared and executed |
| resource fold | The summation of per-block resource counts into the totals of a Program (src/nwqlib/resources/fold.py) |
| Run journal | The log file in which a Run records its rows (src/nwqlib/_run_journal.py) |
| acquisition | Running prepared circuits or a host kernel on a backend and collecting their data |
| qualified | Checked by measurement on the named interpreter and library versions |
| u | The binary64 unit roundoff, 2**-53 |
Shared limits¶
These limits and numerical windows are shared by the Run and by several Methods.
Explicit workflow and reference controls¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
execution.py::ExecutionLimits.max_simulation_qubits / simulator_memory_mb |
20 / 1024 MiB | src/nwqlib/execution.py |
Untuned local defaults. The qubit cap applies to Aer and the local NWQ-Sim backend before any register is allocated. The Slurm route is bounded by its own byte limits instead. A 20-qubit complex128 statevector needs 16 MiB. Aer receives the memory value as max_memory_mb, its cap on quantum-state storage, not a process RSS bound. Increase either for an explicitly intended larger simulation. |
_prepared_execution.py::Run.wait(poll_interval=1.) |
1 second | src/nwqlib/_prepared_execution.py |
Untuned pause between two resume passes. Each pass refreshes every pending provider job once, so the interval sets the status-request rate. It has no scientific meaning. Revisit for a provider rate limit or a latency requirement. |
execution.py::ExecutionLimits.max_direct_amplitudes |
65,536 | src/nwqlib/execution.py |
Largest declared direct magnitude/phase state preparation that the Run synthesizes, 16 qubits, the same default as prepare_qiskit. Its synthesis grows as q2*q. The Run checks every native block of the Plan before the first preparation, so an over-limit Plan fails before any charge, while planning and estimates are unaffected. "Explicit input operation defaults" lists the preparations it covers and the admissions that bound the others. Raise it with extend_limits for an explicitly intended larger direct preparation. |
execution.py::ExecutionLimits.max_synthesis_work |
1,000,000,000 | src/nwqlib/execution.py |
Largest total work of the exact dense syntheses that the Run makes while it prepares circuits, over all its preparations, in the units of _dense_synthesis.dense_synthesis_size. NWQ-Sim at optimization levels 0 and 1, Nexus, IonQ with its QIS gateset at optimization levels 0 and 1, IBM at optimization levels 0 and 1, and Aer with a noise model whose gate basis omits the unitary instruction synthesize the dense unitaries of a prepared circuit through the Run's synthesis cache (Run._exact_dense_unitaries), which keys each matrix by its shape, dtype and bytes. Each distinct matrix is synthesized once while the Run is open, and the Run reserves each new synthesis on the preparation's charge before the first one starts. A backend called directly with the Run, outside its preparations, is held to the rest of this limit but reserves nothing, so it neither reads nor adds to the cache. Lowering in a Run also reserves the syntheses of a controlled transformed block and Qiskit's control of them (gatewise_control_size). A noiseless Aer target applies the matrices and synthesizes nothing. On an Apple-silicon Mac with one thread (measurements under Circuit-free synthesis laws), the default takes 0.65 to 1.1 s when spent on syntheses on 8 and 9 qubits, and up to 15 s, at 1.6 to 15 ns per unit, when spent on syntheses on 4 to 7 qubits. It admits one dense unitary on 8 qubits (5.2e8 units, 0.57 s) and refuses one on 9 (4.0e9 units, 2.6 s). It also bounds the bytes of each synthesis, which grow as M2 while its work grows as M3, so at the default one synthesis holds at most 21.1 MB of working arrays and keeps at most 23.1 MB. The cache keeps each circuit, and the 16 M**2 bytes of its key, until the Run is closed or the Run object is released. No construction byte limit covers the cache. Planning and the subroutine builders charge the kept bytes of their own syntheses only, and those syntheses replace their dense unitaries during construction, before a backend sees the circuit. Every cached synthesis was reserved against this limit, and its kept law and key take at most 2.11 bytes per reserved unit on two qubits, 0.79 on three and 0.42 on four or more, so the default bounds the cache near 2.1 GB, 0.79 GB and 0.41 GB. A QLS inverse of a Hermitian matrix of 64 padded coordinates makes two syntheses of 7.0e7 units, once per Run for all its preparations. A reopened Run starts with an empty cache, and the syntheses that it makes again are reserved again. The reservation follows the preparation charge, because the dense unitaries of a circuit are known only after the Plan is lowered. A refused preparation stays charged against max_total_circuits and leaves its experiment unprepared. Raise the limit with extend_limits for an intended larger total. |
execution.py::ExecutionLimits.max_total_circuits / max_total_shots |
512 / 1,000,000 | src/nwqlib/execution.py |
The circuit cap applies separately to the number of circuit preparations and to the number of acquisition attempts. The shot cap bounds the total shots reserved, including those of uncertain attempts. Failed, reserved and uncertain acquisitions stay charged. Increase either cap for an explicitly intended workload, not to hide a failed admission. |
Shared byte default¶
src/nwqlib/_limits.py::DEFAULT_MAX_BYTES is the single editable default for known input/workspace arrays, amplitude access, stored Results, native inspection and compiler/model transport. Its value is 10,000,000,000 bytes (decimal 10 GB), sized for local research workloads. A demonstration-scale default such as 64 MiB rejects ordinary research inputs, for example a 67,403,776-byte source workload. Byte allowance is separate from work counts, timeouts and process RSS. Explicit smaller caller budgets still apply. Defaults are bound at module import, so restart Python after editing this file.
Byte-accounting conventions¶
src/nwqlib/evidence/_work.py::BOOKKEEPING_BYTES is H0 = 65536, an explicit engineering allowance for scalar bookkeeping, ndarray headers, iterators and fixed sort stacks on the checked 64-bit CPython/NumPy stack. Variable populations of Python objects are charged separately. It is not a universal interpreter or process-RSS theorem. Allocator fragmentation, library thread pools and native synthesis are accounted for separately. src/nwqlib/evidence/_work.py::integer_object_bytes is L(b) = 32 + 4 ceil(max(1, b)/30), a conservative charge for one Python integer of at most b magnitude bits on this stack: CPython uses thirty magnitude bits per four-byte digit, and the 32 covers its header above the observed 28-byte minimum. Numerical arrays use eight-byte float64, int64, uint64 and intp entries, sixteen-byte complex128 entries and one-byte Boolean entries. Stored payload, incremental workspace, peak simultaneous storage and archive size are distinct quantities. B_held denotes other already admitted allocations that stay live in the phase being checked, a shared buffer counted once. The energy-shift endpoint capture and relation laws and the Pauli plan classification gate below use these conventions. Revisit on another interpreter or NumPy stack, or when measurements qualify these allowances.
Exact evidence integer representation¶
src/nwqlib/evidence/_work.py::DEFAULT_MAX_INTEGER_BITS is 4096. ExactArithmetic checks prospective numerator and denominator widths, including unreduced cross products, before each operation. Every finite binary64 input fits individually, but composed operations may not. The guard can conservatively reject an operation whose reduced result would fit. It performs no clipping or approximation and is not a CPU, RSS or cumulative-work bound. Revisit with a concrete larger exact-scalar domain. See Error evidence internals.
Numerical guards and tolerances¶
Probability windows of exact execution¶
NUMERICAL_RELATION_RTOL = 1e-12 in src/nwqlib/_validation.py is a dimensionless input convention with three uses: the tolerance at the endpoint one for host-kernel probabilities and masses that have no derived window of their own, the floor of the circuit and expm_multiply windows below, and subset-mass comparisons. It is not a forward-error bound or a PSD certificate. Unnormalized subset masses scale the relative window by max(1, whole_mass). Revisit for changed precision or units, without enlarging a window in response to a failed check.
exact_probability_window(G, n, per_operation=c) in src/nwqlib/_validation.py is the admitted gap between one and the total of an exact binary64 statevector probability population, max(NUMERICAL_RELATION_RTOL, (c*G + 2**(n+1) + 10)*u). G is the native operation count in the preparation record (on Aer, the native evolution operations of a coherent readout, without save instructions), n the circuit width and u = 2**-53. PreparedArtifact.probability_window evaluates it for one preparation record and takes c from the record's target: AER_STATEVECTOR_ROUNDOFF_PER_OPERATION, about 6957, for Aer 0.17.2 and NWQSIM_CPU_SV_ROUNDOFF_PER_OPERATION, about 164, for NWQ-Sim CPU/SV. Any other target uses Aer's constant, the larger one.
Each constant is a first-order worst-case bound on how much one native instruction can change the squared state norm through binary64 roundoff. It is derived, not fitted. It charges the instruction its worst inexact fold into a fused matrix, its worst applied block and the entry error of a parameterized gate matrix, and doubles the sum for the squared norm. The comment above the constants in src/nwqlib/_validation.py derives each term from Higham's gamma_n = nu/(1-nu) (Accuracy and Stability of Numerical Algorithms, 2nd ed., SIAM 2002, doi:10.1137/1.9780898718027, Lemma 3.1) and the inspected simulator sources. An m-term complex inner product in any summation order has error at most sqrt(2)gamma_(2m) times the sum of its |products|. Above 14 qubits, its default fusion threshold, Aer folds instructions into fused matrices on at most five qubits. Each fused matrix starts from the identity, so its first fold is exact and G instructions give at most G - 1 inexact folds and at most G applied blocks. Aer's worst fold is a dense five-qubit instruction folded after another operation, about 2896 in 2-norm per unit u. Its worst application is a dense five-qubit fused block, 512, and its gate-entry term is about 70. The NWQ-Sim runner sends only U and CX gates, and NWQ-Sim's fusion (efd0226, include/circuit_pass/fusion.hpp) builds gates on at most two qubits, so its terms are about 45, 23 and 14. The readout term 2(n+1) + 10 covers summing at most 2N amplitude products into a probability total or Pauli expectation for N = 2n, by any summation tree or a compensated sum, and one multiplication of the state by a global phase. QHD pools the bins of several exact chunks with running sums and adds QHDAnalysis.pooled_summation_roundoff, (M + 2)u for M pooled bin values, to the largest chunk window before it compares its masses. On the default QCELS schedule of the eight-qubit H4 problem on Aer (nine qubits with the ancilla, qiskit-aer 0.17.2 on macOS arm64), circuits of up to 2,659,881 native instructions missed normalization by up to 6.1e-12, about 0.02Gu, which a fixed 1e-12 window rejects. Their windows reach about 2.1e-6, and no miss exceeded 3.6e-5 of its window.
The derivation assumes cos, sin and the complex exponential of an imaginary argument are accurate to one ulp. The preparation record's probability_window_exclusions lists what it does not bound, and the numeric window is the same with or without exclusions. For Aer the list names supplied matrix, state or channel data (unitary, diagonal, multiplexer, initialize, set_statevector and noise instructions), whose own unitarity or normalization defect is an input property. It also names mid-circuit measure and reset, which renormalize the state, control flow, whose executed instruction count is data dependent, and a phase-sum gate whose |parameters| sum exceeds PHASE_SUM_LIMIT, as "<name> phase sum". A dense unitary, kraus or quantum_channel on seven or more qubits also exceeds the per-instruction charge. These operations are listed wherever they occur in the native circuit, including after the readout. An amplitude readout of either simulator lists "amplitude-derived masses", and so does a trajectory of either simulator with a reduction point. A registered reducer whose own arithmetic bounds those masses resolves the label (Reducer.state_error_resolutions, section "Projected scalar reduction"). NWQLib forms the masses recorded with an amplitude observation through BLAS NRM2 (stable_vector_norm in reduce_amplitudes), whose summation algorithm depends on the installed BLAS, and QHD also decodes the saved amplitudes with NumPy absolute values and sums. Neither mass route is derived here. The QHD tie window (below) bounds only the decoder's individual point probabilities. An Aer release other than 0.17.2 lists "unchecked qiskit-aer version". An NWQ-Sim runner whose revision does not descend from efd0226 lists "unchecked NWQ-Sim revision", and a descendant whose derivation sources differ from efd0226 lists "changed NWQ-Sim roundoff sources" (_ROUNDOFF_SOURCES and _ROUNDOFF_BASE_REVISION in src/nwqlib/backends/nwqsim.py). An optimization level other than the adapter's default (0 for NWQ-Sim and IonQ, 1 for IBM and Nexus) lists "optimization_level", also on the record of a counts readout. The error numbers of such a record are a reference, and its operation count describes the compiled circuit. An empty tuple means the whole execution lies inside the derivation, and None, for host kernels and for counts at the default optimization level, means no assessment. Revisit the constants when the Aer or NWQ-Sim version changes its fusion, kernels or gate-matrix construction, or when a library path needs a bound for supplied matrices.
The QHD completeness checks and the bound above one use the circuit window (exact_probability_window) for probabilities, counts and Pauli expectations, QPE's saved trajectory expectations included, and the saved-state window (saved_state_probability_window) for amplitude-derived masses. An unavailable saved-state correction leaves those masses no finite upper bound. The exact LCHS and QLS scalar reduction checks its complete native norm against the producing preparation record's saved-state error (PreparedArtifact.saved_state_error, the qualified native-state error propagated through the envelope of the host phase correction of the saved state), after resolving its amplitude-derived masses label, when available and otherwise against the same window propagated through that correction (saved_state_probability_window), in both cases plus its host mass allowance (section "Projected scalar reduction"). Aer supplies its stored Python factor. NWQ-Sim supplies the factor reported by its selected runner during preparation and transmitted in the execution request. The same finite envelope applies to either represented factor. Multiplication by (1.0, ±0.0) has a zero envelope under the host arithmetic assumptions. The native first-order allowance remains part of the resulting saved-state estimate. A record alone does not carry its circuit size, so probability arrays, PauliValue and branch masses only require nonnegative or finite values, and QPE raw_mean only requires a magnitude above one. ObservationChunk.validate_unit_bound(receipt) requires a probability total (the math.fsum total saved when the arrays are published), Pauli magnitude or branch mass of at most one plus the preparation record's window. It runs where an observation is created from backend output and where a standalone Result is saved (saved_evidence._validate_recorded_data). QPE bounds raw_mean with the same window. QHD analysis (QHDAnalysis.validate_analysis_masses) and QLS classical analysis bound their reported masses by the window of the producing executions. That is NUMERICAL_RELATION_RTOL for a host kernel, except the QHD classical kernel, whose recorded expm_multiply window (next paragraph) applies.
Host-kernel state budgets¶
expm_multiply_state_error(calls, start=s) in src/nwqlib/_validation.py is the first-order 2-norm error budget delta of a host kernel that evolves a unit initial state of length D by SciPy's expm_multiply, which is Algorithm 3.2 of Al-Mohy and Higham (SIAM J. Sci. Comput. 33 (2011) 488–511, doi:10.1137/100788860). For each call with a Hermitian generator G it takes a bound N on ||G - (trace(G)/D) I||_1 and charges c (expm_multiply_call_roundoff), and delta = (s + sum(c + N) + 5*P)*u for P final phase products. The start term s bounds the construction error of the initial vector (src/nwqlib/algorithms/qhd/initial_state.py::restricted_state_error). state_mass_window(delta, D) turns delta into the mass window max(NUMERICAL_RELATION_RTOL, 2*delta + delta**2 + (D + 4)*u*(1 + delta)**2), and the QHD tie window below uses the same delta. The docstrings derive each term to first order in u under two assumptions that SciPy's own error control makes. Mathematics gives the per-call charge, the derivation, the start terms of the initial states and a measured comparison. The QHD classical kernel computes N from its stored schedule and support tables (method._host_generator_norms) and records the mass window as the probability_window argument of its KernelApplication, the counterpart of the instruction count that a circuit's preparation record holds.
The QHD split-step kernel has its own first-order budget, split_step.state_error, relative to the exact split-step product with the stored step weights, the exact kinetic eigenvalues at the grid's binary64 spacing and the exact sum of the support tables without the objective constant, whose global phase method._constant_phase restores after the probabilities are read. With delta = (s + sum_k rho_k)*u for the same start term s as above, each step charges rho_k = 2*d*5*L + (d + 2)*5 + 15*|dt*a_k|*sum_j max E_j + (M + 1)*|dt*b_k|*W in units of u, where L is the transform level count (split_step.transform_levels), M the number of support tables and W = sum_S max |v_S|. The terms are the transform assumption TRANSFORM_ROUNDOFF, one GLOBAL_PHASE_STATE_ROUNDOFF per phase multiplication, and the formation of the kinetic angles fl(fl(dt*a)*E) with E within EIGENVALUE_ROUNDOFF units and of the potential angles fl(fl(dt*b)/2*V_0) with the table sum V_0 formed in M - 1 roundings, since |exp(i*a) - exp(i*b)| <= |a - b|. The same state_mass_window gives its mass window, and probability_difference_window of this budget is the ceiling of its tie window. The kernel selects with the smaller budget that split_step.evolve observes on its trajectory, which charges each kinetic phase 15*|dt*a_k|*||E_j z|| for the normalized mode populations z after that axis's forward transform and each potential half |dt*b_k|/2*(2*||V_0 z|| + max(M - 1, 0)*W) on its own input, with each moment capped at its maximum above. The angle terms dominate when a kinetic integral is large. With ShiftedCubicSchedule(s=1e-3), the integrated rule and a five-point grid of spacing 1/2, the first kinetic angles reach 3e7 rad, the computed state from the kinetic ground state differs from a 30-digit evaluation of the same product by 4.3e-10, the observed budget is 3.6e-9 and the Plan's budget is 5.0e-8, while the transform and multiplication terms alone give 1.9e-14 (Python 3.12.14 and SciPy 1.18.1 on macOS arm64, with the objective x³/8 - x²/2 + x/4 + 3/8 on the interior grid of [-1.5, 1.5], three steps to T = 0.75 and a 30-digit mpmath evaluation).
QHD most probable point and tie window¶
The QHD most probable point is the lexicographically smallest valid grid point with positive probability whose computed probability p satisfies M - p <= w for the computed maximum M (method._summarize). The window w bounds how far the computed difference of two probabilities can lie from the difference in the model the path evaluates. For a unit state and a computed state within delta in 2-norm, up to a phase, the operator |i><i| - |j><j| has norm one, so the difference moves by at most 2*delta*sqrt(p_i + p_j) + delta**2 <= 2*delta + delta**2. Evaluating each probability with relative error e adds e*(1 + delta)**2, which gives probability_difference_window(delta, e) in src/nwqlib/_validation.py. method._readout_window chooses delta and e by path. The classical kernel uses e = 5u (ABSOLUTE_SQUARE_ROUNDOFF) with the delta that it publishes as its readout_state_error scalar, the observed split-step budget above or the Plan budget of the other flavors (method._host_state_error). Its KernelApplication records the window of the Plan budget as tie_window_ceiling (method._host_tie_window), and the recorded window must not exceed it (method._executed_host_window). Exact native chunks take delta from PreparedArtifact.state_error, which is native_state_error(G, c) = (c*G/2 + 5)*u, half of each doubled per-instruction constant plus one global-phase product, when the target has a derived constant c, the instruction count G is known, the preparation record assessed its probability_window_exclusions and the caller resolves every entry. Otherwise the preparation record reports the reason and QHD records no window and compares the computed probabilities exactly. Native probabilities use e = 2u (COMPONENT_SQUARES_ROUNDOFF). For C averaged chunks the window is mean_c E_c + rho_C*mean_c (1 + 2u)(1 + delta_c)**2 with E_c = probability_difference_window(delta_c, 2u), rho_1 = 0 and rho_C = C*u/(1 - C*u), because a grid point receives at most C contributions of at most C roundings each. Stored amplitudes use e = 5u, and the QHD decoder resolves the preparation record's "amplitude-derived masses" label for these point probabilities only. Their delta is PreparedArtifact.saved_state_error, the same budget propagated through the envelope of any host phase correction of the saved state (statevector_roundoff), and a preparation record whose host correction was not assessed records no window. Counts are pooled as integers and compared exactly, so w = 0. Evaluated from these formulas, not measured, w is 3.2e-13 for one variable with K = 3 and two host steps, 9.0e-12 and 1.3e-11 for a 3-by-3 grid with 80 host steps at T = 1 and 10, 3.2e-8 for the two-step K = 39 grid above, and 2.8e-11, 7.6e-11 and 1.5e-10 for exact Aer probabilities of 36, 98 and 190 native instructions. Each is a first-order worst-case bound, not a measured error. Revisit with the constants it reuses, or when a readout sums several amplitudes into one bin.
The mode representative is the first observed valid grid point with positive weight in lexicographic order whose computed probability deficit from the maximum over that same population passes the numerical tie window. Its computed probability can therefore be below the maximum. Read the reported deficit and tie window alongside the point. If the window is unavailable, the implementation compares the computed probabilities exactly and records why the numerical allowance is unavailable. With counts, the comparison uses empirical frequencies with a zero window and gives no sampling-error guarantee.
mode_status is resolved when the tie rule admits only one positive observed valid point, unresolved when it admits another, and unavailable when there is no positive valid outcome or no derived window. A zero deficit can still have unresolved status.
If every pairwise difference of the computed probabilities has error at most the window, Proposition 45 bounds the model-probability deficit of the selected point.
Roundoff constants¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
AER_STATEVECTOR_ROUNDOFF_PER_OPERATION |
2*(2896.3 + 512 + 70.3), about 6957, in units of u | src/nwqlib/_validation.py |
Derived first-order change of the squared state norm per Aer 0.17.2 native instruction: worst five-qubit fold, worst five-qubit applied block and gate-entry term (paragraph on exact_probability_window above). Revisit when Aer changes its fusion, kernels or gate matrices. |
NWQSIM_CPU_SV_ROUNDOFF_PER_OPERATION |
2*(45.3 + 22.6 + 14.1), about 164, in units of u | src/nwqlib/_validation.py |
The same bound for NWQ-Sim CPU/SV at efd0226 and its descendants with the same derivation sources (_ROUNDOFF_SOURCES in src/nwqlib/backends/nwqsim.py: value type and fusion switch, circuit gate records, gate-matrix storage, fusion pass, U/CX matrices and CPU/SV kernels), whose runner sends only U and CX and whose fusion builds gates on at most two qubits. Revisit when upstream changes those files. |
EXPM_MULTIPLY_TERMS_MAX / EXPM_MULTIPLY_THETA_MAX |
55 / 9.9 | src/nwqlib/_validation.py |
SciPy's m_max and the double-precision theta_55 of Al-Mohy and Higham's Table 3.1 (doi:10.1137/100788860), which fix the largest Taylor degree and the substep size of expm_multiply. Paper and SciPy constants mirrored for expm_multiply_call_roundoff, not tuning choices. Revisit when SciPy changes its m_max, theta table or algorithm. |
GATE_ENTRY_ERROR |
10, in units of u | src/nwqlib/_validation.py |
Entrywise relative error of a parameterized gate matrix under one-ulp cos, sin and complex exponential: at most 5u for Aer's constructions and 9.83u for NWQ-Sim's U gate. Revisit if a simulator composes more factors per entry. |
PHASE_SUM_LIMIT |
4*pi | src/nwqlib/_validation.py |
Largest sum of absolute parameter values of an Aer phase-sum gate (u, u2, u3, cu, cu2, cu3, mcu, mcu2, mcu3) inside the derivation. It admits three angles reduced to [-pi, pi]. A larger sum is listed as an exclusion instead of widening every window. Revisit if library constructions produce larger unreduced angles. |
GLOBAL_PHASE_STATE_ROUNDOFF |
5, in units of u | src/nwqlib/_validation.py |
2-norm change of a unit state multiplied by one global phase exp(iphi). A one-ulp phase entry contributes 2u and the complex product sqrt(2)gamma_2, 4.83u in total, rounded up. exact_readout_roundoff charges it twice for the squared norm, native_state_error once per execution and expm_multiply_state_error once per final phase product. Revisit with the complex exponential or complex-product model. |
COMPONENT_SQUARES_ROUNDOFF / ABSOLUTE_SQUARE_ROUNDOFF |
2 / 5, in units of u | src/nwqlib/_validation.py |
Relative error of one probability evaluated from one amplitude, as used by the QHD tie windows and the host mass window. re*re + im*im (Aer full-register probabilities, NWQ-Sim runner probability bins) makes two products and one nonnegative addition, gamma_2. np.abs(z)**2 (QHD host kernel and amplitude decoder) is (1 + 2u)**2 (1 + u) - 1 to first order, assuming NumPy's complex absolute value is accurate to 2u for normal nonzero magnitudes. Revisit when a readout sums several amplitudes into one bin or NumPy changes its complex absolute value. |
Finite search frontier envelope¶
nwqlib.search.scan defaults to at most 256 candidates, 4096 objective values, 4096 forecast rows, 1000000 ordered coordinate comparisons and 4096-bit exact integers. For N rows and k objectives, admission bounds N*k values and N*(N-1)*k coordinate comparisons before evaluation. Rows that prove incomparable or dominated early still count toward those finite limits. SearchSelection applies the same value, pair and integer controls when validating stored values. The bounded quadratic frontier keeps strict-dominance ties. These controls are operation-size ceilings, not scientific tolerances or CPU/RSS predictions. Increase the relevant explicit limit for a larger supplied population, or replace the frontier algorithm if its quadratic cost becomes material. Existing ExactArithmetic checks intermediate rational products, and the profile code bounds the model expansion before prediction. Stored rescoring checks the supplied forecast population without reevaluating its models. Its original scan counters are not rewritten as new work.
Inputs¶
These limits govern reading and converting the matrices, operators, circuits and fermionic input that a Problem supplies.
Explicit input operation defaults¶
Three explicit classical operations on admitted inputs have default ceilings. They are policy defaults that give an operation called without explicit limits a finite size. They are not counts derived from a workload, accuracy settings or CPU-time predictions. The caller can change each one. An operation above its ceiling is rejected, and the ceiling is never raised automatically.
src/nwqlib/problems/inputs.py::prepare_qiskit(max_direct_amplitudes=65_536)limits direct magnitude/phase synthesis to 216 amplitudes, that is 16 qubits. Its classical arithmetic grows as q*2q, about 106 operations at the limit, and its circuit has up to2*(2**q-2)CX, 131,068 at q = 16 (_preparation_laws.direct_preparation_cx_bound). It does not limit compact product or occupation preparations or the dimension of an admitted input. The value issrc/nwqlib/_limits.py::DEFAULT_MAX_DIRECT_AMPLITUDES, shared withlower_qiskit(max_direct_amplitudes=...)andExecutionLimits.max_direct_amplitudes. Selection, planning and the circuit-free estimate accept larger direct preparations, because the direct CX law needs no synthesis. A Run applies its own limit instead of this default to every block that declares aqiskit.directpreparation, which means its constructor synthesizes that state with the direct magnitude/phase tree ofsrc/nwqlib/subroutines/state_preparation/direct.py(src/nwqlib/blocks/selection.py::direct_preparation_amplitudes). These are the native PREP leaf throughprepare_qiskitwith its controls and adjoints, the ADAPT query reference and the LCHS vector PREP (src/nwqlib/algorithms/lchs/native.py::construct_preparation). The Run checks every native block of the Plan before the first preparation and passes its limit through lowering to the constructors that callprepare_qiskit. Three undeclared uses of the same tree are bounded by their own admission instead. The LCU coefficient PREP ofsrc/nwqlib/subroutines/lcu/core.pyis admitted by_admit_lcu,16*P*(a+1)work againstmax_workand 128 bytes per address for P = 2a addresses, which admits up to 218 addresses at the defaultmax_workof 108 and 221 at the QLS default of 109. The band-address PREP ofsrc/nwqlib/subroutines/block_encoding/banded.pyis admitted bysrc/nwqlib/subroutines/block_encoding/core.py::_banded_plan,64*2**a*(q+16)bytes and32*2**a*q**2work, which with at most 2q distinct bands allows up to 213 addresses at amax_workof 108 and 216 at the QLS default of 10**9. The LCHS source-branch preparations inside SELECT (src/nwqlib/algorithms/lchs/native.py::_source_preparations) have the padded dense LCHS dimension D, which the dense input admission (src/nwqlib/algorithms/lchs/selection.py::select_dense,64*D**2 + 128*Dbytes) bounds before planning continues, at most 8192 amplitudes at the defaultmax_bytes. When the dissipative part L is nonzero, the eigensolve (D**3 <= max_spectral_workfor the unpadded dimension) lowers this to 512 amplitudes by default. MPS preparations are not this tree and follow their MPS envelopes. The explicit preparations of QPE and ADAPT verification keep this default. Revisit for a workflow that needs a larger direct preparation outside a Run.src/nwqlib/operators/inputs.py::OperatorInput.matvec(max_products=1_000_000_000)limits the admitted work of one action, D² for dense storage, nnz + D for CSR/CSC and the grouped-action law ofsrc/nwqlib/operators/_pauli.py::pauli_action_requirementsfor Pauli terms. The default is a fuse against runaway planning work, sized so that no example or test workload in the repository reaches it. For N Pauli state coordinates, L stored terms and b state columns, the shared action allowance is(L+g) + 3*L + N*(L+g*b+b) + N*b + L*N*b, where g bounds the number of distinct flip masks. The 3L preprocessing visits are shared by all columns. The grouped pass visitsN*(L+g*b)term-phase and group-column entries, with metadata and output initialization charged separately. Admission also reserves the complete ordered overflow retry and its output reset. Label inputs addL*(q+1)decoding work for width q._matvec_requirementsuses one column andg = min(L, N)from metadata. AtN=2**20, the term allowance follows this full work law and the caller'smax_products. These are conservative kernel/pass units, not measured CPU work. The same 109 is the default of several other explicit classical work ceilings (DEFAULT_MAX_BLOCK_WORK,DEFAULT_MAX_LCU_WORK,FactorizedOperatorProduct.expandand the MPS contractions).DEFAULT_CONVERSION_WORKis 108. Methods that act repeatedly pass their ownmax_products. Revisit for an explicit action that needs more.src/nwqlib/operators/_pauli.py::PauliTerms.group(max_comparisons=1_000_000)is the running comparison budget of first-fit grouping. Grouping admits each candidate tile before evaluating it.comparison_countcounts all candidates in evaluated tiles. For L input terms it is at most L(L−1)/2. A refusal names the actual grouping work and the next-tile minimum, which is necessary to continue, then reports L(L−1)/2 as the value sufficient for complete grouping in this phase. Grouping can succeed below that envelope when its actual comparisons fit. Later phases using the same option are admitted separately. General commuting grouping charges actual member-pair tests and can reach that quadratic count even with one group. An all-Z table of 68,000 terms forms one group with 67,999 comparisons and is admitted at the default cap, while 68,000 terms that each form a different group need68000*67999/2=2,311,966,000comparisons and a smaller budget refuses them. Its library callers pass their own budgets and the name of their field for the refusal to name. The QWC grouping of a sampled quantum ADAPT plan's inputs insrc/nwqlib/algorithms/gcim/adapt_inputs.pypasses the Method'smax_products, the QWC partition of a sampled FixedGCIM plan insrc/nwqlib/algorithms/gcim/fixed_basis.py::sampled_groupspassesFixedGCIM.max_analysis_work, the grouped sampled readout ofsrc/nwqlib/algorithms/expectation.py::_measured_programpassesExpectationMethod.max_classical_products, and the sampled readouts of QLS and LCHS, which group throughsrc/nwqlib/_quantum_readout.py::qwc_groups, forwardlimit_name="QLS.max_work"(src/nwqlib/algorithms/qls/quantum.py::_readout) andlimit_name="LCHS.max_readout_work"(src/nwqlib/algorithms/lchs/quantum.py::lchs_readout). Classical and exact quantum ADAPT plans read no readout groups and form none. The default applies only to direct calls. Revisit for a workload whose first-fit group count makes the comparison count impractical.
Input cost controls lists the byte laws checked beside them.
Supplied circuit snapshot allowances¶
src/nwqlib/blocks/_qiskit_intake.py::snapshot_circuit charges a census against max_bytes before it copies or shares each object of a supplied circuit. Shared objects are charged once. The census has two kinds of terms.
Stored data is charged at its size. An ndarray parameter costs its array bytes. Its one-byte-per-entry finiteness mask is freed before the array is copied and is never larger than the copy, so the copy's reservation covers it. A scalar parameter costs one complex128, 16 bytes, or the byte length of a larger integer. A control modifier adds the bytes of its count-bit control state, which the SDK may form as 2**count. A UCGate costs 8 bytes per control index. A StatePreparation whose amplitudes are all Python floats or complex numbers has them admitted together, with the same 16 bytes per amplitude and one finiteness check of their complex128 array. A StatePreparation also costs 16 bytes per parameter for the constructor argument that its inverse is rebuilt from, which shares the admitted amplitude list, and a label-form StatePreparation one more byte per label character.
The rest of the snapshot is Python and SDK objects, whose size the byte convention excludes. A census of stored data alone would let a long circuit of parameterless gates pass any limit, so each object also receives a fixed nominal allowance. A definition circuit costs 32 bytes plus 16 per qubit, and a stored instruction 16 bytes plus 8 per qubit operand, one 64-bit index per operand, with barriers included. A gate object costs 32 bytes, including an immutable SDK gate that the snapshot shares instead of copying. An annotation costs 32 bytes plus 8 per modifier slot, and each modifier 32 bytes. These are policy allowances that make the census grow with the number of copied objects and operands. They are not measured object sizes, SDK synthesis work or RSS. src/nwqlib/problems/inputs.py::state_input reuses the 16-byte per-qubit allowance for its circuit-width check. Revisit the nominal allowances if the census must predict memory, which needs measured sizes of the SDK's stored circuits and instructions. Revisit the data terms when the supported native classes or their stored payloads change.
Eigen input conversion and classical preparation¶
src/nwqlib/algorithms/_eigen_inputs.py::AUTO_DENSE_PAULI_DIMENSION=16 is the existing small explicit-dense convenience limit shared by Lanczos, FixedGCIM, ADAPT, the quantum ExpectationMethod path and the LCHS observable readout (src/nwqlib/algorithms/lchs/quantum.py::observable_terms and src/nwqlib/algorithms/lchs/selection.py::plan_lchs, row "dense observable dimension at most 16" below). It is not a Problem dimension limit and does not restrict compact Pauli inputs. src/nwqlib/algorithms/expectation.py::_terms reads this constant for an explicit dense observable and offers no explicit conversion option, so a larger observable is supplied as Pauli terms or evaluated with classical execution. The metadata descriptor in src/nwqlib/algorithms/expectation_metadata.py repeats the value 16 in its limitation text, which must be edited together with the constant. DEFAULT_CONVERSION_WORK=100_000_000 bounds q*D² traffic in an explicitly selected dense block transform before allocation, where D includes padding. Revisit with a concrete conversion workload and measured implementation change.
Gershgorin normalization uses component scaling before row magnitudes, zeroes the diagonal magnitudes before the reduction, gives empty compressed rows or columns zero radius, and uses the outward allowance k*eps/(1-k*eps) with k = 8*n+16, where n counts a row's stored off-diagonal values, including explicit zeros, and eps = 2**-52 is twice the unit roundoff, so the allowance is wider than gamma_k. It covers sequential or pairwise positive summation under round-to-nearest binary64 arithmetic, gradual underflow and a one-ulp hypot assumption. The final half-width covers both outward-rounded endpoints about the stored center. This is a conservative numerical allowance, not a directed-rounding proof. Scalar inputs bypass this margin exactly. No eigensolve estimates the enclosure. Gershgorin frame construction admits 32*N+256*D additional array-data bytes and a 65,536-byte fixed bookkeeping allowance. Here N includes every stored entry for CSR/CSC, including explicit zeros, and equals D squared for dense storage. The input matrix is excluded. The array law covers the simultaneous magnitude, coordinate, reduction and endpoint arrays. The fixed allowance is qualified for the declared native stack and is separate from process RSS and external library workspace. The scale expression peaks at three float arrays per stored value, and the sparse major-coordinate and mask arrays are released before the row endpoints are assembled. Revisit the fixed allowance when the interpreter, NumPy/SciPy or the allocation schedule changes.
preparation_requirements counts product expansion as 2*D-2 amplitude products and a supported k-qubit standard gate as at most D*2**k products. Custom or opaque circuit action remains unknown. Three-vector recurrence admission includes its original matvec envelope, affine-copy storage, m actions and scalar reductions before preparation. These laws count known arrays/products, not process RSS or SDK decomposition time. Revisit when the numerical kernels change.
Explicit EnergyEndpoint capture in src/nwqlib/evidence/energy_shift.py::_operator_payload admits the borrowed operator storage plus the scalars and characters of the table in the EnergyEndpoint record that its two callers, the Lanczos and FixedGCIM energy_endpoint, build from the returned table, and the table's part of the two JSON copies that an identity hash makes. The record's fixed fields are not counted. A matrix endpoint costs 32+2*(6+2*digits+48) bytes per scanned entry, with N=D² scanned entries for stored dense input and N=nnz for CSR/CSC. The 32 bytes are the two indices and two binary64 parts of each (row, column, real, imag) tuple, which record validation and model_dump share. Each JSON copy of the entry has at most 6 punctuation characters, two indices of digits = len(str(D-1)) digits and two binary64 reprs of at most 24 characters. A Pauli endpoint costs 3q+68 bytes per term, the largest of three peaks. The first holds the materialized label, its complex coefficient and the kept real part (q + 24). The second holds the kept label and real part and the label copy of record validation (2q + 8). The third holds the record's label and real part and two JSON copies of at most q + 30 bytes. Both laws add 4 bytes for the brackets of the two copies. A matrix capture also admits its extraction arrays, B_extraction = 41N + 24D + H0 (dense input omits 24D): for CSR/CSC, arange(D), diff(indptr) and a possible intp repeat-count conversion need at most 24D bytes, the repeated major-coordinate vector 8N, its keep mask N, and the gathered row, column and values at most 32N. Dense input needs at most D² mask bytes plus 32 bytes per surviving entry. The complete matrix capture law is B_operator + N{32 + 2[6 + 2 digits(D-1) + 2F]} + 4 + 41N + 24D + H0 with F = 24. Endpoint capture admits coordinate expansion, zero filtering and gathered values in addition to its tuple and JSON payloads. Extraction preserves the input representation's established entry order. The exact relation check (_relation_bytes) admits linear coordinate, value, grouping and NumPy sort workspace before allocation. A matrix relation of U = N_base+N_target stored entries admits 192U + H0 bytes: the 160U logical envelope of converted, concatenated and sorted coordinates, permutation, group maps, union values, masks and gathers, plus 32U of NumPy 2.5.2 lexsort copy and merge buffers. Coordinates beyond int64 borrow the endpoint's own Python integers, so no coordinate objects are boxed. A Pauli relation of N = N_base+N_target+1 contributions on q qubits admits [16 max(1,q)+256]N + H0: four Unicode label populations of at most 4 max(1,q) N bytes and 256N for values, grouping and sort workspace. Endpoint terms are unique per word, because every public Pauli operator input is coalesced before an endpoint reads it. Missing matrix diagonals are counted without materializing them. H0 is the byte-accounting convention above. The capture laws exclude Python object overhead, that is tuple, list, dict and set slots and object headers, except the scalar bookkeeping and ndarray headers that H0 covers. The relation laws charge NumPy sort workspace (the 32U lexsort buffers of a matrix relation and the sort part of 256N for a Pauli relation), the scalar bookkeeping and ndarray headers in H0. They exclude other Python object overhead. The check walks original storage, never a sparse D² realization. Revisit if endpoint storage or the NumPy grouping and sort implementation changes.
Fermionic input from XACC¶
XACC input (src/nwqlib/subroutines/xacc.py) supports at most four ladder operators per term (two-body Hamiltonians), bounding JW expansion to at most 16 branches per coalesced term. Revisit for an explicit higher-body workload and resource contract. Caller-supplied num_modes bounds written indices before expansion. XACC normal ordering keeps every nonzero accumulated float64 coefficient. The final Pauli coefficient_cutoff defaults to zero and is explicit when pruning is wanted. Existing fermionic-pool callers keep their 1e-12 defaults (GENERATOR_COEFFICIENT_CUTOFF), applied after stable coalescing. Normal-order and Pauli coalescing buckets keep O(contributions) scalar workspace until each stable sum is complete, rather than only the final O(unique labels) coefficients. Pruning individual contributions would change the kept operator. Revisit storage only with an equally accurate summation method and an explicit workspace budget.
Operator and input representation sizes¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| Original Hermitian input / native unitary predicates | Exact Hermitian equality. Native unitarity uses the entrywise absolute atol of is_unitary, 1e-8 by default |
src/nwqlib/problems/records.py, src/nwqlib/subroutines/_matrix_checks.py |
Generic Eigenproblem rejects a non-Hermitian original input. The caller can explicitly symmetrize it. Native unitarity and computed projected-Gram roundoff have separate domains. |
subroutines/_matrix_checks.py::is_unitary |
atol=1e-8, rtol=0 |
src/nwqlib/subroutines/_matrix_checks.py |
Untuned entrywise window for U†U-I in explicit dense-unitary builders. Unitarity fixes the operator scale, so the tolerance is dimensionless. This dense predicate is not used to check the Hermitian input that structured QSP assumes. Revisit condition: Precision, admitted matrix dimensions or a derived arithmetic error bound. |
_PAULI_TILE |
1024 rows or amplitudes | src/nwqlib/operators/_pauli.py |
Tile length C of the label decoder, the candidate blocks of PauliTerms.product and PauliTerms.group and the coordinate tiles of apply_terms. For N state coordinates, b columns and L stored terms, put t = min(N, C) and supply a bound g on the number of distinct flip masks. pauli_action_requirements admits a*N*b + 24*L + 16*g + 8 + max(9*L, (64+32*b)*t) bytes, with a = 33 for an already complex128 input and a = 49 when conversion is possible. Grouped tiles use at most (64+16*b)*t explicit scratch bytes and an ordered overflow retry uses (64+32*b)*t. The preceding AND/popcount frontier uses 9L bytes and is released before either tile phase. Label inputs add 32L packed-array bytes that remain live through preprocessing and action. The metadata allowance counts logical payloads, excluding Python object overhead and RSS. Changing C changes storage, not labels, order or scientific records. Revisit with a changed workspace budget or kernel representation. |
operators/refinement.py::_scalar_frontier |
32 integer slots of 3174+bit_length(count) bits |
src/nwqlib/operators/refinement.py |
A sum of count finite binary64 values has a numerator of at most 2098+bit_length(count) bits over the common denominator 2**1074. Two more bits cover the Gershgorin radius and the diagonal offset, and a cross comparison adds 1074. At most 29 such integers are alive at once, the 22 integers of eleven live Fractions in the sparse Gershgorin scan and seven temporaries of one CPython Fraction addition or subtraction. The Pauli scan peaks at 19. Thirty-two slots cover both. Integers inside C-level arithmetic and Python object overhead are excluded. Revisit condition: A changed scan algorithm, exact-scalar representation or CPython Fraction arithmetic. |
algorithms/_eigen_inputs.py |
(64+q)*D**2 bytes |
src/nwqlib/algorithms/_eigen_inputs.py |
Four complex128 matrix-sized buffers plus q bytes per potential Pauli label cover the explicit padded conversion's known data. This excludes Python object overhead and vendor workspace. The caller opts into larger or sparse-to-dense conversions. Revisit condition: Transform storage or label representation changes. |
Programs and lowering¶
These limits bound the Programs that Methods build, the Qiskit circuits built from them and the circuit-free resource counts.
Shared Program admission limits¶
src/nwqlib/ir/expressions.py::AdmissionLimits defaults to 16,384 kept definitions, expressions, parameters, registers, classical values and signatures (the ir/records.py::Program row of Record identity JSON gives its reason), depth 128 (also the hard recursive lifecycle ceiling), 100000 kept field slots and 100000 expression/lifecycle work steps per check, and 4096 bits per integer. These guard finite metadata admission. Repeat counts and range lengths are not expanded. The limits admit the builtin Methods' selected Programs at the sizes measured below while bounding accidental graph/context explosion. Revisit with a concrete larger kept graph and its memory/context inventory, without raising them merely to admit dynamic instruction materialization. max_steps checks kept inventory separately from evaluation/lifecycle work. The latter includes each distinct context's width scan, context/cache collection slots and correlation-member processing, charged before that work. Identical partitions are reused. Genuinely different branch partitions traverse their finite member incidences once. Per-check interning of immutable correlation sets avoids duplicate member storage and repeated full-set equality comparisons per wire. These work units bound kept admission collections. They are not CPU-cycle or measured byte totals.
QLS(max_admission_steps=...) sets max_steps of the quantum QLS Program. Its default, 1,000,000, is ten times the shared max_steps default, because the admission work of a sampled QLS Program grows with the number of its distinct measured registers, as the QLS sampled count registers row states. Planning checks the largest of the kept field slots and the admission work of Program validation, preparation and native lowering of the selected Programs, and a complete measured refusal names a value that admits them (QLS guide). FixedGCIM, ExpectationMethod and LCHS have the same field with the same default for their Programs. A refusal there names the field, the refused stage and its count, either the complete kept-field-slot inventory or the admission work reached when the check stopped, a lower bound. Neither promises a value that admits the later admission, preparation and lowering stages. The shared default bounds the Programs of every other Method.
The selected resource fold reuses Program's AdmissionLimits for its kept inventory, definitions, depth and integer growth. Its work limit is max(max_steps, FOLD_WORK_PER_ADMISSION_STEP * a), where a is the admission work the Program check just measured and FOLD_WORK_PER_ADMISSION_STEP = 1 + len(COUNTS) = 24 in src/nwqlib/resources/records.py. The factor comprises one traversal unit plus one unit for each of the 23 count metrics, 24 in all. The fold repeats the admitted traversal and carries one cost per count metric, so a fold proportional to its Program fits, while a fold that grows faster is rejected. Measured with Qiskit 2.5.2 on macOS arm64, the fold's total work, its admission included, was 4.7 to 5.9 units per admission step for the QLS Programs of 2D Laplacian grids 4 x 4, 8 x 4 and 8 x 8 (inverse polynomials of degree 59, 103 and 239) and 12.1 to 13.5 for H4 QCELS Programs with the default schedule and with tau = 0.2 and longest times 10 and 40, in the default and CX resource contexts. A fold limited to the Program's own max_steps would reject Programs that admit, plan and execute normally, for example the degree-103 QLS Program of an 8 x 4 grid under the shared default max_steps. WorkloadEstimate admits work_units <= 24 * max_steps, the largest limit an admitted Program can reach. The fold's combined traversal, expression output and cache bookkeeping are charged against this limit. It checks selected law integer-growth bounds before power-of-two construction and rejects excessive kept expression definitions/depth before returning a WorkloadEstimate. These are bounded metadata work units, not CPU or RSS limits. No additional default production benchmark or materialization budget is implied. The registered native CX adapter reserves the integer bit length of 32 times max(1, factor) times 2exponent before evaluating the registered law. Controlled SELECT has leading coefficient 24q + 8 <= 32max(1, q). Controlled direct PREP is <= 16*2q, including its uniform fast path. The five reserved factor bits therefore bound these current kernels. This conservative admission can reject a near-limit result whose tighter actual integer would fit. Revisit the envelope if a registered native law changes. Both kept-metadata walkers reserve collection lengths before copying their children, so the same max_steps envelope bounds frontier admission as well as later scalar visits. The small test of this reservation uses 129 already-supplied scalar entries against 128 slots. It is not a memory allocation benchmark.
Selected block lowering limits¶
src/nwqlib/blocks/lowering.py::lower_qiskit defaults to max_operations=100_000 dynamic IR visits, counting Repeat multiplicity, and to max_qubits=4096 and max_clbits=4096 for the total declared quantum and classical register widths. It checks both widths and then counts the visits before Qiskit is imported, so an oversized selected construction rejects before any register or native gate exists. The same counting pass compares each reachable declared direct state preparation with max_direct_amplitudes, whose default, reason and covered preparations are under "Explicit input operation defaults". Emission then walks the same visits, appending one instruction for each BlockCall and one for each register qubit of a Measure or Reset, so the visit limit bounds the emission loop and, up to those register widths, the instruction count of the lowered circuit. Before emission controls the realized base of a controlled transformed block, lowering adds the work of its exact syntheses and of Qiskit's control of them to a running total, which max_synthesis_work=1_000_000_000 bounds. That default equals ExecutionLimits.max_synthesis_work, and a Run charges the same work to its own limit instead.
These values are policy defaults, not counts derived from a construction. They give one explicit lowering a finite size when the caller sets no limit. The operation default equals the instruction ceilings of explicit native inspection (inspect_resources in src/nwqlib/_prepared_execution.py and inspect_circuit_resources in src/nwqlib/backends/inspection.py), QASM export (src/nwqlib/io/qasm.py::export_qasm) and NWQEC compilation (src/nwqlib/backends/nwqec.py). A Run lowers each preparation through _lower_qiskit without passing max_operations, so every Run uses the 100,000-visit default and has no control to raise it. Its width limit is min(4096, max_simulation_qubits) for a local simulator and 4096 otherwise. The limits bound one lowering, not a backend capacity. Revisit them when a selected construction that a Run must execute needs more than 100,000 visits, which also needs a Run-level control, or registers wider than 4096.
Block construction sizes¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
blocks/selection.py::select_zero_reflection |
(q+1)*2**q work |
src/nwqlib/blocks/selection.py |
Conservative logical construction envelope for the selected diagonal reflection, including address-dependent synthesis. It is neither a minimum gate count nor measured SDK runtime. transform_block adds the same law at width q+1, (q+2)*2**(q+1), to the base's (q+1)*2**q for a controlled reflection. Revisit condition: The selected reflection constructor changes. |
blocks/selection.py::_signed_pauli_requirements |
m*(q+16)+32*2**a bytes and m*(q+1)+(q+1)*(a+1)*2**a work |
src/nwqlib/blocks/selection.py |
For m terms on q qubits and a index qubits, the first byte term is the PauliTerms.labels law for the labels that the native constructor materializes during lowering. The 32 bytes per label slot bound the selection's m coefficients and amplitude array together with either two float64 temporaries or the tobytes copy. The coefficients are frozen as a read-only float64 array before the amplitudes are built, so coefficients (8m), amplitudes (8P) and at most two float64 temporaries (16m) give 24m+8P <= 32P. Coefficient freezing peaks at 16m before amplitudes exist and amplitude freezing at 8m+16P <= 24P, in sequential phases. This law belongs to the selection, not to the archive. The work is the reported construction_work size law, in table-entry visits. It counts reading m labels and coefficients and, for each of the q per-qubit UCG tables and the sign diagonal, one pass that builds the table padded to 2a entries and one pass per address bit. For a UCG table that pass is the dependency test. For the diagonal it is one halving level, whose differences are lowered as an RZ multiplexor with a Walsh butterfly. The projected copy, the UCGate padding and checks, the diagonal's phase wrapping and the butterflies add a bounded number of passes of at most 2a entries per table and per address bit, so the law counts these visits up to a small constant factor, not exactly. These are logical allowances, not measured SDK synthesis work. Revisit condition: The SELECT constructor, its dependency projection or its coefficient preparation changes. |
blocks/encoding.py::_select_planned_encoding |
32*size*max(1,n) |
src/nwqlib/blocks/encoding.py |
Conservative table-construction work for structured encodings, with size=2**a*max(1,n). The coefficient 32 is an untuned arithmetic allowance. Dense dilation has its separate cubic law. Revisit condition: A changed constructor or measured operation accounting. |
blocks/_archive.py::write_blocks |
signed Pauli SELECT and readout data 16mw + 16m + 8P bytes |
src/nwqlib/blocks/_archive.py |
With m terms on q qubits, w = ceil(q/64) and P = 2a padded amplitudes, the archive stores the packed operator (16mw + 16m) and one shared amplitude file (8P), plus four NPY headers of at most 256 bytes and the actual graph and manifest JSON. Loading derives the coefficients from the restored operator. Shared SELECT/readout payloads and an already written operator are counted once by object identity. With contiguous arrays the NumPy 2.5.2 NPY writer holds at most one data chunk of min(nbytes, 224), so the numerical write-phase peak is B_held,other + B_op + 8m + 8P + min(max(8mw, 16m, 8P), 2**24) + H_archive, where H_archive, the graph serialization and writer bookkeeping, is not fixed for an arbitrary graph. write_blocks takes no byte limit. Revisit condition: The block archive layout or the NPY writer changes. |
core/planning.py::Plan._resolve_selection and Plan._selected_construction |
one recent entry per operation, two slots per Plan | src/nwqlib/core/planning.py |
Each slot keeps its latest successful pair and requested bindings, and recognizes canonical bindings through the stored Program. A miss replaces that slot before resolving the next point. This bounds the number of cached graphs while preserving adjacent reuse of a point's resolution and construction. Evicted points repeat their normal admitted resolution work. The cap is an engineering policy and does not bound the bytes or CPU cost of an arbitrarily large point. Revisit condition: A qualified prospective per-point byte inventory permits admission against the Method's remaining max_bytes, or measured reuse justifies a larger byte-admitted working set. |
Direct native output allowances¶
src/nwqlib/io/qasm.py::export_qasm admits at most 100,000 supplied top-level native instructions by default, before SDK export, matching the population used by native inspection. Its default emitted UTF-8 text allowance uses DEFAULT_MAX_BYTES. Expanded definitions, QASM ASTs and SDK workspace remain unknown. Revisit operation admission when supporting a different native inventory domain.
Circuit resource calibration¶
Resource work is selected explicitly as a formula, representative sample or prepared inspection. The matched-workload measurements below document calibration fixtures. They do not select a resource tier or establish accuracy for a new input.
The independent PF checks are in tests/test_lchs_resource_structural_law.py. test_product_formula_law_matches_recursive_built_select compares the law with the CX count of the built SELECT after decompose(reps=10), and the transpiled checks in the same file use basis cx,u, optimization level 0 and transpiler seed 7. These fixture counts do not authorize circuit construction at a new workload scale.
The PF accuracy rows span address widths 3, 4, and 5, formula-occurrence counts 2, 3, and 10, and measured/block/analytic CX triples 34/34/34, 90/90/90, and 422/422/426. The largest analytic/measured ratio is 426/422 = 1.00948.... Positive-amplitude PREP pairs use 2*max(0,2**a-2) CX. The generic complex system preparation adds at most 2*max(0,2**n-2). The two-level rows saturate this law, while the heat row uses a cheaper system-state fast path.
QLS and LCHS-QSP estimates follow the actual selected query classes, controls, projector predicates, PREP and reconstruction scales. Explicit representative sampling and prepared inspection consume that same selection. The current resource tests compare structural relations and matched native inventories. Historical compiler counts are not calibration for a changed construction or a portable accuracy interval.
Clifford and T angle sets¶
src/nwqlib/resources/records.py::CLIFFORD_ANGLES and T_ANGLES are the stored conventional angles that name a one-qubit phase or single-axis rotation. A phase or rotation by 0, ±pi/2, ±pi, ±3pi/2 or ±2pi is a Clifford gate, and one by ±pi/4 is one T gate up to Clifford gates, in each case up to global phase. The resource fold's primitive recipes (fold._Fold.primitive) and the QHD rotation laws (src/nwqlib/algorithms/qhd/resources.py::rotation_population and _binary_population) classify every emitted angle with these sets. The comparison is exact on the binary64 value, with no reduction modulo 2 pi and no tolerance, so a nearby arbitrary angle is never counted as a Clifford or T gate. An angle convention, not a tuned value. Revisit when a new primitive or rotation gate uses another angle convention.
Runs and archives¶
These sizes and windows govern a Run's journal, its readouts and its saved files.
Local journal scalar admission¶
nwqlib._run_journal admits each JSON row delta using actual escaped UTF-8 string lengths, integer decimal lengths, exact punctuation and a 32-byte finite-binary64 envelope. It traverses Record fields and rejects nesting deeper than 128. FrozenArray values use their existing dtype, shape and base64 JSON form. Admission computes its encoded size from array metadata and includes the shape list in the nesting limit. The stored-data limit bounds encoded bytes, while the producing method admits its live arrays and serialization workspace. The float allowance exceeds the longest decimal representation of a finite binary64 value, while the depth cap is an untuned recursion boundary. Actual encoded sizes are checked again. These size bounds do not estimate CPU or RSS. Acquisition reservation uses the same admission and adds ExecutionLimits.max_completion_metadata_bytes=65536 for variable metadata. This untuned default reserves storage without allocating buffers. Revisit it for larger provider annotations or changed record schemas. Host KernelApplication records that a kernel declares in SelectedKernel.application_bytes are reserved separately and do not spend this allowance. SQLite's default single-row length limit is 1,000,000,000 bytes, independent of NWQLib's 10 GB total-data allowance and subject to the installed SQLite build.
The Run journal, the archive readers and the backend adapters use these further size constants. None of them changes which data are stored.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
_run_journal.py::LocalJournal.commit and _prepared_execution.py::Run._saved_payload, binary block size |
1 MiB, or the whole payload under a smaller data cap | src/nwqlib/_run_journal.py, src/nwqlib/_prepared_execution.py |
Keeps each stored value far below SQLite's per-value limit and bounds the buffer of each chunked read. Hydration must read with the same block size. Revisit condition: A changed SQLite value limit or read buffer requirement, in both sites together. |
_run_archive.py::load and _prepared_execution.py::Run._saved_manifest, run.json read cap |
1 MiB | src/nwqlib/_run_archive.py, src/nwqlib/_prepared_execution.py |
The manifest is the first file a Run load parses. It names the selection file and does not contain the Plan, so its size does not grow with the problem, and the cap refuses a manifest from a defective writer before it is parsed. Revisit condition: A manifest that legitimately grows beyond it. |
_run_journal.py::LocalJournal, current limits row read cap |
4096 bytes | src/nwqlib/_run_journal.py |
The encoded ExecutionLimits record is a few hundred bytes. It is read first because an extended max_data_bytes decides how large later rows may be. Revisit condition: A larger limits schema. |
artifacts.py::_check_finite_blocks, finiteness-check block at array publication (ArtifactStore._publish), and artifacts.py::probability_readout, block of its finiteness and nonnegativity checks and of its math.fsum total |
_FINITE_CHECK_ENTRIES = 65536 entries, or one leading-axis slice when that slice is longer |
src/nwqlib/artifacts.py |
Bounds the Boolean mask of np.isfinite (and of values >= 0) to 64 KiB, and the list of Python floats that math.fsum reads to one block, instead of one entry per array entry. Untuned. Any block length gives the same accept or reject decision and the same correctly rounded total. Revisit condition: A measured publication cost that a different block length reduces. |
_prepared_execution.py::_merged_associations, provider association check |
32 bytes per item | src/nwqlib/_prepared_execution.py |
A cheap check before _merged_associations builds its lookup tables. The stored items already fit the data cap at more than 32 bytes each (an attempt UUID alone has 36 characters), so the check rejects only an oversized provider report, before the loop walks it. Revisit condition: A different attempt identity format. |
Adapters reading saved or returned JSON text (backends/ibm_runtime.py, ionq.py, nexus.py) |
4 * len(text) |
src/nwqlib/backends/ibm_runtime.py, src/nwqlib/backends/ionq.py, src/nwqlib/backends/nexus.py |
UTF-8 encodes at most 4 bytes per character, and the text has at most len(text) JSON nodes. The value is checked against the Run's data cap before parsing. It bounds the text, not the memory of the parsed objects. Revisit condition: A parse whose object memory must also be bounded. |
Adapters writing or reading a saved native payload (backends/ibm_runtime.py, ionq.py, nwqsim.py) |
2 * max_input_bytes before a QPY dump, 2 * len(payload) before a saved payload is decoded |
src/nwqlib/backends/ibm_runtime.py, src/nwqlib/backends/ionq.py, src/nwqlib/backends/nwqsim.py |
Untuned allowance for the encoded bytes plus one payload-sized copy (the getvalue() bytes of the QPY buffer, the stream that qpy.load reads, or the text that json.loads decodes). Revisit condition: A changed serializer or decoder that keeps more copies. |
backends/ibm_runtime.py::sampler_decode_bytes, scalar Sampler PUBs with 0<=w<=64 and at least one named register |
max(J_join, S*(Q+B)+max(32*S,25*S+24*min(S,2**w)+16)), with Q=sum(ceil(w_r/8)), B=ceil(w/8) and J_join=S*max(9*Q+w,Q+w+(9 if w%8 else 1)*B)+96*bool(w%8) |
src/nwqlib/backends/ibm_runtime.py |
Live ndarray element bytes of the selected packed input, Qiskit 2.5.2 padded join, uint64 packing/gather, and NumPy 2.5.2 sorted unique/count extraction. The returned pair uses 16*m bytes. Padding's 96-byte term covers its two-axis intp specification arrays. Earlier PUB temporaries are released before the next check. Python objects, public records, other PUBs, transport and fixed native workspace are outside this array-payload law. Revisit condition: A changed join or padding algorithm, gather expression, unique implementation, intp width, output representation or generator lifetime. |
backends/ibm_runtime.py::sampler_decode_bytes, scalar Sampler PUB string route above 64 classical bits |
S*max(9*Q+w,Q+w+(9 if w%8 else 1)*B)+96*bool(w%8), with S shots, w=sum(w_r), Q=sum(ceil(w_r/8)) and B=ceil(w/8) |
src/nwqlib/backends/ibm_runtime.py |
Bounds live ndarray element bytes of the selected packed inputs and Qiskit 2.5.2 join. Concatenation keeps padded unpacked bases alive, and padded repacking keeps its input bits and packed output. The 96-byte term covers the two-axis intp padding specification. The subsequent row views, Python byte/string conversion, count dictionaries and classical permutation add no ndarray payload. Python object memory, public records, other PUBs, transport and fixed native workspace are outside this bound. Revisit condition: A changed join, padding or get_counts algorithm, intp width, output representation or lifetime. |
Readout and completion allowances¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| Detached outcome completion growth, NWQ-Sim CPU/SV and NWQ-Sim Slurm trajectories and single-endpoint probability readouts | Fresh outcome growth (161,168) JSON bytes, or (92,99) with an already revised event |
src/nwqlib/_prepared_execution.py::detached_completion_growth |
The event adds a parent, completed status, finish time and observation identity. The already committed status row is replaced, and results_consumed=True saves one byte. The result is independent of the point count. Local UUID point envelopes cost 58 bytes each. SlurmLauncher._locator admits positive uint32 job IDs of at most 10 ASCII digits, so NWQSimSlurmBackend.native_job_id_length returns 10 and the job and empty values envelope costs at most 32 JSON bytes per point. For K Pauli or probability points on the fresh success path, the sufficient completion allowance is 32K + 168 bytes without a forecast or 32K + 99 bytes with an already revised event. Revisit this bound when the admitted job-ID format changes. Admission uses the maximum, 168 bytes, or 99 bytes with a forecast. Acknowledgement and status growth use separate cumulative data admission. Revisit on changes to journal encoding, event fields, revision parents, timestamp format, detached commit order, or backend timing and association fields. |
| Probability trajectory JSON headers | Maximum over legal dense and sparse encodings | src/nwqlib/_prepared_execution.py::_probability_point_metadata, _probability_header_credit |
Each point's data reservation funds its ProbabilityArrays record, publication-row growth and payload-reference rows using the journal field-tree bound. Sparse entry counts are bounded by (2**q - 1)//(W + 1), with W = ceil(q/64). Completion credits at most that point's funded amount and only included net JSON bytes. Shared payload rows receive one credit. Binary values have the separate 8*2**q reservation. Revisit when the record schema, encoding choice or publication lifecycle changes. |
| Projected reducer observable payload | P=32L+H+64+J, J=512+64w+(n+40)M for float coefficients |
src/nwqlib/_quantum_readout.py::projected_requirements, readout_requirements |
M counts all declared JSON rows and L counts nonzero packed rows. H counts the decoded lists, labels, numbers and validated tuple view on the checked 64-bit CPython stack. Integer coefficients use their magnitude-bit count plus two characters in place of 32. The 96L packing frontier covers kept-row tuples and mutable arrays. The conservative scaled reserve is 3P+48L, with H0=65536 for fixed bookkeeping. Original parameter text, parsed objects and immutable packed bytes have separate lifetimes and are charged as distinct representations. This bounds known objects and buffers, not process RSS. Revisit when parsing, snapshots, numeric representation or the interpreter changes. |
amplitudes.py::admit_materialization |
16*N+192*M bytes |
src/nwqlib/amplitudes.py |
A complex128 native vector has N entries. M requested entries require bounded gather, normalization and physical-scaling buffers. The 192-byte per-output envelope covers simultaneous known arrays. Wider mass projection adds 16*W+8*k bytes for a contiguous slice or 80*W+8*k for an indexed slice. This is admission for requested amplitude readout, not a scientific-property check. Revisit condition: The actual reduction buffers or native representation change. |
Runtime seed range¶
MAX_RUNTIME_SEED = 2**32 - 1 in src/nwqlib/core/planning.py bounds the backend seed recorded in RuntimeOptions. Method and backend PCG64 streams derive from the chosen SeedSequence. Static draws are stored in order, and resume restores the existing state and pending draw. There is no arithmetic offset or wraparound. Revisit with a changed backend seed domain.
Readout count representation¶
_validation.MAX_COUNT, also named nwqlib.execution.MAX_COUNT, is 2**63 - 1. This is a representation boundary, not a statistical or workload tolerance. ObservationChunk.histogram() returns counts as int64 weights. A counts declaration requires 1 <= shots <= 2**63 - 1, and an ObservationChunk whose counts are not each at most that maximum, with an exact total equal to its returned shots and within its requested shots, is rejected on every construction and load path. An int64 sum of one validated chunk's counts therefore cannot overflow. Totals across acquisitions need an overflow-safe accumulation. Revisit with a wider integer weight representation.
Selected amplitude mass consistency¶
Amplitude acquisition and the amplitude and host-kernel results of QLS use the shared nwqlib._validation.NUMERICAL_RELATION_RTOL convention, written w below. Concrete negative masses reject. For algorithm mass a and physical mass p, subset admission requires p-a <= w*max(1,a). A Plan selecting the whole branch uses the two-sided difference. Raw values remain unchanged. When its preparation record supplies a saved_state_probability_window, an amplitude acquisition also requires each normalized mass to be at most one plus that window (ObservationChunk.validate_unit_bound). The available window is never below w. An unassessed saved-state correction leaves only the nonnegativity check. Host-kernel QLS masses must be at most 1+w. The window w is a numerical input convention, not a simulator error theorem, recovery scale or verification threshold. The comparison adds no norm evaluation. An exact scalar QLS result publishes the masses of its saved reduction, whose subset relation carries the host mass allowance of _quantum_readout.validate_saved_masses, and a sampled QLS result compares nested integer counts. On a Plan that selects the whole branch, both publish one mass for the two populations.
Projected scalar reduction¶
_quantum_readout.readout_requirements admits the exact projected reduction with scan tiles of at most 1024 amplitudes (tile <= 1024) and the bookkeeping allowance H0 = 65536 above. _quantum_readout.pair_value accepts saved scalar exponents in [-8192, 8192], a derived envelope for finite binary64 inputs and admitted lengths below 2**51, not a precision cutoff. The projected kernel resolves only the producing preparation record's amplitude-derived masses label. Its own host allowance is a_X = gamma(2*n_X+1) with U_X = 32*n_X*eta*M_hat_X/(1-(2*n_X+1)*u) for each declared population, where u=2**-53 and eta=2**-1074. With the resulting qualified native-state budget delta, validate_saved_masses uses the complete-mass window 2*delta+delta**2+a_F*(1+delta)**2+U_F. If another native assumption remains unavailable, it uses the probability-window convention omega+a_F*(1+omega)+U_F. The registry supplies the resolved budget at acquisition, and LCHS and QLS use the same resolution at publication. Revisit these with a different mass kernel or a wider native state domain.
Binary inference representation¶
nwqlib.evidence.binary.MAX_BINARY_COUNT is 2**53-1. This is a representation boundary, not a statistical or workload tolerance. Binary64 exactly represents all admitted integer counts and the n+1 endpoint in the selected anytime schedule. Accumulation checks that boundary before adding another population. Beta posterior construction additionally checks that positive prior/count inputs and tail probabilities survive binary64 addition/complement. Revisit only with another explicitly supported numerical representation. The BinaryReadoutMitigation.minimum_contrast value is a caller-selected conditioning threshold (default 0.05), not an inferred calibration-quality guarantee.
nwqlib.evidence.binary.MAX_BETA_SHAPE_TOTAL is 5e11. When the posterior shape total a+b exceeds it, the equal-tail Beta interval is unavailable, and the posterior mean and variance remain available. Below it the lower endpoint is SciPy's betaincinv(a, b, alpha/2) and the upper endpoint is betainccinv(a, b, alpha/2), which inverts the complement 1-I_x(a,b) directly because rounding 1-alpha/2 changes the upper tail by up to 2-54. The accuracy criterion is an endpoint error of at most 1e-3 posterior standard deviations. For a near-normal posterior an endpoint error of eps standard deviations changes the tail mass alpha/2 by a relative amount of about eps*phi(z)/Phi(-z), where z > 0 solves Phi(-z) = alpha/2. That ratio is 2.3 at alpha/2 = 0.025 and 8.4 at 2-54, the smallest admitted tail, so for such a posterior the criterion keeps each tail mass within 1 percent of alpha/2.
The threshold is measured. The reference integrates the Beta density with mpmath 1.3.0 at 50 digits over one-standard-deviation panels and refines each quantile by safeguarded Newton steps. Where mpmath.betainc also runs, the two agree to better than 1e-26 in relative tail mass. For Beta(a, 1), whose lower quantile is (alpha/2)**(1/a), the reference matches that closed form within 1e-22 standard deviations up to a = 1e13. The fixed grid takes, at each a+b, zero fractions n0/n = 0.5, 0.3, 0.7, 0.1, 0.9, 1e-3 and 0.999, the counts n0 = 0, n0 = 10, n1 = 0 and n1 = 10, the uniform prior and, for n0 = 0, also the Jeffreys prior, with tails alpha/2 = 0.025, 1e-6, 1e-12 and 6e-17 and both endpoints. The dense sample adds 300 zero fractions, 150 log-uniform in [1e-6, 0.5] and their mirror images, with the same tails and the uniform prior. SciPy 1.18.1, NumPy 2.5.2 and Python 3.12.14 on macOS arm64 gave these largest endpoint errors, in posterior standard deviations:
| a+b | Fixed grid | Dense sample |
|---|---|---|
| 1e10 | 4.1e-7 | |
| 1e11 | 1.4e-5 | |
| 5e11 | 2.6e-4 | 4.4e-4 |
| 1e12 | 4.3e-4 | 1.2e-3 |
| 2e12 | 5.1e-4 | 2.0e-3 |
| 5e12 | 1.1e-3 | |
| 1e13 | 4.1e-3 |
The error grows about in proportion to a+b, and the largest dense-sample errors are about 10(a+b)u with u = 2-53. That scale is consistent with rounding of order (a+b)u in the exponent aln(x) + bln(1-x) of the incomplete beta prefactor, an explanation inferred from the scaling alone. 5e11 is the largest measured total whose samples stayed below 1e-3, with a factor of 2.3 to spare. On Ubuntu aarch64 with SciPy 1.18.1, a retest found a largest endpoint error of 2.58e-4 posterior standard deviations at a+b = 5e11, so the same 1e-3 criterion holds there. A sampled maximum does not bound every input. Revisit when SciPy changes its incomplete beta inverse, or with a kernel that evaluates I_x(a,b) in extended precision.
Saved JSON and analysis-origin admission¶
ArchiveFiles.read_json checks the file-byte allowance of a folder opened with one before reading text, rejects duplicate keys and nonfinite numbers, and refuses with ValueError a file whose container nesting exceeds the recursion limit of the standard recursive JSON parser, which is the nesting boundary of archive reading. The Run journal's 128-level nesting limit applies when a row is written. This is a metadata parser boundary, not binary payload authentication or a bound on SDK deserialization. Revisit it only for a concrete supported archive requiring deeper nesting.
capture_analysis_origin(max_dependencies=64) bounds the complete supplied list or tuple of package names, including duplicates, before deduplication and version lookups. The fixed nwqlib and pydantic names are additional. External producers can select a larger explicit bound for an identified dependency inventory. These are installed metadata queries, not imports or scientific execution.
Record identity JSON¶
src/nwqlib/core/records.py::Record.content_id encodes the whole record as JSON text and hashes its UTF-8 bytes. model_dump shares the record's scalar objects, so the hash holds two copies of the JSON beyond the record's own data. A FrozenArray field is the exception, because its serializer creates a new base64 string for each array occurrence in the dump. The record does not know its caller's max_bytes, so no check applies there. The table lists the records whose JSON grows with the problem size, the check that bounds each one and whether that bound also covers the identity JSON.
| Records | Size grows with | Bounded by | Identity JSON against that bound |
|---|---|---|---|
evidence/energy_shift.py::EnergyEndpoint (terms, matrix_entries) |
M Pauli terms, or nnz or D² scanned matrix entries | _operator_payload |
Covered. The capture law counts the table's scalars and the table's share of both JSON copies, at most 6+2*digits+48 bytes per entry or q+30 per term plus 2 bracket bytes per copy. The record's fixed fields are not counted. |
ir/records.py::Program with its nodes, settings and bindings, and the selections of blocks/records.py::SelectedConstruction, one per Program signature, held by the construction of every Plan |
Definitions, signatures, measurement settings and bindings | src/nwqlib/ir/expressions.py::AdmissionLimits, 16,384 definitions and 100,000 kept field slots |
A fixed envelope independent of D and M. It binds Plans whose definitions grow with the terms. A sampled Expectation Plan with 100-shot counts at q = 12 forms nearly one qubit-wise commuting group per term in the following measured prefix. Initialize rng = numpy.random.default_rng(1) once and draw 400 successive labels with "".join(rng.choice(list("IXYZ"), 12)), remove duplicates in first-seen order and remove the all-I label. Use successive standard-normal coefficients from an independent default_rng(2), the all-zero occupation reference and Plan seed 7. Each group experiment allocates one 12-qubit register, prepares the state, applies the group's own basis block, measures the whole register into one 12-bit classical value and releases the register. At every measured prefix the validator charged 42G + 33 admission-work units for G groups. The 60-term prefix has 57 groups and uses 2,427 units, the 61-term prefix has 58 groups and uses 2,469, and the full 400-term prefix has 329 groups and uses 13,851. The max_definitions default of 16,384 is the smallest power of two with at least twofold margin over the combined inventory of 5,951 definitions, expressions, parameters, registers, classical values and signatures (5,761 of them definitions) in the Program of a sampled FixedGCIM Plan built as follows. Initialize rng = numpy.random.default_rng(2) once and add labels "".join(rng.choice(list("IXYZ"), size=12)) to a set until it holds 200 distinct labels. Sort the labels and give them, in sorted order, successive rng.normal() coefficients from the same generator. Use the occupation basis states 1, 2, 3 and 5 as 12-bit strings, 2,000 shots and Plan seed 7. That Program is refused at 4,096. The Plan JSON grows by about 4.8 KB per random term between the 60- and 126-term prefixes, mostly in each group's basis selection and experiment definitions, and the admitted 400-term Plan holds 1.7 MB of JSON. These figures were measured with Python 3.12.14 and Qiskit 2.5.2 on macOS arm64. |
Pauli-term reconstructions and readout specs: ExpectationReconstruction, FixedGCIMReconstruction, QPEReconstruction, LCHSReconstruction, QLSReconstruction, AdaptReconstruction with AdaptPoolMember, and core/planning.py::ObservableEstimateSpec and ObservationSpec |
M terms of an admitted Pauli operator, and for ADAPT its pool terms and kept commutator rows | The operator's input admission, which adds S=6q+360 bytes per term (src/nwqlib/operators/inputs.py::_pauli_identity_bytes) to q+4L per label term, 4L per mask term and 5L per SparsePauliOp term. The Fermion mapping and the final Pauli expansion of FactorizedOperatorProduct include S in their laws. ADAPT checks a running total in src/nwqlib/algorithms/gcim/adapt_inputs.py, because Pauli admission does not see the commutators. It charges S per Hamiltonian and pool term and, for the other fields of each pool member, three copies of the sum of 2048 bytes, 128 per fermionic term and 16 per ladder entry or orbital index. It also charges four copies of the reconstruction's remaining JSON, measured from its actual metadata with ASCII escaping, collection delimiters and the identity wrapper (_reconstruction_fixed_json). A quantum Plan's commutator tables are arrays and a sampled Plan's readout groups are index tuples. array_record_bytes charges these with the Pauli-row and pool-member allowances as other live bytes and the reconstruction and table metadata as J_fixed (section "GCiM work and byte limits"). |
Covered for QPE, ADAPT, Expectation without sampled counts, the labels of grouped sampled Expectation and QLS with vector or exact scalar output. LanczosReconstruction holds no per-term data. It stores the scalar frame and the identity of the selected Pauli table, and the bound Plan derives its SELECT readout table (src/nwqlib/algorithms/lanczos/readout.py::LanczosReadout) from that table after planning and after loading. S is three copies of J=2q+120 bytes, the two JSON copies of the identity and one copy that bounds the records' own label and coefficient data. J holds the widest per-term layout, a PauliTerm or PauliCoefficient record of q + 88 bytes with a 24-character binary64 repr plus the label (q + 3) and coefficient (25) again in the readout spec of a provider-estimate Expectation and 4 bytes of list brackets. The measurements below, made with Qiskit 2.5.2 on macOS arm64, count the identity JSON of the whole Plan for random Pauli labels with standard normal coefficients and occupation reference states. From q = 12 to q = 40, where J grows from 144 to 200, one JSON copy grew per term by 96 to 123 bytes for QPE, 112 to 168 per ADAPT Hamiltonian term, 110 to 166 for exact Expectation, 129 to 185 for provider-estimate Expectation and, for 100-shot sampled Expectation of distinct all-Z labels that form one group, 109.6 to 165.6, each within J at its q, and by 44 bytes per term of a Pauli A for QLS with vector output at q = 12 and 20. With the spin-adapted pool, the exact quantum ADAPT reconstructions of H4 at q = 8 (the eigenvalue example, 45,896 commutator rows) and LiH at q = 10 (the notebook's ACTIVE_SPACE = (2, 5), 200,910 rows) serialize to 4.9 and 24.9 MB of model_dump_json text against checked totals of 25.5 and 133 MB, measured with Python 3.12.14 and Qiskit 2.5.2 on macOS arm64. The reconstructions' fixed fields, a few KB each, are not counted. At q = 12, FixedGCIM with two occupation basis states and LCHS on a periodic stencil with product-formula evolution add 8.9 to 12.8 KB per Pauli term of the observable or Hamiltonian, mostly in the Program and its selections. That growth far exceeds J, so AdmissionLimits (row above) bounds them instead, as it bounds the group experiments of sampled Expectation. The largest admitted Plans of these kinds, for FixedGCIM and LCHS, hold 341 and 57 terms and 4.4 and 0.6 MB of JSON. QLS on PeriodicStencil(12, mass=1.0, diffusion=0.25) with the all-zero occupation RHS, and with the labels and coefficients of the sampled Expectation measurement in the row above, binds each term of an exact QuadraticForm or NormalizedExpectation output once, in the parameters of its one reduction. Its Plan JSON grew by 38.5 bytes per term between the 40- and 80-term prefixes, within J. Sampled QLS Program size also depends on the distinct measured-register maps. Allocate, Release, coherent evolution and one-site basis changes are shared. Each distinct position/wire pair supplies one Measure definition, and each distinct measured register supplies one Sequence. Their additional record allowance and count-table bounds are stated in the QLS sampled count registers row. The actual Program is admitted against its kept field slots and lifecycle work. |
execution.py::ObservationChunk, ObservationView, ExecutionTrace and the contribution lists of core/analysis.py::Result |
Acquisitions, their count bins, and their probability array manifests and summaries | The output reservation made before each acquisition (encoded readout declaration, declared values, eight bytes per possible outcome of a probability array, and max_completion_metadata_bytes), for a trajectory each point's one-point declaration, fixed chunk header and values with one metadata allowance (_prepared_execution.trajectory_reservation), against ExecutionLimits.max_data_bytes, the item and payload admission before native preparation (_prepared_execution._admit_readout) and max_total_circuits |
Proportional. A chunk's identity is computed before its journal row is written, so the reservation, not the row admission, bounds it. Arrays are separate artifacts whose manifests hold digests only. |
qhd/constrained_records.py::AugmentedLagrangianRecord with its ALIteration rounds |
Rounds times the constraints and variables of the problem | src/nwqlib/algorithms/qhd/constrained.py::_admit_layer_work. Each admitted round evaluates f and every constraint at a point of d coordinates for at least (1 + m)(1 + d) units with m constraints, and stores at most 6 m + 2 d + 91 JSON values outside a nested refinement under inequality_form="phr", counting nested identities, versions and unknown-count reasons (70 to 92 values for one to four inequality constraints on one to three variables), and at most 29 m + 3 d + 100 for m >= 1, or 3 d + 102 for m = 0, with the inequality representation of the other forms (augmented-Lagrangian record size) |
Proportional. The record holds at most 91 JSON values per admitted unit of layer work. Under the other forms the ratio of the round bound to the round's layer work, (29 m + 3 d + 100)/((1 + m)(1 + d)), is at most 33 for m, d >= 1, and (3 d + 102)/(1 + d) is at most 52.5 for m = 0 and d >= 1. The inner Results of the rounds are separate records bounded by their own rows. |
qhd/constrained_records.py::AugmentedLagrangianRecord with nested BoxRefinementResult records |
R stored rounds, m kept constraints, d variables, n_r completed levels and b_r split levels in round r | For solver-produced records with integer entropy, one-entry round spawn keys, at most one QHD missing-data reason per round and level spawn keys of length at most two, R <= options.max_iterations, n_r <= refinement.max_levels and b_r <= min(n_r, refinement.max_splits) separately in each round, with b_r = 0 when splits are disabled. Write H for the scalar leaves outside iterations, including the QHD configuration, root refinement options and aggregate resources, and g_r for the length of both Gaussian parameter vectors, or zero for other initial states. If H_other excludes the root QHD field as well, H = H_other + 34 + 2*g_0, where g_0 is the root Gaussian length or zero for other initial states. The scalar-leaf bound is H + R*(6*m + 2*d + 91) + sum_r(2*g_r + 147 + n_r*(14*d + 77) + b_r*(8*d + 13)), summing over nested refinements that exist. The direct round envelope under inequality_form="phr" is 6*m + 2*d + 91, including at most 49 resource leaves. Under the other forms each round has its own bound, at most 29*m + 3*d + 100 for m >= 1 or 3*d + 102 for m = 0, and the level terms of its refinement count the d + s variables of its inner problem with s slack variables (augmented-Lagrangian record size). _admit_layer_work bounds the layer's evaluation reservation. Each completed level's QHD running total includes at least d*K units, where K is qhd.num_grid_points. For d >= 1 and K >= 2, 14*d + 77 <= 46*d*K without a split and 22*d + 90 <= 56*d*K with a split. |
Structurally proportional, with root and zero-level configuration terms explicit. Leaves include nulls and exported record identities, and exclude object keys and containers. The hash wrapper adds two scalar leaves to the hashed fields. Identity JSON bytes are unbudgeted because text lengths, including aggregated unknown-count reasons, and integer digit counts are not bounded by this law or qhd.max_bytes. Inner QHD Results are separate records. |
qhd/refinement_records.py::BoxRefinementResult |
n completed levels, b split levels, d variables and the length g of both Gaussian parameter vectors, or g = 0 for other initial states | options.max_levels bounds n, and b <= min(n, options.max_splits), with b = 0 when splits are disabled. At most 2*g + 147 + n*(14*d + 77) + b*(8*d + 13) scalar leaves for level spawn keys of length at most two, or 2*g + 147 + n*(14*d + 76) + b*(8*d + 13) for standalone refinement. The fixed part includes 88 leaves for total and stopped resources and 34 + 2*g for QHD. The QHD contribution is 15 direct scalar settings, including encoding and kinetic_model, 3 identities, 5 schedule leaves, 4 + 2*g initial-state leaves and 7 BinarySynthesis leaves. The binary-synthesis record is exported for both encodings. Each completed level admits at least d*K running-total work, where K is qhd.num_grid_points. For d >= 1 and K >= 2, 14*d + 77 <= 46*d*K without a split and 22*d + 90 <= 56*d*K with a split. |
Structurally proportional, with the fixed configuration term present even at zero levels. Leaves include nulls and exported record identities, and exclude object keys and containers. The hash wrapper adds two scalar leaves to the hashed fields. Identity JSON bytes are unbudgeted because scalar strings and integers have variable encoded lengths. Live QHD Results are outside the record identity. |
qhd/records.py::SupportValues and QHDReconstruction.initial_amplitudes |
E total support entries, up to D = K**d for a full support, and d K one-variable grid entries |
QHD._admit_symbolic_work with the byte laws of method._support_table_bytes |
Covered. QHD support tables and initial amplitudes are admitted by QHD._admit_symbolic_work before their construction and first identity computation. Their JSON grows with the total support entries and one-variable grid entries. The allowance includes the portable base64 tree and encoder workspace beside the live arrays, using the byte laws in _support_table_bytes. Each enclosing record counts every array occurrence it serializes. |
QHD grids and marginals, QLS and QSP polynomial coefficients and phases, LCHS quadrature and Duhamel nodes, Lanczos moments, FixedGCIM pencils, WorkloadEstimate, PlanEstimate and telemetry rows |
Grid points, degree, nodes, Krylov dimension, basis size or assessments | Their registered caps and checks: QHD selected operation sizes, max_degree=256, the LCHS quadrature work and byte checks (for Duhamel nodes duhamel_nodes**3 <= max_quadrature_work in src/nwqlib/algorithms/lchs/host.py), the Krylov degrees, max_basis_size=64, 24 * max_steps, max_assessments=4096 and max_comparisons=100000 |
Fixed or method-level caps independent of D. |
Revisit when a record type starts holding data that grows with the problem, or when a per-term Pauli record grows past J.
Augmented-Lagrangian record size¶
Count the scalar JSON leaves of the exported qhd/constrained_records.py::AugmentedLagrangianRecord, each number, Boolean, string or null once, including every nested record's schema, parent and content hashes, and excluding object keys and containers. For solver-produced records with integer entropy, one-entry round spawn keys and at most one QHD missing-data reason per round, a round outside its nested refinement under inequality_form="phr" has at most 6 m + 2 d + 91 leaves for m kept constraints and d variables, the registered bound. Its ALEvaluation has at most 5 m + 2 d + 24, its mode_status included, its ALResources at most 49 (18 counts, 14 nullable with one reason pair each, and 3 content hashes), and its other fields, a null refinement and a null representation included, at most m + 18. Under the other forms, a round's InnerRepresentation with s slack variables and T trials has at most d + m + 10 s + 10 T + 12 leaves. Replacing its null placeholder adds at most d + m + 10 s + 10 T + 11 leaves. The evaluation's slack fields add at most 2 max(s - 1, 0) leaves, because both remain null when s = 0. The round therefore has at most 6 m + 2 d + 91 + d + m + 10 s + 10 T + 11 + 2 max(s - 1, 0) leaves. With s, T <= m, this is at most 29 m + 3 d + 100 for m >= 1, and at most 3 d + 102 when m = 0.
A nested refinement with n completed levels and b split levels adds at most 2 g + 147 + n (14 d + 77) + b (8 d + 13) (refinement record size), with d replaced by the d + s variables of a round with s slack variables, where g is the common length of a GaussianState's center and widths, or zero for other initial states, and level spawn keys have at most two entries. Let H count the leaves outside iterations. If H_other excludes the root qhd field as well, then H = H_other + 34 + 2 g_0, where g_0 is the root Gaussian length or zero for other initial states. H includes the problem, configurations, preprocessing, limits, scales, aggregate resources and remaining root fields. With R stored rounds, the bound is H + R (6 m + 2 d + 91), or H plus the sum of the round bounds above under another form, plus the nested-refinement bounds above, summed over refinements that exist. Adding each full refinement to the round bound overcounts its null placeholder by one.
There are at most options.max_iterations rounds and each nested history has at most refinement.max_levels levels, with b <= min(n, refinement.max_splits) separately in each round and b = 0 when splits are disabled. An unsplit level has at most 14 d + 77 leaves and a split level at most 22 d + 90. Each completed level's running total counts at least d K work units, where K is qhd.num_grid_points. For d >= 1 and K >= 2, these level bounds are at most 46 d K and 56 d K, respectively. Root fields and each refinement's fixed term 2 g + 147 are accounted for separately, including zero-level refinements. These counts give no byte allowance. Text such as the joined unknown-count reasons and integer digit counts have no fixed length, so content-hash JSON bytes are not budgeted by this rule or qhd.max_bytes.
Refinement record size¶
Count the scalar JSON leaves of the exported qhd/refinement_records.py::BoxRefinementResult, each number, Boolean, string or null once, including every nested record's schema, parent and content hashes and every allowed unknown-count reason pair, and excluding object keys and containers. An unsplit level has at most 14 d + 75 + s leaves for d variables and s entries in its spawn key. Its dimension-dependent fields contribute 14 d from box and next_box (2d each), intervals (2d), the center and widths of a best-point initial_state (d each), and spacing, point_indices, point, unit_point, axis_masses and region_axis_counts (d each). Its 24 other scalar fields include region_count, mode_status, the null split and scalar split_declined, and initial_state adds at most its kind and 3 identities. There are 3 identity leaves and at most 44 resource leaves, from 15 counts, 3 resource identities and at most 13 reason pairs. A StallSplit has 8 d + 14 leaves, replacing one null and adding 8 d + 13. Standalone refinement has s = 1 and the augmented-Lagrangian layer s = 2, so the unsplit bounds are 14 d + 76 and 14 d + 77, and the split bounds are 22 d + 89 and 22 d + 90, respectively.
The two selected region count fields contribute d + 1 scalar leaves for counts. Exact readout uses two nulls, which also fit this bound for d >= 1. The separately computed report dictionary is outside the exported record and its content identity.
The fields outside levels add at most 2 g + 147 leaves. This counts 11 for the eight scalar result fields and three identities, 14 for the options, 88 for total and stopped resources, and 34 + 2 g for QHD. The QHD term includes 15 direct scalar settings, including encoding and kinetic_model, 3 identities, 5 schedule leaves, 4 + 2 g initial-state leaves and 7 BinarySynthesis leaves. The binary-synthesis record has 4 settings and 3 identities and is exported for both encodings. Here g is the common length of a GaussianState's center and widths, or zero for other initial states. With n completed levels, b split levels and spawn keys of length at most two, the bound is 2 g + 147 + n (14 d + 77) + b (8 d + 13). Use 76 in place of 77 for standalone refinement. options.max_levels bounds n, and b <= min(n, options.max_splits), with b = 0 when splits are disabled. Each completed level counts at least d K running-total work, where K is qhd.num_grid_points. For d >= 1 and K >= 2, 14 d + 77 <= 46 d K without a split and 22 d + 90 <= 56 d K with a split, since 46 d K - 14 d >= 78 d >= 77 and 56 d K - 22 d >= 90 d >= 90. The fixed fields remain when no level completes. These counts give no byte allowance because text lengths and integer digit counts vary. Identity JSON bytes are unbudgeted, and live QHD results are outside the record's content hash.
Backends¶
These defaults and size limits belong to the backend adapters and the provider services they call.
Provider and compiler policies¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
_prepared_execution.py::LOCAL_SIMULATORS |
{"qiskit_aer", "nwqsim"} |
src/nwqlib/_prepared_execution.py |
The backend kinds that simulate on this machine, local Aer and a locally launched NWQ-Sim runner. The Run bounds their circuit width by ExecutionLimits.max_simulation_qubits before native work, and on them and on host execution prepare(plan, settings="all") prepares every static setting, because their preparation is local lowering with no remote compilation or provider charge. Revisit condition: A new local simulator, or a decision on preparing every setting on remote backends (Limitations and open work). |
backends/nwqsim.py, slurm.py, ionq.py, ibm_runtime.py and nexus.py, optimization_level defaults |
0, 0, 0, 1 and 1 | src/nwqlib/backends/nwqsim.py, src/nwqlib/backends/slurm.py, src/nwqlib/backends/ionq.py, src/nwqlib/backends/ibm_runtime.py, src/nwqlib/backends/nexus.py |
The compiler optimization levels whose preparation records carry no optimization_level exclusion. A user may select any level from 0 to 3. Another level adds that exclusion. A preparation record carrying it has no state_error() bound, and the code that uses it takes its documented fallback. No error or resource accounting of the optimization is attempted. Revisit condition: An error derivation that covers an optimized lowering. |
backends/ionq.py, request timeout |
30 seconds | src/nwqlib/backends/ionq.py |
Untuned connect/read timeout for each HTTP request, not a total request or job deadline. Revisit condition: The selected service and network workload. |
backends/ionq.py, shots |
1–1000000 | src/nwqlib/backends/ionq.py |
The shots range of the single-circuit and multi-circuit job types in IonQ's v0.4 create-job reference, which gives minimum 1 and maximum 1,000,000. A selected device may accept fewer shots. Revisit condition: A changed API version or a device-specific limit. |
backends/ionq.py::_json_bytes |
max_data_bytes//3 |
src/nwqlib/backends/ionq.py |
Untuned serialization admission factor leaving room for simultaneous request representations and journal metadata. JSON text and outgoing UTF-8 bytes coexist. The factor is not a proof of Python-object or requests/SDK peak memory. Revisit condition: A changed serializer or transport lifetime. |
backends/ionq.py::_request |
read size 64 KiB and parse allowance 5 * size |
src/nwqlib/backends/ionq.py |
Untuned read size. With size bytes already read, each read asks for min(64 KiB, max_response_bytes + 1 - size) bytes, so the reader holds at most max_response_bytes + 1 body bytes, and the extra byte reveals an oversized body. This relies on urllib3 2, which the ionq extra requires, returning at most the requested number of decoded bytes per read. Five times the bytes read is an untuned allowance for the parsed JSON, checked against the Run's data cap before json.loads. It is not a measured object-size ratio. Revisit condition: Measured parse memory or a changed response format. |
backends/ionq.py, bytes per circuit at launch and per item at batch refresh |
128 bytes per circuit at launch and 256 bytes per item at batch refresh | src/nwqlib/backends/ionq.py |
Untuned per-item allowances for the item names and associations those calls build, checked before building them. Revisit condition: A changed item naming or association record. |
backends/ionq.py::ionq_decode_bytes |
57*m + 8*min(m, 2**min(q,c)) for q,c<=64 |
src/nwqlib/backends/ionq.py |
Bounds one nonempty histogram decode's live ndarray elements, with m provider entries, q native qubits and c classical bits. NumPy 2.5.2 unique with inverse indices holds seven eight-byte-per-entry arrays, one Boolean mask and the distinct mapped values. The law includes the returned pair and allows a separate gather even for the 64-bit identity. It excludes Python objects, native workspace, transport, public records and other decodes. Revisit condition: A changed gather, remap, unique implementation, dtype width, output representation or lifetime. |
backends/ionq.py::_decode |
m*(c+16) represented count-key allowance |
src/nwqlib/backends/ionq.py |
One bit-string count mapping has at most m entries of c characters, integer counts of at most seven digits and JSON entry punctuation. This admission applies alongside the narrow array check and is the decoder's only byte admission on the wider Python-integer route. Public CountBin JSON uses the common readout reservation. This term does not bound Python dictionary, string or integer memory. Revisit condition: A changed count domain, key spelling, output representation or Python-heap accounting contract. |
backends/nexus.py, shots per program |
10000 | src/nwqlib/backends/nexus.py |
H2's published job limit keeps a job within its allowed calibration interval. See H2 hardware credit limitations. Revisit condition: A changed provider contract or target. |
backends/nexus.py::refresh per download, and _decode |
shots * (len(bits) + 8) per download, and shots * (4 * width + 8) |
src/nwqlib/backends/nexus.py |
Untuned allowances for one downloaded pytket result and for the outcome tuples and count labels built from it. The download allowance uses the largest item because the program identity is known only after download. Revisit condition: Measured result sizes or a changed result format. |
backends/slurm.py::SlurmLauncher.admit |
12 + len(profile.cluster.encode("ascii")) response bytes |
src/nwqlib/backends/slurm.py |
SlurmLauncher.admit requires max_response_bytes >= 12 + len(profile.cluster.encode("ascii")) before the submission intent and, through launch, before any spool write or scheduler call. The allowance covers a ten-digit supported job ID, the separator, the known cluster and one LF in canonical sbatch parsable output. The normal response cap and Run data admission still apply to every command. Status and accounting responses can require a larger cap. Revisit condition: A changed job-ID domain, cluster-name rule or sbatch acknowledgement form, sbatch output on stderr, or a transport that introduces CRLF. |
backends/slurm.py::SlurmTransport, pipe read, and backends/_auxiliary_process.py, read |
pipe read 64 KiB, and read 8 KiB with a 0.1 s select tick | src/nwqlib/backends/slurm.py, src/nwqlib/backends/_auxiliary_process.py |
Untuned I/O granularity. The byte caps and deadlines are enforced independently of these values. Revisit condition: A measured latency or throughput need. |
backends/nwqec.py, epsilon |
1e-10 per gate or 1e-2 otherwise |
src/nwqlib/backends/nwqec.py |
Untuned requested synthesis precision defaults under distinct compiler policies. The request is recorded, and its name does not certify the realized circuit's error. Revisit condition: The intended synthesis accuracy or compiler policy. |
| Nexus execute-job programs | 300 | src/nwqlib/backends/nexus.py::admit_batch |
The documented Nexus job-size limit. Each program carries its own shots and cost cap. This limit counts program items, independently of the 10,000-shot per-item H2 policy. |
| IonQ multi-circuit job | 5,000 circuits and 150,000 gates | src/nwqlib/backends/ionq.py::admit_batch |
The limits of IonQ's v0.4 create-job reference. A multi-circuit job holds at most 5,000 circuits and 150,000 gates in total across them. NWQLib counts one gate per entry of each circuit's IonQ JSON gate list, and the reference does not define the count further. The reference gives no gate limit for a single-circuit job, and NWQLib applies the same 150,000-gate total to one as its own cap. Revisit when the API version or the reference changes. |
Aer simulator¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
AER_EXECUTION_DECOMPOSE_REPS |
10 | src/nwqlib/backends/qiskit_aer.py |
Largest number of decomposition passes that turn nested non-native instructions, such as StatePreparation and controlled custom gates, into Aer-executable gates. Lowering stops earlier once only native classes, names and supported representations remain. Metadata counts actual passes including terminal admission. Selected explicit-analysis preparation paths still use the fixed budget. Revisit when Qiskit/Aer instruction support changes. |
| Aer STATEVECTOR trajectory count | 1 | src/nwqlib/backends/qiskit_aer.py |
Exact readouts, a trajectory schedule of several points included, run one explicitly selected pure-state trajectory, including measurement-conditioned execution. This is separate from SHOTS measurement budgets and the resource shots field. |
| Aer job identifier length | 36 ASCII characters | src/nwqlib/backends/connection.py::AerBackend.native_job_id_length |
Aer 0.17.2 generates each job identifier with str(uuid.uuid4()). Its JSON string costs 38 bytes. The trajectory job and empty values envelope cost 58 bytes per point. Admission uses the 1,113-byte upper bound on synchronous completion-row growth without a forecast, or 1,044 bytes when the event already has a revision parent (src/nwqlib/_prepared_execution.py::synchronous_completion_growth). Revisit when the job format or completion lifecycle changes. |
| Aer statevector thread cap | max(os.cpu_count(), OMP_NUM_THREADS) |
src/nwqlib/backends/qiskit_aer.py::_aer_state_threads |
Positive cap passed to Aer as max_parallel_threads, resolved once per open Run, recorded in the preparation metadata and reapplied on restoration. The trajectory memory check uses the same value for the transient per-worker probability and index arrays of a general marginal save. Aer takes the minimum of this cap and its automatic OpenMP maximum, so the cap can reduce parallelism. Revisit when a workload needs more state-update threads or when Aer changes its thread configuration. |
NWQ-Sim runner guards¶
The NWQ-Sim runner reserves 32 bytes per serialized numeric value when admitting its JSON output envelope, which holds counts and Pauli values. Probability marginals and amplitudes are binary files admitted by their exact bytes. This conservatively covers the bundled nlohmann serializer's binary64 and uint64 spellings, in addition to separately counted key and punctuation bytes. Revisit when the numeric representation or serializer changes. Actual serialized bytes are checked again before writing. This does not claim a bound on JSON object overhead or process RSS.
The NWQ-Sim request writer (src/nwqlib/backends/nwqsim.py::_streamed_gate_bound) admits each compact JSON gate before writing it: at most 134+digits(q) bytes for a U gate and 38+digits(q0)+digits(q1) for a CX gate, plus one separator byte after the first gate. The U skeleton {"name":"u","qubits":[],"params":[]} has 36 bytes, one wire, three parameters in the same 32-byte binary64 envelope and two commas. The CX skeleton is one byte longer, with two wires, one comma and no parameters. The fixed request fields and the closing suffix are reserved first. Revisit when the gate fields, their spelling or the float envelope change.
NWQ-Sim's 58-qubit SV and 29-qubit DM guards keep signed 64-bit state indices and native byte products representable. The 63-classical-bit guard keeps the runner's projected count key in its unsigned 64-bit container. MPI shot counts are additionally limited to 2**31 - 1, the public collective's C int count. Global U/CX gates also require the local state dimension 2**q / R to fit that count because their kernels send the full local partition. Local-only gates, global measured bits and global phase do not require this state transfer. MPI rank counts must be powers of two and each rank must own at least two qubits, as required by the public SV_MPI constructor. These representability guards do not promise an allocatable state. Revisit with changed native types or a different collective/kernel contract.
Native known-array byte formulas follow the inspected constructor and sampler allocations: CPU/SV has three double arrays plus one extra double, CPU/DM has two density arrays plus a diagonal cumulative array, MPI/SV has four local double arrays and three full-shot arrays, and GPU SV/DM has two host plus four device state arrays (two have an extra double), with four full-shot arrays. The NWQ-Sim guide records their per-process formulas and 20-qubit scalar projections. Revisit these formulas when changing the dependency allocation layout. They exclude parsed input/circuits, fused gates, communication buffers and additional SDK workspace. They are not whole-process RSS bounds.
The runner disables PURITY_CHECK before including the CPU kernel. That macro adds a full norm scan after each fused gate and only prints a discrepancy. It is not part of state evolution or required readout. No per-gate reference/diagnostic pass is added by NWQLib execution. Revisit if an explicitly requested diagnostic supplies its own compute budget and a use for its result.
IBM primitive preparation¶
IBMRuntimeBackend uses resilience level 0 unless explicitly selected by the caller, so the wrapper does not enable mitigation work by default. Other unreported server settings and provider usage remain unknown. Revisit the option mapping when the supported Runtime API changes.
The layout-mapped observable uses SparseObservable.simplify(0.): zero is an exact-zero policy, not a numerical tolerance. Qiskit 2.5.2's ordinary ObservablesArray coercion simplifies at 1e-8 and can remove nonzero terms. The already validated sparse representation enters its public prevalidated-array constructor. Before mapping, the run admits 16K(q+Q)+128K known data bytes for K terms, q logical and Q native qubits. This covers sparse Pauli/symplectic representations, exact simplification and mapping. It is not SDK RSS or a statevector/dense-operator allocation. Revisit if that representation changes.
Resource forecast and explicit inspection bounds¶
backends.assessment defaults to 4096 point-assessment plus model-prediction rows. It counts the entire requested expansion before any model evaluation. The pure logical fold has no such forecast cap. backends.telemetry defaults to 100000 actual forecast/timing comparison pairs. These are dedicated output caps, and they do not change the scientific model domains.
Explicit circuit inspection defaults to 100000 input/output operations and 10 GB of known option/inventory/copy metadata, with an option nesting ceiling of 64. Its byte envelopes are 64 bytes per operation plus 16 per quantum/classical wire for inventory, or 96 bytes per operation plus the same wire term for an auxiliary compiler copy. src/nwqlib/backends/inspection.py::_options_snapshot charges each transpile option as 6*len+2 bytes for a string of len code points, max(8, bit_length+2) for an integer, 32 for None, a Boolean or a float, and 32+16*len for a list, tuple or dictionary. The total bounds the length of the options written as UTF-8 JSON (ensure_ascii=False, the form of the Run journal and record identities), whose longest form of one code point is the six-byte \u00XX escape of a control character. ASCII-escaped JSON can take 12 bytes for a code point outside the Basic Multilingual Plane and is not bounded by this charge. The basis snapshot charges 16 bytes per entry plus six per name code point, and the returned inventory 64 bytes plus six per code point for each distinct operation name and 256 bytes for its fixed fields. The option, basis and inventory charges measure the content of those values, not Python object memory. The per-operation and per-wire envelopes are untuned allowances for the instruction collections. None of these charges bounds installed SDK payloads or compiler RSS. Revisit with a concrete larger inspected artifact or changed representation. No cap grows automatically. Effective Qiskit optimization level and seed resolve through the installed compiler's configuration rules and are recorded with its version.
Explicit logical compilation and physical projection¶
Explicit compile_logical uses 100000 input/output operations, 10 GB transport/publication bytes and a 60 s child timeout. The operation populations are top-level source instructions and final compiled gates (or synthesis-stage count-only gates), and native intake separately admits nested stored data by bytes. Child stdout/stderr capture is capped at max_bytes. These reuse inspection-scale limits, contain a hung/noisy local dependency, and do not bound its internal RSS or synthesis work. No retries or cap increases are automatic.
Explicit estimate_physical reuses the 10 GB and 60 s process boundary. Its fixed QDK 1.32.3 model uses GateBased error 1e-4, gate and two-qubit gate time 100 ns and measurement time 500 ns, SurfaceCode distances 3, 5 and 7, the published Litinski19 table (Litinski, arXiv:1905.06903v3), PSSPC with 20 T states per rotation and ccx_magic_states=False, LatticeSurgery slowdown 1.0 and UnionBound max_error=0.01, without graph, post-processing or cache. These are disclosed finite model assumptions, not calibrated optima. The SurfaceCode formulas are QDK's, and Estimate fault-tolerant resources states which parts of them the papers QDK cites support. Counts-model rotation/CCZ/CCiX zeros require the admitted Clifford+T source. The times and error budget are explicit caller controls, and expanding the model set requires a new decision. The QDK time unit is ns, and public entries multiply by 1e-9 to report seconds.
Expectation¶
Expectation classical work and physical output¶
ExpectationMethod.max_classical_products=100_000_000 bounds known matvec products plus 2*D reduction products and the shared preparation_requirements count before preparation. Known bytes reserve preparation plus the shared matvec envelope and 32*D reduction/input workspace. Opaque SDK preparation cost remains unknown. These are logical populations, not RSS or a preemptive timeout.
Quadratic-form recovery uses PhysicalScale mantissa/exponent with the existing 4096-bit ExactArithmetic boundary to avoid an intermediate floating norm square. The sampling allocation replaces C by ||psi||²*C in the integer-safe Hoeffding calculation and keeps its absolute, sampling-only scope. No physical total-error or relative-error guarantee is introduced. Revisit either limit only for a concrete larger input/action or exact-scalar representation workload.
ExpectationMethod.sampling_shots evaluates the shots per QWC group 2*(C/epsilon)**2*log(2*L/delta), where L counts the nonzero nonidentity labels, from exact rationals in Decimal arithmetic with 80 significant digits. Division and multiplication round upward, and the correctly rounded logarithm is moved up by one unit in the last place, so the computed value exceeds the exact one at any precision. The precision sets only the size of that excess. At 80 digits the relative excess is below 1e-78, so for admitted counts below 2**53 the returned ceiling exceeds the exact ceiling only when the exact value lies within 1e-62 below an integer. The exact value is never an integer, because the logarithm of a positive rational other than one is transcendental. Revisit if the admitted count domain grows beyond binary64 integers.
src/nwqlib/operators/inputs.py::_scaled_observable_requirements admits one storage-preserving scaled vector action with B_scaled = B_action + 32*N + 3*P + 48*K bytes and W_scaled = W_action + 2*N + 18*K work. N is the state dimension, P the original operator payload bytes, and K the stored dense-entry, sparse-nnz or Pauli-term count. B_action and W_action come from _matvec_requirements, including the possible input conversion. For L single-word Pauli terms, g = min(L, N) and t = min(N, _PAULI_TILE), its action allowances are 49*N + 24*L + 16*g + 8 + max(9*L, 96*t) bytes and (L+g) + 3*L + N*(L+g+1) + N + L*N work. The additional 3P+48K bytes cover invocation-local scaled storage, component scans and the immutable Pauli snapshot. The 18K work covers up to four maximum/reduction visits, two ldexp visits, eight mask visits, one initial copy and three immutable Pauli-array copies per entry. The 32N bytes and 2N work cover vector framing and reduction. Pauli coefficient scaling occurs once per operator, so its 3P+48L storage and 18L work are independent of the number of state columns. Classical expectation adds state-preparation allowances. The power-of-two exponent is derived from the largest original real/imaginary component and saved with the scalar frame, and recovery combines exponents before binary64 conversion. These counts cover logical buffers and traversal, excluding Python object overhead and SDK RSS. Revisit when the action, storage-scaling or reduction kernels change.
Lanczos¶
Chebyshev Lanczos defaults¶
Sampled execution uses one coherent physical SELECT setting per odd degree and one reflection setting per unknown even degree. Its cached label-controlled product-basis transform is counted as SELECT_basis once per odd setting and per corresponding shot when X/Y rotations are present. Exact probabilities read every requested degree as a point of one walk trajectory. An even degree is read through PREP^dagger, the saved index marginal and PREP, and an odd degree through the basis change B, the saved index and system marginal and B^dagger. A tail with d native operations and an inverse with the same count add 2d operations to the preparation record's G, or d_forward+d_inverse when inverse lowering has a different count. The final observation needs no inverse when nothing follows it, so each odd point executes the transform once and each odd point before the last point executes its adjoint once. Diagonal I/Z-only SELECT requires no basis transform. The uncompressed input has O(n L) two-by-two table entries for L Pauli labels and n system qubits. This bounds table size, not construction or transpilation time. Address-bit dependency analysis can require O(n L log L) bookkeeping, and native compilation has additional unmeasured cost. This is an explicit implementation/cost contract, not a fitted numerical constant or a claim of minimum compiled gate cost. Revisit with an independently justified readout implementation and its shot, gate, and noise costs.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
krylov_dimension |
min(8, original_dimension) when omitted |
src/nwqlib/algorithms/lanczos/method.py |
Bounded projected subspace. Neither the dimension nor the initialization proves ground-state overlap. |
| Per-setting stage minimum | 2 | src/nwqlib/algorithms/lanczos/numerical.py |
At least two requested observations per pilot or main setting give an empirical variance when they are returned. It does not guarantee how many observations the backend returns. |
| Sensitivity pilot fraction | Explicit caller value in (0,1) | src/nwqlib/algorithms/lanczos/workflow.py |
One pilot and one main stage share the declared total shots, adjusted to the stage minima. There is no default pilot fraction. |
| Sensitivity main floor | min(actual returned pilot shots, uniform main allocation) | src/nwqlib/algorithms/lanczos/numerical.py |
Nonbinding allocations are kept. Binding floors are reserved first, and the remaining surplus of the donor settings is distributed, with ties broken by stable index. This heuristic keeps a setting from being starved of shots and is not a confidence bound. |
overlap_cutoff / overlap_cutoff_policy / overlap_noise_multiplier / overlap_failure_probability |
None / empirical / 1 / .05 | src/nwqlib/algorithms/lanczos/method.py, src/nwqlib/algorithms/lanczos/numerical.py |
Exact moments use the 1e-12 numerical floor. Sampled default regularization uses an empirical Gram RMS with adjustable multiplier 1, without a coverage or energy-error claim. The conditional Hoeffding Gram bound (doi:10.1080/01621459.1963.10500830) is reported separately. The explicit confidence policy uses twice that bound, so by Weyl's inequality the kept eigenvalues of the population Gram matrix exceed the bound. The Lanczos guide states the shot requirements and assumptions. Revisit the exploratory threshold and confidence level for the intended workload. |
| Gram input allowance | m*1e-12 deterministic. Zero for sampled admission | src/nwqlib/algorithms/lanczos/numerical.py |
Entry resolution delta gives Gram spectral resolution at most m*delta. This is an admission policy, not native error certification. |
| Automatic pilot separation | kept minimum and kept/discarded gap above 2*rho_S | src/nwqlib/algorithms/lanczos/numerical.py |
Two moving endpoints can each move by one empirical RMS. A fixed cutoff uses a one-RMS boundary margin. |
| Explicit recurrence normalization admission | 1e-12 | src/nwqlib/algorithms/lanczos/numerical.py |
Window for a supplied normalized vector, separate from eigenvalue accuracy. The input that supplies a native preparation is responsible for its normalization. |
| Projected analysis envelope | 512*m*m+128*m bytes; 32*m*m+16*m*m*m work |
src/nwqlib/algorithms/lanczos/numerical.py::admit_projected_analysis, called by src/nwqlib/algorithms/lanczos/method.py::Lanczos.plan and src/nwqlib/algorithms/lanczos/numerical.py::reconstruct |
Known index/pencil/eigenvector/output storage and projected solve. The envelope depends only on m, so planning checks it against max_bytes and max_analysis_work before any acquisition. reconstruct checks it again against the limits of the Method that runs the analysis, which may differ from the planning Method. When the Plan acquires moments under SensitivitySampling, planning also admits the pilot with the same byte law and twice the work, one share for its solve and one for its analytic derivative, which runs after the solve and fits the same byte envelope. These are logical size laws, not CPU/RSS. |
| Pauli census and readout constructor | 65536 + M*(16*W+56) bytes for M source rows and W uint64 words per mask |
src/nwqlib/algorithms/lanczos/readout.py::LanczosReadout |
Checked before the census allocates. The data law covers simultaneous support, coefficient, index and sign arrays and the constructor's owning row copy. The input Pauli table is excluded. The fixed 65,536-byte allowance covers headers and bookkeeping on the qualified stack, including empty and small tables. It is not a bound on process RSS. Revisit when the interpreter, NumPy or the census allocation schedule changes. |
| SELECT/K copies with mask ingestion | J*(80*W+88+S(q)) + (q+8)//8 + 1024 + 65536 bytes for J selected rows, W words per mask and S(q) = src/nwqlib/operators/inputs.py::_pauli_identity_bytes(q) |
src/nwqlib/algorithms/lanczos/method.py::centered_operator_bytes, _centered_operator |
One check before the copies are allocated covers the four caller-owned arrays, J*(16*W+24) data bytes plus 256 bytes each, live throughout ingestion, the ingestion envelope J*(4*(16*W+16)+S(q)) + (q+8)//8 and 65,536 bytes of other fixed bookkeeping. The ingester then receives the remaining limit. The supplied table and readout are excluded. The 256-byte per-array allowance exceeds the 112-byte rank-one and 128-byte rank-two ndarray headers on the checked stack. It is an engineering allowance, not a NumPy guarantee, and does not bound process RSS. Revisit for other array ranks, allocators or interpreter and dependency versions. |
MAX_READOUT_ENTRY_BITS |
20 | src/nwqlib/algorithms/lanczos/method.py |
Planning admits at most 2**20 entries in one exact probability marginal or one histogram of a fixed per-setting shot count. Each exact trajectory point reads one marginal. Exact probability readout is therefore limited to 20 measured qubits. SensitivitySampling plans at its two-shot stage minimum, so the Run's data limit bounds its stage histograms. An engineering cap on returned data size, not an accuracy parameter. Revisit for a wider exact readout. |
LCHS¶
LCHS defaults and error allocation¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| LCHS construction tolerance | .01 | src/nwqlib/algorithms/lchs/method.py |
Untuned starting budget for the kernel-integral approximation and quadrature. Input norm, PSD recovery and other error sources determine its effect in physical units. |
| LCHS kernel exponent and component split | beta=.75, half the construction tolerance per component | src/nwqlib/algorithms/lchs/providers.py |
Beta lies in ACL arXiv:2312.03916v2 Sec. 2.2's practical range. Equal budgets separately constrain the cutoff tail and k quadrature. Neither allocation controls every physical-output error component. |
Composite-Gauss truncation_multiplier |
1.0 | src/nwqlib/algorithms/lchs/providers.py |
Multiplies the cutoff obtained from the finite ACL Eq. (186) tail budget (arXiv:2312.03916v2). Values at least one preserve the allocated tail bound. The default uses the unexpanded certified cutoff. |
| Constant-source time quadrature | 8 nodes | src/nwqlib/algorithms/lchs/method.py |
Untuned finite Gauss-Legendre rule for the Duhamel integral. It sets work and approximation order, without a default Duhamel error certificate. |
| Low–Somma shift parameter | c=1 | src/nwqlib/algorithms/lchs/providers.py |
A working point also used in Low–Somma arXiv:2508.19238v2, Figure 1. The published formula accepts other positive c, whose error and cost must be assessed for the selected workload. Planning refuses a c whose coefficient 1-norm makes binary64 rounding exceed LCHS_ROUNDING_EPSILON_FRACTION of the tolerance. |
lchs/method.py |
approximation_tolerance=.01 |
src/nwqlib/algorithms/lchs/method.py |
The default pair splits this tolerance equally between tail and k quadrature. QSP synthesis and each automatic Trotter branch separately receive 0.1 times the tolerance. These are component allocations, not a total physical-output budget. The constant-source quadrature is registered above. resolve_lchs_coefficient_plan in src/nwqlib/algorithms/lchs/providers.py uses the same default. Revisit condition: The requested component accuracy and available branch budget. |
LCHS_QSP_EPSILON_FRACTION (lchs/native.py) |
0.1 | src/nwqlib/algorithms/lchs/native.py |
Jacobi-Anger allocation as a fraction of the LCHS epsilon. The circuit stage bounds apply the path-specific physical coefficient-state prefactor once to the compensated compiled-QSP bound. |
LCHS_TROTTER_EPSILON_FRACTION (lchs/native.py) |
0.1 | src/nwqlib/algorithms/lchs/native.py |
Per-node [CSTWZ] budget (doi:10.1103/PhysRevX.11.011020) before physical application weighting. Intentional Pauli-pruning cost is deducted first, and budgeted selection uses the Pauli-triangle structural coefficient once. Fixed-step paths do not evaluate a coefficient by default. |
LCHS_ROUNDING_EPSILON_FRACTION |
.1 of approximation_tolerance, compared with u*alpha |
src/nwqlib/algorithms/lchs/providers.py::_admit_coefficient_rounding |
Planning refuses a coefficient table whose 1-norm alpha makes the binary64 rounding scale u*alpha (u = 2-53) of the finite sum exceed this fraction of the construction tolerance. Each term of the sum is rounded at least once, and the exact solution of the shifted problem has norm at most the input norm, so a large alpha means cancellation with rounding of order u*alpha times the input norm, plus up to (N-1)*u*alpha from recursive summation of N terms (Higham, Accuracy and Stability of Numerical Algorithms, 2nd ed., doi:10.1137/1.9780898718027, Eq. (4.4)). The LCHS error model reports floating-point error as unknown, so the kernel and quadrature bounds, which describe exact arithmetic, are only published when this rounding is small against them. The fraction equals the allowance of each realization stage, LCHS_QSP_EPSILON_FRACTION and LCHS_TROTTER_EPSILON_FRACTION. It matters for low_somma_f2, whose alpha is about exp(c)*erfc(1/(2*gamma)) (Low and Somma, arXiv:2508.19238v2, Theorem 2). At c = 50 this is about 1.2e16, so u*alpha is about 1.3. Revisit with a derived bound for the complete floating-point term. |
LCHS_ELLIPSE_STRIP_MARGIN |
.1 | src/nwqlib/algorithms/lchs/providers.py |
Keeps the Bernstein ellipse in the analytic strip at distance .1 from kernel singularities. It bounds the kernel by 1/(C_beta*d), not a fit to sampled errors. Revisit only with a justified analytic bound and disclosed node/slot cost. |
DEFAULT_PSD_TOLERANCE / psd_tolerance default |
1e-12, with decision window psd_tolerance * max(abs(lambda_min), abs(lambda_max)), that is psd_tolerance * norm(L, 2) |
defined in src/nwqlib/algorithms/lchs/solution_error_budget.py. Used by src/nwqlib/algorithms/lchs/time_independent_terms.py, src/nwqlib/algorithms/lchs/quantum.py, and LCHS option records |
One scale-aware numerical PSD policy for public decomposition, host/quantum selection and formula planning. Values inside the window are accepted unshifted. Values below it raise or shift. A backward-stable Hermitian eigensolver returns eigenvalues within p(n)*u*norm(L, 2) of the exact ones (LAPACK Users' Guide, 3rd ed., Sec. 4.7), so the window is relative to that norm, and 1e-12 covers p(n) up to about 9000. The window has no absolute term, so the decision does not depend on the time unit. Replacing A by r*A and T by T/r, for r > 0, scales the window and the shift by r. An admitted unshifted violation lets norm(expm(-A*T), 2) reach at most exp(psd_tolerance*norm(L, 2)*T) (ACL arXiv:2312.03916v2, Lemma 21, Eq. (162)). The comparison is constant-time after the selected existing spectral estimate. Revisit if supported evolution horizons reach norm(L, 2)*T of order 1/psd_tolerance or a user requires formal PSD proof. |
MPS threshold defaults / num_layers |
1e-14 / 2 | src/nwqlib/algorithms/lchs/method.py |
Selected TT-SVD threshold and layered-circuit controls. The fidelity of the built circuit requires a separate comparison. |
lchs/time_independent_terms.py::_PAULI_COEFFICIENT_ATOL |
1e-12 | src/nwqlib/algorithms/lchs/time_independent_terms.py |
After forming H+kL, coefficients with magnitude at or below this absolute cutoff are removed and a finite upper bound on their summed moduli is recorded (src/nwqlib/subroutines/trotterization/error_budget.py::upper_dropped_mass). It also admits negligible imaginary parts in identity/angle processing and in the classical rotation sequence of src/nwqlib/algorithms/lchs/host_pf.py. Units are those of the combined coefficient. The cutoff is absolute rather than relative to operator scale. Revisit against the selected operator scale and transformation. |
lchs/time_independent_terms.py |
64*eps*comparison_scale |
src/nwqlib/algorithms/lchs/time_independent_terms.py |
Untuned relative comparison window between an existing structural coefficient and an explicitly computed dense coefficient. It introduces no default reference calculation. Revisit condition: A changed coefficient derivation or precision. |
lchs/quantum.py::observable_terms and lchs/selection.py::plan_lchs |
dense observable dimension at most 16 | src/nwqlib/algorithms/lchs/quantum.py, src/nwqlib/algorithms/lchs/selection.py |
Both read src/nwqlib/algorithms/_eigen_inputs.py::AUTO_DENSE_PAULI_DIMENSION, the small explicit-dense convenience limit that Expectation also applies. A quantum observable is read from its Pauli terms, by the exact projected reduction or by one counts setting per qubit-wise commuting group, and only a small dense matrix is converted to Pauli terms automatically. Larger observables are supplied as Pauli terms. Revisit condition: An explicit, work-bounded dense-to-Pauli conversion option for LCHS observables. |
LCHS work and byte limits¶
LCHS dense Pauli decomposition uses the block-transform and output-conversion law derived in _admit_pauli_decomposition. For d = 2^q, its current string-table reservation is (34*q+28)*d**2 + 2*H(q) work units and (104+2*q)*d**2 + T(q) logical payload bytes. With h = floor(q/2), T sums j*4**j and H sums (j+1)*4**j over j = h for even q, or j = h and q-h for odd q. The transform count includes overflow-safe averaging. The output count includes label production, coefficient records and dictionary hashing. The byte count follows sequential L/H transforms and their live arrays and tables. Revisit the law when the transform's allocation lifetimes or output representation changes. The shared work and byte defaults are 100,000,000 and 10,000,000,000 respectively.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
lchs/method.py construction limits |
dense SELECT slots 256; SVD, spectrum, quadrature and SELECT work each 100,000,000; QSP degree 256 and evaluations 20,000; selected product-formula steps 100,000; max_admission_steps 1,000,000 |
src/nwqlib/algorithms/lchs/method.py |
Bounded selected construction/preprocessing, with a separate shared byte cap. max_trotter_steps caps the budgeted per-node step counts and the periodic Strang count that trotter_synthesis_tolerance selects, before the periodic construction law is charged. max_admission_steps is the Programs' AdmissionLimits.max_steps, with ten times the shared default (Shared Program admission limits). Work is an operation-size law, not CPU or RSS. Revisit for a deliberately larger construction, preserving before-work admission. |
lchs/refinement.py::LCHSRefinement limits |
dense and structural work each 100,000,000; node evaluations 4096; steps 100,000 | src/nwqlib/algorithms/lchs/refinement.py |
Explicit refinement only, over the already selected model. Counts bound this requested operation, and no diagnostic loop runs automatically. Revisit with its workload and the evidence kept for it. |
lchs/verification.py::LCHSVerification limits |
dense work 100,000,000; node evaluations 4096; PF operations 100,000 | src/nwqlib/algorithms/lchs/verification.py |
Whole-call reference caps. IVP additionally requires explicit tolerance and RHS-call limits. These limits do not certify reference accuracy. |
lchs/quantum_resources.py::sample_resources limits |
13 qubits 100,000,000 build work 100,000 inspected operations |
src/nwqlib/algorithms/lchs/quantum_resources.py |
Untuned policy caps of one explicit representative inventory, which simulates nothing. The qubit cap refuses a Plan wider than 13 qubits before any representative is built, and it bounds every representative circuit that Qiskit constructs or, on request, transpiles. That SDK work and memory lie outside the build-work proxy. No array or work law of this inspection gives 13. A dense exact SELECT builds no representative. It is reported by the value-independent structural upper-bound record of its synthesis and control route (quantum_resources.dense_representative_envelope), an integer evaluation charged no build work. On the dense route the qubit cap, the PREP and readout build work and the operation cap remain. A dense-dilation QSP child adds the synthesis of its unitary on q + 1 qubits and Qiskit's control of it in the controlled query and, with two active children, in its generator branch, where the whole-matrix route charges the whole-matrix synthesis of its controlled dilation on q + 2 qubits instead. Product-formula charges grow with the label count times 2**a and with the number of stored angles (template occurrences times repetition blocks times 2**a). QSP, PREP and readout charges grow with 2**a, 2**q and the width. The build work bounds the summed gate and table size proxy, and the operation cap bounds all inspected representative operations. None of the caps bounds SDK runtime or process RSS. Revisit for a deliberately wider inspection, such as a product-formula Plan above 13 qubits, or when the default build-work cap changes. |
| LCHS readout work cap | max_readout_work=2_000_000_000 |
src/nwqlib/algorithms/lchs/method.py, src/nwqlib/algorithms/lchs/quantum.py |
Inclusive sum of evaluated sampled QWC comparisons and registered exact projected-reduction input visits. max_select_work=100_000_000 separately funds SELECT construction. The readout default covers the reducer envelope at width 20 with up to 513 coalesced kept terms, subject to the other limits. The cap is an admission count, not a runtime estimate. Revisit when the grouping or reduction law, or the supported default execution domain, changes. |
_ELLIPSE_CANDIDATE_WORK |
32 scalar units | src/nwqlib/algorithms/lchs/providers.py |
Conservatively accounts for one candidate's fixed arithmetic/log/exp/order calculation including its single rounding correction. The composite-Gauss provider first completes its scalar panel search. It stops when the minimum node count of every later panel count exceeds the best rule found. It then checks max_quadrature_work against the search charge, 32 units per selected node and the cubic Gauss-rule construction charge, before allocating the node arrays. A refusal names the complete charge of that selected rule. With M candidates and the selected rule of total = 2mQ nodes, the charge is 32M + 32*total + Q**3, checked once after the search, so LCHS.max_quadrature_work admits the search charge plus the selected rule's construction, not every scalar probe. The search tries at most floor(Q1/2) candidates after its first valid Q1. Not a CPU/RSS bound. Revisit if candidate computation changes. |
| LCHS fixed byte and work envelopes | 256 bytes and 32 scalar work units per k-node in the src/nwqlib/algorithms/lchs/providers.py quadrature builders, 8192 bytes per SELECT slot in src/nwqlib/algorithms/lchs/time_independent_terms.py::_build_product_formula_select_plan, 128*K*(u+1)*(q+1) + 8192*K bytes for the K cached combined product-formula nodes with u union labels in src/nwqlib/algorithms/lchs/time_independent_terms.py::_node_cache_bytes and 16 scalar work units per application and node for a step selection from a cached coefficient, 128*(u+1)*(q+1) bytes for the aligned L/H union rows, which bounds their cached rows and construction containers under the 64-bit CPython container allowance for q >= 1, and, per application and quadrature position, 8192 bytes plus the magnitude-width charge of every integer field for the host step records in src/nwqlib/algorithms/lchs/host_pf.py::host_position_record_reserve, with an 8500-bit step count for a budgeted selection, 8192 bytes per representative inventory in src/nwqlib/algorithms/lchs/quantum_resources.py::sample_resources, and 8192 bytes added in src/nwqlib/algorithms/lchs/verification.py::_admit_publication |
src/nwqlib/algorithms/lchs/providers.py, src/nwqlib/algorithms/lchs/time_independent_terms.py, src/nwqlib/algorithms/lchs/host_pf.py, src/nwqlib/algorithms/lchs/quantum_resources.py, src/nwqlib/algorithms/lchs/verification.py |
Untuned per-item envelopes for Python tuples, dictionaries, records and per-node scalar evaluations that the array terms of the same law do not describe. They keep admission before work. They are not measured allocations, operation counts or RSS bounds. Revisit condition: A change in the stored record contents or per-node evaluation, or a workload whose admission approaches the byte or work cap. |
lchs/time_independent_terms.py::_build_product_formula_select_plan |
angle-table construction, 8E data bytes |
src/nwqlib/algorithms/lchs/time_independent_terms.py |
With B compressed repetition blocks, W template occurrences, S padded slots and E = BWS, the float64 table uses 8E data bytes during construction and storage. The fill loop checks finite values and canonicalizes signed zero, then a private FrozenArray ownership transfer publishes a read-only view of the same buffer. Optional affine validation adds 24S data bytes for three temporary float64 row arrays. Admission adds the applicable array terms to selection_bytes, which includes the padded k/time/coefficient arrays and declared record allowances. These known array-data terms do not bound process RSS. Revisit condition: A change in producer validation, buffer ownership, dtype, or affine-check allocation lifetimes. |
lchs/time_independent_terms.py::_spectral_host_requirements and _spectral_branch_requirements |
the dense exact LCHS route, W_host and branch work |
src/nwqlib/algorithms/lchs/time_independent_terms.py |
Host: W_host = D + K_eig(8D**3 + 34D**2) + K((r+1)D**2 + (2r+3)D + 6AD) for K nodes, A applications and r distinct input vectors, with K_eig = K, one for an exactly zero L and zero for identity actions, and phase bytes max(32D**2+16D, 48D**2+32D, 16D**2+(160+32r)D) beside the live inputs. Explicit branches: K_eig(8D**3 + 34D**2) + B(D**3 + 2D**2 + 3D) for B nonzero-time branches plus the selected synthesis, control and preparation work, with the pre-synthesis frontier 64D**2 + 24D, the live L and H, cached eigensystems 16D**2 + 8D each and completed circuits. One Hermitian eigensystem per node, one matvec per input and node, one final action per node and one dense product per explicit branch. The queried LAPACK workspace Q_V(D) of NumPy's zheevd is not charged. The laws count the known NumPy arrays and are declared logical workspaces, not complete workspace caps. These units are admission proxies, not timings or equal-cost CPU operations. Revisit condition: A changed eigensolver, branch formation or host accumulation, or a deployment that qualifies its LAPACK workspace. |
Explicit RK45 verification tolerance¶
LCHSVerification._reference_parameters rejects rtol below 100 * sys.float_info.epsilon. SciPy's integrate._ivp.common.validate_tol otherwise replaces such a request with this float64 floor. Rejecting before reference work keeps the selected and recorded numerical tolerance identical to the one used. This is the existing RK45 dependency's domain, not an accuracy certificate. Revisit if the selected solver or supported native dtype changes. The comparison is strict, so the floor itself is accepted, and tests/test_lchs_verification.py checks that a smaller requested tolerance is rejected when the options are constructed, before any acquisition.
LCHS circuit-free synthesis laws¶
These laws describe the selected synthesis model and are separate from measured compiler output, like the shared circuit-free synthesis laws. Revisit them when the construction or compiler changes, using independent matched-workload evidence.
The compiled QSP SELECT law (src/nwqlib/algorithms/lchs/compiled_selection.py::compiled_select_cx_projection) prices the leaf that native.construct_select builds with Qiskit 2.5.2. Qiskit's add_control unrolls a gate to its basis and controls every basis gate, so each gate is priced by its census kind at the controls it receives (_controlled_kind_cx). Under one control a CX costs 6 CX, an X, Y, Z or H 1 and every other one-qubit gate 2. The parity qubit controls each QSP pass, and with two generator children the combine qubit first controls each branch, so a branch gate is controlled twice in succession. The second control unrolls the output of the first, so the gate costs the one-control price of its one-control census (_ONE_CONTROL_UNROLLED_CENSUS). That is 52 CX for a CX, 16 for an H, 22 for an RY or a U, 18 for a P and 2 for the branch phase, where two controls at once would charge a CX 14. The children are counted as the SELECT unrolls them. A dense dilation is controlled through the exact synthesis of src/nwqlib/subroutines/_dense_synthesis.py, with the gate bounds of _dense_synthesis.dense_synthesis_gate_census on n + 1 qubits. On the whole-matrix route of dense_control_route a dense child beside another child is the controlled dilation on n + 2 qubits (controlled_synthesis_gate_census), which the parity qubit controls once, and the branch's RY multiplexor stays twice controlled. A Pauli child has one U gate for each entry of each dependency-projected UCG on r address bits, 2**r - 1 core CX and, for r >= 1, a completion diagonal of 2**(r+1) - 1 RZ and 2**(r+1) - 2 CX. A banded child has its QFT pair, its UCRZ tables and its address diagonal. The projector flips use LCHS_QSP_PROJECTOR_FLIP_CX below, and each reflection is an X with N - 1 controls from MCX_CX_BY_CONTROLS, for N = b + 3 ancillas. Transpiled built leaves (basis cx,u, optimization level 0) with dense, banded and Pauli children on one to four system qubits, one to three address bits and two to seven block ancillas stayed at or below the law, with ratios from 0.82 to 0.97. The census constructs no matrices or circuits during estimation. The value is recorded as an estimate because another Qiskit version can control or lower the gates differently. Revisit when the QSP evolution, qiskit_compat.controlled or the Qiskit version changes, and whenever tests/test_lchs_resource_structural_law.py fails.
A direct preparation is priced at the costliest construction that _build_normalized_state_preparation can take (compiled_selection._direct_preparation_paths). X gates prepare a basis state and H gates equal amplitudes, at most n gates either way. The prefix-uniform fast path has the elementary slot envelope of _uniform_superposition_gate_slots in src/nwqlib/_preparation_laws.py, which includes the open CH/CRY control wrappers. The magnitude tree has 2**n - 1 RY and 2**n - 2 CX, and a complex state adds a phase diagonal with as many RZ and CX. Each construction is priced at the control level the SELECT adds, and the largest price counts. A QSP child counts its positive coefficient PREP and the inverse. A constant-source input (compiled_selection._controlled_direct_preparation_cx) adds one controlled global phase. Basis, equal, uniform-prefix and general complex states on one to three qubits with 1, 2, 3, 4 and 8 controls stayed at or below the price (test_controlled_source_preparation_bounds_every_construction).
The LCHS dense_exact SELECT law (src/nwqlib/algorithms/lchs/native.py::dense_branch_select_cx) charges each physical branch, on the gate-wise route of dense_control_route, the exact synthesis of its q-qubit unitary with every gate controlled on the a address bits, as qiskit_compat.controlled and Qiskit 2.5.2's add_control build it. On the whole-matrix route, which the default "auto" takes for a = 1, a branch on q >= 2 qubits is one whole-matrix synthesis on q + a qubits, at most (25/96) 4**(q+a) - 2**(q+a) + 4/3 CX, which a Haar-random branch reaches, and a constant-source branch includes its preparation in that synthesis. The gate census of _dense_synthesis.dense_unitary_circuit is bounded by (25/48) 4**q - (3/2) 2**q + 2/3 CX, 7*4**(q-2) U, (3/8) 4**q - (3/2) 2**q RZ and (4**q - 16)/24 H gates for q >= 2, and one U for q = 1. Controlled on a bits, a CX costs the X with a + 1 controls and an H the X with a controls from MCX_CX_BY_CONTROLS. An RZ becomes a multi-controlled RZ (MCRZ) of 2 CX for a = 1 and 2*_dirty_mcx_cx(ceil(a/2)) + 2*_dirty_mcx_cx(floor(a/2)) CX otherwise (src/nwqlib/subroutines/hamiltonian_evolution/pauli_evolution.py). A U costs 2 CX for a = 1 and otherwise two MCRZ, one multi-controlled RY (MCRY, 12 CX for a = 2, 20 for a = 3 and the MCRZ count from a = 4) and _mcphase_cx(a), and the branch phase costs _mcphase_cx(a). These are the counts without ancillas. For a from 1 to 8 with one or three idle qubits, clean or already used, no gate kind lowered to more (test_controlled_gate_costs_match_qiskit_and_bound_every_context). On the gate-wise route, transpiled built SELECT leaves of LCHS plans with q = 1..4 and a = 1..3 (basis cx,u, optimization level 0) stayed at or below the law, with ratios from 0.77 to 1.0 and equality at q = 1 and at q = 2, a = 1. The QSD leading term (23/48) 4**(q+a) per branch is below the built count, for example 491 against 752 CX in total for the four branches at q = 2 and a = 2. The value is recorded as an estimate because another Qiskit version can control or lower the gates differently. For a constant-source layout on the gate-wise route the preparation inside each branch is priced separately by compiled_selection._controlled_direct_preparation_cx. Built leaves with a constant source on one to three system qubits and two to eight address bits stayed at or below the total, with ratios from 0.61 to 0.94. Revisit when src/nwqlib/subroutines/_dense_synthesis.py, qiskit_compat.controlled or the Qiskit version changes, and whenever tests/test_lchs_resource_structural_law.py fails.
The branch-controlled product-formula SELECT law (src/nwqlib/algorithms/lchs/select_synthesis.py::_branch_controlled_product_formula_resource_law) prices every gate of every branch at the a address controls (compiled_selection._controlled_kind_cx). Qiskit 2.5.2 synthesizes an occurrence of a one-qubit Pauli label as one RZ, RX or RY, and a longer label as an H pair for each X factor, an SX and SXdg pair for each Y factor, 2*(support - 1) parity CX and one RZ. A parity CX costs the X with a + 1 controls, an H the X with a controls, an SX the MCPhase on a + 1 qubits, an RZ the MCRZ, and a lone RX or RY Qiskit's "noancilla" MCRY, a Gray-code circuit of 3*2**a - 4 CX for a <= 3 and the MCRZ count from a = 4. Transpiled built leaves with one to three system qubits and one to seven address bits, Lie and Suzuki formulas, with and without a constant source, stayed at or below the law, with ratios from 0.57 to 1.0 and equality for one-qubit labels. The loosest ratios, at three system qubits and four or five address bits, come from Qiskit lowering the multi-controlled X gates with idle system qubits as ancillas, for example 42 CX for a parity CX under four address controls with one idle qubit, against the 84 of the law. The phase contribution of branch-controlled LCHS SELECT uses the same _mcphase_cx count as QHD's MCPhase provider. A phase on an a-bit address is a phase gate controlled on a-1 bits. Qiskit's definition emits MCRZ operations with decreasing control counts. Each MCRZ with at least two controls uses two dirty-MCX pairs, and the one-control MCRZ is a CRZ with two CX. A generic diagonal-table formula does not price this constructor, for example at six address bits it gives 62 CX while this MCPhase decomposition has 84. The calculation uses scalar control counts without constructing a gate or a phase table.
LCHS_QSP_PROJECTOR_FLIP_CX in src/nwqlib/algorithms/lchs/compiled_selection.py gives the CX count of a QSP projector flip, the X gate with k open controls on the block ancillas, after Qiskit adds the parity control of the pass. That control makes add_control unroll the flip through its definition, not through the synthesis that MCX_CX_BY_CONTROLS records, and control each gate once more, so an entry charges 6 CX for each CX of the definition, 1 for each X or H and 2 for each other one-qubit gate. The counts are 8, 56, 122, 336 and 746 CX at k = 1 to 5 and 5092 at k = 10. From k = 5 the definition is Qiskit's synth_mcx_noaux_v24, whose CX count is _mcphase_cx(k + 1) and whose P, H and RZ counts have no published closed form, so the table stores the values. It covers 32 block ancillas. A Pauli child on q system qubits has at most 2q address qubits, so a generator exceeds the table only from 16 system qubits, where the outer classification gate of _compiled_qsp_part_plans already refuses at the default max_bytes (its 96*d**2 matrix reserve alone is 412,316,860,416 bytes), and compiled_select_cx_projection then raises. test_qsp_projector_flip_table_matches_installed_qiskit recomputes every entry with the installed Qiskit and compares k up to 5 with the transpiled build. Revisit when that test fails after a Qiskit upgrade or when the projector phase changes its construction.
QPE¶
QPE defaults¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| QPE static schedule defaults | QCELS 16 times/grid 4096, SPE 32 paired draws/grid 4096, RFE 97 paired draws/49 frequencies | src/nwqlib/algorithms/qpe/method.py |
Untuned finite workloads. SPE and RFE make one query per distinct drawn (power, quadrature) and keep its draw multiplicity in the estimate. Exact readout evaluates each query once on one controlled trajectory. Shot-mode SPE and RFE request 64 and 194 times shots coherent shots, and a query drawn m times acquires m times shots. At RFE's defaults the expected number of distinct settings is about 84.7. K sets RFE phase spacing 2π/K and maximum power K-1. These defaults do not establish a confidence theorem. |
| SPE Fourier filter | degree d=11, beta=6 | src/nwqlib/algorithms/qpe/method.py, src/nwqlib/algorithms/qpe/numerical.py |
WBC arXiv:2110.12071v2 PDF Eq. (A2), HTML Eq. (17), gives the coefficients. Maximum frequency is 2d+1. With epsilon1=epsilon2=eta/8 and epsilon3=eta/2, Theorem 3 has two premises. Its beta premise, beta >= max(W(2/(piepsilon3^2))/(4sin(delta)^2), 1) with W the principal Lambert W function, gives the smallest transition half-width delta that beta supports, and PDF Eqs. (A6)-(A7), HTML Eqs. (21)-(22), give the degree premise that spe_filter_domain checks for d. The resulting filter bound is 3eta/8, leaving eta/8 below the decision threshold for sampling error. A finite sample count is not automatically sufficient. The overlap lower bound is a caller premise. |
QCELS finite complex search (QCELS_BRACKET_EVALUATIONS) |
default maximum power 10, effective G=max(requested grid, 8*power span+1), 64 evaluations per local bracket, 32 scalar work units per complex point and evaluation | src/nwqlib/algorithms/qpe/numerical.py, src/nwqlib/algorithms/qpe/method.py |
Dimensionless finite search that keeps the original grid candidates. Operation-size bounds and grid resolution are neither coverage nor a global-optimum theorem. No QCELS uncertainty interval. The 62 golden-section iterations after two initial points shrink a bracket of two grid cells by a factor of about 1.1e-13. At the default grid the final width, about 3.4e-16, is below the binary64 spacing 4.4e-16 near pi. The work law and the saved-fit check use at most (1 + QCELS_BRACKET_EVALUATIONS)*G evaluations. |
QCELS and RFE signal resolution (_signal_vanishes) |
count data: exact zero. Exact data: sqrt(2)*mean_window*W + gamma_(2K)*sum_k w_k*|z_k| for K samples with positive integer weights w_k summing to W |
src/nwqlib/algorithms/qpe/numerical.py (qcels, rfe) |
A power carries a signal only when the sum of its samples exceeds what the admitted means and the summation can produce from an exact zero. mean_window is the window that admitted each exact mean, PreparedArtifact.probability_window for exact probabilities and NUMERICAL_RELATION_RTOL for a host kernel. The QPE Method passes it from src/nwqlib/algorithms/qpe/method.py::_mean_window, the largest preparation-record window of a Run with exact probabilities, NUMERICAL_RELATION_RTOL for classical execution and None for count data. The rounding term follows Higham, 2002, doi:10.1137/1.9780898718027, Lemma 3.5 for the complex additions and the standard model (2.4) for the real weights. A repeated power in count data sums several means, and an exact cancellation among them can leave a roundoff residue that counts as a signal. The decision sets QCELS aliasing and flatness and RFE flatness. Revisit with those windows. |
QPE_INTERVAL_CONFIDENCE |
0.95 | src/nwqlib/algorithms/qpe/numerical.py, src/nwqlib/algorithms/qpe/records.py::QPEInterval.level |
Nominal level of the RWPE moment-matched Gaussian interval, mean plus or minus 1.96 standard deviations. It is a model level, not frequentist coverage, and the single-phase Gaussian model can fail for mixed spectra. Revisit together with QPEInterval.level, which stores the same value. |
| RWPE Gaussian prior and horizon | mean=0 radians, std=pi radians, 14 steps | src/nwqlib/algorithms/qpe/method.py, src/nwqlib/algorithms/qpe/controller.py |
A Gaussian centered on the principal phase branch with width pi is a broad starting model. Fourteen one-bit updates are an untuned finite workload. The exact model width shrinks by sqrt((e-1)/e) per step, while the basic walk has a finite exploration range. Revisit the prior and horizon together for the intended spectral population and longest evolution time. |
| QPE automatic time | margin .9, principal radius pi, SPE cap pi/3 | src/nwqlib/algorithms/qpe/method.py, src/nwqlib/algorithms/qpe/numerical.py |
The .9 factor is an untuned interior margin. QCELS, RFE and continuous-time RWPE use the principal spectral radius pi. SPE uses min(pi/3, pi/2-delta) when its filter transition delta is available, preventing evaluation in the periodic wrap transition. An unavailable filter proof remains visible. |
QPE minimum_overlap |
SPE's declared eta, otherwise .9 when omitted | src/nwqlib/algorithms/qpe/records.py::QPEVerification |
Criterion for an explicit nominal projector comparison. SPE checks the lowest cluster, and other estimators check the largest prepared cluster. Revisit with the desired scientific target. This does not alter acquisition or establish total error. |
| QPE controlled_power_error_budget | 1e-3 | src/nwqlib/algorithms/qpe/method.py |
Per-power allowance for the complete structural subtotal of a product-formula power: pruning, the product formula at the represented step, represented-time displacement and identity-phase and rotation-angle formation. Native numerical/synthesis error remains unknown. |
qpe/method.py::_QPEMethod.pauli_pruning_rtol |
1e-12 | src/nwqlib/algorithms/qpe/method.py |
Relative-to-largest-coefficient cutoff after coalescing. Only nonidentity terms are pruned, and the identity remains in the phase. Removed L1 mass enters the selected power-evolution allowance. Revisit for the requested operator scale/error budget. |
| QPE group_atol | 1e-8 by default | src/nwqlib/algorithms/qpe/records.py |
Explicit nominal numerical clustering tolerance, in the Problem's energy unit for a Hamiltonian target and in radians of eigenphase angle for a unitary target. Not a proof of exact degeneracy. |
| QPE identity phase | exact zero only | src/nwqlib/algorithms/qpe/method.py |
Each emitted increment P(-(p_k-p_(k-1)) tau identity) between adjacent static power positions, or P(-power tau identity) at an RWPE time, is kept on the control and omitted only when it is exactly zero. No small-phase cutoff. |
| QPE projector roundoff window | n*eps/(1-n*eps)*max(1, norm(psi)**2) with n = 16*D*k |
src/nwqlib/algorithms/qpe/verification.py::_projector_weight |
Admits a QR projector weight slightly outside [0, 1] for D rows and k cluster vectors, with eps = 2**-52, twice the unit roundoff u. The first factor is therefore Higham's gamma_(2n) = 2nu/(1-2nu) (Accuracy and Stability of Numerical Algorithms, 2nd ed., SIAM 2002, doi:10.1137/1.9780898718027, Lemma 3.1). The factor 16 is an untuned allowance for the inner products, squares, cluster sum and QR orthogonality loss, not a derived bound. Revisit with a derived Householder-QR bound or if a legitimate weight is rejected. |
QPE selected numerical controls¶
The shared QPE-family controls max_power=100000, max_trotter_steps=1000000, max_work=1000000000 and max_bytes=10_000_000_000 bound the selected finite schedule, native step construction, known numerical/controller work and arrays. A dense block charges its matrix arithmetic, D**3 units per product of two D-square matrices (src/nwqlib/algorithms/qpe/powers.py::_dense_block, with the eigendecomposition of a Hamiltonian charged by _linalg_laws.hermitian_eigensystem_work), and the exact synthesis of its controlled matrix on n+1 qubits (_dense_synthesis.controlled_synthesis_size, see DEFAULT_MAX_LCU_WORK). A unitary input's polar base is charged once per Plan, 9*D**3 + 8*D**2 work units (src/nwqlib/algorithms/qpe/powers.py::polar_base), and a unitary block's bytes include every cached binary square of that base, (96+16J)*D**2 plus the synthesis bytes and 256(J+1)+65536, for the largest square index J of the Plan's exponents. Each block is admitted on its own, and a dense block on seven system qubits needs about 2.3e8. The default max_work is a fuse against runaway planning work, sized so that no example or test workload in the repository reaches it. The exact trajectory constructs one block per distinct gap between consecutive powers, and the Plan's total over its blocks is not compared with max_work. src/nwqlib/subroutines/qpe/coherent.py::build_coherent_qpe_circuit takes the same max_work and max_bytes defaults and charges its m powers of a D-square unitary m*D**3 work plus m times the work of controlled_synthesis_size(log2(D)+1), and 96*D**2 bytes plus the working bytes of one synthesis and the kept bytes of all m, before its unitarity check. They are separate from the work a Run submits and do not bound SDK or LAPACK CPU/RSS. QPEVerification uses the same work/byte defaults for its explicit reference. Revisit for a concretely selected larger method workload. No implicit reference or environment work expands these defaults.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| QPE common-step exact arithmetic admission | max_work=1_000_000_000, 1075-bit width chunks, V=32L+320K+128+32K²+4Kk, B=65536+192L+2048K+(2L+13K+64)(32+4⌈b/30⌉) bytes |
src/nwqlib/algorithms/qpe/powers.py::_common_recheck_law, _census_coefficient_bits, _census |
Exact static trajectories only. Sampled and RWPE Plans select each power independently and reserve no common-step work. One shared coefficient/angle reduction and one exact running phase prefix for K distinct powers and L kept nonidentity terms, with k=bit_length(max(1,K)). The integer envelope b uses actual input exponents, supplied rational widths, a candidate-count bound that includes refused counts, and binary64 publication boundaries. With W=None, the coefficient width w is the pre-census width bound of _census_coefficient_bits, computed from the term count and input exponents before any census structure exists. Work is V*c²*(1+⌈log2 c⌉) with c=⌈b/1075⌉. This charge is reserved in both census admissions of an exact trajectory. Every census candidate needs its work plus the reservation within max_work, and its bytes and the reserved bytes each within max_bytes. An empty kept generator has no census and is admitted directly with its exact W=0. The 32K² term covers integer-key collision comparisons. Bytes use the qualified 64-bit CPython logical-object convention. This is an admission proxy, not a timing or process-RSS bound. Revisit when the arithmetic schedule, scalar representation or input domain changes. |
| QPE independent exact recheck admission | max_work=1_000_000_000, 1075-bit width chunks, V_T=32L(K+1), V_S=320K+128, V_I=128K_a |
src/nwqlib/algorithms/qpe/powers.py::independent_recheck_work, _independent_powers |
Independent sampled QPE and RWPE powers admit their exact parameter recheck before forming the coefficient Fractions. For K selected powers and L kept nonidentity terms, the recheck prices 32L(K+1) term visits, 320K+128 scalar visits and 128K_a candidate-inversion visits, with each group weighted by its derived integer-width factor ceil(b/1075)^2*(1+ceil(log2(ceil(b/1075)))). K includes zero powers processed by the recheck. These are counted arithmetic units, not measured time. Census, schedule, construction and analysis keep their own admission laws. K_a counts nonzero powers when L > 0 and is zero for an empty kept generator. Revisit when the recheck's arithmetic schedule, scalar representation or input domain changes. |
QPE controlled Suzuki step law¶
These laws describe the selected synthesis model and are separate from measured compiler output, like the shared circuit-free synthesis laws. Revisit them when the construction or compiler changes, using independent matched-workload evidence.
A QPE controlled second-order Suzuki step (src/nwqlib/algorithms/qpe/powers.py::suzuki_step_cx) costs 4*sum_j w_j CX, where w_j is the number of nonidentity symbols of Pauli term j. Qiskit 2.5.2's PauliEvolutionGate.control controls only the central RZ of each one-term rotation, so a controlled rotation of support w has a parity chain and its inverse, 2(w-1) CX, and a controlled RZ of 2 CX, and the step applies every term twice. The count is exact for lowering to cx,u at optimization level 0 before routing. The step also records 2L selected logical operations for L terms, and the one-qubit identity phase and feedback RZ record one operation and zero CX. The shared common step's laws bind its step_time, an independently selected sampled or RWPE step's laws also bind its power and steps, and Repeat supplies the executed step count. Transpiled counts matched these laws on 11 one-term Pauli rotations with supports one to four, each at three angles with and without control, on the identity-phase, feedback and Hadamard gates, on five selected steps and on whole small QCELS and RFE workloads with 8 and 10 settings (Python 3.12.14 and Qiskit 2.5.2 on macOS arm64, lowering to cx and u at optimization level 0 without a coupling map). Revisit when Qiskit changes its Pauli-evolution synthesis or construct_suzuki_step changes its schedule.
QLS¶
QLS 1/x fit and phase pipeline¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
QLS_TARGET_MARGIN |
1e-3 | src/nwqlib/algorithms/qls/constants.py |
Rescale margin over the shared rigorous Chebyshev norming bound, which leaves each QLS phase target with norming bound 1/(1 + 1e-3). Every sampled QLS phase target at this margin converges, but L-BFGS does not always reach the tolerance, and the Newton start of src/nwqlib/subroutines/qsp/phases.py then supplies the phases. With the default epsilon_inv 0.01, every sampled target that needed it was a qsvt_inverse target with polynomial_kappa between 5.01 and 5.72 (degrees 29 to 33). These targets were at 112 of the 157 polynomial_kappa values sampled from 5.00 to 5.78 in steps of 0.005, and at 12 of the 387 admitted values of a 400-point geometric grid from 1.01 to 50. The kernel-reflection targets of the shortcut solvers converged from the L-BFGS start at all 261 admitted values of a 300-point geometric grid from 1.01 to 80. At other epsilon_inv values, on geometric grids of 100 values from 1.01 to 60, one to three qsvt_inverse targets with polynomial_kappa between 4.28 and 6.74 needed the Newton start at each epsilon_inv of 0.1, 0.03, 1e-3 and 1e-4. For kernel reflection, 22 of the 84 admitted values needed it at epsilon_inv 1e-3, all between 1.29 and 18.1, and none of 100 did at 0.1. All sampled targets come from src/nwqlib/algorithms/qls/host_planning.py::select_polynomial with the QLS default limits and were solved as src/nwqlib/algorithms/qls/quantum.py::plan_quantum solves them (Python 3.12.14, NumPy 2.5.2 and SciPy 1.18.1 on macOS arm64). No Plan of the QLS examples or guide reaches the Newton start. The margin costs 0.1% of success amplitude. Revisit if a fit change makes a QLS phase solve fail at this margin. |
QLS_POLYNOMIAL_KAPPA_FLOOR |
1.01 | src/nwqlib/algorithms/qls/constants.py |
Separate polynomial-domain floor. Actual alpha/sigma_min and original condition remain unchanged, including a condition number of exactly 1. The polynomial requires a parameter greater than 1, and 1.01 also avoids kernel-reflection cancellation near 1. Revisit if the domain requirement of either polynomial construction or the kernel-reflection coefficient fit changes. |
QLS_SPECTRAL_PREMISE_RTOL |
1e-12 | src/nwqlib/algorithms/qls/constants.py |
Relative window comparing selected alpha/kappa with already-computed singular endpoints. Consistent with the existing dense-dilation SVD normalization guard, with no absolute floor. Raw endpoints and the selected alpha are kept, and automatic kappa has the mathematical minimum 1. Numerical consistency only, not a certified spectral or physical-error bound. Revisit with precision or spectral-estimator changes. |
QLS_INVERSE_DEGREE_LAW_FACTOR (C_law) |
1.25 | src/nwqlib/subroutines/qsp/inverse.py |
Engineering tripwire ceil(C_law*kappa*log(kappa/eps)) on the inverse-fit degree (inverse_degree_law_bound), motivated by the [CKS] d = O(kappa log(kappa/eps)) law (arXiv:1511.02306v2), not a finite guarantee that a passing fit exists below it. Revisit when the fit construction or certificate grid changes. Historical measured slack does not qualify a new fit. |
QLS_FIT_LSQ_NODE_MULTIPLIER |
4 | src/nwqlib/subroutines/qsp/inverse.py |
Overdetermined LSQ grid (the exactly-determined grid invites fit-the-grid minima), mirroring the QSP solver grid-multiplier rationale. |
QLS_FIT_CERTIFICATE_GRID_PER_DEGREE |
25 | src/nwqlib/subroutines/qsp/inverse.py |
N=25d affine Chebyshev-zero nodes on [1/kappa,1]. The residual has degree d+1, and the shared norming code inflates its sampled maximum by sec((d+1)pi/(2N)). Even symmetry covers the negative interval. This is a binary64 evaluation of an exact-arithmetic bound, not interval certification. Revisit with a changed residual degree or grid. |
_DOUBLING_FACTOR_OVER_TRIPWIRE |
8 | src/nwqlib/subroutines/qsp/inverse.py |
Runaway stop of the doubling degree search at eight times the QLS_INVERSE_DEGREE_LAW_FACTOR tripwire degree, capped by max_degree and rounded down to odd. A failure reports the finite search scope and does not prove that a larger degree or a different fit would fail. Revisit when the fit construction or the certificate grid changes. |
qsp/inverse.py::inverse_candidate_cost |
peak_bytes of 8*(n*(d + 1) + 3*n*u + 10*n + u + 3*(d + 1) + 8*q) bytes for one candidate of odd degree d, with u = (d + 1)/2 odd unknowns, n = 4*u least-squares nodes and q = 25*d certificate nodes |
src/nwqlib/subroutines/qsp/inverse.py |
Bytes and work counted for one inverse-fit candidate before it is computed. peak_bytes counts 8-byte slots: n*(d + 1) for the Chebyshev-Vandermonde matrix from chebvander, whose odd columns are a view that the scaling turns into the design matrix, 3*n*u for three n x u slots, two for the design matrix and the copy that lstsq factorizes and one to spare, 10*n for length-n vectors (reference and affine nodes, their temporaries, the right-hand side and its copy inside lstsq), u + 3*(d + 1) for the lstsq solution, the coefficient table and its copies inside chebval, and 8*q for certificate-grid vectors (the nodes, the Clenshaw recurrence of chebval and the residual temporaries). The terms are summed although the Vandermonde matrix is released before lstsq runs, so peak_bytes is an allowance above the peak of these arrays. work is the work law of numpy.linalg.lstsq for an n x u matrix (row _linalg_laws.py::least_squares_work) for the dense least squares, 4*n*(d + 1) for the Vandermonde recurrence and design scaling, and 8*q*(d + 1) for the Clenshaw evaluation and residual on the certificate grid. The caller checks peak_bytes per candidate and sums work over all candidates. |
QLS Dalzell shortcut pipeline¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
shortcut eta choice |
epsilon_inv / sqrt(2) |
src/nwqlib/subroutines/qsp/shortcut.py |
The paper's known-norm choice at t = ||x|| (Dalzell arXiv:2406.12086v2, pp. 5-6). |
| norm-grid detection threshold | 0.725 |
src/nwqlib/algorithms/qls/norm_search.py |
Dalzell arXiv:2406.12086v2, Sec. 5.1, separates a good log-grid candidate (>= 0.8 success) from candidates more than ln(2) away (< 0.65) with midpoint margin. |
| norm-grid trials per candidate | ceil(100 ln(20 |T|)) |
src/nwqlib/algorithms/qls/norm_search.py |
Dalzell arXiv:2406.12086v2, Sec. 5.1, log-grid detector formula. Stored as trials_per_candidate in the explicit grid model's metadata. This prospective trial count is not an acquired shot population. |
| norm-sequence trials per candidate | ceil(100 ln(20*4))*(1+log2(kappa_rounded)-j) |
src/nwqlib/algorithms/qls/norm_search.py |
Dalzell arXiv:2406.12086v2, Eq. (24), with four candidates per visited step and the Sec. 5.3 overhead. Recorded as trials_per_candidate in the actual step records and trials_per_candidate_formula in sequence metadata. It is not a measured coverage guarantee. |
PROJECTOR_ROW_TILE |
64 | src/nwqlib/algorithms/qls/numerical.py |
Row-tile length of the rank-one update that forms G_t in place. It bounds the outer-product temporary at 16 r a bytes with r = min(a, 64), the B_G term of the classical workspace law in host_planning.selected_work. Tiling changes storage, not the per-entry multiply and subtract. The value is untuned. Revisit if the rank-one temporary becomes a measured peak. |
qsp/shortcut.py::kernel_reflection_cost |
peak_bytes of 8*(3*n*(d + 1) + 12*n + 4*(d + 1)) bytes for degree d = 2*ell with n = max(2001, 8*(d + 1)) fit nodes |
src/nwqlib/subroutines/qsp/shortcut.py |
Bytes and work counted for computing the kernel-reflection polynomial of one plan. peak_bytes counts 8-byte slots: 3*n*(d + 1) for the Chebyshev-Vandermonde matrix built inside chebfit, its column-scaled copy and the copy that lstsq factorizes, which are alive together, 12*n for node, filter and target vectors, including the Clenshaw recurrence vectors of chebval, and 4*(d + 1) for coefficient vectors. work is the work law of numpy.linalg.lstsq for an n x (d + 1) matrix (row _linalg_laws.py::least_squares_work) for the dense least squares and 8*n*(d + 1) for the Vandermonde matrix, the degree-ell filter evaluation and the scaling. LAPACK workspace is excluded. |
Explicit QLS verification criteria¶
These constants in src/nwqlib/algorithms/qls/constants.py set the numerical and heuristic windows of explicit QLS verification. They are separate from NUMERICAL_RELATION_RTOL record admission and are not certified forward-error or finite-sample confidence bounds.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
QLS_SPECTRAL_COVERAGE_TOLERANCE |
1e-9 | src/nwqlib/algorithms/qls/constants.py |
Dimensionless lower/upper spectral-domain deficit window. Preserves ordinary auto-kappa rounding. Change only with an independently justified criterion. |
QLS_EQ17_ABSOLUTE_WINDOW |
2e-12 | src/nwqlib/algorithms/qls/constants.py |
Absolute additive window for automatic Eq. (17) comparisons (Dalzell arXiv:2406.12086v2) after both phase endpoints are widened. Does not modify an explicitly selected tolerance. Revisit with a derived roundoff bound for the Eq. (17) endpoints. |
QLS_INVERSE_BINOMIAL_VARIANCE_FLOOR |
1e-12 | src/nwqlib/algorithms/qls/constants.py |
Floor on the predicted per-trial Bernoulli variance before division by returned shots, so a predicted rate of exactly 0 or 1 keeps a positive shot allowance. Used only when the reference rate is in [0,1]. No coverage guarantee. Revisit if the shot term adopts a finite-sample confidence method. |
The automatic shot term is four times the predicted-rate standard deviation. Dalzell's Eq. (17) shot term (arXiv:2406.12086v2) has no variance floor, and QLS_EQ17_ABSOLUTE_WINDOW keeps the combined Eq. (17) allowance positive. Default shortcut direction tolerance is 5*epsilon_inv, a heuristic. Verification receipts (VerificationReceipt) keep raw discrepancy, deterministic and shot allowances separately. An automatic ratio is a distinct heuristic fact. Revisit these choices through their scientific criterion, never by changing a value merely to pass a failed result.
QLS construction and readout limits¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
qls/method.py::QLS and qls/verification.py::QLSVerification limits |
QLS: max_degree 256, max_qsp_evaluations 20,000, max_work 1,000,000,000, max_bytes 10,000,000,000 and max_admission_steps 1,000,000. QLSVerification: max_work 1,000,000,000 and the same max_bytes |
src/nwqlib/algorithms/qls/method.py, src/nwqlib/algorithms/qls/verification.py |
Explicit construction and reference allowances, each checked before the work it bounds. max_degree caps the selected polynomial and its fit search, max_qsp_evaluations the objective, Jacobian and verification evaluations of the phase solve, max_work the size laws of the original SVD or Hermitian eigendecomposition, spectral selection, dense completion, encoding construction, the exact synthesis of controlled dense queries, polynomial fits, the classical model and the comparisons of sampled Pauli grouping, each step against the whole cap, and max_bytes the peak live arrays of each of these steps and of the phase solve. The exact projected reductions of one Run are charged together against max_work. With one control, the default dense_control_route synthesizes the forward and the adjoint controlled dilation, twice controlled_synthesis_size(log2(p)+2) for padded dimension p. With the two controls of shortcut_dilation it takes the gate-wise route, twice the sum of dense_synthesis_size(log2(p)+1) and gatewise_control_size(*gatewise_control_counts(log2(p)+1, 2)). The default max_work therefore admits a controlled dense dilation of at most 64 padded coordinates (4.2e8 units with one control, 2.4e8 with two) and refuses 128 (3.2e9 and 1.4e9). The Hermitian inverse queries the dilation without controls and keeps the completion bound of at most 256 coordinates. max_admission_steps is the quantum Program's AdmissionLimits.max_steps, with ten times the shared default (Shared Program admission limits). The verification caps bound its reference solve and spectrum. They are not accuracy requirements, elapsed-time limits or RSS bounds. Revisit for a deliberately larger selected workload, keeping admission before work. |
| QLS sampled count registers | r_g measured bits per setting, c=max_g r_g shared classical bits |
src/nwqlib/algorithms/qls/quantum.py::_counts_bits, _selection, _program |
A group measures its support and valued selectors. Samples and physical-prefix mass settings measure all coordinates. A table has at most min(shots,2**r_g) nonzero outcomes stored with c-bit keys. Generic readout admission conservatively reserves min(shots,2**c) CountBin JSON records with the existing prototype bound. The decoder checks zero unused suffix bits and maps each parity mask into the measured order. Its uint64 route requires c<=64. Distinct measurement maps add the qualified Program allowance 1024*(M_p+V+T), where M_p counts distinct position/wire pairs, V distinct ordered maps and T their summed lengths. Actual Program field slots and admission work remain subject to max_admission_steps. The admission work of a sampled Program grows with the number of its distinct measured registers. One measured example uses the following observable. Initialize rng = numpy.random.default_rng(11) once, draw labels with "".join(rng.choice(list("IXYZ"), size=8)) until 160 distinct labels other than the all-I label exist, sort them, and then draw one standard-normal coefficient per sorted label with rng.normal() from the same generator. With 100 shots and Plan seed 5, a NormalizedExpectation of this observable for A = diag(linspace(1, 2, 256)) and b of standard-normal entries drawn with default_rng(3) forms 108 groups with 27 distinct registers, and its Program needs 130,560 units, against 60,425 when every setting measures all wires. The 160-term prefix of the sampled Expectation measurement in the src/nwqlib/ir/records.py::Program row, on PeriodicStencil(12, mass=1.0, diffusion=0.25) with the all-zero occupation RHS, a NormalizedExpectation or QuadraticForm output, 100 shots and Plan seed 7, forms 149 groups with 111 distinct registers, and its Program needs 339,214 units, against 82,729 when every setting measures all wires. Both fit the QLS default max_admission_steps of 1,000,000. These figures were measured with Python 3.12.14 and Qiskit 2.5.2 on macOS arm64. Revisit when the count representation, layout or Program object convention changes. |
GCiM¶
ADAPT-GCiM defaults and tolerances¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
Adapt-GCiM theta default |
pi/4 | src/nwqlib/algorithms/gcim/adapt.py |
Fixed angle in Zheng et al. (2024), arXiv:2312.07691v3, Methods. Explicit optional BFGS may later change it, including to zero. |
ADAPT-GCiM t_auto_fraction / t_user |
0.20 / 10 | src/nwqlib/algorithms/gcim/adapt.py, src/nwqlib/algorithms/gcim/adapt_acquisition.py |
Zheng et al., arXiv:2312.07691v3, Methods, uses T=min(0.2*unselected_count, T_usr). An integer flat count reaches a fractional T at its ceiling. The default user cap 10 is the paper's H4 setting and is adjustable for other pools. With max_iterations=8, a pool larger than 42 members cannot reach this flat threshold. The controller checks the iteration cap first and stops at the eighth selection, so the last flat check follows the seventh. The flat count is then at most seven, and the stop needs 0.2*(P-7) <= 7 for a pool of P members (src/nwqlib/algorithms/gcim/adapt_acquisition.py::_stopping_reason, GCiM guide). Count small changes from the reference energy and reset the count after a large or unavailable change. This heuristic is not a physical error certificate. |
ADAPT-GCiM max_iterations |
8 | src/nwqlib/algorithms/gcim/adapt.py, src/nwqlib/algorithms/gcim/adapt_acquisition.py::_stopping_reason |
Finite cap on selection rounds, which bounds the working basis at 2*min(8, pool size) states and the number of pool screenings. Zheng et al., arXiv:2312.07691v3, Methods, stop by the flat counter and state no iteration cap. Measured with Python 3.12.14 on macOS arm64 with exact evaluation, the spin-adapted pool and θ = π/4, eight selections reach chemical accuracy for both of the following molecules. The H4 chain of the eigenvalue example (2 Å, 8 qubits, 66 pool members) reaches FCI to roundoff at the eighth selection and is 14 mHa above it after the seventh. LiH at 1.6 Å with active space (2, 5) (10 qubits, 160 members) is 0.46 mHa above FCI after the eighth selection and at FCI after the ninth. With the cap at 32, the largest that max_basis_size=64 admits, the flat counter stops them at selections 18 and 19, ten after each energy reached FCI, with 36 and 38 basis states instead of 16. For H4 under quantum execution with exact probabilities, that raises the pair circuits from 136 to 666 (formula of the eigenvalue notebook, Appendix A). A cap equal to the pool size cannot be planned for either pool unless max_basis_size rises to 132 or 320. With t_user=10 the cap keeps the flat counter from stopping pools larger than 42 members (row above). The cap does not certify convergence. Revisit when a default workload needs more than eight selections or when the flat counter should stop larger pools by default. |
Adapt-GCiM energy_change_tolerance / gradient_norm_floor |
1e-6 / 1e-8 | src/nwqlib/algorithms/gcim/adapt.py, src/nwqlib/algorithms/gcim/adapt_acquisition.py |
Flat-counter change tolerance and surrogate gradient-norm exit floor. Neither establishes total output accuracy. |
ADAPT operational residual_norm_tolerance |
1.6e-3 | src/nwqlib/algorithms/gcim/adapt.py, src/nwqlib/algorithms/gcim/adapt_acquisition.py |
Explicit THEORY stopping criterion ||H psi - E psi|| <= tol in the declared energy unit. Its chemical-accuracy motivation applies when that unit is Hartree. It is not a ground-energy certificate or a supplied-reference comparison tolerance. |
ADAPT explicit diagnostic residual_tolerance |
1.6e-3 | src/nwqlib/algorithms/gcim/adapt_verification.py |
Selected post-hoc residual classification. It does not change controller termination or the selected output accuracy. SHOTS does not claim the deterministic spectral interval. |
ADAPT input_symmetry_tolerance |
1e-12 | default: src/nwqlib/algorithms/gcim/adapt.py. Admission: src/nwqlib/algorithms/gcim/adapt_inputs.py |
Relative Frobenius correction window for the anti-Hermitian projection (A - A^dagger)/2 of each pool generator. The Hamiltonian gets no window and must already be exactly Hermitian. Its Pauli coefficients must be real, and the classical matrix route keeps the original Hermitian matrix. Zero requires an exactly anti-Hermitian generator. The actual correction is recorded and the processed generator is the target. Because the window is relative, a generator with a large Frobenius norm, for example from an identity term, admits a larger absolute correction. Revisit for changed precision or an application requiring an absolute correction budget. |
ADAPT pauli_coefficient_cutoff |
1e-14 | default: src/nwqlib/algorithms/gcim/adapt.py. H/commutator preparation: src/nwqlib/algorithms/gcim/adapt_inputs.py |
Absolute approximation in the operator's units. After duplicate-label accumulation and symmetry projection, keep a real coefficient iff abs(c) > cutoff. Zero disables nonzero-term pruning. This does not scale with operator size or identity shifts and is not a universal accuracy guarantee. Choose it relative to target accuracy and inspect the reported dropped L1 mass. |
| ADAPT supplied-reference comparison tolerance | explicitly required positive reference_tolerance |
src/nwqlib/algorithms/gcim/adapt_verification.py |
No automatic THEORY/STATEVECTOR/SHOTS reference tolerance. The selected comparison checks the recorded processed Ritz scalar against a supplied framed scalar in their declared unit. It does not establish ground identity or total output accuracy. |
GCiM sampled enclosure diagnostic (SAMPLED_PENCIL_ENCLOSURE_RTOL) |
1e-8 relative | src/nwqlib/algorithms/gcim/fixed_basis.py, src/nwqlib/algorithms/gcim/adapt_acquisition.py, src/nwqlib/_projected_eigensolver.py |
Compare the finite sampled Ritz value with the processed Pauli L1 enclosure. Scale the distance by the largest endpoint magnitude and keep its absolute distance. A violation is a plausibility diagnostic, does not stop the controller or veto previous-energy/flat-counter updates, and is not a statistical or spectral certificate. Actual solver refusal is separate. Revisit for changed precision or conditioning. |
gcim/sector.py |
imaginary-part window 1e-10 |
src/nwqlib/algorithms/gcim/sector.py |
Untuned absolute admission window for a spin expectation on the already normalized input. Rescaling the original state does not change this expectation. It is not a dimension-uniform arithmetic bound. Revisit condition: Larger spin sectors, precision or a derived summation-error policy. |
| ADAPT Taylor-model gradient regression tolerance | absolute 2e-11, relative 0 | tests/test_adapt_shift_science.py |
Compares optimization.normalized_taylor_energy_gradient for iX then iZ at angles 3/16 and -5/16, a zero first angle and degree 1, with exact rational ascending-power Taylor matrices on a unit-scale two-level example. It is a small-case regression threshold, not a bound for arbitrary chains or a derivative of the step-count selector. |
Fermionic generators and chemistry input¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
GENERATOR_COEFFICIENT_CUTOFF |
1e-12 absolute | src/nwqlib/subroutines/fermionic_pool.py |
Removes binary64 residues of exactly cancelling terms after all contributions to a normal-ordered term or Pauli label are summed. The smallest kept Pauli coefficient is 0.0255 in the six-orbital default pool and at least 0.088 in the UCCSD, QEB and CEO pools of six occupied and six virtual spin orbitals. Revisit for a generator family whose genuine coefficients can approach the cutoff, or for another precision. |
| Generator occupation planning limits | at most 8 active modes and block dimension 8 | src/nwqlib/subroutines/fermionic_pool.py, src/nwqlib/algorithms/gcim/optimization.py |
Bounded local spectrum planning for default generators. Shared-index circuit compilation separately allows at most 6 modes and dimension 5. Revisit only with a new supported generator family and its cost model. |
| Generator repeated-eigenvalue merging | 16 * eps * max_block_dimension * spectral_scale |
src/nwqlib/algorithms/gcim/optimization.py |
Engineering window for numerical splitting of repeated eigenvalues in bounded Hermitian blocks. This is not a certified eigenvalue error bound. Revisit if a new generator family has smaller distinct spectral gaps. |
fermionic_pool.py, GENERATOR_TAYLOR_DEGREE |
18 and ceil(2*abs(theta)*sum(abs(c))) steps |
src/nwqlib/subroutines/fermionic_pool.py |
The Pauli triangle bound gives step norm at most 1/2. Each exact-arithmetic Taylor remainder is at most exp(.5)*.5**19/19! < 2.6e-23. Accumulated step error and floating-point arithmetic need separate accounting. No operator-norm solve selects the steps. src/nwqlib/subroutines/fermionic_pool.py::_taylor_steps computes the count for both the exponential and ADAPT's classical work laws (src/nwqlib/algorithms/gcim/adapt_records.py::_exponential_action_work), so the two cannot differ by a step. Revisit condition: Different precision, larger step populations or a requested total action-error bound. |
fermionic_pool.py::apply_spin_squared |
fixed DEFAULT_INPUT_BYTES and 1_000_000_000 products |
src/nwqlib/subroutines/fermionic_pool.py |
Default ceilings for an action that takes no caller limit. The S^2 action is a product of two ladder sums with q/2 terms each. Revisit condition: A caller-selected limit. |
fermionic_circuits.py::_require_default_double |
8192 bytes and 4096 work | src/nwqlib/subroutines/fermionic_circuits.py |
Untuned fixed charge, checked against the caller's max_bytes and max_products, for re-normalizing one default double of at most six base terms with four ladder operators each. The mapping of the result is charged separately by its own law. Revisit condition: A larger generator family. |
gcim/adapt.py::generator_pool_slot_bound |
B(q,T) = max(6q*max(1,T), 24q+6192) logical slots per generator, or 6q*max(1,T) when every pool generator has commuting terms |
src/nwqlib/algorithms/gcim/adapt.py |
Family-level envelope for the routes of fermionic_circuits.build_generator_circuit, with q system qubits and the largest stored Pauli term count T of the pool: 6q per commuting term, 5(2q+11) for pair/split, 6(2q+516) and 12(2q+516) for singlet and triplet four-distinct doubles, and at most 24q+2544 for the shared-index route with at most six active modes and blocks of dimension five. The quantum query envelope is 2*W_reference + 2*l*(1+B) + 8q + 7 with l = min(max_iterations, pool size). It counts appended logical instructions, not native synthesis, and does not establish support for an arbitrary custom noncommuting generator, which planning checks separately. Revisit condition: A new compiler route or supported generator family. |
GCiM chemistry coefficient_cutoff default |
1e-12 | src/nwqlib/algorithms/gcim/chemistry.py |
Pauli/Fermion term pruning at machine precision in the optional PySCF/OpenFermion input path. |
ORBITAL_PHASE_RELATIVE_THRESHOLD |
1e-8 relative | src/nwqlib/algorithms/gcim/chemistry.py |
Sign convention for RHF orbitals: the first AO coefficient whose magnitude exceeds this fraction of the orbital's largest magnitude is made positive. Coefficients that vanish by symmetry carry roundoff near 1e-16 of that scale, far below the threshold, so the chosen coefficient and its sign do not depend on the LAPACK build. Fixed-angle ADAPT-GCIM depends on these signs. Rotations within degenerate orbitals are not fixed. Revisit for molecules with degenerate orbitals or a changed precision. |
| H2 chemistry-table cross-check tolerance | 1e-8 | tests/test_gcim_chemistry.py |
High-precision hardcoded H2 electronic JW table at R = 0.7414 Angstrom. Pipeline comparison subtracts nuclear repulsion from the identity term before comparing electronic coefficients. |
GCiM work and byte limits¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
gcim/adapt.py |
pool cap 1024 and basis cap 64 | src/nwqlib/algorithms/gcim/adapt.py |
Untuned workload ceilings. Pool admission precedes expansion, and the basis cap precedes projected matrix construction. Neither cap establishes convergence or bounds native SDK workspace. Revisit condition: An explicitly intended larger pool or projected basis. |
gcim/fixed_basis.py::FixedGCIM, ADAPT and AdaptVerificationOptions limits |
max_basis_size=64, max_experiments=100_000, max_analysis_work=1_000_000_000 (row "Sampled FixedGCIM analysis admission"), max_classical_products=100_000_000, max_admission_steps=1_000_000. max_products=1_000_000_000 of ADAPT and AdaptVerificationOptions |
src/nwqlib/algorithms/gcim/fixed_basis.py |
Untuned workload ceilings, checked before pair expansion, QWC grouping, the grouped analysis pass, the projected solve, classical actions or verification. max_admission_steps is the Programs' AdmissionLimits.max_steps, with ten times the shared default (Shared Program admission limits). They bound derived counts, not CPU time, and establish no accuracy. Native ADAPT residual and sector verification charges each decomposed k-qubit gate of the circuits it simulates 1 + 4**k + 2**q * (2**k + 2) operations on q qubits. Whether the default admits it depends on those circuits and the reference. With an occupation reference, a residual check of any selection of up to eight spin-adapted generators at 8 qubits needs at most 91.1 million operations before the Hamiltonian action, which adds the shared Pauli action law (pauli_action_requirements). The max_products default is a fuse against runaway planning work, sized so that no example or test workload in the repository reaches it. Full-state checks at 20 to 25 qubits must stay an explicit choice. Revisit condition: An explicitly intended larger basis, matrix-element population or classical workload. |
| Fixed GCIM projected size law | 512*B*B+256*E selected bytes and 32*B*B+16*B*B*B projected work |
src/nwqlib/algorithms/gcim/pencil.py (_projected_solve_bytes, _projected_solve_work), used by src/nwqlib/algorithms/gcim/fixed_basis.py, src/nwqlib/algorithms/gcim/adapt.py and src/nwqlib/algorithms/gcim/adapt_acquisition.py |
B actual basis states and E selected FixedGCIM settings (exact pair reductions, or grouped sampled settings), checked before expansion. 512*B*B is an untuned allowance of 32 complex128 B-by-B arrays for the projected solve, and 256*E is FixedGCIM's untuned record allowance for each setting, not a derived Python-heap bound. A sampled Plan adds G*n+8L logical basis and index bytes for its group description. ADAPT has no matrix-element settings in this law. It checks both laws, the work against max_products, for its largest basis of 2*max_selections states at planning (src/nwqlib/algorithms/gcim/adapt.py), so a run that might stop at a smaller basis is still refused when that largest solve does not fit. It checks both again before each projected analysis and each reanalysis with a new cutoff (src/nwqlib/algorithms/gcim/adapt_acquisition.py::_matrix_from_observations). No CPU/RSS guarantee. |
Classical FixedGCIM and ADAPT identity offset (_mean_diagonal, _offset_free_requirements, _shifted_sparse) |
c = trace(A)/d removed before the projection. Extra work d*d + d for dense input. For CSR/CSC input with z stored entries, h missing diagonals and index itemsize I, persistent bytes 16z when h = 0 and the pattern is shared, else (16 + I)(z + h) + I(d + 1), plus the vectorized scan and union scratch, setup work W_setup = 4z + 6d + 3h and n_s*h extra action work for n_s columns. Before the scan h is bounded by d |
src/nwqlib/algorithms/gcim/fixed_basis.py, src/nwqlib/algorithms/gcim/adapt_inputs.py, src/nwqlib/algorithms/gcim/adapt_acquisition.py::_offset_free_column |
The mean diagonal is the mean eigenvalue, so ||A - cI||_2 <= lambda_max - lambda_min whatever the identity offset of A, and the solve adds c back after the ill-conditioned projection. Dense row blocks are refilled in one buffer of at most n_s rows. A Run creates its shifted canonical CSR/CSC payload once, shares its original pattern only when every diagonal is present, and charges that payload throughout later actions. ADAPT keeps it in the Run's context for every later action. The ADAPT per-query charge includes the setup work conservatively. The docstrings of _offset_free_requirements and _shifted_sparse state the laws. Revisit when the classical projection kernel or the sparse representation changes. |
Exact FixedGCIM and ADAPT pair overlap allowance (_exact_overlap_allowance) and pair reduction (pair_reducer.pair_work, pair_bytes, TILE = 1024) |
Allowance max_i sum_(j != i) E^S_ij from each off-diagonal pair's preparation-record saved-state error (PreparedArtifact.saved_state_error), unavailable when any is unresolved. The row sums are accumulated exactly and converted upward. Work (L+F) + 3L + N(L+F+1) + N(1+L) + 2N per pair (N for L = 0), bytes 17N + max(9L, 96t) + 24L + 16F + 8 + 65536 with F = min(L, N) and t = min(N, 1024) |
src/nwqlib/algorithms/gcim/pair_reducer.py, src/nwqlib/algorithms/gcim/fixed_basis.py, src/nwqlib/algorithms/gcim/adapt_acquisition.py |
The work covers the attempted grouped pass, its reserved ordered retry, group construction and the two complex contractions. The bytes cover one action output and a finite mask, the larger of group construction and tile scratch, group metadata and the fixed heap allowance. The strided branches are views, and the saved state and the Plan's resident Pauli table are admitted by their own checks. FixedGCIM admits its summed pair work against max_classical_products at preparation, before the first pair's native work, ADAPT each matrix stage's new pair work against max_products before the stage's first query, and both each point's workspace against max_bytes before acquisition. The tile length is an engineering choice, not a scientific threshold. |
| ADAPT packed classical Pauli actions | B_A=33d+24L+16g+8+max(9L,96t), W_A=2dL+dg+2d+4L+g, g=min(L,d), t=min(d,1024) | src/nwqlib/operators/_pauli.py::pauli_action_requirements, src/nwqlib/algorithms/gcim/adapt_records.py, src/nwqlib/algorithms/gcim/adapt_acquisition.py | q is the system-qubit count, d=2**q, and L counts the action's stored rows including zeros and duplicates. The single-column law uses complex128 states and includes preprocessing, the grouped pass and a possible ordered retry. H0, generator and insertion rows are packed once per classical Plan. Let L0 count the nonidentity Hamiltonian rows and T_j the rows of generator j. For R=L0+2 sum_j T_j newly packed rows, their logical payload is 32R bytes, construction reserves 64R beside other live storage, and packing charges R(q+4) pass units. Input tables and Python headers are admitted separately. Revisit if masks, snapshots or the action workspace change. |
| ADAPT Taylor and derivative action work | W_exp=s[d+18(W_A+2d)]+10d, with s from _taylor_steps | src/nwqlib/algorithms/gcim/adapt_records.py::_exponential_action_work, _energy_gradient_work | Each Horner action uses the shared packed Pauli allowance. Derivative Horner terms have two actions and reverse terms have one. The energy/gradient law sums the complete forward, normalization, tangent, reverse and Hamiltonian work. Invocation work is admitted as a whole before execution. Packing is a separate once-per-Plan charge. These are logical pass allowances, not CPU-time or accuracy bounds. |
| ADAPT classical action workspace | 16(8ell+2k+12)d+24k vector/scalar bytes for an energy-gradient query, or 16(8ell+12)d otherwise, plus packed tables, the full-H payload, max_A(B_A-32d), reference and offset-free setup storage | src/nwqlib/algorithms/gcim/adapt_acquisition.py::bind_classical, src/nwqlib/algorithms/gcim/adapt_verification.py | ell=max_selections, k is the query chain length, and d=2**q on q system qubits. The frontier already contains action input/output vectors. The remaining shared allowance includes the finite mask, action metadata and maximum preprocessing/retry scratch. Verification uses its own streamed vector frontier and adds the action remainder for each applicable classical generator or residual-H action, along with saved arrays and native gate matrices where applicable. Revisit when live vectors or shared action buffers change. |
| Sampled FixedGCIM groups and moments | 2(M-D)max(1,G)+DG settings for an acquired pencil, shots per setting | src/nwqlib/algorithms/gcim/fixed_basis.py sampled planning and group analysis | M is the required pair count including D diagonal pairs, G is the number of nonempty QWC groups, and n is the system-qubit count. Groups use separate X/Y ancilla settings off diagonal and system-only measurements on a diagonal. The current whole-register construction has c=n+1 shared classical bits, with an unused zero phase bit on diagonals. Generic CountBin JSON admission reserves min(shots,2**c) bins. Group analysis admits the sum of the decoded arrays cached on all chunks and the largest live moment workspace under its logical payload law. Physical-H variance includes the overlap–Hamiltonian sample-mean covariance of the designated overlap group. One-shot variances are unavailable, and negative raw variance estimates are flagged without clipping. Binary64 moment fields require finite representable moments. The scalar-identity shortcut acquires no pencil. Revisit when the measured-register representation, count codec, reduction schedule or pooling rule changes. |
| Sampled FixedGCIM analysis admission | W_plan=b² sum_g W_g(m)+Lc+32b²+16b³ and B_plan=8(w+1)b²Gm+80m+8max_g L_g+8b²(L+16G)+(n+16)L+512b² |
src/nwqlib/algorithms/gcim/fixed_basis.py::sampled_analysis_requirements, FixedGCIM.plan, read_groups |
For b trial states, n system qubits, c=n+1 classical bits, w=ceil(c/64), G nonempty QWC groups with sizes L_g, L=sum_g L_g and s shots per setting, put m=min(s,2c). Each group occurs b² times. The one-table work is W_g(m)=m[c+28+(4w+4)L_g]+L_g[(c-1)w+1]. Planning checks W_plan against max_analysis_work and B_plan against max_bytes before circuit-block selection. The common readout check enforces m on returned entries, including zero-count bins. The bound covers one acquisition per setting and conservatively includes the unused diagonal phase bit. Each analysis pass checks its actual chunk populations before decoding. Bytes are logical cached-array, workspace and scalar payloads, and work uses the group reducer's kernel/pass convention. Stored JSON, Program records and grouping have separate admissions. The scalar-identity shortcut has no grouped pass. The bound is conservative, so the default max_analysis_work is 109. On the three-qubit fixture of tests/test_gcim_production.py::test_grouped_analysis_is_admitted_before_setting_construction (18 settings at 64 shots) it plans 17,088 units where the analysis of one sampled run (seed 104) needed 12,384, about 1.38 times, and on the 14-qubit fixture of test_default_analysis_work_admits_a_fourteen_qubit_grouped_plan in the same file, 80 random Pauli words (G = 60), three product states and 540 settings, it plans 118,714,800 units at 4,096 shots and 237,416,880 at 8,192, both above 10**8. The same default also caps the QWC grouping comparisons of sampled_groups and the projected solve. Sampled FixedGCIM checks grouping, the projected solve and grouped analysis against max_analysis_work. For basis size b, L nonidentity terms, n system qubits and s shots per setting, write c=n+1, w=ceil(c/64) and m=min(s,2^c). The grouped-analysis charge is b²{m[(c+28)G+(4w+4)L]+L[(c−1)w+1]}+Lc+32b²+16b³, where G is the selected group count. Before grouping, substituting G=L gives a sufficient retry envelope. These are phase ceilings, so the required common limit is their maximum. Before grouping, planning refuses when the projected solve or the G=1 analysis charge exceeds the limit and names that maximum (src/nwqlib/algorithms/gcim/fixed_basis.py::sampled_analysis_work). Revisit when the acquisition population, cardinality contract, layouts or reduction lifetimes change. |
gcim/adapt_verification.py::verify |
B_states + reference_action_bytes + 16*4**k_max + B_other + B_saved bytes |
src/nwqlib/algorithms/gcim/adapt_verification.py |
Known-buffer admission of the streamed residual and sector checks. B_states = 16d*max(F(l), 8*[residual], 32*[sector]) + d with F(0) = 3, F(1) = 5 and F(l) = max(6, l+2) full vectors for l selections (_state_frontier), where the residual and sector allowances of 8 and 32 vectors are explicitly identified allowances, not measured minimal peaks. 16*4**k_max covers the widest lowered native gate matrix from the census that charges the evolution work, B_other is the shared action remainder max_A(B_A-32d) of the ADAPT classical action workspace row, with the packed tables, on a classical Plan, or that of the H action when a quantum Plan requests residual verification, and B_saved adds once the distinct arrays the saved Result already owns. Python object headers, allocator overhead and SDK decomposition storage are excluded, and the reference synthesis keeps its own admission. Revisit condition: A change of the shared Pauli action law or of the packed action tables, or a new verification check. |
gcim/adapt_inputs.py::array_record_bytes |
B_other_live + 2B + 4J + 16I bytes |
src/nwqlib/algorithms/gcim/adapt_inputs.py |
Charge of the ADAPT reconstruction's known record data and serialization forms. Each [H, A_i] table of a quantum Plan is an int64 (rows, q) letter-code array and a float64 (rows,) coefficient array (fixed_basis.PauliArrays, 8q+8 raw bytes per row). A sampled Plan stores each screening or energy readout family as I integer indices into U labels, divided among G nonempty QWC groups. B is the raw array bytes. J includes their base64 data and dtype, shape and key headers, the enclosing reconstruction metadata and collection delimiters (_reconstruction_fixed_json), each table record's fixed JSON (_commutator_record_fixed_json), and I*(digits(max(U-1,0))+1) + 2G + 2 for each screening or energy index family. The reconstruction metadata allowance includes both empty group lists and bounds supplied text by its actual escaped JSON length. Each freshly constructed table adds 145 spaced JSON characters outside its two array headers, including an empty table. 2B covers two raw copies, 4J three simultaneous JSON or base64 text forms and one transient encoding form, and 16I two logical index sequences. The other live bytes are the 6q+360-per-term Hamiltonian and generator charge plus the pool-member allowances in the record identity table. Classical Plans charge the same reconstruction metadata without commutator arrays or index payloads. Python object headers, allocator overhead and SDK workspace are outside this law. Revisit condition: Another record field, metadata source, label dtype or codec, or an identity encoder that streams. |
gcim/adapt_inputs.py::label_cache_bytes |
B_kept + 65536 + max(X_decode, X_build) |
src/nwqlib/algorithms/gcim/adapt_inputs.py |
Per-Plan decoded commutator rows, sorted label tuples, active-label frozensets and cache keys use the object-population formulas at the helper. The allowance includes variable headers, slots, hash-table capacity and construction overlap on 64-bit CPython 3.12.14. Preparation adds it once to the reconstruction's array and encoding charge. The cache's strings are shared by its label collections, and its contents are rebuilt after loading. This qualified object allowance does not bound total process RSS. Revisit condition: A different interpreter, object layout, decoder, cache capacity, key population, ownership lifetime or measured construction peak outside the allowance. |
gcim/adapt_inputs.py::_commutator_arrays |
tiles of _COMMUTATOR_TILE = 2**14 candidate pairs; B_tile = (64W+128)b, B_coalesce(C) = (80W+128)C + 65536, W_j = P(W+1) + C(12W+16) + (2W+2)C ceil(log2 max(2,C)), plus pauli_array_conversion_requirements (B_conv, W_conv) |
src/nwqlib/algorithms/gcim/adapt_inputs.py |
Commutator construction tests bounded packed pair tiles, stores only anticommuting products and coalesces the complete survivor population before publication. One running work total covers every pool member against max_products, and byte admission includes completed member tables (8J_i(q+1) each), the current survivor population and its sort, conversion and snapshot workspace. The tile length is an engineering choice, not a scientific threshold. These are accounting envelopes, not exact elapsed-work counts or heap bounds. Revisit condition: Another survivor layout, sort, conversion or tile schedule. |
QHD¶
QHD Method and refinement defaults¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| QHD Method defaults | num_grid_points=2, num_steps=1, total_time=1, schedule=QuadraticSchedule(gamma=0.3), coefficient_rule="midpoint", trotter_order=2, boundary="dirichlet", encoding="one_hot", initial_state=KineticGroundState(), initial_state_preparation="structured", kinetic_model="finite_difference", binary_synthesis=BinarySynthesis() (min_cx for both diagonals, relabel, exact QFT), theory_flavor="schrodinger" |
src/nwqlib/algorithms/qhd/method.py, src/nwqlib/algorithms/qhd/schedules.py, src/nwqlib/algorithms/qhd/initial_state.py, src/nwqlib/algorithms/qhd/binary.py |
The smallest legal grid and step count with unit time, so a default solve uses two qubits per variable. The Dirichlet boundary is the one that admits K = 2 with the default one-hot encoding, because the one-hot periodic grid needs an even K of at least 4 (method.QHD._periodic_grid). The one-hot encoding is the established circuit. The finite-difference kinetic model is the one that every route applies (method.QHD._kinetic_route), and the Schrodinger flavor has no splitting error. Gamma 0.3 is the illustrative schedule of the QHD guide. Any positive gamma makes the kinetic-to-potential weight ratio decay to zero, as Leng et al. require after their Eq. (1) (arXiv:2303.01471v1). The midpoint rule stays the default because a convergence check of a fixed model, with the quadratic schedule and both cubic schedules at s = 1, found both rules second order with errors within 1% of each other, and the midpoint rule needs no interval integrals. The check evolved the one-dimensional two-mode cosine of Liu et al., arXiv:2607.16996v1, on an eight-point periodic grid from the kinetic ground state to T = 1, with the split-step flavor at 64, 128 and 256 steps against an adaptive DOP853 solution (Python 3.12.14, NumPy 2.5.2 and SciPy 1.18.1 on macOS arm64). Revisit the coefficient rule after a check at small s, where the rules differ most. CubicSchedule and ShiftedCubicSchedule have no default s, because s is a model parameter that the caller chooses. Apart from the initial state below, none of these defaults is an accuracy choice, and scientific use selects them explicitly. Revisit them when a default workload is chosen for a scientific purpose. The initial state is the ground state of the kinetic term, whose weight relative to the potential, a(t)/b(t), is largest at t = 0 and decays afterwards. On a Dirichlet grid it has no excited kinetic component, whose phase would depend on the step count (docs/algorithms/qhd.md, "Initial states and preparation"). Under the quadratic schedule both terms are present at t = 0, so it is not an eigenstate of H(0) and gives no adiabatic guarantee. On the periodic grid it equals the uniform state. initial_state_preparation="structured" emits each encoding's own structured construction under that one recipe name, one H per qubit on the binary encoding, and on the one-hot encoding over the Dirichlet grids the kinetic ground state costs the same 3(K - 1) CX per variable as the uniform state, because both have full positive support. In a comparison with UniformState() on two-variable Dirichlet interior grids, with the classical Schrodinger model, the quadratic schedule with gamma 0.3, midpoint weights, T = 10, 200 steps and exact readout, it gave the smallest best-point distance in all six standalone refinement runs with the most probable point (double well, anisotropic quadratic and Ackley, each with gains 1 and 8, at most ten levels) and a much better refined constrained Rastrigin point (distance 0.0013 against 0.034 under most_probable, three levels per augmented-Lagrangian round). It is not better everywhere. The refined most_probable run on a mixed-scale problem took 5 rounds instead of 3, and the unrefined mode_or_mean run on the unit disk ended farther from the optimum (distance 0.023 against 0.00079). The split-step flavor with 3,200 steps gave the same refined constrained Rastrigin distances, so the Schrodinger and split-step flavors share the default. These runs were made with Python 3.12.14, NumPy 2.5.2, SciPy 1.18.1 and SymPy 1.14.0 on macOS arm64, and the QHD guide lists their remaining settings (docs/algorithms/qhd.md, "Measured evidence"). Revisit the initial state if a comparison at resolved step counts on another problem class favors a different start. UniformState(), the start of Leng et al.'s Algorithm 1, step 3, stays an explicit option. The binary synthesis defaults are exact: min_cx takes the cheaper of two exact syntheses per table, relabel saves the swap layers without changing the operator, and the QFT keeps every controlled phase. |
QHD schrodinger_infidelity_tolerance / ir_product_infidelity_tolerance |
0.1 / 1e-9 | src/nwqlib/algorithms/qhd/records.py |
Explicit selected state-comparison criteria. The binary IR floating-point budget describes state operations on computed phase arrays and selected stored QFT angles. Its first-order arithmetic model excludes phase-array formation and the discrepancy between a reconstructed Walsh array and the exact action of the stored rotations. The kept-state phase-admission allowance includes R_W, but R_W is not propagated into IR readout accuracy, the mode tie window or verification uncertainty. No automatic reference or process guarantee. Revisit with the intended state/schedule and numerical accuracy. |
qhd/records.py::QHDVerification tolerances |
objective gap 1e-12, Schrodinger infidelity 0.1, IR product infidelity 1e-9, minimum 0 |
src/nwqlib/algorithms/qhd/records.py |
The default comparison tolerances are 1e-12 for an objective difference, 0.1 for Schrodinger infidelity and 1e-9 for IR-product infidelity. The default minimum_tolerance=0.0 counts equality of evaluated binary64 objective values. These thresholds define comparison decisions and evaluated success mass. They do not certify exact objective values or equality between the computed and mathematical feasible sets. Revisit against the requested physical comparison and discretization. |
rotation_threshold default |
0 radians | src/nwqlib/algorithms/qhd/validation.py::DEFAULT_ROTATION_THRESHOLD |
Preserves every computed nonzero angle independently of objective or time scale. Explicit positive thresholds are approximation choices, and the raw dropped angles are recorded. Revisit only with a declared physical error allocation. |
| QHD grid admission | spacing h finite and positive, h*h and h**2 normal, (4 h) h finite and 0.25/(h*h) normal, and binary64 coordinates finite and strictly increasing, the periodic grid from lower to below upper, the interior grid strictly inside, the endpoint grid from lower |
src/nwqlib/algorithms/qhd/grid.py::OneHotGrid.__post_init__ |
Derived, not tuned. The code that uses the grid computes the kinetic coefficients 1/(h*h), -0.5/(h*h), -w/(4.0*h*h) and 1/h**2 in binary64. An overflowing denominator makes a coefficient zero, an underflowing h*h divides by zero, and a subnormal value keeps fewer than 53 significant bits. The range is about 1.5e-154 <= h <= 3.3e153. Strict increase rejects duplicate encoded points, as on [1e16, 1e16 + 4) with K = 4. The endpoint grid's last coordinate is not required to equal upper, which rounding misses for many ordinary boxes, for example [-1, 1] with K = 50. The coordinate pass streams K values per variable and is admitted with the d K initial-state work. Revisit if the kinetic coefficients or coordinates move to another precision. |
| Box refinement defaults | max_levels=3, mass_threshold=0.99, max_no_improve=2, point_rule="most_probable", box_rule="centered" |
src/nwqlib/algorithms/qhd/refinement_records.py::BoxRefinement |
Three levels is the default of the refinement prototype written by the authors of Wu et al., arXiv:2605.12066v1. The threshold 0.99 per axis is the eta of that paper's Sec. V (the scripts of its constrained experiment use 0.9999). In exact arithmetic, marginal masses of at least 0.99 give the joint box at least max(0, 1 - d/100) of the level's valid mass (refinement._joint_mass_bound). With counts these are empirical masses conditional on a valid outcome, not a bound for the underlying distribution. Two levels without improvement is the setting of the paper's experiment scripts. The centered rule lets each kept grid point represent the cell centered on it. On the interior Dirichlet grid the first point lies at a + h, so the left-endpoint cells of the paper's Eq. (14) would move the lower bound inward even when the interval starts at the first point. Revisit the threshold after comparing thresholds on the paper problems. The most probable point is the candidate rule stated in the paper's Sec. V, the default for exact and sampled readout alike, and it keeps a point that the computed distribution defines. With the default kinetic initial state it matched best_observed on exact refined constrained Rastrigin and on the standalone double-well and Ackley comparisons, and on sampled refined constrained Rastrigin it tied twice and gave a better point once in three seeds. With the uniform initial state best_observed gave better points on refined constrained Rastrigin and on sampled standalone Ackley, so it stays an explicit option. These comparisons were made in the environment that the refinement module docstring states, with classical Schrodinger evolution over T = 10 in 200 steps and max_work = max_bytes = 10**10 per solve. The standalone runs used the problems and grids of the scaling comparison in the next row, gain 8, at most ten levels and max_no_improve=2. Each augmented-Lagrangian round on constrained Rastrigin refined at most three levels at gain 8 with max_no_improve=2, for at most 15 rounds from penalty 1 at normalized tolerance 1e-6. Sampled runs drew 256 outcomes per level from each computed distribution (seeds 7, 19 and 43) by composing one-level classical solves, since classical execution takes no shots. Within a level, best_observed chooses the least evaluated value of the objective that level solved. A refined round chooses the completed level with the least recorded relative L_k value, with earlier levels winning exact ties. The refinement boxes still depend on the distribution. The comparison used the classical Schrodinger model on two-variable problems and does not establish that either rule is better in general. |
| Box refinement scaling and potential_gain | scaling="search_model", potential_gain=None, which selects DEFAULT_POTENTIAL_GAIN = 8 for the search model |
src/nwqlib/algorithms/qhd/refinement_records.py::BoxRefinement, BoxRefinement.gain |
The search model uses the same form of unit-grid kinetic operator and level-normalized potential and is invariant under translations, positive coordinate rescalings and additive constants of the objective (refinement._level_problem). The physical model solves the original objective and has no gain, so it rejects an explicit potential_gain. In a comparison on two-variable Dirichlet interior grids with the classical Schrodinger model, T = 10, 200 steps, the quadratic schedule with gamma = 0.3, the uniform initial state, threshold 0.99, the most probable point, max_levels=10, max_no_improve=10 and max_work = max_bytes = 10**10 (seed 7, Python 3.12.14, NumPy 2.5.2, SciPy 1.18.1 and SymPy 1.14.0 on macOS arm64), the physical model stopped with an unchanged box after 7 levels on the double well of the QHD guide with K = 12 and after 3 on the anisotropic quadratic (x - 3/10)**2 + (y - 2/5)**2 + x y/2 on [0, 1] x [-1/2, 3/2] with K = 12, where the search model kept shrinking for ten levels. On the two-variable Ackley function of Wu et al., arXiv:2605.12066v1, Eq. (16), on [-5, 5]**2 with K = 32, both reached ten levels, and the physical model ended closer to the minimizer than the search model with gain 1 and farther than with gain 8. That minimizer is the binary64 point (0.961528396769231, 0.3696594771697266) of examples/generators/qhd_scientific.py. The paper's Sec. VI.A reports this seed-123 shift as x* ≈ (0.962, 0.370), and the point is the fourth of the successive uniform(-1, 1, size=2) draws from numpy.random.RandomState(123) that the authors' experiment script, which the paper does not link, makes one per test function, the Ackley draw after those of the quadratic, Rosenbrock and Rastrigin functions. No gain is best for every problem (refinement._level_problem shows why kappa has no monotone effect). With the kinetic initial state and otherwise the same settings, gain 8 gave a smaller best-point distance than gain 1 after ten levels on all three standalone problems: double well 5.0e-6 against 2.2e-4, anisotropic quadratic 3.0e-6 against 3.9e-4, and Ackley 3.0e-8 against 1.4e-5. Its evolution work was 8 to 13 percent higher in the uniform-start runs. With best_observed and the uniform initial state, gain 1 did better on the exact-readout double well (0.00112 against 0.00360) and on one of three sampled Ackley seeds, so both controls stay configurable. The sampled runs, in the same environment, used max_no_improve=2 and drew 256 outcomes per level from each computed distribution (seeds 7, 19 and 43) by composing one-level classical solves. Search refinement uses a fixed unit-grid kinetic operator and a level-dependent normalized potential. Its Hamiltonian has the form a(t) T_u + kappa b(t) (F-c)/E. The levels share this form, but their sampled potentials can differ. Physical refinement keeps physical coordinates, so shrinking a box can increase the kinetic scale through the inverse square of the grid spacing. The step count must resolve both the kinetic and potential evolution. A larger gain can increase potential phases and product-formula error, so the default gain does not establish that a chosen step count is adequate. The gain-8 convergence evidence uses classical Schrodinger evolution. For a circuit-product comparison, also check the product approximation and the numerical-reference limitations of ir_product. Only gains 1 and 8 were compared, and the comparison makes no claim for other schedules, dimensions or problem classes. |
| Box refinement width floor | each rounded midpoint of two adjacent grid coordinates strictly between them, and the first and last strictly inside the faces | src/nwqlib/algorithms/qhd/refinement.py::_resolved |
Derived, not tuned. On a resolved side every grid point has its own binary64 coordinate, and the next box is not empty (_next_box). A side of n grid intervals reaches the floor at a width of the order of n ulp(M), with M the largest coordinate magnitude. It keeps coordinates distinct, not objective values. Revisit if coordinates move to another precision. |
| Box refinement resolution threshold | a positive exact sum of the support-table ranges at most sum_S ulp(max_i abs(T_S(i))) stops the level as unresolved_objective, a zero sum as flat_objective |
src/nwqlib/algorithms/qhd/refinement.py::_resolution_stop |
A resolution rule, not a proof of flatness. At most one unit in the last place of each table's largest value, the variation cannot be told from evaluation error, and dividing by E would turn it into a potential of order one. A varying function can stop, such as exp(2**-52 x) on the grid {0, 1}, and a level whose evaluation error exceeds the threshold passes, so RefinementLevel.conditioning is reported, not bounded. For normal values the threshold is at most 2 u sum_S max abs(T_S) with u = 2**-53. Revisit if table values move to another precision, or if an evaluation-error bound for the table expressions becomes available. |
| Box refinement stall split | stall_split="none", max_splits=1 |
src/nwqlib/algorithms/qhd/refinement_records.py::BoxRefinement |
A split gives up the discarded region's probability on the evidence of two objective values, and a narrow, deep minimum in that region can reverse their ranking, as can a broad basin when the distribution is spread out and each side's most probable point lies far from its minimum (refinement._stall_split), so the default keeps the box_unchanged stop. A budget of one split allows one such discard per refinement. The double-well example in that docstring made one split under budgets 1 and 3 and lowered the best-point error from 0.0826 to 0.0108 in ten levels (in the environment and with the seed that the refinement module docstring states). Revisit after comparing the split on the paper problems. |
| Box refinement level initial state | level_initial_state="configured", level_gaussian_width=None, which selects DEFAULT_LEVEL_GAUSSIAN_WIDTH = 1/6 with "best_point_gaussian" |
src/nwqlib/algorithms/qhd/refinement_records.py::BoxRefinement, DEFAULT_LEVEL_GAUSSIAN_WIDTH |
A Gaussian at the best point so far carries its location into the next box but adds momentum components that the kinetic ground state does not have. On (x - 3/10)**2 over [0, 1] with K = 16, the split-step flavor, T = 1, the quadratic schedule with gamma 0.3, gain 8, max_levels=6 and max_no_improve=6 it ended at best-point error 0.00668 against 0.0534 for the kinetic ground state at every level, and with the default max_no_improve=2 it stopped after three levels at 0.112. On the gain-1 double well of the QHD guide with K = 12, T = 10, 200 steps, a uniform first level and max_levels = max_no_improve = 10 it stopped at 0.116 against 0.0062 for the uniform and 0.00022 for the kinetic start at every level, and after a kinetic first level it reached 0.00137. The kinetic start at every level gave the smallest error in all six cases of that two-variable comparison, with either first level for the Gaussians (refinement._best_point_gaussian), so the default keeps QHD.initial_state. These runs were made with seed 7, Python 3.12.14, NumPy 2.5.2, SciPy 1.18.1 and SymPy 1.14.0 on macOS arm64. The width is a fraction of each side of the level box, the unit-coordinate width of the search model. 1/6 is the only width compared, a tested value and not an optimum. Revisit with a comparison of widths or problem classes that shows when the Gaussian helps. |
| QHD selected region coverage confidence | total run failure probability alpha = 0.05, half for each method | src/nwqlib/algorithms/qhd/_coverage.py::DEFAULT_FAILURE_PROBABILITY |
A predeclared confidence policy for default print/report. Standalone refinement uses H = max_levels, and AL uses H = max_iterations * max_levels. Hoeffding covers all candidate marginal intervals, CP all candidate joint boxes, and the report chooses the larger available lower bound. Each concerns the level's backend-sampled distribution conditioned on valid decoding. CP unavailability keeps the original Hoeffding half-budget. The user can set alpha in report or pass None. Revisit if the default confidence policy or sampling model changes. The derivation and binary64 evaluation limits are in Proposition 49. |
| Box refinement split confidence | false-valley allowance alpha = 0.01 per refinement | src/nwqlib/algorithms/qhd/refinement.py::_FALSE_VALLEY_ALLOWANCE |
A confidence policy, not a derived number. With counts, a valley is admitted for a stall split only when the smaller flanking peak's count exceeds the valley's by more than sqrt(2 N log(2 H d K/alpha)) (N valid shots, H max_levels, d variables, K grid points), which by Hoeffding's inequality for each cell's empirical mass (doi:10.1080/01621459.1963.10500830, Theorem 1, Eq. (2.3), p. 15, with the lower tail of Eq. (1.4), p. 13) and a union bound over cells and levels keeps the probability of admitting any valley that sampling cannot resolve at most alpha within one refinement for independent shots (refinement._valley_admission). It concerns the population that the backend samples, so correlated shots, drift between pooled batches or a systematic device error against the ideal model need a different claim. In the augmented-Lagrangian layer each of R rounds runs its own refinement, so over the run the union bound gives at most R alpha. A split discards probability, so an unresolved sampled valley must not cause one. A smaller alpha declines more real valleys at a given shot count. Revisit if sampled splits are compared on the paper problems. |
| Box refinement improvement test | strict decrease of the best relative objective F - C | src/nwqlib/algorithms/qhd/refinement.py::refine_box |
C is one exact constant for the whole refinement, the constant term of the first level's decomposition (refinement._tabulated_objective), so a large additive constant of F cannot round the objective's variation out of the comparison when every coefficient of F is an exact SymPy number. A binary64 Float coefficient makes SymPy round each constant term at the size of that constant before the subtraction. Exact comparison replaces the relative threshold (best - current)/max(abs(best), 1e-30) > 1e-16 of the Wu et al. arXiv:2605.12066v1 experiment scripts. For abs(best) >= 1e-30 the two agree, because adjacent binary64 numbers differ by a relative 2**-52/(2 - 2**-52), about 1.11e-16, or more. |
| QHD number-projector lowering costs | structured provider minimum of the native phase-diagonal construction and multi-controlled-phase CX laws | src/nwqlib/subroutines/hamiltonian_evolution/pauli_evolution.py |
Circuit-free routing selects the exact provider and prices that same realized nonzero-angle artifact. Exact-zero projector blocks are omitted before routing and cost zero, independently of the rotation threshold. A direct structured diagonal-provider call with angle zero emits no gate. Literal costs and transpiled equality are pinned beyond support 3. |
QHD augmented-Lagrangian defaults¶
The fields belong to src/nwqlib/algorithms/qhd/constrained_records.py::AugmentedLagrangian unless another site is named. Tolerances, penalty tests and multiplier bounds act on normalized quantities (docs/algorithms/qhd.md, "Constrained problems"). Revisit the rows marked "Study" when a study on the released QHD model compares these choices on problems with known references.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
initial_penalty, penalty_growth, max_penalty, max_iterations |
1, 2, 1e9, 15 | src/nwqlib/algorithms/qhd/constrained_records.py |
Wu et al. arXiv:2605.12066v1, Sec. VI.B (the authors' prototype stops after 10 rounds and has no penalty cap). Study. |
reduction_ratio |
0.25 | src/nwqlib/algorithms/qhd/constrained_records.py |
The authors' prototype. Birgin and Martinez, doi:10.1137/1.9781611973365, Algorithm 4.1, only require 0 < tau < 1. Study. |
objective_scale, equality_scales, inequality_scales |
1 for every scale | src/nwqlib/algorithms/qhd/constrained_records.py |
A dimensionless problem. A problem with mixed units needs caller scales. Automatic gradient scaling, such as the book's Eq. (10.5), is not implemented. |
DEFAULT_INEQUALITY_FEASIBILITY_TOLERANCE, the value of feasibility_tolerance=None |
1e-9 with inequalities only, none with an equality | src/nwqlib/algorithms/qhd/constrained_records.py |
The constraint tolerance of Wu et al., Sec. VI.B. It classifies the normalized residual of a returned point and claims no accuracy of the solution. A grid point strictly inside the feasible set has zero violation, a grid point on the boundary can have a rounding residual that a positive tolerance accepts, and for x² + y² - 1 on the interior grids of [-1, 1]² with K = 16, 32 and 64 the smallest positive residuals, 1/289, 1/1089 and 9/4225, lie far above both 1e-9 and 1e-6. With the default point rule the tested refined runs on the unit disk (K = 16 on [-1, 1]², round 7) and on constrained Rastrigin (K = 32 on [-5, 5]², round 10) stopped at the same points and rounds with either value, under the Schrodinger flavor at 200 steps and the split-step flavor at 3,200 steps. These runs were made with Python 3.12.14, NumPy 2.5.2, SciPy 1.18.1 and SymPy 1.14.0 on macOS arm64, with classical execution on the Dirichlet interior grid, the kinetic ground state, QuadraticSchedule(gamma=0.3), total time 10 and refinement in every round with the search model, gain 8, at most three levels and mass threshold 0.99. A caller who accepts a small positive violation, as the off-grid mean of mode_or_mean can have, sets a larger tolerance explicitly. The residual that a grid can reach on an equality depends on the grid and the constraint, so the caller sets it. Revisit if a problem class with the default point rule stops at different points under 1e-9 and 1e-6. |
complementarity_tolerance=None |
The resolved feasibility tolerance | src/nwqlib/algorithms/qhd/constrained_records.py |
The book's Eqs. (10.7) and (10.8) share one tolerance. |
multiplier_bounds=None |
No safeguard, stated in the result | src/nwqlib/algorithms/qhd/constrained_records.py |
The book gives no bounds valid for every problem. Its convergence proofs use the boundedness that Step 4 of Algorithm 4.1 gives the multipliers of every round (p. 35, and the proof of Theorem 5.1, pp. 41–42). Its analysis of the penalty in Chapter 7 also assumes that the true multipliers lie inside the bounds, away from lambda_min, lambda_max and mu_max (Assumption 7.7, Sec. 7.5, p. 64), which no default can know. Without that assumption the global convergence results of Chapter 6 still hold, but the penalty may no longer stay bounded (p. 64). The docstring of src/nwqlib/algorithms/qhd/constrained_records.py::MultiplierBounds states these conditions. |
equality_multipliers, inequality_multipliers |
Zeros | src/nwqlib/algorithms/qhd/constrained_records.py |
The authors' prototype. The book's Algorithm 4.1 starts the multipliers inside the safeguard, so a run refuses equality_multipliers=None for a problem with equalities when the equality interval of multiplier_bounds excludes zero. |
absorb_bounds |
False | src/nwqlib/algorithms/qhd/constrained_records.py |
Absorption changes the grid and the dynamics, and only the endpoint grid can represent an optimum on an absorbed boundary. |
stationarity |
False | src/nwqlib/algorithms/qhd/constrained_records.py |
Request the diagnostic with options=AugmentedLagrangian(stationarity=True). Read it from result.last.evaluation.stationarity when result.last is present. If the value is unavailable, result.last.evaluation.stationarity_unavailable gives the reason. It differentiates f, h and g, charged as layer work, and does not enter the default feasibility-and-complementarity stopping rule. |
inequality_form |
phr |
src/nwqlib/algorithms/qhd/constrained_records.py |
The PHR form for every kept inequality. slack and auto are explicit choices (Proposition 54 of docs/mathematics.md, src/nwqlib/algorithms/qhd/constrained.py::_round_problem). A comparison of the three forms kept this default because the automatic rule has no solution-quality criterion (QHD guide). Revisit when a restricted class has the quality evidence that ROADMAP's narrower-eligibility item names. |
inner_point |
most_probable, for exact and sampled readout |
src/nwqlib/algorithms/qhd/constrained_records.py |
Wu et al., arXiv:2605.12066v1, Sec. V, where each refinement level takes the most probable grid point, and the same default as BoxRefinement.point_rule, whose row under "Box refinement defaults" gives the comparison behind it and its limits. A refined round compares the recorded relative values of its inner objective, L_k under the PHR policy and L(x, s) with an inequality representation, at its completed levels and chooses the least one, with the earlier level winning an exact tie. Within a level, best_observed chooses the least evaluated value of the objective that level solved. A search-scaled level solves its normalized search objective, while a physical level solves the inner objective. Comparisons between levels use recorded relative values of the inner objective. With exact evaluations and full observation of the first product grid, best_observed gives an inner objective value no larger than that grid's inner minimum. Its projection has PHR value at most the first original grid's PHR minimum plus the recorded first-grid slack error bound, under Proposition 54's cap, mesh and arithmetic premises. With table error at most epsilon_T in inner-objective units and relative-comparison error at most epsilon_R, the additional allowance is 2 epsilon_T + 2 epsilon_R. The comparison term is unnecessary with one level. These numerical error bounds are premises that the API does not certify. Partial observation, another point rule, or a numerically flat level does not supply the full-grid premise. mode_or_mean needs off-grid evaluations. |
termination |
feasibility_and_complementarity |
src/nwqlib/algorithms/qhd/constrained_records.py |
The book's Eqs. (10.7)–(10.8). The complementarity test can keep the outer loop running when both a normalized slack and its tentative normalized multiplier exceed the complementarity tolerance. feasibility follows Wu et al.'s code, with <= where their code compares strictly. With the default kinetic initial state and the most probable or best observed point, continuing past this stop to a stop on the penalty measure of the book's Eq. (4.9) found the same best feasible points on five exact refined problems and six sampled runs. Where the two stops differed, continuing took more rounds and a larger final rho, for example 11 rounds and rho 512 against 10 and 256 on constrained Rastrigin. With the uniform start or mode_or_mean some runs did improve by continuing. Neither stop certifies stationarity or optimality. The comparison ran with Python 3.12.14, NumPy 2.5.2, SciPy 1.18.1 and SymPy 1.14.0 on macOS arm64, with classical Schrodinger execution on the Dirichlet interior grid, QuadraticSchedule(gamma=0.3) with midpoint coefficients, total time 10 and 200 steps, normalized tolerance 1e-6, and refinement in every round with the search model, gain 8, at most three levels, max_no_improve=2 and mass threshold 0.99. The five exact problems and the six sampled runs are those that the comment at AugmentedLagrangian.termination lists. The split-step flavor gave the same finding on exact constrained Rastrigin at 3,200 steps, on the exact combination problem at 320 steps and on the three sampled constrained Rastrigin seeds. |
penalty_update |
on_insufficient_decrease |
src/nwqlib/algorithms/qhd/constrained_records.py |
The book's Algorithm 4.1, Eq. (4.9). every_iteration follows Wu et al.'s code. |
GRID_MINIMUM_MAX_WORK, GRID_MINIMUM_MAX_BYTES for qhd/constrained.py::constrained_grid_minimum |
1,000,000,000 and 10 GB | src/nwqlib/algorithms/qhd/constrained_records.py, src/nwqlib/algorithms/qhd/constrained.py |
The defaults of QHDVerification, for the same kind of explicit grid evaluation. constrained_grid_minimum evaluates the original functions in admitted C-order slabs of b points and keeps no D-entry table. It charges W_min = D (d + N_f + sum_j N_j + 8 C + 12) units for D = K**d grid points, C kept constraints and tree sizes N_f and N_j (node counts), in a declared node/coordinate/reduction-input unit. The constraints add normalization, absolute or positive-part/max operations, finite checks and the final feasibility comparison, and the fixed 12 visits fund objective validation, mask/selection and running-minimum updates. With sequential constraint evaluation a sufficient slab rule is B_min(b) = 8 d K + H_eval + max_j W_eval,j(b) + (8 d + 96) b bytes, whose rate covers coordinate/index construction, the objective and violation arrays, finite-real conversion, absolute/maximum/normalization temporaries, a feasibility mask and a where result, and each constraint is scaled in place on its private buffer. Here W_eval,j(b) = 8 b (N_j + 2) is the evaluator allowance of f and each constraint (row "Support-table evaluator allowance"), an engineering allowance, not a derived bound, on the assumption that the NumPy-printed callable holds at most one slab-sized temporary per expression node, plus the output and one broadcast input, and H_eval = H0 is an allowance until measured. The slab b is the largest that max_bytes admits after the coordinate axes and headers, and these defaults bound the same two arguments. The running best value and index, the feasible count and the chosen value are streamed. The grid_minimum reference of src/nwqlib/algorithms/qhd/verification.py charges D (d + N_f + 12) units and 8 d K + 8 D + H_eval + W_eval,f(b) + (8 d + 96) b bytes for its one array table of the original objective, evaluated in slabs of b points, with W_eval,f(b) = 8 b (N_f + 2) and H_eval = H0 + 256 (N_f + 2 d + 16) (src/nwqlib/algorithms/qhd/verification.py::_grid_minimum_bytes). constrained_grid_minimum(result) admits its work, evaluates f, h and g on all K**d grid points, and finds the least evaluated objective among grid points that pass the computed feasibility test. Its signed difference subtracts this freshly evaluated minimum from the result's stored best objective. That difference can be negative, including when the stored point belongs to the evaluated feasible grid, because the two values can have different rounding errors. An off-grid or infeasible stored point gives no grid-optimality conclusion. If no grid point passes computed feasibility, indices, point, objective and gap are None. The gap is also None when the result has no best point. The helper does not support a result that uses box refinement. Revisit for a deliberately larger reference grid or a changed evaluation. |
QHD numerical guards¶
Numerical guards and tolerances derives the probability windows, the split-step state budget and the tie window that use these constants.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
TRANSFORM_ROUNDOFF |
5, in units of u times the level count L | src/nwqlib/algorithms/qhd/split_step.py |
Qualification constant of the assumption ||fl(T x) - T x|| <= 5 u L ||x|| for SciPy's orthonormal FFT, inverse FFT and DST-I, with L = ceil(log2 K) for the FFT and ceil(log2(2 (K + 1))) for DST-I, covering normalization, twiddles, Bluestein lengths and the transforms of the real and imaginary parts of a complex DST. SciPy 1.18.1 evaluates all three with ducc0.fft (scipy.fft._duccfft, SciPy 1.18.0 release notes, "scipy.fft improvements"). Neither SciPy nor ducc0 publishes such a bound, so it is a stated qualification assumption about that backend, not a derivation, and the state budget that uses it is first order in u. 80-digit mpmath sums on three unit inputs each, at ten FFT and inverse-FFT lengths and eight DST-I lengths between 2 and 127, including primes and a DST-I logical length with a large prime factor, measured at most 0.80 in units of u L, for the forward FFT at K = 2 (Python 3.12.14, NumPy 2.5.2, SciPy 1.18.1 and mpmath 1.3.0 on macOS arm64), and test_scipy_transforms_meet_the_qualification_constant in tests/test_qhd_split_step.py keeps that canary. Revisit when SciPy changes its FFT backend or the canary fails. |
EIGENVALUE_ROUNDOFF |
13, in units of u | src/nwqlib/algorithms/qhd/split_step.py |
Relative error of one kinetic eigenvalue 2 (g(x)/h)**2 with x in [0, pi/2] and g = sin or the identity: 3u for x, 5u for sin(x) under one-ulp sin and condition number at most 1, one rounding for the division by h and doubled error plus one rounding for the square. Derived, not tuned. Revisit with the eigenvalue formula or the sine model. |
| QHD binary product budget: H / controlled phase or diagonal / start vector | 3 / 5 / 2, in units of u | src/nwqlib/algorithms/qhd/theory.py::binary_product_state_error |
2-norm charge of one operation of the classical binary product, relative to the computed phase arrays and selected stored QFT angles. An H gate rounds the complex sum, the real product with the scale fl(sqrt(1/2)) and the scale itself. A controlled phase or phase diagonal is GLOBAL_PHASE_STATE_ROUNDOFF per entry. The start vector 1/sqrt(D) has one square root and one division. _validation.state_mass_window turns the budget into the mass window, and probability_difference_window into the tie window. The binary IR floating-point budget describes state operations on computed phase arrays and selected stored QFT angles. Its first-order arithmetic model excludes phase-array formation and the discrepancy between a reconstructed Walsh array and the exact action of the stored rotations. The kept-state phase-admission allowance includes R_W, but R_W is not propagated into IR readout accuracy, the mode tie window or verification uncertainty. Revisit when the kernel changes its arithmetic. |
QHD one-hot product budget: hopping / projector (ONEHOT_HOPPING_STATE_ROUNDOFF / ONEHOT_PROJECTOR_STATE_ROUNDOFF) |
6 / 10, in units of u | src/nwqlib/algorithms/qhd/theory.py::onehot_product_state_error |
First-order 2-norm charge of one executed block of the direct one-hot kernel, relative to the analytic block at its stored angle. A hopping forms cos and sin within 2u each, an error matrix of norm at most 2 sqrt(2) u, plus sqrt(2) u for the real scalings and u for the addition, (1 + 3 sqrt(2)) u < 6u. A projector forms its active factor from two phase entries and one complex product and multiplies the state once, (4 + 4 sqrt(2)) u < 10u on the active slice and at most 5u elsewhere. The budget is u (start + 6 B_H + 10 B_P + 5) over every executed occurrence, under normal round-to-nearest arithmetic and one-ulp sine, cosine and complex exponential. A zero-angle block is skipped exactly and has no charge. Revisit when the kernel changes its arithmetic. |
qhd/binary.py::qft_error_bound outward factor and AQFT total factor |
1 + 8u and 1 + 4u |
src/nwqlib/algorithms/qhd/binary.py |
The sum of omitted controlled-phase norms has relative error below 5u (binary64 pi, one-ulp sine, one product, one fsum rounding), so 1 + 8u rounds each QFT bound outward, and 1 + 4u covers the two products of the total 2 N_s d e_F. Derived, not tuned. Revisit with the sine accuracy assumption or another format. |
qhd/schedules.py::_ATANC_ONE_BELOW |
2**-27 |
src/nwqlib/algorithms/qhd/schedules.py |
Below this z, the quadratic kinetic integral takes atan(z)/z as 1, whose relative error z**2/3 is at most u/6, and it avoids dividing a subnormal atan(z) by z. Derived, not tuned. Revisit condition: Another floating-point format. |
qhd/validation.py::_NORMAL_MIN, _NORMAL_MAX and the range rule _normal_range |
2**-1022 and (2 - 2**-52) 2**1023 (sys.float_info.min, max) |
src/nwqlib/algorithms/qhd/validation.py |
A multiplication, division or power-of-two scaling that a normal relative-error argument uses must have exact result zero or magnitude in this range, zero only where the operands or the producer's structure give zero. Then rounding errs by at most u relatively, with no inexact underflow, and a power-of-two scaling is exact. Only a rounded result at either end needs the exact value. QHD planning checks its durations, weights, coefficients, angles and identity products where each is formed and before pruning. Stored table values, the objective constant and initial amplitudes need only be finite, and a contribution formed from them whose exact result would be nonzero and below 2**-1022 is omitted and charged to the error record (validation._underflows, records.QHDRangeOmissions). Additions may leave exact subnormal scratch, and error bounds, rounded upward, may be subnormal. Exact format constants. Revisit condition: Another floating-point format. |
qhd/validation.py::_ANGLE_UNIT_BITS |
1074 | src/nwqlib/algorithms/qhd/validation.py |
Every finite binary64 number is an integer multiple of 2**-1074, so angle_units writes each dropped rotation angle as an exact integer count of that unit, and the running pruning total sums them without rounding. The exact sum, halved for binary Rz angles, plus the exact charges of contributions omitted below the normal range, is converted to binary64 once and moved to the next number upward when the conversion lands below it, so QHDReconstruction.pruning_error_bound is never below the exact bound and is 0.0 when nothing is removed. Exact, not tuned. Revisit condition: Another floating-point format. |
qhd/schedules.py::_UNIT_MEAN_NODES |
four panels of the 16-point Gauss–Legendre rule | src/nwqlib/algorithms/qhd/schedules.py |
The cubic schedule's kinetic integral averages its positive transformed integrand over four equal panels of [0, 1] with 16 nodes each (_GAUSS_LEGENDRE_16, within 0.46 ulp of the 50-digit rule). The panel half-width 1/8 and the distance sqrt(3)/2 of the integrand's poles from [0, 1] bound the truncation error by 1.9e-23 relative, below 1.7e-7 u, so roundoff sets the integral's 48u bound. An interval costs 64 or 128 integrand evaluations. The integrated weights of CubicSchedule(s=1) for 10,000 steps to T = 10 took 0.08 s (Python 3.12.14 on an Apple M3 Max). The integrated coefficient rule charges it as 129 work units per step. Revisit condition: A tighter accuracy target, a step count where this cost matters, or another integrand. |
qhd/schedules.py::QuadraticSchedule.kinetic_integral_roundoff, CubicSchedule.kinetic_integral_roundoff and CubicSchedule.kinetic_integral_range |
32, 48 and (1e-8, 1e8, 1e4) |
src/nwqlib/algorithms/qhd/schedules.py |
First-order relative error constants C of the computed kinetic integral A, |A_hat - A| <= C u A + 2 tau with u = 2-53 and tau = 2-1074, derived in the kinetic_integral docstrings and the module docstring. The quadratic constant holds for gamma > 0 over the whole finite binary64 range, and the cubic one on s in [1e-8, 1e8] and 0 <= t <= 1e4, the range that kinetic_integral_range records. evolution_bounds._coefficient_term uses C for the first-order coefficient residual of the QHD error record, which is then an estimate, and outside the cubic range it reports the residual as unavailable. Derived, not tuned. Revisit condition: A changed kinetic_integral. |
qhd/schedules.py::interval_work |
2 for QuadraticSchedule and ShiftedCubicSchedule and 129 for CubicSchedule |
src/nwqlib/algorithms/qhd/schedules.py |
Planning work of the integrated coefficient rule per step (QHD._admit_symbolic_work): one closed form for A and one for B, or for the cubic schedule at most 128 integrand evaluations of A (64 Gauss nodes on each side of t = s**(1/3)) and one closed form for B. The midpoint rule charges 2 units per step for its two point values. Revisit condition: A changed integral routine. |
qhd/initial_state.py::_EXPONENT_RELATIVE, _EXPONENT_ABSOLUTE, _SUBNORMAL_ULP, _OUTWARD, the overflow test z_min <= 2**1020 and the separation factor 1 + 16u of GaussianState._variables |
5u/(1 - 10u), 2-1073, 2-1074, 1 + 2**-40, 2**1020, 1 + 16u |
src/nwqlib/algorithms/qhd/initial_state.py |
Constants of the finite Gaussian and product-state bounds. A computed exponent has five rounding factors, so gamma_5/(1 - gamma_5) = 5u/(1 - 10u) bounds its relative error against the computed value, and 2-1073 covers gradual underflow. One ulp below the normal range is 2-1074, the error of exp there and the representable allowance for every subnormal rounding, including the half-ulp 2-1075 that binary64 cannot represent. _OUTWARD lifts each computed bound above its exact value, covering fewer than 256 roundings of at most 2u with relative sensitivity at most 2 and the argument rounding of exp. Error norms use math.hypot, which stays finite whenever the norm is representable, and an exp, expm1 or hypot result beyond the binary64 range becomes an infinite bound, which the cap absorbs. An overflowed exponent exceeds 21022 exactly, which leaves its amplitude below 2-1074 when the minimum is at most 21020. When every exponent overflows, a nearest distance that every other rounded distance exceeds by the factor 1 + 16u is exactly nearer by the factor 1 + 8u, so the other amplitudes lie below 2**-1074. Derived, not tuned. Revisit condition: Another floating-point format. |
QHD selected operation sizes¶
QHD.max_work=1_000_000_000 bounds known symbolic/table/action work, including every expm_multiply call of a classical host evolution charged with _linalg_laws.expm_multiply_requirements, and QHD.max_bytes=10_000_000_000 bounds known arrays before selection or explicitly chosen construction. QHDVerification uses the same defaults for its separate explicit reference invocation. They exclude unknown SymPy/SciPy/Qiskit internal work and process RSS. The max_work default is a fuse against runaway planning work, sized so that no example or test workload in the repository reaches it. Revisit the defaults when a concrete larger workload is selected. Shared ExecutionLimits separately govern acquisition and stored results. Mathematics gives the charges these limits bound, for the support tables, the step rows and block occurrences, the initial state, the augmented-Lagrangian layer and each classical kernel law, with measured examples. Before any table work planning admits the evaluation of the initial state and its stored payload, one unit per grid point of each variable and 320 d K + (8 d K + 120 d + 40) + (8 d K + 120 d + 64 + 204 d) + 4096 (d + 1) bytes (src/nwqlib/algorithms/qhd/method.py::_initial_state_bytes, QHD.plan). The coefficient 320 per grid point (_INITIAL_STATE_BYTES) covers the evaluator's Python float lists and float64 vectors, and 4096 (d + 1) (_INITIAL_STATE_OBJECT_BYTES) is an untuned allowance for Python and NumPy bookkeeping. This allowance is measured and specific to the implementation and runtime, Python 3.12.14, NumPy 2.5.2 and 64-bit CPython on macOS arm64, not a derived bound. Requalify it when the evaluator, the stored representation, the object lifetimes or the runtime changes.
QHD uses explicit caller planning/construction/materialization/publication envelopes in place of automatic resource tiers. Native exact readout keeps its full 2*(dK) population. Explicit classical scalar execution keeps no array unless requested. QHD fidelity reduction roundoff uses the window of _numerics.normalized_fidelity_with_window, derived in that function. A negative raw infidelity inside that window is disclosed and adjusted to zero, with the raw value kept.
A QHD work/byte admission refusal reports the stage, its admission charge and the exceeded limits. For a complete charge, raising each exceeded limit to the displayed amount clears that admission check. Other planning or execution checks can still refuse. At a lower-bound check, counting may have stopped early. Setting a limit to the displayed amount can therefore refuse again at the same stage. If resources permit, raising each exceeded limit by a larger factor, such as twice the displayed amount, gives progress without guaranteeing that the next attempt succeeds.
Symbolic expansion and support-table admission can report lower bounds. An explicit verification request has its own QHDVerification work and byte limits, separate from the QHD Method's limits. A verification work count can also be truncated, while its displayed byte allowance is already the complete declared allowance.
Compact binary planning admits its compilation work before BinaryModel or compile_binary_steps runs. The charge includes kinetic setup, identity enclosures, each potential table's absolute mean and identity radius, Walsh normalization losses, requested synthesis trials, omission charges and final numerical reductions. A requested dense block with N entries costs at most 2N+7 units. A Walsh-capable request costs at most 7N+80 units, including a min_cx trial that ultimately selects dense synthesis.
With b=log2(K), n_S=b*len(S), N_S=2**n_S and C(b)=10K+2bK+b**2+6b+16, the additional compilation charge is
where f=1 for first order and f=2 for second order. Use L(N)=2N+7 for a dense request and L(N)=7N+80 for a Walsh-capable request. The symbolic, table and initial-state charges are added once.
These are explicit scalar/element arithmetic units, with one extra unit for each nextafter used by numerical bounds. Exact rational operations count as operations without pricing their integer bit complexity. Representation conversions, count bookkeeping and dependency internals are excluded. Classical schrodinger and split_step do not compile the compact product and do not pay this compilation term.
The separate classical binary_product_sizes work model charges (n+1)2**n for a Walsh transform. That term excludes the normalization-loss arithmetic charged by native construction, so it has a different scope from the native (n+4)2**n term.
Native binary construction counts scalar arithmetic and array-element arithmetic, including phase wrapping, Walsh transforms and normalization-loss calculations. One addition, subtraction, multiplication, division, absolute value, negation, square, or declared elementary-function call costs one work unit. A complex exponential or angle call counts as one declared unit. The charge includes arithmetic evaluated along the admitted construction path, while comparisons and arithmetic used only by range predicates, indexing, copies, permutations, records and SDK internals are outside this unit.
For b=log2(K), define
The upper charge also includes the declared gate-emission envelope. The native upper charge is
Here B ranges over emitted block occurrences, including both potential halves in a symmetric product, N_B=2**n_B, and c_B is 3 for a dense block and 5 for a Walsh block. For a potential support S, n_S=b*len(S). The support sum counts each needed Walsh table setup once. R_d counts below-range dense phase-table entries, R_w counts below-range Walsh rotations, and P counts threshold-dropped Walsh rotations. All count occurrences across the selected blocks. Structured preparation uses W_prep=d*b, the Qiskit preparation allowance uses d*b*K, and a resource-only construction has no preparation charge.
The binary byte allowances are _BLOCK_BYTES per compiled block occurrence, 64 bytes per table entry for the values, Walsh coefficients, string weights and one construction's arrays. Native circuit construction is admitted by the row qhd/native.py::_binary_native_bytes (table "Budgets and mechanical bounds"). The classical binary product charges its element operations (including the same (n + 1) 2**n per support-table Walsh transform except under a dense-diagonal potential, and d C(b) for the kinetic tables that each model construction builds), (32 + 16 i_start) D bytes for the in-place state operations: one live complex state of 16D, 8D more for the half-state copy of an H gate, a contiguous reordered copy during a QFT swap that gives 32D, and 16D for the start vector that the caller keeps live through the call (i_start = 1), and 64 bytes per table entry (src/nwqlib/algorithms/qhd/theory.py::binary_product_sizes). Before a binary model builds any array, it admits against QHD.max_bytes the caller's live data, its own arrays (16 bytes per table entry, 8 bytes per entry of one popcount array per distinct qubit count, E - 1 mask bytes per table of E entries, and the metadata allowance H_model of the row qhd/binary.py::_model_metadata), one current construction per table (16 bytes per entry plus the metadata allowance h = 4096 + 16|S| + 8 L(4200 + ceil(log2 E))), and the larger of one new construction (80 bytes per entry of the largest table plus h, and for a Walsh-capable potential the normalization-loss allowance of that row) and the phase-use workspace (64 bytes per entry of the largest table). It builds each construction once per construction key and keeps it in the whole-model cache only when its payload and h fit the remainder, and otherwise keeps only the latest construction of each table (src/nwqlib/algorithms/qhd/binary.py::BinaryModel.construction, _model_reservation). The classical binary product holds the same 48D bytes of state buffers, because each diagonal's reshaped view is released right after its in-place multiplication, 16 (sum_potential E_t + K + E_max) bytes of phase buffers and the Plan's source records (src/nwqlib/algorithms/qhd/method.py::_binary_source_reservation) in that admission (src/nwqlib/algorithms/qhd/theory.py::run_binary_product). Planning admits the model with its source records and compiler residents held, and native construction with the source records and the circuit graph held (rows qhd/method.py::_binary_source_reservation and qhd/native.py::_binary_native_bytes). circuit_resources and run_resources admit the binary rotation census against the Plan's QHD.max_work and, together with the model, QHD.max_bytes: D_dense(E) = H0 + 112E + (16 + L(n)) floor(E/2) bytes of angle formation for a dense table of E = 2**n entries, H0 + (1024 + 4 L(b_C)) U bytes for the Python population of at most U distinct magnitudes, and 2N ceil(log2 max(2, N)) + 8N comparison visits for grouping N angles (src/nwqlib/algorithms/qhd/resources.py::_census_sizes). Refinement admits the save and load of each level's tables.json against the level's QHD.max_bytes (src/nwqlib/algorithms/qhd/refinement.py::_level_table_phase_bytes). A version-3 refinement level file stores its format, unscaled support tables, table-evaluation count and rational reference offset. Resume reuses the tables and count, and obtains the reference offset from the first level's file. A null offset follows the symbolic recomputation path. src/nwqlib/algorithms/qhd/refinement.py::_level_json_bound prices the fixed ASCII skeleton, table-record/base64 envelopes, evaluation-count digits and rational-offset digits without serializing array entries. The file-byte bound is 77+T(specs)+D(evaluations)+R(offset). The census, construction and level-file allowances are qualified for 64-bit CPython 3.12.14, NumPy 2.5.2, SymPy 1.14.0 and Pydantic 2.13.5, and they are not process-RSS bounds. The other binary allowances are untuned. Revisit them when a measured binary workload approaches the limits.
Budgets and mechanical bounds¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
SCRATCH_BYTES_PER_POINT / SCRATCH_BYTES_PER_AXIS |
1024 bytes per grid point of one axis / 4096 bytes per variable plus 4096 | src/nwqlib/algorithms/qhd/split_step.py |
Measured transform-scratch allowance of the split-step byte law, beyond its counted arrays. SciPy 1.18.1's transform backend, ducc0.fft (scipy.fft._duccfft), keeps plans and fiber buffers in native memory that tracemalloc does not see. The native observation is the growth of the process high-water mark (ru_maxrss of resource.getrusage) across an in-place transform pair along either axis of a K-by-K array in a fresh process with one transform worker, less the traced allocations. It was at most 688 bytes per point over two runs at K = 1008, 1009, 1023 and 1024, which include DST-I lengths whose logical length has large prime factors. A high-water mark grows only past the earlier peak and in whole pages, so an allocation below an earlier peak does not show, and these figures are samples, not an upper bound on the backend's scratch. The kernel fixes one transform worker (workers=1 in split_step._dst and _transform), the setting of the measurement. The measuring script called scipy.fft alone, so this figure does not depend on NWQLib's code. Beyond the counted arrays tracemalloc saw at most 7034 bytes of Python and NumPy objects (array headers, the einsum buffer of the kinetic moments, per-axis index and phase temporaries) at d = 1, rising to 17340 at d = 8, below 4096 (d + 1) in every case, after the process's first scipy.fft call had loaded about 1.6 MB of modules and caches once. The cases were three-step evolutions on both boundaries and both kinetic models, with d = 1 to 3 and K = 4 to 128 and with d = 4 to 8 and K = 2 and 4, at most 300,000 points, in two runs of the kernel that forms the scaled kinetic and potential moments. The two runs differed by up to 1426 bytes in one case. SciPy 1.18.1 and NumPy 2.5.2 on macOS arm64. Revisit when SciPy changes its FFT backend, when the kernel changes its temporaries or when an admission approaches max_bytes. |
_STEP_BYTES / _BLOCK_BYTES |
768 bytes per schedule step / 4096 bytes per compiled block occurrence | src/nwqlib/algorithms/qhd/method.py |
Planning allowances set above measured sizes. _STEP_BYTES is admitted with the initial state before the expansion, and _BLOCK_BYTES in the final table admission, both before the compiler forms any step row or block (QHD._admit_symbolic_work). A Plan keeps per step its (t, a, b) row in the compiler and in QHDReconstruction.step_weights, and each row's share of the two identity JSON copies of the reconstruction and the Plan. The slope of the traced peak of plan() between 2,000 and 4,000 steps, after one warm-up Plan, was 404 to 409 bytes per step for the split-step flavor and 492 to 497 for the Schrodinger flavor, over d = 1 and 2, K = 4, the three schedules and both coefficient rules. The Schrodinger slope includes about 80 bytes per step for a generator norm that planning held in the traced code and that method._generator_norms does not keep, since it yields the norms one at a time. A compiled block occurrence, the compiler's block, its QHDBlock or QHDBinaryBlock record and its identity JSON, traced 3,257 to 3,506 bytes on the one-hot encoding and 2,446 to 2,468 on the binary encoding (the ir_product slope between 500 and 1,000 steps less the split-step rows, per block of a step, over d = 1 to 4, K = 2 to 16 and projector supports of up to 4 qubits). Python 3.12.14, Pydantic 2.13.5 and NumPy 2.5.2 on macOS arm64. Revisit when the rows, the block records, their identity encoding or the Python runtime changes. |
_SCALAR_BYTES |
1024 bytes per published host-kernel scalar | src/nwqlib/algorithms/qhd/method.py |
Readout allowance of the classical kernel's admission, set above a measured size (QHD._host_construction), one charge for each of the d K + 3 d + 13 ScalarValue records that _execute_theory returns (_scalar_names). After one warm-up call, the traced memory still held by the kernel's output grew by 594 bytes per scalar between K = 1,024 and 8,192 at d = 1 and averaged 648 bytes per scalar at K = 65,536. Small grids also keep about 20 KB of fixed Python objects, which this per-scalar allowance does not cover and _KERNEL_CALL_BYTES in the next row does. Python 3.12.14, Pydantic 2.13.5 and NumPy 2.5.2 on macOS arm64. Revisit when the kernel's outputs or the record implementation changes. |
_KERNEL_CALL_BYTES |
262,144 bytes per classical kernel call | src/nwqlib/algorithms/qhd/method.py |
Fixed allowance of the classical kernel's admission, set above a measured size and charged once per call (QHD._host_construction). It covers Python objects of bounded total size that no size law counts: the fixed objects of the readout and of the returned records and, on the Schrodinger flavor's expm_multiply calls, the 2-element tuples of SciPy's sparse shapes that CPython keeps on its free list, at most 2,000, which tracemalloc counts as about 112 KB. The kernel reads its generator norms and compiled blocks one at a time (method._generator_norms, native.raw_blocks), so none of this grows with the step or block count. After one warm-up call, the traced peak of _execute_theory exceeded the workspace that the kernel's size laws admit, for the split-step flavor less its whole transform-scratch allowance, by at most 156,775 bytes over the Schrodinger, split-step, one-hot IR-product and binary IR-product kernels. The cases were d = 1 and 2 with K = 2 to 64 and 2 to 2,048 steps, with the grid size times the step count at most 216 for the one-hot IR product up to 222 for the split-step flavor, the initial states, keep_state and both coefficient rules at 32 and 512 steps, d = 3 with K = 2 and 4, d = 4 with K = 2, and up to 32,768 steps for the Schrodinger and binary IR-product kernels at d = 1. Traced with Python 3.12.14, NumPy 2.5.2, SciPy 1.18.1 and Pydantic 2.13.5 on macOS arm64. A second trace of the same cases without the one-hot IR product at K >= 32 gave at most 155,938 bytes. The one-hot IR-product cases were traced with an earlier kernel that evaluated each block with expm_multiply, and the direct block actions of theory._run_ir_product have not been traced against this allowance. Revisit when the kernel, its outputs, SciPy or the Python runtime changes. |
| Schrodinger table and assembly charges | gamma_(T-1) for T ≥ 1 tables, zero for T = 0, gamma_3 for relative assembly, U=eta*((r+1)*abs(dt)+r) for assembly underflow, and factor 2 for centering |
src/nwqlib/algorithms/qhd/method.py::_schrodinger_step_bounds |
Here eta=2**-1074 and r=1+2d. The first broadcast addition to zero is exact. Finite binary64 sums satisfy the relative model even when their result is subnormal. Each assembly leaf has at most three roundings. The absolute errors of at most r kinetic products and one potential product propagate through the addition and final scaling, and the final r products add their own error. The bound eta/2*((r+1)*abs(dt)*(1+u)**2+r) is at most U for u=2**-53. A symmetric perturbation and its exact trace shift have combined norm at most twice the perturbation norm. The column-major COO fill and SciPy 1.18.1's ordered CSR duplicate reduction make the stored diagonal range zero. The scalar kinetic bound uses the actual h*h coefficients, including doubled periodic K=2 links. Requalify when the fill, conversion, assembly order or arithmetic changes. The SciPy state model remains conditional and first order. |
| Restricted kinetic stencil construction work | W_stencil=28S+4M+(37d+12)D+40(d+1) logical visits |
src/nwqlib/algorithms/qhd/theory.py::restricted_kinetic_work, restricted_kinetic_sparse, src/nwqlib/algorithms/qhd/method.py::restricted_sizes |
D=K**d, S is the raw COO slot count and M the final CSR entry count. Raw row, column and value writes and their separate integer-index checks cost 6S. Base, length, prefix, index, mask and gather/scatter passes cost at most (37d+3)D. The qualified SciPy 1.18.1 conversion uses at most 22S+4M+9D+5 visits for index checks and possible casts, counting conversion, sorted/canonical scans, duplicate reduction and possible compaction. At most 32(d+1) scalar shape, stride, coefficient and loop-bookkeeping operations plus the five endpoints fit within 40(d+1) for d>=1. One unit is an element or record visit in a named pass. Column-major fill gives sorted CSR rows. Admission uses the boundary-free upper counts S=3dD and M=(1+2d)D. Requalify after changes to fill order, NumPy Boolean indexing and indexed assignment, conversion, dtype checks or the work-unit convention. |
| Fixed-pattern Schrodinger setup and fill work | 5z+6D+1 once for setup and 2z+5D per fill |
src/nwqlib/algorithms/qhd/method.py::restricted_sizes, _generator_workspace, _fill_generator |
NumPy 2.5.2 repeat setup prices arange, diff, an optional count cast, count validation/summation and group traversal, plus z repeated outputs. Pattern equality visits z entries. flatnonzero counts and scans z Boolean entries and its branchless path can write z indices. The empty CSR initializes D+1 pointers. Each fill has two z-entry multiplications and D-entry potential multiplication, diagonal gather, addition, scatter-index validation and scatter writes. The integer get checks within its gather and the set has a separate validation pass. Requalify when these NumPy paths, pattern lifetimes or the logical-visit convention change. |
compiler.H0 |
65,536 bytes | src/nwqlib/algorithms/qhd/compiler.py |
Explicit allowance for scalar bookkeeping, ndarray headers, iterators and fixed sort stacks on the checked 64-bit CPython/NumPy stack, charged once in the fixed headers H_S of each support table's chunk charge, once in the table metadata H_tables and once in the serialization headers H_json,arrays (method._support_table_bytes, QHD._admit_symbolic_work). It is not a universal interpreter or process-RSS theorem. Requalify when NumPy or the Python runtime changes. |
| Dense QHD probability summary | H0+max(17D+H_decode,25D+8dK) incremental bytes, plus the borrowed 8D input bytes |
src/nwqlib/algorithms/qhd/method.py::_summarize, summary_sizes |
D is the dense population and P its positive count. For C=min(P,4096), H_decode=64C(d+1)+512(d+1) when P is positive, and zero otherwise, prices the chunk's intp coordinate arrays, Python coordinate and position lists, integer objects and descriptors on the qualified 64-bit stack. Each chunk's iterator is released before constructing the next chunk, and the position array is released before selectors and mean gathers. H0 covers fixed bookkeeping. The 25D+8dK+H0 form applies when H_decode<=8D+8dK. Other held inputs are additional. This law excludes sparse-point and object-count routes, allocator arenas and process RSS. Revisit with the decoder chunk, object convention or allocation lifetimes. |
compiler.OBJECT_HEADER_BYTES |
256 bytes per inventoried array or wrapper object | src/nwqlib/algorithms/qhd/compiler.py |
Charged in H_S for each of the N_S + 2 arrays that the evaluator allowance assumes (compiler.support_workspace). |
| Support-table metadata and serialization allowances | H_tables = 65536 + 6144 T + (24 + L(b_d)) M + 256 d bytes and H_json,arrays = 65536 + 2048 (T + d) + (8 + L(b_d)) M bytes for T tables with M total support indices in d variables, b_d = bit_length(max(0, d - 1)) and L(b) = 32 + 4 ceil(max(1, b)/30) |
src/nwqlib/algorithms/qhd/method.py::_support_table_bytes |
Qualified object allowances for 64-bit CPython 3.12.14, NumPy 2.5.2 and Pydantic 2.13.5, not universal interpreter bounds. H_tables covers two SupportValues populations at 2048 bytes per record (an inventory of 1396 bytes: the object, field dictionary and field-name set, measured at 1080 bytes, the wrapper, a hash, a cached identity, three extrema and the support-tuple header), 2048 bytes per table for the shared owning-array and view headers, the compiler's wrapper, the source support tuple and up to six construction and cache mapping entries at 256 bytes each, 24 bytes per support slot in at most three tuples, their integer referents and 256 bytes per coordinate-array header. H_json,arrays covers each serialized table's two portable dictionaries, its support and shape lists and string headers, each initial vector's array dictionary, and the support-list slots with their integers. Requalify when the record fields, their lifetimes or the runtime change. |
| Support-table evaluator allowance | W_S(b) = 8 b (N_S + 2) bytes for a chunk of b entries of a support whose NumPy-printed callable has N_S expression nodes |
src/nwqlib/algorithms/qhd/compiler.py::support_workspace, which src/nwqlib/algorithms/qhd/constrained.py::_support_tables also uses, and W_eval,j(b) of src/nwqlib/algorithms/qhd/constrained.py::constrained_grid_minimum for f and each kept constraint |
Engineering allowance, not a derived bound. It assumes that the NumPy printer's evaluation holds at most one chunk-sized float64 temporary per expression node, plus the output and one broadcast input. A primitive with complex intermediates or hidden storage lies outside it. On the shifted Ackley objective of examples/generators/qhd_scientific.py extended to d = 4 and 5 variables with K = 16, one support of N_S = 95 and 115 nodes, the traced peak of one evaluator call was 131,968, 1,573,664 and 25,166,624 bytes at b = 4,096, 65,536 and 1,048,576, about 24 bytes per entry against 8 (N_S + 2) = 776 and 936 (Python 3.12.14, NumPy 2.5.2 and SymPy 1.14.0 on macOS arm64). Requalify when the printer, NumPy or the supported expressions change. |
_FROZEN_ARRAY_BYTES |
204 bytes per stored initial-amplitude vector | src/nwqlib/algorithms/qhd/method.py |
The FrozenArray wrapper beyond its owning array's data and header: the object with its two slots, 48 bytes, and the separate read-only view header that it keeps, 112 bytes, measured with sys.getsizeof on 64-bit CPython 3.12.14 and NumPy 2.5.2, and its cached hash, a Python integer of at most 63 magnitude bits charged 44 bytes. Charged once for each distinct stored rank-one vector, one per variable, in _initial_state_bytes, which with the owning arrays gives the stored-population allowance 8 d K + 324 d + 64. Revalidated wrappers share the existing array storage. Requalify when FrozenArray's storage or the runtime changes. |
Grid coordinate validation within _initial_state_bytes |
16K array bytes for d = 1, 25K−1 for d ≥ 2 |
src/nwqlib/algorithms/qhd/grid.py::OneHotGrid.__post_init__, src/nwqlib/algorithms/qhd/method.py::QHD.plan |
int64-to-float64 conversion holds two K-entry arrays. From the second axis, the previous float64 coordinates and K−1-byte failure mask coexist with that conversion. This scratch fits the existing 320dK coefficient, and grid validation precedes initial-state evaluation. The existing 4096(d+1) term supplies the qualified fixed-object allowance. The array law is derived, while the object allowance is specific to CPython 3.12.14 and NumPy 2.5.2. Requalify when the coordinate producer, loop lifetimes or runtime changes. |
QHD_DECODING_PROBABILITY_CHUNK_SIZE |
4096 | src/nwqlib/algorithms/qhd/decoding.py and the kept-state stream in src/nwqlib/algorithms/qhd/refinement.py |
The decoder reuses up to 32 KiB of float64 probability storage plus its bounded masks and indices. Refinement uses 4096 as a ceiling, with an admitted actual chunk c and a separate incremental allowance 65536+1024d+(16d+128)c bytes. The 16d term permits old/new coordinate overlap and 128 prices indices, complex gathers, probability arrays and consumer masks/selections. Dense readout is preferred whenever its complete byte and work laws fit. A failed dense admission does not waive streaming admission. Joint mass uses one pass and paired split modes use two. Requalify the object constants when NumPy, the runtime, chunk lifetimes or the consumers change. The chunk size changes neither the probabilities nor the tie rule. |
WRAPPED_PHASE_CACHE_ENTRIES and wrapped_phase_workspace_bytes |
64 scalar entries per owner, 65536+40E bytes for a reached owner with largest wrap array E = 2**s |
src/nwqlib/subroutines/hamiltonian_evolution/pauli_evolution.py::wrapped_projector_phase, src/nwqlib/algorithms/qhd/method.py::_admit_table_phases, src/nwqlib/algorithms/qhd/resources.py::_admitted_onehot_population |
Fixed storage policy for separate potential-compiler and per-call rotation-census caches. Keys are (angle.hex(), s) and values are Python floats. Hits preserve the stored bits and do not refresh insertion order. A miss computes the original 2**s array-shaped wrap and evicts the oldest entry after insertion when the cap is exceeded. Call order determines reuse, and every call can miss. Miss work is proportional to 2**s, with current QHD diagonal providers using s = 4 through 7. Insertion permits 65 entries before oldest-entry eviction. One miss holds the float64 phase array and two complex128 arrays, giving the derived 40E numerical peak. The fixed term is a qualified cache/object allowance on CPython 3.12.14 and NumPy 2.5.2. Compilation releases the cache before identity serialization and shares the largest of the table, wrap and identity workspaces. A later census admits its source records, magnitude population, cache and one miss together with caller-held buffers. Revisit the cap when measured wrap cost warrants a different storage allowance. Requalify when cache capacity, key representation, provider policy, NumPy expression or object lifetimes change. |
| One-hot wrapped-phase census inventory | B_source + 65536 + (1024+4L(bit_length(max(1,C)))) U + B_wrap, plus caller-held bytes |
src/nwqlib/algorithms/qhd/resources.py::_onehot_wrap_census_bytes |
U bounds distinct magnitude keys and C bounds total rotation/T multiplicity from the stored provider metadata and structured preparation. L(b)=32+4 ceil(max(1,b)/30) is the existing CPython integer allowance. The qualified population rate covers Counters, temporary projector slots, sorting and result validation. Source tables and stored vectors use the existing metadata inventories, with step/block allowances and a register-width-dependent support-position extension. The cache and one array-shaped miss coexist with the census. The gate applies when a stored block uses the diagonal provider and can raise the minimum bytes of planning or inspection. These named-allocation allowances exclude process RSS and arbitrary symbolic or SDK internals. Requalify when the census populations, record representation or runtime changes. |
qhd/binary.py::_model_metadata and _model_reservation |
Model bytes 16A+sum_distinct_n 8*2**n+sum_t(E_t-1)+H_model |
src/nwqlib/algorithms/qhd/binary.py |
With R=d+T phase tables of E_t entries and n_t=log2(E_t), A=sum E_t, M total potential support positions and u distinct widths, model bytes are 16A+sum_distinct_n 8*2**n+sum_t(E_t-1)+H_model. H_model=65536+sum_t[4096+8L(4200+n_t)]+[16+L(bit_length(max(0,d-1)))]M+64d+256u+H_QFT, where L(v)=32+4ceil(max(1,v)/30). For b=log2(K), m the admitted AQFT cutoff (m=min(aqft_cutoff,b-1), with m=b-1 for None), C_p=m*b-m*(m+1)//2, w the optional swap count (w=b//2 for swaps and zero for relabel) and g=b+C_p+w, H_QFT=256+32g+56b+64w+192C_p+L(bit_length(max(0,b-1)))*(b+2C_p+2w). Latest constructions use sum(16E+h). The build reservation is the maximum of 80E+h and, for every Walsh-capable potential, [192+2L(2200+2n)]E+h for the simultaneous normalization-loss Fractions. Here h=4096+16s+8L(4200+n), with s the table's support size. Phase use reserves 64E_max. These are named-payload and qualified object allowances on CPython 3.12.14, NumPy 2.5.2, SymPy 1.14.0 and Pydantic 2.13.5. Optional cache entries share the remainder after caller residents and the complete baseline. Revisit condition: A change to QFT tuple construction, table/omission representation, synthesis scratch, cache lifetimes or the qualified runtime. |
qhd/method.py::_binary_source_reservation and QHD.plan |
Source bytes 8sum_potential E+H_tables+8dK+324d+64+768N+4096B+(32+L_d+12digits(max(0,d-1)))J |
src/nwqlib/algorithms/qhd/method.py |
Source tables, their metadata, stored initial vectors, schedule rows and compiled block records are reserved once. With B=N(d+fT), J=N(d+fM), f=1 or 2, H_tables=65536+6144T+(24+L_d)M+256d and L_d=L(bit_length(max(0,d-1))), source bytes are 8sum_potential E+H_tables+8dK+324d+64+768N+4096B+(32+L_d+12digits(max(0,d-1)))J. Planning adds 8dK+120d+40 and 16 bytes per grouped symbolic-term list position. It releases its model before prospective quantum inspection and admits B_plan+B_census+B_local. Classical binary ir_product planning also admits B_source + 48D + 16(sum_potential E_t + K + E_max) + B_local, the baseline used by theory.run_binary_product, where D=K**d, E_max=max(K,E_1,...,E_T) and B_local=model+latest+max(build,use) uses the selected cutoff, swap policy and potential-Walsh setting. Array identity encoding has the separate check B_plan+I. Symbolic expression payloads and process RSS keep their separate scope. Revisit condition: A change to source records, grouped-term containers, support serialization, compiler/model lifetimes or the row/block qualifications. |
qhd/native.py::_binary_native_bytes and construct_qhd |
2*(65536+4096C+512Q+2048G+32A_q+64V)+W_native bytes |
src/nwqlib/algorithms/qhd/native.py |
Both QHD._select_binary_native during planning and construct_qhd before allocation admit the same complete native byte phase, and the selected workspace records that total. The native graph and source-record allowances replace the 128-byte-per-block charge. Before creating a binary QHD circuit, admit its source records, model baseline and 2*(65536+4096C+512Q+2048G+32A_q+64V)+W_native, where C counts created circuit containers, Q their total local wire positions, G gate positions, A_q qubit arguments and V scalar parameters. The helper derives these counts from each stored block, including dense nested definitions and the selected initial preparation. W_native=65536+(320+L(n_max))*E_max+256*n_max*(1+digits(max(0,d*b-1))) prices one block's padding, wrapping, Gray-code and Walsh-list scratch. The rates are qualified engineering allowances for CPython 3.12.14 and Qiskit 2.5.2. They exclude later SDK synthesis, shared immutable SDK-singleton definitions and process RSS. Revisit condition: A changed builder, gate/definition copy behavior, parameter representation, SDK version or measured workload outside the checked range. |
QHD rotation law¶
These laws describe the selected synthesis model and are separate from measured compiler output, like the shared circuit-free synthesis laws. Revisit them when the construction or compiler changes, using independent matched-workload evidence.
The one-hot QHD rotation law (src/nwqlib/algorithms/qhd/resources.py::rotation_population) counts the rotations of Qiskit 2.5.2's definitions of the emitted gates. The fixed parts are those of synth_mcx_n_dirty_i15 inside MCPhaseGate: 7 T gates for two controls, 15 phases of ±pi/8 for three controls and 8 k - 2 T gates for k ≥ 4 controls (_mcx_rotations). The law and its test test_projector_rotations_match_the_qiskit_definitions in tests/test_qhd_resources.py compare these with the expanded definitions. Revisit when that test fails after a Qiskit upgrade. T_PER_PRECISION_BIT = 3 in the same module is the leading coefficient of the typical Ross–Selinger T count 3 log2(1/eps) of one Rz synthesized to operator error eps (arXiv:1403.2975v3), used with intercept zero because neither that result nor NWQEC 0.1.2 supports a finite intercept. The resulting T estimate is labeled estimate. On one small QHD circuit it was 392.9 against 348 T and T-inverse gates compiled by NWQEC 0.1.2 at the same budget (docs/algorithms/qhd.md, "Fault-tolerant resources", measured with Python 3.12.14 and Qiskit 2.5.2 on macOS arm64). Revisit with a qualified synthesis law or a calibration that has a consumer.
Subroutines¶
These constants belong to the shared subroutines: exact dense synthesis, QSP, product formulas, block encodings, state preparation, the projected eigensolver and the linear-algebra work laws.
Circuit-free synthesis laws¶
The following laws describe the selected synthesis model. They are separate from measured compiler output. Revisit them when the construction or compiler changes, using independent matched-workload evidence.
The classical cost of one exact synthesis _dense_synthesis.dense_unitary_circuit on m qubits, with M = 2m, is registered by _dense_synthesis.dense_synthesis_size as 113 M**3/4 + (m**2 + 16 m + 512) M**2 work units, 320 M**2 + 65536 working bytes and 352 M**2 + 16384 kept bytes. The units are those of the dense-completion law of DEFAULT_MAX_BLOCK_WORK, h3 for a product of h-square matrices and 8 h3 for an SVD or Schur factorization. A generic block-ZXZ step on an s-square block costs 113 (s/2)3 units and passes four half-size blocks on, so the recursion stays below (113/4) M3. Each of at most M2/16 two-qubit blocks is allowed 8192 units for its A.2 work, which stays below 7200 units also with the blocks that a failed batch of _dense_synthesis._apply_a2 discards, and the elementwise passes stay below (m2 + 16 m) M2. The kept law allows 256 bytes for each of at most 11 M**2/8 instructions. Kept circuits, rebuilt from their gate lists in a fresh process, held 124 to 173 bytes of resident memory per instruction for m = 3 to 9, and the kept law was 1.7 to 5.6 times that memory for m = 2 to 9 and 2.0 and 2.2 times at m = 8 and 9. The traced working peak stayed at or below 0.54 of its law for m = 1 to 9, and 0.45 for m = 6 to 9, with Haar-random, near-identity and block-diagonal inputs. A Haar-random input took 0.12 s at m = 7, 0.57 s at m = 8 and 2.6 s at m = 9 on an Apple-silicon Mac with one thread, 0.65 to 1.1 s per 1e9 units at m = 8 and 9. _dense_synthesis._A2_BATCH = 32 is the largest number of consecutive two-qubit blocks whose Weyl frames the A.2 loop computes in one stacked call. It keeps the arrays of one batch near 100 KB, and without a cap the synthesis at m = 9 took the same time. qiskit_compat.dense_synthesis_widths lists the syntheses that a rewrite would make, without making them. QLS charges the forward and the adjoint synthesis of a controlled dense query, which a Run builds once each, to max_work and max_bytes at planning (src/nwqlib/algorithms/qls/host_planning.py::_admit_query_synthesis). LCHS charges one synthesis per physical dense_exact branch with address bits and at least two system qubits, and the syntheses of a dense-dilation QSP child (src/nwqlib/algorithms/lchs/compiled_selection.py::compiled_select_dense_syntheses) to max_select_work and max_bytes. The QSP evolution and joint-generator builders check their own max_work and max_bytes, the MPS builder checks its at most layers * (n - 1) two-qubit syntheses against its max_svd_work and max_bytes, and lowering inside a Run and a backend that lowers a circuit to a gate basis reserve theirs against the Run's ExecutionLimits.max_synthesis_work. lower_qiskit called outside a Run checks them against its own max_synthesis_work. Revisit when _qsd or _apply_a2 changes its steps or when Qiskit changes how a circuit stores standard gates.
_dense_synthesis.AUTO_WHOLE_MATRIX_MAX_CONTROLS = 1, called K, is the largest number of controls for which dense_control_route="auto" synthesizes the whole controlled matrix (controlled_unitary_circuit) instead of controlling each gate of the synthesized unitary (Exact dense synthesis). For an M-square unitary with k controls the whole-matrix synthesis costs about 11 (2**k M)**3 work units by controlled_synthesis_size, about 88 M3 with one control and 704 M3 with two, against about 28 M3 for dense_synthesis_size alone, and its bytes grow as (2**k M)**2, four times per added control. For Haar-random unitaries on 2 to 5 qubits (Qiskit 2.5.2) it took 3.2 to 4.0 times fewer CX than the gate-wise route with one control, 3.6 to 3.8 times fewer with two and 1.9 to 2.1 times fewer with three. With four controls the gate-wise route took fewer CX, 15.5 to 19.0 percent fewer with the random_unitary seeds 10 n + 4 for n = 2 to 5 qubits and 15.5 to 23.9 percent over the seeds 10 n + 4, 10 n + 104 and 10 n + 204, counted after transpile at optimization level 0 with its default qubits_initially_zero=True. With one or two controls it was also faster to build and lower to CX and U gates on 3 to 5 qubits. K is chosen on the admitted classical cost. With one control the whole-matrix work is of the same order as the synthesis alone for a CX saving of 3 to 4 times. With two controls the saving stays near 3.7 times, but the work is about 704 M3, so a controlled QLS query of a dense dilation of 64 padded coordinates would need about 3.2e9 units, beyond the default max_work, where the gate-wise route needs 2.4e8. With two controls the laws also give the whole-matrix synthesis 256 (4 M)**2 working and 176 (4 M)**2 kept bytes, against 320 M**2 and 352 M**2 for the synthesis alone, about 12.8 and 8 times as many for large M (12.65 and 7.98 times at M = 128). Each further control multiplies the whole-matrix work by 8 and its bytes by 4, and the saving falls to about 2 times with three. "whole_matrix" is available for any control count, and the construction limits charge its cost (controlled_synthesis_size). Revisit when controlled_unitary_circuit changes its cost or CX count, when Qiskit's control of a gate changes its counts (CONTROLLED_U_INSTRUCTIONS), or when a workload with two controls needs the CX saving.
On the gate-wise route, when a construction controls a composite gate that holds exact syntheses, Qiskit's add_control unrolls the composite and controls every synthesized gate, once for each place where the synthesis occurs (qiskit_compat.dense_control_counts). This control step has its own law, the gate-wise control law of _dense_synthesis.gatewise_control_counts and gatewise_control_size. For one synthesis on m qubits with k controls, gatewise_control_counts bounds the gates that Qiskit unrolls (the census of dense_synthesis_gate_census plus the global phase), the instructions that it emits and the heavy instructions among them, which hold angles or are Python objects. CONTROLLED_U_INSTRUCTIONS and CONTROLLED_RZ_INSTRUCTIONS store the number of instructions that Qiskit 2.5.2 emits for a U and an RZ gate with 1 to 64 controls, measured because its dirty-ancilla X synthesis has no published closed-form size. A CX and the global phase emit one instruction and an H seven. gatewise_control_size charges 2048 gates + 16 instructions work units, 1024 gates + 96 instructions + 1024 heavy + 65536 working bytes and 96 instructions + 1024 heavy + 16384 kept bytes. The control step does no floating-point linear algebra, so its units were calibrated by time. Each unrolled gate costs one Python call, 6 to 16 us with one control, and a U gate took up to 320 us with up to eight controls (Apple-silicon Mac, one thread, Qiskit 2.5.2). For Haar-random syntheses on 4 to 7 qubits with one to eight controls, Qiskit's control call took 4.4 to 35 ns per unit, and the synthesis took 1.6 to 15 ns per unit of its own law. Both laws feed the same limits, which count work units, not time. Per unit, the control call took 0.45 to 3.3 times as long as the synthesis on 4 qubits and 2.6 to 21 times as long on 7 qubits. With these units the default QLS max_work still admits a controlled dense dilation of 64 padded coordinates. An LCHS dense_exact SELECT of 16 branches on 5 system qubits needs 7.1e7 units for its control steps, 2.5e7 for its syntheses and 8.2e6 for its branch exponentials in the test problem _dense_select_problem of tests/test_lchs_resource_structural_law.py, 1.04e8 in total, so the largest power-of-two branch count that the default max_select_work of 1e8 admits on 5 system qubits there is 8. In circuits of one instruction kind a parameterless standard gate held 40 to 59 bytes of resident memory, a gate with angles 123 to 177, an MCX gate 355 and an MCPhase gate 733. For 4 to 7 qubits and one to eight controls, the kept laws of the synthesis and of this step together were 1.7 to 7.7 times the memory that kept controlled gates held, and the sum of the working and kept bytes of this step was 2.1 to 11 times the peak increase of resident memory during the call, wherever that increase exceeded 0.1 MB. test_gatewise_control_tables_match_installed_qiskit and test_gatewise_heavy_counts_match_installed_qiskit in tests/test_synthesis_admission.py recompute the stored counts. Revisit when those tests fail after a Qiskit upgrade, when qiskit_compat.controlled changes its route, or when a workload needs a control step that these units charge far below its time.
MCX_CX_BY_CONTROLS in src/nwqlib/subroutines/_mcx_counts.py gives the CX count of an X gate with 1 <= k <= 64 valued controls. Entry k is the CX count of NWQLib's controlled(XGate(), k) after Qiskit 2.5.2 lowers it without ancillas to basis cx,u at optimization level 0, which is the logical circuit before routing. Open and closed control states give the same count. The counts are 6, 14, 36 and 84 CX at k = 2, 3, 4 and 5, 464 at k = 10 and 3010 at k = 25, so no per-control constant describes them. For k up to 5 they coincide with Qiskit's synth_mcx_noaux_v24 and for the larger k compared with synth_mcx_noaux_hp24. Neither synthesis has a published closed-form CX count, so the table stores the values. The module imports nothing, so the circuit-free laws of QLS and LCHS read the table without importing Qiskit. Every count read from it is recorded as an estimate, because another Qiskit version can synthesize the gate differently. A gate with more than 64 controls has no CX law. In QLS, _primitive_cx_law in src/nwqlib/algorithms/qls/quantum.py prices each X gate with 2 <= k <= 64 controls in the Program, such as the projector flips and the Q_b' and A_t predicates. A singly controlled X is exactly one CX. Other controlled primitives and controlled queries have no CX law, so only the Hermitian inverse route can reach a concrete CX total, and there the table enters through the projector flips of an encoding with at least two block ancillas. The shortcut circuits and the dilation of a non-Hermitian inverse report an unknown total. The LCHS SELECT laws read the table for the multi-controlled X gates that Qiskit adds when it controls a branch, and for the QSP reflections (native._controlled_gate_cx, compiled_selection.compiled_select_cx_projection). test_multi_controlled_x_cx_table_matches_installed_qiskit_synthesis in tests/test_mcx_counts.py recomputes every entry with the installed Qiskit. Revisit when that test fails after a Qiskit upgrade, or when the gates move to an ancilla-assisted construction with a closed-form count.
The selected MPS outer-PREP slot bound is 3 * layers * max(0, n - 1) CX. In mps_to_circuit, every layer truncates to rank at most two and the last site has right bond one, so at most n-1 local two-qubit unitaries are emitted. _dense_synthesis.dense_unitary_circuit synthesizes each exactly to rounding with at most three CX. Initial preparation and LCU preparation have separate provider/layer inputs, and the latter has multiplicity two. This circuit-free bound describes the selected construction before hardware routing, and early rank-one termination can reduce its realized cost.
Dense synthesis rounding window¶
ROUNDING_WINDOW = 2**-46 in src/nwqlib/subroutines/_dense_synthesis.py is the largest deviation that the exact dense synthesis of controlled gates, of backend lowering and of MPS preparation treats as zero or as an exact special value: a Weyl coordinate from a multiple of pi/4, a one-qubit rotation from the identity, an off-diagonal block entry from zero, or an entry of a block's deviation from a multiple of the identity from zero. A replaced coordinate changes the operator by at most this amount in operator norm, an omitted one-qubit gate by at most 1.5 times this amount, and a dropped block entry changes one entry by at most this amount. The value is 128 unit roundoffs (u = 2**-53). For exactly special inputs these quantities come out of the factorizations at about 1e-16, so the window recognizes them, and a deliberate deviation such as a Weyl coordinate of 1e-12 is far outside it and keeps its gates. A larger window would drop real rotations. A smaller one would only cost CX for special inputs, not accuracy. For U = expm(-i t G) with random Hermitian G, t from 1e-10 to 10 and one to five qubits, the largest entry error measured was 14 u times the dimension, and tests/test_dense_synthesis.py checks 64 u times the dimension. Revisit when the synthesis changes or a workload needs special inputs recognized at a coarser precision.
QSP phase solver and evolution synthesis¶
Each definition includes a rationale comment.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
QSP_SOLVER_GRID_MULTIPLIER |
4 | src/nwqlib/subroutines/qsp/phases.py |
Overdetermined objective grid removes fit-the-grid spurious minima seen at the exactly-determined 1x grid. Revisit if solver cost binds. |
QSP_SOLVER_LBFGS_OPTIONS |
maxiter 2000, ftol 1e-30, gtol 1e-18 | src/nwqlib/subroutines/qsp/phases.py |
The objective is a mean square residual, so a 1e-12 residual needs objective ~1e-24, below scipy's default stopping thresholds. Revisit when the residual tolerance or the objective changes. |
QSP_SOLVER_POLISH_MAX_STEPS / ..._MAX_HALVINGS |
30 / 15 | src/nwqlib/subroutines/qsp/phases.py |
Damped Gauss-Newton iteration on the objective-node residual, which is Newton's method on the Chebyshev coefficients (Dong, Lin, Ni and Wang, arXiv:2307.12468v1, Eq. (3.1)). It polishes each L-BFGS point. When L-BFGS from the start of Dong, Meng, Whaley and Lin, arXiv:2002.11649v2, leaves the residual at or above the tolerance, or nonfinite, it runs from that start. It rejects non-improving steps, so it can only improve its starting point. On the default grid of the sweep below, the Newton start ran on 347 parity solves and converged on each within 10 residual evaluations, including the evaluation at the starting point. Revisit if either use stops above the residual tolerance on a target that passes the max|f| <= 1 check. |
QSP_SOLVER_RESIDUAL_TOLERANCE |
1e-12 | src/nwqlib/subroutines/qsp/phases.py |
Numerical acceptance threshold for the QSP verification-grid residual. It is a solver control, not a total Hamiltonian-evolution error bound. Revisit for a different floating-point precision. |
QSP_TARGET_BOUNDARY_ULPS |
256 | src/nwqlib/subroutines/qsp/phases.py |
src/nwqlib/subroutines/qsp/phases.py solver-domain admission band above 1 for the binary64 value B of the rigorous target supremum bound. A target whose B exceeds 1 + 256*eps*max(1, B) is rejected before optimization. Constant-time after norming. Revisit if the norming implementation or supported floating-point precision changes. |
QSP_EVOLUTION_TARGET_MARGINS (evolution.py) |
(1e-3, 3e-3) | src/nwqlib/subroutines/qsp/evolution.py |
The truncated cosine reaches |f| = 1 at x = 0, and the sine does for tau >= pi/2. Dong, Meng, Whaley and Lin (arXiv:2002.11649v2, Sec. IV.5 and Fig. 13) report that the Hessian condition number at the optimum grows like eta^-gamma with gamma > 1 as max|f| = 1 - eta approaches 1. Dividing by s = max(1, cos bound, sin bound)*(1 + margin) keeps each target's max|f| at most 1/(1 + margin). With the Newton start, every preparation in the sweep below converges at 1e-3, and 3e-3 is a fallback it never uses. The margin costs only the recorded quadratic OAA amplitude deficit ~6(m/2)^2. |
Chebyshev norming grid (phases.py) |
64(d+1) Chebyshev-zero nodes on [-1,1] |
src/nwqlib/subroutines/qsp/phases.py |
The shared admission, scale and residual code uses Ehlich-Zeller Satz 2 (doi:10.1007/BF01111276), Eqs. (12)–(14), with factor sec(d pi/(2N)). At this density its inflation is at most sec(pi/128)-1 ~= 3.01e-4, below one third of the 1e-3 target margin. The implementation evaluates the exact-arithmetic inequality in binary64. |
QSP_EVOLUTION_BESSEL_TAIL_TERMS (evolution.py) |
200 | src/nwqlib/subroutines/qsp/evolution.py |
Adaptive-terminal search horizon beyond min(ceil(tau), max_degree). The first terminal whose finite suffix plus analytic infinite remainder admits a degree is used, and the finite search never replaces the analytic infinite remainder. |
QSP preparation max_degree / max_evaluations / max_bytes defaults (evolution.py::prepare_qsp_evolution) |
256 / 20000 / 10 GB | src/nwqlib/subroutines/qsp/evolution.py |
Explicit preprocessing controls. Evaluations count actual objective/Jacobian and verification calls cumulatively across both parity solves, all starts and attempted target margins. Bytes bound known simultaneous arrays rather than process RSS or vendor workspace. |
qsp/phases.py::_damped_newton |
stop 1e-14 |
src/nwqlib/subroutines/qsp/phases.py |
Untuned maximum objective-node residual at which the damped Gauss-Newton iteration stops, below the default 1e-12 acceptance tolerance (QSP_SOLVER_RESIDUAL_TOLERANCE). The iteration polishes each L-BFGS point, and it also runs from the start point of Dong, Meng, Whaley and Lin, arXiv:2002.11649v2, as the Newton start. In both uses the stop ends the iteration, and neither use replaces the final residual check on the verification grid. Revisit condition: Different precision, or measured convergence of the polish or the Newton start. |
qsp/shortcut.py::_KR_CONSTRUCTION_GRID_POINTS |
2001 |
src/nwqlib/subroutines/qsp/shortcut.py |
Untuned least-squares sample floor. The population is max(2001,8*(d+1)) for degree d, providing an overdetermined fit. Separate norming evidence evaluates the resulting polynomial. Revisit condition: A changed polynomial family, degree range or fit workload. |
python docs/scripts/qsp_phase_solver_sweep.py --evolution runs prepare_qsp_evolution with its default controls at each point of a tau x epsilon grid, and each preparation solves both parity targets. The table compares the solver with and without the Newton start on the default grid tau = 0.05, 0.10, ..., 10 (Python 3.12.14, NumPy 2.5.2 and SciPy 1.18.1 on macOS arm64). With the Newton start, every preparation converges at margin 1e-3, and so do the 74 preparations with tau from 10 to 100 in steps of 2.5 at epsilon 1e-3 and 1e-9. The third column counts the parity solves in which L-BFGS left the objective-node residual at or above the tolerance, so that the Newton start supplied the phases. The last two columns replace the _damped_newton(residual_and_jacobian, start) call in solve_symmetric_qsp_phases by the L-BFGS result, as mutation probe qsp_newton_start_skipped does. Without the Newton start, 340 of the 1000 preparations fail at both margins, and 5 succeed only at the 3e-3 margin.
epsilon |
Preparations | Parity solves that took the Newton start | Raised without the Newton start | Needed the 3e-3 margin without the Newton start |
|---|---|---|---|---|
| 1e-2 | 200 | 41 | 41 | 0 |
| 1e-3 | 200 | 51 | 50 | 1 |
| 1e-4 | 200 | 58 | 58 | 0 |
| 1e-6 | 200 | 82 | 79 | 3 |
| 1e-9 | 200 | 115 | 112 | 1 |
Trotter error bounds and Pauli census¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| Trotter-bound imaginary coefficient window | 1e-12 absolute | src/nwqlib/subroutines/trotterization/error_budget.py::_validated_terms |
Same window as the native Pauli rotation path in src/nwqlib/subroutines/hamiltonian_evolution/pauli_evolution.py. It accepts the residues near 1e-17j that library constructions leave on exactly Hermitian operators and rejects genuinely complex coefficients. It is absolute, not relative to the coefficient scale. Revisit for operators whose coefficients are far from unit scale. |
Pauli census scratch allowance H0 and contraction block |
H0 = 65536 bytes. Block: the largest member of {max(1, E, F)} and the powers of two below it whose byte envelope fits. The public helpers replace a chosen block above 65536, the default block of coefficient_up, by 65536 when the byte envelope at 65536 fits, so their coefficient equals the one computed with the default block |
src/nwqlib/subroutines/trotterization/error_budget.py::census_bytes, choose_census_block, _pauli_bound_coefficient_from_terms; admitted by src/nwqlib/algorithms/qpe/powers.py::_census, src/nwqlib/subroutines/trotterization/error_budget.py::_bound_coefficient_evaluation and, for one structure shared by all distinct LCHS k-nodes, by src/nwqlib/algorithms/lchs/time_independent_terms.py::_lchs_census_choice |
H0 is an explicit engineering allowance for scalar bookkeeping, ndarray headers, iterators and fixed sort stacks on the checked 64-bit CPython/NumPy stack. It is not a universal interpreter or process-RSS theorem, and variable populations of Python objects (chunk objects, partial sums) are charged separately in the same law. The block is a scratch-size choice, not a tolerance. The outward coefficient bound holds for every positive block, and a smaller block does not necessarily lower total memory because each block keeps one partial sum. LCHS borrows the shared census tables for a node keeping all labels. Otherwise it selects each pair/triple table once and remaps it in place with np.take(..., mode="clip"). Boolean row selection includes one eight-byte integer index per selected row. On the checked 64-bit stack, the complete census envelope adds 9*max(E,F) bytes to the shared law, using that law's existing double-table reserve. Node arrays are released between contractions. choose_census_block prefers block 65,536 when the full recomputed envelope fits. The selected block also determines the grouping of outward coefficient reductions. Revisit when the census kernels, NumPy or CPython change their object or temporary-array populations. |
| Pauli census admission of the public Trotter helpers | max_work = 1,000,000,000 (DEFAULT_CENSUS_MAX_WORK) and max_bytes = DEFAULT_MAX_BYTES. Label/term conversion: 128*p*(q+1)+65536 bytes held during the census and 16*p*(q+1) logical preparation work, plus p*q label visits for packing |
src/nwqlib/subroutines/trotterization/error_budget.py::_bound_coefficient_evaluation, used by trotter_bound_coefficient, evaluate_trotter_bound, trotter_error_bound and select_trotter_step_count |
Engineering defaults, not mathematical ceilings. The work default matches the QPE planning-work default, and the byte default is the shared one. The conversion envelope covers the simultaneous rendered-label list, source tuples and complex scalars, and validated tuples and real scalars under the native-object accounting convention of LCHS. The pair stage is admitted with E = P = p(p-1)/2 and no triples before the labels are converted. The second stage uses the actual E and reserves F = N for the full second-order expression before any triple test. Order one and the suffix relaxation charge no nested tests or triples. The complete work expression includes the pair work once. The pair stage is admitted before label conversion. The requested expression is admitted again using the actual pair count before nested tests or contraction. A refusal identifies the failed stage and reports a sufficient complete envelope for the requested order and variant. Its block-one byte value is a checked candidate, rather than a minimum over all block sizes. A byte-fit failure is reported only when no checked candidate fits the current byte allowance. With P = p(p-1)/2, J = p(p-1)(2p-1)/6, w = ceil(q/64) and L0 = 16p(q+1)+pq, the complete sufficient work is L0+(w+1)P at order one, L0+wP+p+2P for the relaxed prefix and L0+w(P+J)+P+J for the order-two exact census. Revisit when the label conversion or the census kernels change their object populations. |
| Trotter dense-comparison tolerance | 1e-12 relative | tests/test_trotterization.py::_DENSE_TAIL_SUM_REL_TOL |
The anchor and seeded tests compare the Pauli-triangle coefficient W_up, an outward upper bound on the exact rational W for either bound variant, with the dense commutator sums of the same Hamiltonian. For these inputs the dense sums' magnitude bound B is below 5 W and their evaluation errs by at most gamma_256 B, so a relative allowance below 1.43e-13 covers the roundoff of both sides and of the comparison, derived in the comment at the constant. The allowance stays eleven orders of magnitude below the 13.64 % to 50 % decrease of W that a lost factor causes on the tight two-qubit anchor. Revisit when the fixtures, their generation or the dense helpers change. |
PAULI_COEFFICIENT_RTOL |
1e-12 | src/nwqlib/subroutines/pauli_decomposition.py |
Scale-covariant pruning of matrix-decomposition roundoff. The removed coefficient mass is propagated and deducted from product-formula error budgets. |
hamiltonian_evolution/pauli_evolution.py::_matched_xx_yy |
imaginary window 1e-12 |
src/nwqlib/subroutines/hamiltonian_evolution/pauli_evolution.py |
Absolute window, in the coefficient's units, below which the imaginary part of an XX or YY coefficient is discarded before the pair uses one XXPlusYYGate. QHD kinetic coefficients are real. A larger imaginary part sends the block to the general Pauli-evolution path. Revisit condition: A relative window or another precision. |
Block encoding, LCU and multiplexors¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| dilation unitarity guard | 1e-10 | src/nwqlib/subroutines/block_encoding/core.py |
float64 SVD roundoff ~1e-14 against O(1) construction defects, four orders of separation. |
| Dense-dilation normalization relative slack | 1e-12 times alpha | src/nwqlib/subroutines/block_encoding/core.py::_DENSE_ALPHA_RELATIVE_SLACK |
Compare a planned alpha with a computed value that is at most the spectral norm in exact arithmetic, the largest entry magnitude or the largest singular value of the already-required SVD, at the operator's scale. Planning, the explicit dense-dilation plan and the constructor use this one rule (_require_alpha_covers). LAPACK's SVD is backward stable, so a computed norm, and an alpha taken from it, can fall a few units in the last place below an entry magnitude or the exact norm. 1e-12 is about 9000 unit roundoffs, which covers that error at the dimensions a dense dilation serves, while a larger deficit is refused with both raw values in the message. Revisit for a different precision or SVD error model. This is a numerical admission tolerance, not an additional reference decomposition. |
banded structure-detection atol |
1e-12 x max(1, max|A|) | src/nwqlib/subroutines/block_encoding/banded.py |
Scaled machine-precision equality test for circulant structure. |
block_encoding/core.py |
64*eps*max(1,n,terms,bands) |
src/nwqlib/subroutines/block_encoding/core.py |
Untuned relative candidate-comparison window for circulant classification. The exact rational residual is separately computed, and automatic routing requires zero residual. The window alone is not a certificate. Revisit condition: Changed accumulation or classification algorithm. |
block_encoding/core.py::_admit_pauli_plan and _circulant_classification_work, rational circulant classification of m kept terms on n qubits with b = min(m, 2**n) candidate bands |
Work C(n,m,b), bytes B_held + 64P(n+16) + [128 L(2J)+512]V + H0 |
src/nwqlib/subroutines/block_encoding/core.py |
Work C(n,m,b) = (n+32)m + 4(n+1)mb + 16b + 2b ceil(log2 max(1,b)) + 32 logical visits (label parsing, band accumulation, the mb term-band trace calls with 4n carry transitions and four rational operations each, band records and the offset sort), added to the 32*P*max(1,n)**2 table work. Bytes B_held + 64P(n+16) + [128 L(2J)+512]V + H0 with J = 4208 + 2n + 2 ceil(log2(b+1)) + ceil(log2(m+b+1)) and V = m+b+n+1. Every finite binary64 band component is a multiple of 2-1074 below 21024, and the squared masses over m terms and b bands fit a numerator width of 4196 + 2n + 2 ceil(log2(b+1)) + ceil(log2(m+b+1)) plus a small fixed allowance, and J adds twelve bits. The factor 128 bounds the simultaneous integer population per V, L(2J) each integer's product width and object, and 512V the parsed tuples, Fraction wrappers, dictionary entries and complex values. There is no mb rational table. B_held is the caller's live payloads (the decomposition's labels and coefficients, a dense input at 16D²), and classify=False omits the classifier terms. The caller charges each live decomposition once as m(n+16) logical label and complex-coefficient bytes, together with each distinct live dense matrix at 16D² and any completed child payloads that coexist. SELECT construction adds 64P(n+16). An active circulant classifier additionally requires [128L(2J)+512]V+H0, including its fixed bookkeeping allowance. With classify=False, or with an empty skipped classifier, that entire classifier allowance is zero. These charges follow the stated logical payload conventions and do not bound process RSS. These units are admission proxies, not timings or equal-cost CPU operations. The LCHS outer gate src/nwqlib/algorithms/lchs/compiled_selection.py::_compiled_qsp_part_plans checks max{B_dec, 96d² + sum_c m_c(n+16) + sum_c 64P_c(n+16) + max_c X_c} bytes at the full envelope of d² terms per child before decomposition (_qsp_outer_bytes, with X the bracketed classifier allowance above), then passes each nonempty child its live payloads as B_held and one remaining work budget across the two decompositions, 32d² pruning, both classifications and tables, dense reconstructions and norms. That outer estimate is temporary until the LCHS compiled-selection gate adopts this law. A checked one-qubit circulant I*DBL_MAX + X*min_subnormal produces a 4196-bit rational numerator, so the 2200-bit estimate is not a bound. Revisit condition: A changed rational classifier or carry recurrence, or measurements that qualify the allowances. |
DEFAULT_MAX_BLOCK_WORK |
1,000,000,000 | src/nwqlib/subroutines/block_encoding/core.py |
Local work ceiling for dense conversion, Pauli classification, SELECT table construction and dense completion. The QSP evolution and joint-generator builders of src/nwqlib/subroutines/qsp/evolution.py take the same value as their default max_work for the exact synthesis of the dense unitaries that controlling their passes or branches makes (_dense_synthesis.dense_synthesis_size). Callers can lower or raise it. Dense completion counts 8D³ for the SVD, D³ for the complement product and 8D³ for the unitarity product of the 2D-square dilation, or 9D³ when the planning SVD is reused. Planning a dense dilation without a normalization charges src/nwqlib/_linalg_laws.py::singular_values_work for the spectral norm np.linalg.norm(A, 2), or 8D³ when the caller computes the norm with a full SVD that the completion reuses. The completion byte law 384D²+32D counts the main completion arrays, listed in _admit_dense_input, as if they were alive together. At most 304D² of them are alive at once. The QLS padded completion uses the difference for its extra singular frames. Planning without completion uses an allowance of 192D²+32D bytes, more than twice the 82D² its largest counted step holds. SDK and LAPACK workspace is not counted. Revisit when the completion's arrays change or when an allowance rejects an input that the counted arrays would admit. |
DEFAULT_MAX_LCU_WORK |
1,000,000,000 | src/nwqlib/subroutines/lcu/core.py |
Local coefficient, PREP and dense SELECT size estimate, checked by _admit_lcu before any matrix conversion or synthesis. N supplied matrices, P padded addresses, D system dimension, M=PD and m=log2(M) contribute 16P(log2(P)+1)+ND³+N(11M³+(m²+5m+256)M²) work units and 128P+64ND²+256M²+65536+N(176M²+16384) bytes. Implicit padding stores no matrices, and the synthesis terms vanish for a single term, which is not controlled. Each controlled branch is synthesized exactly from its M-square controlled matrix (_dense_synthesis.controlled_unitary_circuit), and _dense_synthesis.controlled_synthesis_size derives the terms from its recursion. The work counts the demultiplexing and block-ZXZ factorizations (8h³ for an SVD or Schur factorization of an h-square matrix and h³ for a product, as in the dense-completion law of DEFAULT_MAX_BLOCK_WORK), 4096 units for each of at most M²/16 two-qubit blocks, and fewer than (m²+5m)M² elementwise units. The 256M²+65536 bytes hold the controlled matrix, the factorization arrays and the gate angles of one branch at a time, and each kept branch is allowed 256 bytes for each of at most (11/16)M² instructions plus 16384 bytes of objects. Kept branches, rebuilt from their gate lists in a fresh process, held 115 to 165 bytes of resident memory per instruction for m=3 to 9, and the kept law was 1.7 to 7.0 times that memory for m=2 to 9 and 1.95 and 2.03 times at m=8 and 9. The traced working peak stayed at or below 0.48 of its law for m=2 to 9. One branch took 0.003 s at M=32, 0.011 s at M=64, 0.044 s at M=128 and 0.21 s at M=256 on an Apple-silicon Mac with one thread. The default is a fuse against runaway planning work, sized so that no example or test workload in the repository reaches it. Revisit when the dense SELECT or the exact synthesis changes its construction, or when Qiskit changes how a circuit stores standard gates. |
_multiplexors.py::_affine_angle_tolerance |
gamma_(2k+1) |
src/nwqlib/subroutines/_multiplexors.py |
The window is gamma_m*S with gamma_m=m*eps/(1-m*eps), m=2k+1 for k controls, eps=2**-52=2u and S=abs(offset)+sum_q abs(c_q), which bounds every entry of the affine law. affine_angle_table_values forms an entry as the offset plus at most k coefficients, which takes at most k rounded additions and no products, so the reconstruction differs from the exact law by at most k*u*S/(1-k*u). The window is more than four times that bound, and the remainder absorbs the rounding of the supplied table. Its only producer, the LCHS product-formula SELECT, computes each entry as mu*t*(h+k_j*l)/r for occurrence multiplier mu, elapsed time t, r steps, Pauli coefficients h and l and quadrature node k_j. _attach_affine_pauli_structure passes the affine law with offset (mu*t/r)*h and coefficients (mu*t/r)*l*w_q, where w_q are the power-of-two bit weights of the node grid, and both computations share the rounded mu*t. With at most two controls every partial sum that the reconstruction forms is at most abs(k_j*l)*mu*t/r in magnitude, and counting the remaining roundings to first order in u bounds the disagreement by (6*abs(h)+8*abs(k_j*l))*u*mu*t/r. With one control (k=1) the node products and one-term sums are exact, and the bound is (6*abs(h)+6*abs(k_j*l))*u*mu*t/r. With three or more controls the reconstruction adds the negative sign-bit coefficient last, so a partial sum can exceed abs(k_j*l)*mu*t/r by a factor of up to 2**(k-1)-1. The count is then taken against S, which bounds every partial sum, and the producer's four roundings, the law's two and the reconstruction's k give at most (k+6)*u*S to first order. The window, about (4k+2)*u*S, covers these bounds for every k. It equals the one-control bound only at the node k_j = w_0, when h and k_j*l have the same sign, the producer's three roundings (sum, product and division) are extreme with one sign and the four roundings of the law and its reconstruction are extreme with the other. A larger residual keeps the generic multiplexor under automatic SELECT and raises for an explicit structured request. Revisit condition: Changed arithmetic or supported precision, another table producer, or a rejected table built from exactly affine data. |
State preparation¶
src/nwqlib/subroutines/state_preparation/mps.py::DEFAULT_MAX_SVD_WORK is 100,000,000. Before the first TT-SVD step it bounds the sum of rows times columns times the smaller matrix dimension over all prospective economy SVDs. Prior rank caps affect later matrices, while current truncation does not avoid current full factors. This initial explicit dense-analysis ceiling is adjustable for an intended larger workload and does not measure FLOPs, wall time or RSS. src/nwqlib/subroutines/state_preparation/mps_circuit.py::build_mps_circuit_state_preparation also compares the work of its at most layers * (n - 1) exact two-qubit syntheses (_dense_synthesis.dense_synthesis_size) with max_svd_work, separately from the TT-SVD, and their bytes with max_bytes, before the first synthesis. The layered circuit construction is admitted against the same two limits before scikit_tt starts, by src/nwqlib/subroutines/state_preparation/mps.py::layered_construction_size, which the direct builder and LCHS planning share (layered MPS construction law). The law follows the scikit_tt sweeps, products and mpo.dot calls of src/nwqlib/subroutines/state_preparation/mps_circuit.py::mps_to_circuit, with each TT rank bounded by min(2**min(j, n-j), 4**l b_j) for bond dimensions b_j at layer l. Its untuned 32768 + 4096 n bytes cover the Python objects of the TTs and gate MPOs. Over 180 traced constructions (random, structured, LCHS weight and pulse targets, 3 to 14 qubits, 1 to 4 layers), the law stayed at 1.33 to 230 times the traced SVD and product work and at 1.11 to 3.8 times the traced peak bytes. The large work ratios occur where threshold truncation keeps the actual ranks of a low-rank state far below the bound. At the default limit it refuses random complex states of 12 qubits with 4 layers and of 13 or 14 qubits with 2 layers, and it admits the loader cells and the Gauss weights of the LCHS intro. The existing 10 GB input-byte default separately bounds known simultaneously live arrays. Explicit reference fidelity is computed as |<t,r>|²/(<t,t><r,r>) by _numerics.normalized_fidelity_with_window, so it does not depend on how accurately the reference and the raw tensor were normalized. Its window, ((1+g)/(1-g))²(1+u)/(1-u)-1 with g=gamma_(4D+4), is derived in that function from the evaluation's roundings for dimension D. The raw value is kept, and truncation or SVD error is not treated as roundoff.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| Controlled direct PREP slot bound | 4*(D-1) + 12*max(0,D-2), D=2**n |
src/nwqlib/_preparation_laws.py |
Two trees with D-1 rotations and D-2 CX each. An extra control costs at most 2 CX per rotation and 6 per CX, the textbook Toffoli circuit reproduced by Shende and Markov, arXiv:0803.2316v1, Fig. 1. Global phase adds no CX. The contiguous-uniform fast path has an elementary bound of 41*n-59 for n>=2, inside the complex-tree envelope. Revisit when direct preparation or elementary control synthesis changes. |
state_preparation/mps_circuit.py |
canonical-column tolerance 1e-9 and inter-layer ortho threshold 1e-12 |
src/nwqlib/subroutines/state_preparation/mps_circuit.py |
Untuned local column-admission and tensor-compression tolerances. The local unitary completion uses at most two-qubit blocks. Tensor orthogonalization can discard singular values, so these tolerances do not give a circuit fidelity certificate. Revisit condition: Precision, tensor rank or requested preparation accuracy. |
TT-SVD threshold=1e-14 in state_preparation/mps.py::decompose_state_to_mps, and threshold=1e-14, num_layers=2 in state_preparation/mps_circuit.py::build_mps_circuit_state_preparation and the lcu/core.py PREP builders |
threshold=1e-14, num_layers=2 |
src/nwqlib/subroutines/state_preparation/mps.py, src/nwqlib/subroutines/state_preparation/mps_circuit.py, src/nwqlib/subroutines/lcu/core.py |
Each singular value at or below the threshold is dropped from the TT-SVD of the unit-norm input, keeping at least one per bond. The discarded squared weight is reported. Two disentangling layers are an untuned circuit depth whose circuit error is not evaluated at construction. The lower-level mps_to_circuit defaults to one layer. Revisit condition: Requested preparation accuracy, precision or tensor rank. |
state_preparation/mps.py |
max_products=1000000000 |
src/nwqlib/subroutines/state_preparation/mps.py |
Explicit contraction-work ceiling. Each contraction charges output entries times the contracted rank before tensordot. The default is an untuned resource allowance, independent of quantum accuracy. Revisit condition: A deliberately larger requested materialization. |
Layered MPS construction law¶
src/nwqlib/subroutines/state_preparation/mps.py::layered_construction_size returns upper bounds (work, bytes) of building the layered circuit from MPS cores. The law follows the scikit_tt calls of mps_circuit.py::mps_to_circuit, and of the conversion of the cores to a scikit_tt tensor train, for cores with these bond dimensions, with an upper bound on each TT rank in place of the rank a run reaches. It lives in mps.py, without Qiskit, so that LCHS planning can check the construction without importing an SDK.
Ranks. Write b_j for the bond dimensions of n sites and S_j = min(2**j, 2**(n-j)). A left sweep of TT.ortho sets rank j+1 to at most twice rank j, and a right sweep sets rank j to at most twice rank j+1. Neither sweep raises a rank, and a truncation only lowers one, so after ortho every rank is at most S_j and at most its value before. The MPO of an extracted two-qubit gate on sites s and s+1 has rank 4 at bond s+1 and 1 elsewhere, and mpo.dot multiplies the ranks of the working MPS by those of the MPO. Each bond carries at most one such gate per residual pass, so the working ranks at layer l are at most min(S_j, 4**l b_j).
Work, in the units of the TT-SVD work limit of decompose_state_to_mps:
- An SVD of an m x c core matrix,
m c min(m, c), charged twice because scikit_tt repeats a failedgesddwithgesvd. - A left-sweep step with
k = min(m, c):diag(s) @ vh,k k c, and itstensordotinto the next core of right rank r,2 k c r. - A right-sweep step: the previous core, a
2 a x mmatrix, times U,2 a m k, and timesdiag(s),2 a k k. mpo.dot: four products for each entry pair of MPO and MPS cores, twice the entries of the product TT.- One unit per entry for each copy or scaling of a TT and 64 units for each 4 x 4 SVD of a gate's MPO and each unitary completion.
Each layer copies the working MPS and applies ortho(max_rank=2), norm (a copy and a right sweep of the rank-2 result) and the scalar multiplication (two copies), and completes n local unitaries. Each of the num_layers - 1 residual passes applies n gates, each an MPO SVD, mpo.dot and ortho(threshold=1e-12) of the product.
Bytes, 16 per entry: the target vector, the stored cores, the right-canonical input TT and the rank-2 copies of a truncation are held throughout. On top of them comes the largest of four steps: the initial ortho, a truncation (working and truncated TT and their sweep), mpo.dot (the old working TT, the product TT and one core's broadcast copy) and the ortho of the product (the product TT and its sweep). The peak of a sweep is the larger of an SVD step and a product step. An SVD step holds the reshaped core and the Fortran copy that LAPACK factors, U, s and Vh and zgesdd's work arrays, at most k k + 194 k complex entries for LAPACK block sizes up to 64 (SciPy 1.18.1's optimal size stayed at or below 0.95 of it for m and c up to 4096), k max(5 k + 7, 2 max(m, c) + 2 k + 1) reals and 8 k integers. A product step holds U, Vh, diag(s) with its complex copy and the product arrays. An untuned allowance of 32768 + 4096 n bytes covers the gate MPOs, the Python objects of the TTs and the headers of their core arrays. Under tracemalloc these exceeded the array bytes by up to 31 kB for ten sites (SciPy 1.18.1). Qiskit circuit objects are outside the law, and the two-qubit syntheses have their own limit check. Cores whose bond dimensions are all one form a product state, which the builder prepares site by site without scikit_tt, 64 units per site.
Projected eigensolver¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
DEFAULT_OVERLAP_EIGENVALUE_CUTOFF |
1e-12 | src/nwqlib/_projected_eigensolver.py, consumed by GCiM records and Lanczos resolution |
Numerical null-space policy in Lowdin orthogonalization, not a certified input-error bound. Revisit with the numerical precision policy. |
Deterministic Gram formation allowance (gram_formation_allowance) |
g*trace(S)/(1 - g) with g = sqrt(2)*gamma_(2L) for inner products of length L |
src/nwqlib/_projected_eigensolver.py, used by src/nwqlib/algorithms/gcim/fixed_basis.py (classical, L = d) and src/nwqlib/algorithms/gcim/adapt_acquisition.py (classical, L = 2**n) |
Derived first-order bound on the spectral norm of the error of a computed Gram matrix. Each entry lies within g*||v_i||*||v_j|| (the complex inner-product bound derived in src/nwqlib/_validation.py from Higham, 2002, doi:10.1137/1.9780898718027, Lemma 3.1 and Eq. (3.5)), and with a_i = ||v_i|| the nonnegative bound matrix g*a*a^T has spectral norm g*sum_i ||v_i||**2. By Weyl's inequality a legal rank-deficient basis then has no computed overlap eigenvalue below minus this allowance and the eigensolver roundoff. It is an admission allowance, not a rank cutoff. Revisit if the Gram products leave binary64 or change their length. |
Linear-algebra work laws¶
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
_linalg_laws.py::expm_requirements |
work (10 + s) n**3 + 128 n**2 with s = max(0, ceil(log2 ||M||_1)), bytes 320 n**2 + 65536, 1-norm limit MAX_EXPONENTIAL_NORM = 2**37 |
src/nwqlib/_linalg_laws.py |
Admission of one dense matrix exponential on an n-square matrix M, by scipy.linalg.expm or scipy.sparse.linalg.expm. Both SciPy kernels follow Algorithm 5.1 of Al-Mohy and Higham (2009), doi:10.1137/09074721X, up to five products for the powers that choose the Pade degree, three products and one LU solve for the degree-13 quotient, and at most ceil(log2(||M||_1/4.25)) squarings, so the charged s has two squarings of margin. The quadratic term covers the 79 products of |M| with a vector in the backward-error tests and 34 scaled sums. The bytes cover 20 complex n-square arrays and small objects, and traced peaks of scipy.sparse.linalg.expm stayed at or below 17.2 such arrays for n from 32 to 256. The work law is a size count, not measured time. The LCHS expm and closed_form references use it with the exact 1-norm. The 1-norm limit keeps every power norm that the kernels compute, of M**k up to k = 10 and |M|**k up to k = 27, below 2**999. Above it SciPy 1.18.1's scipy.linalg.expm can take an infinite squaring count, so a reference is unavailable and planning refuses. The shared-index route of src/nwqlib/subroutines/fermionic_circuits.py::build_generator_circuit refuses an angle whose occupation-block exponent exceeds the 1-norm limit before it calls scipy.linalg.expm. Those blocks have at most five rows, so the route charges no work law. Revisit condition: A changed kernel or Pade selection in SciPy, or a workload near the work or byte cap. |
_linalg_laws.py::expm_multiply_requirements |
55 ceil(N/9.9) Taylor products, 968 norm-estimation products above N = 63.36, entries + 7 n units per product, bytes 64 (entries + n) + 512 n |
src/nwqlib/_linalg_laws.py |
Admission of one scipy.sparse.linalg.expm_multiply(G, v) on an n-square sparse G with at most entries stored entries and shifted 1-norm at most N. SciPy 1.18.1 implements Algorithm 3.2 of Al-Mohy and Higham (2011), doi:10.1137/100788860. Its parameter search includes m = 55 and s = ceil(alpha/9.9) with alpha at most N, which bounds the Taylor products. Above the 1-norm 63.36 of condition (3.13) for one vector it estimates the norms of the powers 2 to 9 with onenormest, at most 22 p products with a vector for the p-th power. Counted products stayed at or below 0.78 of the law for tridiagonal and random Hermitian generators of order 16 to 256 and N from 0.05 to 5000. The classical QHD Schrodinger evolution and its explicit fidelity references charge every call (src/nwqlib/algorithms/qhd/method.py::restricted_sizes). The one-hot and binary IR-product kernels make no such call. Revisit condition: A changed SciPy expm_multiply, or a workload near the work cap. |
_linalg_laws.py::seeded_norm_estimates |
seed 0 | src/nwqlib/_linalg_laws.py |
Seed of NumPy's global generator while scipy.sparse.linalg.expm or expm_multiply estimates norms with onenormest. The value has no accuracy meaning. A fixed seed makes the draws, and so the chosen parameters and the rounding of the result, the same for the same input, and the caller's state is restored afterwards (Dependency issues). Revisit condition: A SciPy version whose onenormest takes its own generator. |
_linalg_laws.py::singular_values_work |
ceil(4 d**3/3) + 32 d (d + 1) |
src/nwqlib/_linalg_laws.py |
Work of the singular values of a d-square matrix without its singular vectors, as numpy.linalg.norm(ord=2) and numpy.linalg.svd(compute_uv=False) compute them. Both run LAPACK's gesdd with JOBZ='N', whose bidiagonal reduction takes 16 d**2 (d - d/3) real flops for complex input and 4 d**2 (d - d/3) for real input, ceil(4 d**3/3) multiply-adds in either case. Its Householder vectors, the scaling scan of gesdd and NumPy's copy of the input add about 4 d2 units. The singular values of the bidiagonal matrix come from the dqds algorithm, six units per element of each transform (a division, an addition, a multiply-add, a multiplication and two comparisons). The division count that dlasq2 reports stayed at or below 4.12 d2 for random complex and real matrices and for matrices with clustered or graded singular values of order 4 to 128 (macOS Accelerate LAPACK), so the dqds stage took at most about 25 d2 units, and the 32 d (d + 1) term covers the steps after the reduction and the terms linear in d. The full SVD with both frames, charged 8 d**3 by the dense laws, is not computed. The admissions that charge this law before the call are those of the LCHS Duhamel remainder and the refinement's spectral_norms (_scaling_safe_spectral_norm), the dense commutator norms of the fixed product-formula validation (src/nwqlib/algorithms/lchs/time_independent_terms.py::_fixed_trotter_certificate_records), the QLS spectral selection of a general matrix (src/nwqlib/algorithms/qls/host_planning.py::_spectrum) and the QLS verification reference for a missing spectrum (src/nwqlib/algorithms/qls/verification.py::_verify), and the block-encoding planning norm np.linalg.norm(A, 2) of a dense dilation (src/nwqlib/subroutines/block_encoding/core.py::_dense_norm_work). A dense-dilation plan whose caller supplies a full SVD for the completion keeps the 8 d**3 charge. Revisit condition: Another norm routine or NumPy's SVD driver, or a dlasq2 division count above 4.6 d2. |
_linalg_laws.py::hermitian_eigensystem_work |
8 d**3 + 32 d**2 |
src/nwqlib/_linalg_laws.py |
Work of the eigenvalues and eigenvectors of a d-square Hermitian matrix by numpy.linalg.eigh, which runs LAPACK's zheevd, or dsyevd for real input, with JOBZ='V'. The tridiagonal reduction takes about (2/3) d3 multiply-adds and the back-transformation of the eigenvectors d2 (d - 1). The tridiagonal eigenvectors take at most (2/3) d3 by divide and conquer above order 25 and about 4 d3 units by implicit QL or QR iteration at order 25 and below, at two iterations per eigenvalue. The cubic term 8 d**3 equals the dense laws' charge of a full SVD and exceeds the roughly 6 d**3 of either route, and 32 d**2 covers the lower-order terms. The charge exceeds the 4 d**3 that src/nwqlib/algorithms/qls/host_planning.py::_spectrum charges for eigvalsh without vectors. The QLS classical model charges it for the one original eigendecomposition of a Hermitian A whose factors its inverse or linear norm model reuses (src/nwqlib/algorithms/qls/host_planning.py::original_factor_laws) and for the eigendecomposition of the dilation of G_t (src/nwqlib/algorithms/qls/host_planning.py::selected_work), and a dense QPE power of a Hamiltonian for the shared eigendecomposition (src/nwqlib/algorithms/qpe/powers.py::_dense_block). The QLS classical model also charges 8 d**3 for each SVD with both frames, of G_t and of a general original A whose factors its inverse or linear norm model reuses. Revisit condition: NumPy's eigensolver driver, or a workload whose admission approaches its work cap. |
_linalg_laws.py::least_squares_work |
m c**2 + c**3 + 4 min(c, 25) c**2 + 8 m c |
src/nwqlib/_linalg_laws.py |
Work of numpy.linalg.lstsq for an m x c matrix with m at least 1.6 c and one right-hand side. It runs LAPACK's gelsd, which factors the matrix by Householder QR, c**2 (m - c/3) multiply-adds, and bidiagonalizes the triangular factor, ceil(4 c**3/3), together m c**2 + c**3. It solves the bidiagonal system through its SVD. At order 25 and below that solve accumulates the right singular vectors by QR sweeps, about 4 c**3 units at two sweeps per singular value, and above 25 it uses divide and conquer with leaves of order at most 25, which 4 min(c, 25) c**2 covers. 8 m c covers NumPy's copies, the reflectors applied to the right-hand side and the remaining vector work. The QLS inverse fit charges it with m = 4u rows for u odd coefficients (src/nwqlib/subroutines/qsp/inverse.py::inverse_candidate_cost), and the kernel-reflection fit with m = max(2001, 8 (d + 1)) nodes for d + 1 coefficients (src/nwqlib/subroutines/qsp/shortcut.py::kernel_reflection_cost). Revisit condition: NumPy's least-squares driver, or a fit with fewer than 1.6 nodes per unknown. |
Example instance sizes¶
The instance sizes in examples/ have the rationales listed below. Revisit them when the intended example workload changes.
| Constant | Value | Site | Reason and revisit condition |
|---|---|---|---|
| H4 bond length / basis / pool / iterations / angle | 2.0 Å / STO-3G / spin_adapted_sd / 8 / π/4 |
examples/generators/gcim_lanczos_qpe_eigenvalue_intro.py |
At 2.0 Å CCSD lies 18 mHa below FCI, while ADAPT-GCIM reaches FCI in eight iterations of exact matrix-element evaluation. With exact probabilities one iteration as quantum circuits needs 3 circuits of at most nine qubits and the full run 136 (notebook Appendix A). Each pair circuit is reduced by applying all Pauli terms of H to its saved state, so the notebook runs the circuits on H2 at 0.74 Å instead, 3 circuits of at most five qubits. In the stored notebook run the chemistry setup and the eight exact iterations took 4.0 s, and the notebook states about 2 s for the H2 circuits, on a laptop with Python 3.12.14. |
| Eigenvalue notebook QCELS time step / schedule / longest time | automatic, 0.930 / num_times=64, all powers 0 to P / 40 per hartree |
examples/generators/gcim_lanczos_qpe_eigenvalue_intro.py |
The Hartree–Fock state has squared overlap 0.48 with the ground state, and the default QCELS schedule, powers 0 to 10, leaves an error of 121 mHa. The automatic τ = 0.9π/R uses the row-sum bound R = 3.04 Ha of the dense matrix. With num_times=64 every longest time up to 60 per hartree gives the consecutive powers of Ding and Lin's Eq. (6), and 40 per hartree, powers 0 to 43, reaches 0.42 mHa. Spread schedules reach similar or smaller errors at 40 per hartree, 0.073 mHa with τ = 0.2 and 32 powers, but can lock onto a replica of the ground-state peak. With τ = 0.2 and 32 powers spread over 1 to 190, whose gaps are 6 or 7, a longest time of 38 per hartree returned an energy 5.16 Ha too high, near the replica spacing 2π/(6τ) = 5.24 Ha. The consecutive schedule gave no such failure at the tested longest times from 10 to 60 per hartree. Its error does not fall monotonically, for example 0.42 mHa at 40 and 6.9 mHa at 45 per hartree, and Theorem 1 (arXiv:2211.11973v2, p. 14) does not cover an overlap below 0.71, so the errors are empirical. |
| Plate grid / heaters / inverse tolerance | 4 × 4 / two / 0.01 | examples/generators/qls_linear_system_intro.py |
Sixteen unknowns use 4 system qubits and 9 in total. The condition number 9.5 keeps the inverse polynomial at degree 59. No reflection or rotation of the plate maps the two unequal heaters onto themselves, and b has a component in all 9 eigenspaces of A, so the solve uses the inverse polynomial at every distinct eigenvalue. Two equal heaters at (1, 1) and (2, 2) would reach only 6 of them. |
| Channel points / speed / diffusivity / pulse width / time / Trotter steps | 8 / 1 / 0.01 / 0.15 / 0.25 / 8 | examples/generators/lchs_linear_dynamics_intro.py |
Upwind transport keeps concentrations nonnegative. The default kernel gives 3 system and 9 address qubits. Speed 1 makes one time unit a trip around the channel, and the final time 0.25, a quarter trip, moves the pulse two grid points while it keeps 71% of its norm. The Péclet number speed/diffusivity = 100 makes transport dominate the physical spreading. With grid spacing h = 1/8, the pulse width 0.15 ≈ h/√ln 2 gives a full width at half maximum of 2h, so the neighbors of the peak hold half its value. Plans for final times 0.25, 0.375 and 0.5 and diffusivities 0.005, 0.01, 0.02, 0.04 and 0.0625 keep 9 address qubits, so the qubit budget does not set these two values. Eight Trotter steps are an untuned illustration. The notebook separately compares finite-sum and quantum errors, which need not be ordered by this step count. |
| QHD grid points / total time / schedule / steps | 6 / 10 / quadratic, gamma 0.3 / 80 | examples/generators/qhd_optimization_intro.py |
Two variables at six points use 12 qubits. Without a schedule the probability stays near uniform, and this schedule places 0.79 on the best grid point. |
| QHD objective / bounds | (2x²−1)² + 3x/5 + 2(y−3/10)² + 6xy/5 / ±1.2 in both variables | examples/generators/qhd_optimization_intro.py |
The double well has minima near x = ±0.71, and the tilt 3x/5 makes the valley at negative x deeper, with the global minimum −0.81 at (−0.77, 0.53) and a local minimum 0.57 at (0.66, 0.10). The coupling 6xy/5 places the valleys at different y, so the function does not split into two one-dimensional problems. Projected gradient descent from 500 random starts in the box reaches the global minimum in 54% of runs. The box contains both valleys, and its 6 × 6 grid of interior points includes one 0.08 from the global minimum. |
| QHD constrained objective / constraint / bounds / grid points | (x−1)² + (y−1)² / x² + y² ≤ 1 / [0, 1] in both variables / 4 | examples/generators/qhd_optimization_intro.py |
The continuous solution (1/√2, 1/√2) has f* = 3 − 2√2 ≈ 0.172 and an active constraint with multiplier √2 − 1, so the notebook can relate a grid point to it exactly. The box is the unit square, which does not depend on where the solution lies. Its 4 × 4 interior grid, with coordinates 0.2 to 0.8, uses 8 qubits and contains (0.6, 0.8) and (0.8, 0.6) on the circle, so the constraint is also active on the grid. These two points are the feasible grid minima, with objective 0.2 at distance 0.14 from the continuous solution. The 3 × 3 grid of the same box has no point on the circle. The rounds use the schedule, steps and total time of the unconstrained example. |
| LCHS scientific dimension / D / time / penalty / source rate | 4 / 4 / .005 / 1000 / 298 | examples/generators/lchs_scientific.py |
Four-point heat flow with a constant source and imaginary boundary penalty. D = 4 and the source rate 298 are the values of the heat experiments of Schleich et al., arXiv:2506.21751v1, which place a point source of that strength at the centre element (caption of Fig. 5 and Sec. III.2.1). On four points D/h² = 36, so the paper's final time 1 would reach the steady state. The time .005 gives DT/h² = 0.18, where the slower interior mode keeps 84% of its amplitude and the source injection 1.49 is comparable with the initial norm 0.91, so the initial-state and source branches carry similar weight. The penalty 1000, λT = 5 radians, lies within the paper's numerical range, which stops at 10⁶ (Sec. III.1). With this quadrature it makes the finite-penalty error, 3.8%, comparable with the LCHS error, 2.8%. λ = 300 gives a penalty error of 11%, and λ = 3000 and 10000 make the LCHS error the larger one. The notebook performs one quantum solve and separates its error from the finite-penalty approximation using two independent references. |
| LCHS scientific k qubits / source-time nodes / total-qubit cap | 5 / 4 / 12 | examples/generators/lchs_scientific.py |
The default 32 k nodes and four source-time nodes give 160 weighted branches in 256 address slots, for ten total qubits. The cap is the 12-qubit simulator width of the example notebooks (docs/examples.md). It sets max_dense_select_slots to 1024, so planning refuses a larger dense SELECT. In the stored run the compiled circuit has about 3900 CX gates per weighted branch. |
| QLS scientific tau / time step / norm guess / construction tolerance | 1 / .1 / 4 / .001 | examples/generators/qls_scientific.py |
Two Euler updates and two final-state copies of three collision populations form a 15-coordinate history. Every history block stays close to f(0), and ‖b‖ ≈ ‖f(0)‖, so the norm guess estimates ν = α‖x‖/‖b‖ ≈ √5 α before the run. With α = 1.88 the estimate is 4.20, and the actual ν is 4.18. Dalzell's success probability, about sin²(2θ_t) with θ_t = arctan(ν/t) (arXiv:2406.12086v2, Eqs. (7) and (17)), is 0.998 at t = 4 and 0.205 at t = 1. The relaxation changes the normalized history by about 1%, so the tolerance must be finer than that. With the notebook's construction, the QLS direction error at tolerances .01, .003 and .001 is 0.33, 0.045 and 0.017 of the no-evolution baseline for t = 4, and 1.40, 0.38 and 0.13 for t = 1. |
| QLS scientific default qubits / total-qubit cap | 10 / 12 | examples/generators/qls_scientific.py |
Dense encoding pads the history from 15 to 16 coordinates. Encoding and shortcut ancillas give ten total qubits. The cap is the 12-qubit simulator width of the example notebooks (docs/examples.md). |
| Resource notebook sizes / QLS tolerance / LCHS allowances / QCELS schedule | q = 80, 90, 100 / 0.01 / 0.005 for kernel cutoff and quadrature and 0.005 for the Strang product / τ = 0.005, longest time 0.16, 32 positive powers, grid 256, 1024 shots per setting | examples/generators/resource_estimation_at_scale.py |
The sizes are the 80 to 100 qubits the notebook is meant to show, where a complex128 state vector of the system alone needs 1.9e25 to 2.0e31 bytes. The screened Poisson ring keeps its coefficients as it grows, so κ = 2 and the inverse polynomial has degree 11 at every size. The default vector output plans all four tolerances from 0.1 to 0.0001 at q = 100 within the program admission limit. The LCHS split meets an ideal vector-error target of 0.01 with 53 Strang steps, whose construction charge at q = 100, 69,127,013 units, fits max_select_work. With approximation_tolerance=0.01, a Strang allowance of 0.001 needs 112 steps and is rejected at q = 90 and 100. With τ·3q ≤ 1.5 < π the spectral enclosure excludes phase wrapping of the integer powers. Powers 0 to 32 give 66 settings. The requested 256-point grid (effective 257) keeps the run short, and the default 4096-point grid, at 281,153,664 units, also fits the QPE work limit. The Ising and Heisenberg chains have 200 and 300 Pauli terms at q = 100, below 851 terms, the initial full-envelope ceiling of the census alone at the default QPE work limit. The compiled checks of Appendix E use 4, 6, 8, 12 and 16 system qubits for QLS and LCHS (8 to 25 total qubits) and 4 and 8 spins for the two QPE chains, where the same constructions and accuracy controls compile at optimization level 0. prepare(plan, settings="all") builds all 66 circuits of each QPE plan without submission. The 14 checks took 15 s together in the stored run. Compiling the LCHS circuit on 4 system qubits at level 3 had not finished after 3 min, so only QLS gets a level-3 row. The Qiskit version and machine of that compile attempt were not recorded, so the 3 min explains this choice and is not a timing to compare. In the stored run, with Python 3.12.14 and Qiskit 2.5.2 on an Apple M3 Max laptop, the 36 planning and estimation calls took 1.1 s together. |
| Resource notebook GCiM Hamiltonian / trial states / shots | −Z^⊗q − 0.7 X^⊗q / two product states (cos θ|0⟩ + sin θ|1⟩)^⊗q at θ=π/8 and 3π/8 / 1024 per setting | examples/generators/resource_estimation_at_scale.py, src/nwqlib/algorithms/gcim/fixed_basis.py::_sampled_construction, src/nwqlib/blocks/selection.py, src/nwqlib/_preparation_laws.py |
The two nonidentity terms form G=2 QWC groups. The complete b=2 pencil has b²G=8 settings, four diagonal and four off diagonal, and 8192 shots. An uncontrolled product preparation costs zero CX. A controlled product preparation uses q copies of the one-qubit complex-tree bound, 4q CX. Group basis changes and ancilla rotations are one-qubit operations. Each off-diagonal setting is bounded by 8q CX, so the complete acquisition is bounded by 32q*1024. At q=3, Qiskit 2.5.2 with basis cx,u, optimization level 0, approximation_degree=1.0 and seed 7 compiles the settings to [0,0,12,12,12,12,0,0], or 49152 CX over all shots against the bound 98304. The ratio two follows from one two-CX CRY per real product-state site and is specific to these states and compilation settings. Revisit when the preparation law, readout construction or compiler changes. |
| Resource notebook QHD grids / time / steps / start | K = 8, 16, 32 per variable / 0.001 / one first-order step, midpoint cubic-schedule weights / uniform | examples/generators/resource_estimation_at_scale.py |
The largest one-hot circuit has 96 logical qubits, within the notebook's 80 to 100 qubit band, while binary uses 15. One step from the uniform state gives a small, reproducible workload for comparing the two encodings' circuit constructions, and the quadratic's minimum (1/4, 1/4, 1/4) lies on every grid. At K = 32 the one-hot evolution bound reaches its uninformative cap of 2, and the difference from binary reflects how each construction is bounded, not achieved accuracy. The compiled checks use K = 4 in both encodings and K = 8 in binary, which cover both binary kinetic selections within the default 20-qubit preparation limit. |