Skip to content

Diagonalize the batches concurrently over groups of processes - #54

Draft
Jim Garrison (garrison) wants to merge 1 commit into
mainfrom
concurrent-batch-diagonalization
Draft

Jim Garrison (garrison) wants to merge 1 commit into
mainfrom
concurrent-batch-diagonalization

Conversation

@garrison

Copy link
Copy Markdown
Member

solve_sci_batch diagonalized each subspace in turn over every process. Since qiskit-addon-sqd passes all of an iteration's subspaces in one call, the processes can instead be divided into one group per subspace and the subspaces diagonalized at the same time.

The division happens when there is more than one subspace and at least as many processes as subspaces. Otherwise the previous behavior is unchanged. The returned list is ordered by subspace either way.

What this gives you

This is what the two-round trim SQD schedule needs, and it needs no new parameter. The loop calls the solver once per round: TrimPolicy's screening round hands over every batch in a single call, which is divided into groups here, and its merged round hands over one subspace, which is given every process. Which round is which follows from the number of subspaces, so dividing the processes stays the solver's concern rather than the schedule's — as diagonalize_fermionic_hamiltonian documents for a collective sci_solver.

Details worth a reviewer's attention

The FCIDUMP is still read once. The split happens after it is loaded, over the whole communicator, rather than once per group.

Groups are equal in size, and a remainder is left idle rather than making one group larger. This is not tidiness. A screening round diagonalizes these batches in order to rank them against each other, and a batch solved over more processes than its siblings reaches a slightly different floating-point energy, so the ranking would turn partly on how the processes happened to divide. With 10 processes and 3 subspaces, three groups of three are formed and one process takes no part. Idle processes still enter every collective, or the ones doing the work would wait on them forever.

The wavefunction dump is renamed from wavefunction.bin to wavefunction_<group>.bin. SBD does not return the amplitudes through the binding: it writes them to this file and the solver reads them back to build the SCIState. The directory holding it is created once and broadcast, so every group shares it, and the name used to be a constant — safe only while the batches ran one after another, each write being consumed before the next began. Run concurrently, the groups would write one path at once and each would then read back whichever write landed last, building an SCIState from another group's subspace with no error raised. The batch index distinguishes them, and needs no coordination: every member of a group agrees on it and distinct groups differ. With carryover_type=1 the dump is never requested, so the collision cannot arise there.

Communicators are freed rather than left to the garbage collector, since a long run would otherwise exhaust them. Only the ones Split produced are freed, never the caller's.

A consequence for the process grid

adet_comm_size, bdet_comm_size and task_comm_size describe a single diagonalization, so they are relative to whichever communicator that diagonalization is handed — a group, once the processes are divided, rather than all of them. The rule is divisibility: diag() derives the helper dimension by integer division, h_comm_size = mpi_size / (task_comm_size * base_comm_size), and TaskCommunicator then checks that multiplying the four back recovers mpi_size, which fails exactly when that division truncated. So adet × bdet × task has to divide the communicator evenly.

Nothing about that rule changed; which communicator it applies to did. A config sizing the grid from MPI.COMM_WORLD was only ever shorthand for "every process performing this solve", and the world stopped being that set. With 6 processes and 3 subspaces each group holds 2, so a grid asking for 6 no longer divides it — while 2, or the default 1, would. A caller cannot size the grid to the group in any case, not knowing how the processes were divided, and one config has to serve both rounds. The new tests therefore leave the grid at its default.

Testing

_split_for_batches is arithmetic over (size, rank, num_batches), so the whole decision table — group sizes, leader selection, idle ranks, contiguity — is checked in a single process for counts up to eleven, including the uneven divisions. Two MPI tests then exercise the real split: one diagonalizes several subspaces both divided and one at a time over every process and requires the energies to agree, and one runs the two-round shape repeatedly in one process lifetime, which is what would expose a leaked communicator or a dump read back by the wrong round.

Verified on the CPU backend under OpenMPI, passing at one through seven processes. Not yet exercised: the GPU backends, and more than one node.


This pull request was drafted by Claude Opus 5 under my guidance.

solve_sci_batch diagonalized each subspace in turn over every process. Because
qiskit-addon-sqd passes all of an iteration's subspaces in one call, the processes
can instead be divided into one group per subspace and the subspaces diagonalized
at the same time.

The division happens after the FCIDUMP is loaded, so the Hamiltonian is still read
once over the whole communicator rather than once per group. Each group then runs
_solve_sci_core over its own communicator, so a result is produced on each group's
rank-0 process rather than only on the world's. The group leaders exchange their
results and the collected list is broadcast, ordered by subspace index, so every
process returns the same thing as before.

Groups are equal in size, and any leftover processes are given no group rather than
making one group larger. This is not tidiness: a schedule such as TrimPolicy
diagonalizes these batches in order to rank them against each other, and a batch
solved over more processes than its siblings reaches a slightly different
floating-point energy, so the ranking would turn partly on how the processes
happened to divide. Those idle processes still enter every collective, or the ones
doing the work would wait on them forever.

This is what trim SQD needs, and it needs no parameter. qiskit-addon-sqd's loop
calls the solver once per round: TrimPolicy's screening round hands over every
batch in one call, which is divided here, and its merged round hands over a single
subspace, which is given every process. Dividing the processes is the solver's
concern rather than the schedule's, which is what diagonalize_fermionic_hamiltonian
documents for a collective sci_solver.

The wavefunction dump is renamed from wavefunction.bin to wavefunction_<group>.bin.
SBD does not return the amplitudes through the binding: it writes them to this file
and _solve_sci_core reads them back to build the SCIState. The directory holding it
is created once and broadcast, so every group shares it, and the name used to be a
constant -- safe only while the batches ran one after another, each write being
consumed before the next began. Run concurrently, the groups would write one path at
once and each would then read back whichever write landed last, building an SCIState
from another group's subspace with no error raised. The batch index distinguishes
them: every member of a group agrees on it and distinct groups differ, so the names
cannot collide and no coordination is needed. With carryover_type=1 the dump is
never requested, so the collision cannot arise there.

Communicators are freed rather than left to the garbage collector, since a long run
would otherwise exhaust them. Only the ones Split produced are freed, never the
caller's own.

One consequence for the process grid. adet_comm_size, bdet_comm_size and
task_comm_size describe a single diagonalization, so they are relative to whichever
communicator that diagonalization is handed -- a group, once the processes are
divided, rather than all of them. The rule is divisibility: diag() derives the
helper dimension by integer division, h_comm_size = mpi_size / (task_comm_size *
base_comm_size), and TaskCommunicator checks that multiplying the four back recovers
mpi_size, which fails exactly when that division truncated. So adet * bdet * task
has to divide the communicator evenly.

Nothing about that rule changed; which communicator it applies to did. A config
sizing the grid from MPI.COMM_WORLD was only ever shorthand for "every process doing
this solve", and the world stopped being that set. With 6 processes and 3 subspaces
each group holds 2, so a grid asking for 6 no longer divides it -- while 2, or the
default 1, would. A caller cannot size the grid to the group anyway, not knowing how
the processes were divided, and one config has to serve both rounds. The new tests
therefore leave the grid at its default instead of using the _sbd_config helper,
which sizes it from the world.

Tested on the CPU backend under OpenMPI. _split_for_batches is arithmetic over
(size, rank, num_batches), so the whole decision table -- group sizes, leader
selection, idle ranks, contiguity -- is checked in a single process for counts up
to eleven, including the uneven divisions. Two MPI tests then run the real split:
one diagonalizes several subspaces both divided and one at a time over every
process and requires the energies to agree, and one runs the two-round trim shape
repeatedly in one process lifetime, which is what would expose a leaked
communicator or a dump read back by the wrong round. The suite passes at one
through seven processes. Not yet exercised: the GPU backends, and more than one
node.

Assisted-by: Claude Opus 5
@garrison Jim Garrison (garrison) added this to the 1.8.0 milestone Oct 9, 2026
@garrison Jim Garrison (garrison) added the enhancement New feature or request label Oct 9, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant