Repository navigation
Diagonalize the batches concurrently over groups of processes - #54
Draft
Jim Garrison (garrison) wants to merge 1 commit into
Draft
Jim Garrison (garrison) wants to merge 1 commit into
Jim Garrison (garrison) wants to merge 1 commit into
Conversation
solve_sci_batch diagonalized each subspace in turn over every process. Because qiskit-addon-sqd passes all of an iteration's subspaces in one call, the processes can instead be divided into one group per subspace and the subspaces diagonalized at the same time. The division happens after the FCIDUMP is loaded, so the Hamiltonian is still read once over the whole communicator rather than once per group. Each group then runs _solve_sci_core over its own communicator, so a result is produced on each group's rank-0 process rather than only on the world's. The group leaders exchange their results and the collected list is broadcast, ordered by subspace index, so every process returns the same thing as before. Groups are equal in size, and any leftover processes are given no group rather than making one group larger. This is not tidiness: a schedule such as TrimPolicy diagonalizes these batches in order to rank them against each other, and a batch solved over more processes than its siblings reaches a slightly different floating-point energy, so the ranking would turn partly on how the processes happened to divide. Those idle processes still enter every collective, or the ones doing the work would wait on them forever. This is what trim SQD needs, and it needs no parameter. qiskit-addon-sqd's loop calls the solver once per round: TrimPolicy's screening round hands over every batch in one call, which is divided here, and its merged round hands over a single subspace, which is given every process. Dividing the processes is the solver's concern rather than the schedule's, which is what diagonalize_fermionic_hamiltonian documents for a collective sci_solver. The wavefunction dump is renamed from wavefunction.bin to wavefunction_<group>.bin. SBD does not return the amplitudes through the binding: it writes them to this file and _solve_sci_core reads them back to build the SCIState. The directory holding it is created once and broadcast, so every group shares it, and the name used to be a constant -- safe only while the batches ran one after another, each write being consumed before the next began. Run concurrently, the groups would write one path at once and each would then read back whichever write landed last, building an SCIState from another group's subspace with no error raised. The batch index distinguishes them: every member of a group agrees on it and distinct groups differ, so the names cannot collide and no coordination is needed. With carryover_type=1 the dump is never requested, so the collision cannot arise there. Communicators are freed rather than left to the garbage collector, since a long run would otherwise exhaust them. Only the ones Split produced are freed, never the caller's own. One consequence for the process grid. adet_comm_size, bdet_comm_size and task_comm_size describe a single diagonalization, so they are relative to whichever communicator that diagonalization is handed -- a group, once the processes are divided, rather than all of them. The rule is divisibility: diag() derives the helper dimension by integer division, h_comm_size = mpi_size / (task_comm_size * base_comm_size), and TaskCommunicator checks that multiplying the four back recovers mpi_size, which fails exactly when that division truncated. So adet * bdet * task has to divide the communicator evenly. Nothing about that rule changed; which communicator it applies to did. A config sizing the grid from MPI.COMM_WORLD was only ever shorthand for "every process doing this solve", and the world stopped being that set. With 6 processes and 3 subspaces each group holds 2, so a grid asking for 6 no longer divides it -- while 2, or the default 1, would. A caller cannot size the grid to the group anyway, not knowing how the processes were divided, and one config has to serve both rounds. The new tests therefore leave the grid at its default instead of using the _sbd_config helper, which sizes it from the world. Tested on the CPU backend under OpenMPI. _split_for_batches is arithmetic over (size, rank, num_batches), so the whole decision table -- group sizes, leader selection, idle ranks, contiguity -- is checked in a single process for counts up to eleven, including the uneven divisions. Two MPI tests then run the real split: one diagonalizes several subspaces both divided and one at a time over every process and requires the energies to agree, and one runs the two-round trim shape repeatedly in one process lifetime, which is what would expose a leaked communicator or a dump read back by the wrong round. The suite passes at one through seven processes. Not yet exercised: the GPU backends, and more than one node. Assisted-by: Claude Opus 5
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
solve_sci_batchdiagonalized each subspace in turn over every process. Since qiskit-addon-sqd passes all of an iteration's subspaces in one call, the processes can instead be divided into one group per subspace and the subspaces diagonalized at the same time.The division happens when there is more than one subspace and at least as many processes as subspaces. Otherwise the previous behavior is unchanged. The returned list is ordered by subspace either way.
What this gives you
This is what the two-round trim SQD schedule needs, and it needs no new parameter. The loop calls the solver once per round:
TrimPolicy's screening round hands over every batch in a single call, which is divided into groups here, and its merged round hands over one subspace, which is given every process. Which round is which follows from the number of subspaces, so dividing the processes stays the solver's concern rather than the schedule's — asdiagonalize_fermionic_hamiltoniandocuments for a collectivesci_solver.Details worth a reviewer's attention
The FCIDUMP is still read once. The split happens after it is loaded, over the whole communicator, rather than once per group.
Groups are equal in size, and a remainder is left idle rather than making one group larger. This is not tidiness. A screening round diagonalizes these batches in order to rank them against each other, and a batch solved over more processes than its siblings reaches a slightly different floating-point energy, so the ranking would turn partly on how the processes happened to divide. With 10 processes and 3 subspaces, three groups of three are formed and one process takes no part. Idle processes still enter every collective, or the ones doing the work would wait on them forever.
The wavefunction dump is renamed from
wavefunction.bintowavefunction_<group>.bin. SBD does not return the amplitudes through the binding: it writes them to this file and the solver reads them back to build theSCIState. The directory holding it is created once and broadcast, so every group shares it, and the name used to be a constant — safe only while the batches ran one after another, each write being consumed before the next began. Run concurrently, the groups would write one path at once and each would then read back whichever write landed last, building anSCIStatefrom another group's subspace with no error raised. The batch index distinguishes them, and needs no coordination: every member of a group agrees on it and distinct groups differ. Withcarryover_type=1the dump is never requested, so the collision cannot arise there.Communicators are freed rather than left to the garbage collector, since a long run would otherwise exhaust them. Only the ones
Splitproduced are freed, never the caller's.A consequence for the process grid
adet_comm_size,bdet_comm_sizeandtask_comm_sizedescribe a single diagonalization, so they are relative to whichever communicator that diagonalization is handed — a group, once the processes are divided, rather than all of them. The rule is divisibility:diag()derives the helper dimension by integer division,h_comm_size = mpi_size / (task_comm_size * base_comm_size), andTaskCommunicatorthen checks that multiplying the four back recoversmpi_size, which fails exactly when that division truncated. Soadet × bdet × taskhas to divide the communicator evenly.Nothing about that rule changed; which communicator it applies to did. A config sizing the grid from
MPI.COMM_WORLDwas only ever shorthand for "every process performing this solve", and the world stopped being that set. With 6 processes and 3 subspaces each group holds 2, so a grid asking for 6 no longer divides it — while 2, or the default 1, would. A caller cannot size the grid to the group in any case, not knowing how the processes were divided, and one config has to serve both rounds. The new tests therefore leave the grid at its default.Testing
_split_for_batchesis arithmetic over(size, rank, num_batches), so the whole decision table — group sizes, leader selection, idle ranks, contiguity — is checked in a single process for counts up to eleven, including the uneven divisions. Two MPI tests then exercise the real split: one diagonalizes several subspaces both divided and one at a time over every process and requires the energies to agree, and one runs the two-round shape repeatedly in one process lifetime, which is what would expose a leaked communicator or a dump read back by the wrong round.Verified on the CPU backend under OpenMPI, passing at one through seven processes. Not yet exercised: the GPU backends, and more than one node.
This pull request was drafted by Claude Opus 5 under my guidance.