OSC-SM: Notified RMA Implementation with Set notify and Bounds - #28
Open
joe-explr wants to merge 24 commits into
Open
OSC-SM: Notified RMA Implementation with Set notify and Bounds#28joe-explr wants to merge 24 commits into
joe-explr wants to merge 24 commits into
Conversation
Signed-off-by: Joseph Schuchart <joseph.schuchart@stonybrook.edu>
This commit adds notification support to the OSC SM component by implementing the put_with_notify, get_with_notify, rput_with_notify, and rget_with_notify functions. These functions perform the same operations as their non-notify counterparts but also increment notification counters after the data transfer completes. The changes include: - Added function pointer types for notify variants in osc.h - Added function prototypes in osc_sm.h - Implemented the notify functions in osc_sm_comm.c - Updated the module template to register the new functions - Removed TODO comments that have been addressed Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
put_with_notify get_with_notify Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
put_with_notify
get_with_notify
Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
…for a single and multi rank window. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
…eph.antony18@gmail.com>
… <jajoseph.antony18@gmail.com>
MPI-5.1 names the error class for an invalid notification index MPI_ERR_RMA_NOTIFICATION. Rename the placeholder used by the notified RMA work to match the standard. The class was never registered with the error code subsystem, so MPI_Error_string() and MPI_Error_class() did not know about it; add the missing CONSTRUCT_ERRCODE()/OBJ_DESTRUCT() pair in errcode.c. Also fix the binding generator's ERROR_CLASSES list, where the entry was inserted without a trailing comma and so was silently concatenated with the following 'MPI_ERR_TYPE' element rather than added as a class of its own. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Add the two remaining notification-management procedures from MPI-5.1 section 12.6.1: MPI_WIN_SET_NUM_NOTIFY is a blocking, synchronizing collective that sets the number of notification counters attached at the calling MPI process to exactly num_notifications and resets all of them to zero. MPI_WIN_GET_NUM_NOTIFY is local and returns the number of counters attached at target_rank. Both are wired through the osc framework as new module entry points, so components that do not implement them return MPI_ERR_UNSUPPORTED_OPERATION rather than crashing. The osc/sm implementation carves a fixed per-rank counter region out of the shared segment at window creation, which is therefore the effective MPI_WIN_NOTIFICATION_NUM_UB; a request beyond that capacity is rejected with MPI_ERR_ARG. Each rank publishes its own attached count into the shared segment, so the collective needs only a barrier -- no counts have to be exchanged -- and MPI_WIN_GET_NUM_NOTIFY is a plain shared-memory read. Also add the missing put_notify/get_notify entries to interface_profile_sources, which were omitted when those two procedures were introduced. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Three correctness problems in the osc/sm notified communication path: 1. The notification counters were plain uint64_t but are incremented concurrently by remote origins with opal_atomic_add() and polled by the local rank. Type them opal_atomic_int64_t so that the reads in MPI_WIN_GET_NOTIFY_VALUE are atomic and cannot be hoisted out of a caller's polling loop. 2. The notification index was validated *after* the data movement, so an erroneous call had already overwritten the target window (or, for get, the origin buffer) by the time the error was returned. MPI-5.1 section 12.6.1 makes referencing an out-of-range counter erroneous at initiation, so hoist the check above ompi_datatype_sndrcv() in all four notified operations. The check is factored into a helper, which also fixes MPI_GET_NOTIFY returning OMPI_ERR_BAD_PARAM instead of MPI_ERR_RMA_NOTIFICATION. 3. The get paths used opal_atomic_rmb() before incrementing the target's counter. The notification tells the target that the get has read the window, so the constraint is load-before-store, which a load-load fence does not express; opal_atomic_add() is relaxed and adds no ordering of its own. Use a full opal_atomic_mb(). In MPI_WIN_GET_NOTIFY_VALUE the barrier was likewise placed before the counter load, where it ordered nothing; move it after so that it gives the acquire semantics the caller needs. MPI_WIN_RESET_NOTIFY_VALUE also gains the trailing barrier for the same reason. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Adds the six remaining notified operations from the MPI-5.1 draft -- accumulate, get_accumulate, and the four request-based forms -- so that osc/sm covers all eight defined in section 12.3, and sizes the notification counters from window info rather than a compile-time constant. The counter region was a fixed 16 entries per MPI process carved out of the shared segment, and MPI_WIN_SET_NUM_NOTIFY rejected anything larger. That conflicts with the mpi_assert_max_num_notify info key, whose default of 0 the standard defines as "the implementation does not assume any limit on the number of notification counters". The reservation now comes from that key when one is given, and otherwise from a new osc_sm_num_notify_counters MCA parameter. A request beyond the reservation relocates the counters to a dedicated shared segment instead of failing; when the key was given it is a hard bound, since the window was sized on the strength of that assertion. Growth is collective and runs inside MPI_WIN_SET_NUM_NOTIFY, which the standard already defines as a blocking synchronizing collective. Every process agrees on the new layout through an allgather of the requested counts, and on whether the attach succeeded through an allreduce, so a failure at one process cannot leave others incrementing counters that nobody reads. The published count stays clamped to the current allocation until the larger one exists, so a failed growth cannot leave behind a count that would admit writes past the end of the region. A barrier separates the attach from the unlink, because attach opens the backing file by name and the broadcast does not tell rank 0 that the other processes are finished with it. Each process now caches a per-target pointer to the counters, making the lookup on the path of every notified operation a single indexed load -- cheaper than the previous base-plus-offset arithmetic -- so the ability to relocate the region costs the hot path nothing. Also corrects the reset at the end of component_select(), which zeroed the whole node state and so wiped the notification fields it had just written, and removes a stray double semicolon in ompi_osc_sm_fetch_and_op(). Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
MPI-5.1 section 12.2.6, Table 12.1 caches three attributes on every window: MPI_WIN_NOTIFICATION_NUM_SB, the number of notification counters the implementation supports efficiently; MPI_WIN_NOTIFICATION_NUM_UB, the upper bound on that number; and MPI_WIN_NOTIFICATION_VALUE_UB, the upper bound on a counter value. Without them a program has no portable way to ask how many counters it may request, since MPI_WIN_GET_NUM_NOTIFY reports how many are attached rather than how many are available. The values come from a new osc_win_get_notify_bounds entry point on the osc module, queried once per window while it is configured. A component that does not implement notified communication leaves the entry point NULL and all three attributes read zero, which is the honest answer for such a window and is consistent with its notified operations returning MPI_ERR_UNSUPPORTED_OPERATION. For osc/sm the bounds follow the reservation: with an mpi_assert_max_num_notify assertion both NUM_SB and NUM_UB are that value, since the window was sized for exactly it; without one, NUM_SB is what was reserved and nothing bounds NUM_UB short of the notification index type, because the counters grow on demand. The keyvals are appended after MPI_FT so that the existing predefined values stay put, with matching entries in mpif-values.py to keep the C and Fortran numbering identical. The predefined-keyval bitmap is already bounded by MPI_ATTR_PREDEFINED_KEY_MAX and needed no change. Table 12.1 types VALUE_UB as MPI_Count *, and the attribute machinery has no MPI_Count slot -- every other predefined attribute is integer- or address-valued. It is stored as an MPI_Aint, which is the same width wherever Open MPI runs now that 32-bit environments are unsupported; the reasoning is recorded at the call site so the choice does not later read as a type error. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
Covers counter management, all eight notified operations in their blocking and request-based forms, notification index errors, growth past the reserved capacity, the mpi_assert_max_num_notify bound, and the three notification window attributes. A second test forces osc/rdma and checks that every notified entry point reports MPI_ERR_UNSUPPORTED_OPERATION without moving any data, and that the attributes read zero. That is the contract which lets components that do not implement the chapter remain untouched, so it is worth testing directly rather than assuming. Both are single-process tests wired into make check, so the shared segment growth path is exercised only in its single-rank form, where the counters are a plain heap allocation. The collective path -- segment creation, broadcast, attach, the status allreduce and the barrier before unlink -- needs a multi-rank test that this harness cannot host. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
The notified communication code carried long block comments that restated the standard at length and recorded design deliberation. Reduce them to short notes that say what the code does and cite the relevant MPI-5.1 section, so the comment density matches the surrounding osc/sm sources. No functional change. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com>
joe-explr
commented
Aug 12, 2026
| @@ -0,0 +1,151 @@ | |||
| /* | |||
| * Copyright (c) 2026 Joseph Antony. All rights reserved. | |||
| * $COPYRIGHT$ | |||
Author
There was a problem hiding this comment.
I will remove this , not sure what to fill here
MPI_WIN_GET_NUM_NOTIFY takes target_rank as a nonnegative integer in the group of the window, but the frontend did not validate it. The osc/sm and osc/ucx backends both range-check it and return MPI_ERR_RANK, so this was not a crash, but argument validation belongs in the frontend under MPI_PARAM_CHECK, consistent with the notified communication operations and with MPI_Win_shared_query. MPI_PROC_NULL is deliberately not accepted here: unlike the notified communication operations, the spec specifies target_rank as nonnegative. Signed-off-by: Joseph Antony <jajoseph.antony18@gmail.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements the MPI-5.1 notified communication chapter (§12.6) for the
osc/smone-sided component, together with the framework plumbing, Cbindings, window attributes, and error class the feature requires.
Notified RMA lets an origin attach a notification to an RMA operation:
the target's notification counter is incremented only after the data
movement has completed at the target, so the target can poll a counter
instead of participating in a synchronization epoch.
Public MPI interface
bigcount
_cform (16 entry points):MPI_Put_notify,MPI_Get_notify,MPI_Accumulate_notify,MPI_Get_accumulate_notify, and the request-returningMPI_Rput_notify,MPI_Rget_notify,MPI_Raccumulate_notify,MPI_Rget_accumulate_notify. Each takes an additionalint notification_idximmediately beforewin.MPI_Win_set_num_notify,MPI_Win_get_num_notify,MPI_Win_get_notify_value,MPI_Win_reset_notify_value.MPI_ERR_RMA_NOTIFICATION(value 83), registeredin
ompi/errhandler/errcode.candompi/include/mpif-values.py.MPI_WIN_NOTIFICATION_NUM_SB,MPI_WIN_NOTIFICATION_NUM_UB,MPI_WIN_NOTIFICATION_VALUE_UB(keyvals 13–15), appended after the existing
MPI_WIN_*values sothat no established attribute value shifts.
C bindings
12 new
.c.ingenerator templates underompi/mpi/c/(one per notifiedoperation and window procedure), wired into
ompi/mpi/c/Makefile.am,with
MPI_ERR_RMA_NOTIFICATIONadded toompi/mpi/bindings/ompi_bindings/consts.py. No generated output ishand-edited.
osc framework
ompi/mca/osc/osc.hgains the 8 notified operation slots plus 5 modulefunction pointers:
osc_win_get_notify_value,osc_win_reset_notify_value,osc_win_set_num_notify,osc_win_get_num_notify, andosc_win_get_notify_bounds. A componentthat does not implement notified communication leaves them NULL.
Window setup
The three notification attributes are cached on every window at
creation — querying
osc_win_get_notify_boundswhen the componentsupplies it, and caching zeros when it does not, which is an honest
report that no notification counter can be attached to such a window.
MPI_WIN_NOTIFICATION_VALUE_UBis stored viaompi_attr_set_aintbecause the attribute machinery has no
MPI_Countslot and every otherpredefined attribute is integer- or address-valued. This is safe while
MPI_Ainttracks the pointer width; a comment inompi/win/win.crecords that a 32-bit revival would need a real
MPI_Countslot inattribute_value_t.osc/sm implementation
int64_treserved inline in the main sharedsegment (16 per rank by default, tunable via the
osc_sm_num_notify_countersMCA parameter), reached through aper-process
notify_bases[]pointer array so that each notifiedoperation costs a single indexed load regardless of where the counters
currently live.
node_states[]: theattached count, the reserved capacity, and the segment
offset. The last is an offset rather than a pointer because the
shared segment is mapped at a different address in every process.
MPI_Win_set_num_notifygrows on demand into a separatelyallocated overflow segment: allgather the requests so the grow/no-grow
decision is identical everywhere, bcast the new segment descriptor (an
empty
seg_namesignals failure so that no process hangs in a latercollective), allreduce on attach success before any shared state is
mutated, raise the attached count only once the space behind it
exists, barrier, then unlink and drop the old mapping. The allocation
never shrinks.
mpi_assert_max_num_notifyinfo key turns the reservation into ahard cap rather than a trigger for reallocation, and is reported back
through
MPI_Win_get_infoonly when one was actually given.store-store fence; get and the whole accumulate family use a full
barrier, because there the notification asserts that a read of the
target window happened — a load-before-store constraint that a
one-sided fence cannot express.
opal_atomic_add()is relaxed andcontributes no ordering of its own.
count before any data moves, so an erroneous call cannot leave the
target window (or, for get, the origin buffer) modified before the
error is reported.
never bumped for an operation that failed.
MPI_NO_OPis explicitlynot treated as failure: the target window was still read into the
result buffer, which is an access the notification must cover.
MPI_Win_reset_notify_valueuses a singleopal_atomic_swap_64sothat increments arriving between a read and a separate zeroing cannot
be lost.
Instrumentation
Four new SPC counters:
OMPI_SPC_PUT_NOTIFY,OMPI_SPC_RPUT_NOTIFY,OMPI_SPC_GET_NOTIFY,OMPI_SPC_RGET_NOTIFY.Tests
Two new programs under
ompi/test/general/, both wired intomake check, with the binaries added to.gitignore.win_notify.ccovers counter management, blocking and request-basednotified operations, notification-index error cases, on-demand counter
growth, the
mpi_assert_max_num_notifyassertion path, and the threewindow attributes. All 8 notified operations and all 4 management
procedures are exercised.
win_notify_unsupported.cverifies that a window whose component doesnot implement notified communication returns
MPI_ERR_UNSUPPORTED_OPERATIONfrom every notified entry point ratherthan calling through a NULL function pointer.
Open question for reviewers
osc/smpre-attaches its reserved counters, whereasosc/ucxstarts atzero and requires
MPI_Win_set_num_notifybefore any counter may bereferenced. MPI-5.1 §12.6.1 does not state what the initial attached
count is, so neither behavior is provably wrong — but the divergence
means a program that omits
MPI_Win_set_num_notifyworks onsmandfails on
ucx. Worth settling before this lands.