This commit is contained in:
Peter Boyle committed 2026-09-29 13:46:09 -04:00
1 parent 3bd5e8883f
commit 3a3a20b8c9
11 files changed
+4506 -41

No files matched your search

+47 -38
View File
@@ -206,49 +206,58 @@ Tursa/nsight-sys
--------------------------------------------------------------------
============================================================================
2026-08-28 libfabric memory-registration-cache "memhooks" monitor -- DEFECTIVE
(libfabric issue #11451, filed by PB, HPE JIRA opened)
Piz Daint
============================================================================
STATUS: the ALPS/aarch64 case below is DEMONSTRATED (issue + both fixes verified
there). The Frontier case is a HYPOTHESIS as of 2026-08-28: the NO_TRANSLATION
failure is observed, its attribution to the memhooks MR cache is by mechanism and
by analogy, and the fix is UNTESTED here. Test ladder, one knob per run, NRHS=12:
(1) FineSloppyComms=0 (2) FI_MR_CACHE_MONITOR=kdreg2 (3) FI_MR_CACHE_MAX_COUNT=0
(4) --disable-accelerator-aware-mpi build.
REVISIT this entry with the result; if (2) does not cure it, remove kdreg2 from the
job scripts' justification (it stays as OLCF's recommendation regardless).
2026-08-28 22:18, job 5372414: step (2) DID NOT cure it -- kdreg2 set (and the
environment already carried FI_MR_CACHE_MAX_COUNT=786432), sloppy comms OFF, still
NO_TRANSLATION at outer step 60 (Couter, Waitall count=2).
CORRECTION 2026-08-28 23:xx (fi_mr(3) man page): FI_MR_CACHE_MONITOR governs SYSTEM
memory only; device (HMEM_ROCR) registrations are monitored by
FI_MR_ROCR_CACHE_MONITOR_ENABLED=0|1. Grid's comms window is hipMalloc'd device
memory, so the kdreg2 step tested nothing relevant -- the ALPS PLT bug (#11451) and
the Frontier NO_TRANSLATION are the same *library* but different monitors.
Attributed by gdb on a hung rank (5371826): main thread in MPI_Waitall <- Grid
0x5d539c3 (StencilSendToRecvFromComplete/CommsComplete class), i.e. the HALO
EXCHANGE; the two aborts' request counts (16 = fine Dhop, 2 = coarse PaddedCell
direction) agree. Next: FI_MR_CACHE_MAX_COUNT=0 (all memory) and
FI_MR_ROCR_CACHE_MONITOR_ENABLED=0, one per cell.
Symptom (Frontier, x86, Cray MPICH 8.1.x, ROCm 7.2): device-buffer MPI fails
with
MPI_Waitall ... MPIDI_OFI_handle_cq_error: OFI poll failed
(ofi_events.c:MPIDI_OFI_handle_cq_error: Input/output error - NO_TRANSLATION)
on a LIVE, never-freed hipMalloc buffer (the sloppy-comms compressed halo buffer,
a static deviceVector) once the process does sustained hipMalloc/hipFree churn
(MemoryManager eviction at NRHS>=12, EvictAll/DropCache). Translations cached by
the provider go stale.
Symptom (ALPS/CSCS, aarch64 Grace/H200, cray-mpich 8.1.32, libfabric 1.22): the
https://github.com/ofiwg/libfabric/issues/11451
SEGFAULT in first MPI_Comm_dup on more than one node, after MPI_Init
Cause: symptom (ALPS/CSCS, aarch64 Grace/H200, cray-mpich 8.1.32, libfabric 1.22): the
memhooks monitor intercepts munmap by PATCHING THE PLT at MPI_Init
(ofi_memhooks_start -> ofi_write_patch with a garbage data_size); the write overruns
munmap@plt into the NEXT PLT entry (MPI_Comm_dup in Grid's binary), leaving
`br x15` where an adrp belongs -> segfault on first MPI_Comm_dup. Diagnosed with
a hardware watchpoint on the PLT entry during MPI_Init; see the issue for the
munmap@plt into the NEXT PLT entry (MPI_Comm_dup in Grid's binary)
leaving
`br x15` where an adrp belongs ->
Diagnosed with a hardware watchpoint on the PLT entry during MPI_Init; see the issue for the
gdb transcript. Reproduce on one node with MPICH_SINGLE_HOST_ENABLED=0.
Fix (either): export FI_MR_CACHE_MONITOR=kdreg2 (kernel-driven invalidation; OLCF's
own recommendation, NOT the default)
export FI_MR_CACHE_MAX_COUNT=0 (no registration cache at all)
Fix (either):
export FI_MR_CACHE_MONITOR=kdreg2
OR
export FI_MR_CACHE_MONITOR=disabled
export FI_MR_CACHE_MAX_COUNT=0
On ALPS this is a CORRECTNESS requirement for MPI on Slingshot, not a tuning.
============================================================================
Frontier
============================================================================
a) FI_MR_CACHE_MONITOR=kdreg2 leads to Runtime MPI errors with "NO_TRANSLATION"
Symptom (Frontier, x86, Cray MPICH 8.1.x, ROCm 7.2): device-buffer MPI fails with
MPI_Waitall ... MPIDI_OFI_handle_cq_error: OFI poll failed
(ofi_events.c:MPIDI_OFI_handle_cq_error: Input/output error - NO_TRANSLATION)
on a LIVE, never-freed hipMalloc buffer
b) export FI_MR_CACHE_MONITOR=disabled
export FI_MR_CACHE_MAX_COUNT=0
Leads to INCORRECT RESULTS
https://github.com/ofiwg/libfabric/issues/12773
https://github.com/ofiwg/libfabric/issues/12775
Fix:
export FI_HMEM_ROCR_USE_DMABUF=0
Every systems/Frontier job and sourceme now sets kdreg2 on OLCF's recommendation;
whether it resolves the Frontier NO_TRANSLATION is the pending test above.