Files
Grid/systems/WorkArounds.txt
T
2026-09-29 14:58:18 -04:00

375 lines
12 KiB
Plaintext

The purpose of this file is to collate all non-obvious known magic shell variables
and compiler flags required for either correctness or performance on various systems.
A repository of work-arounds.
Contents:
1. Interconnect + MPI
2. Compilation
3. Profiling
************************
* 0. MPI IO
************************
--------------------------------------------------------------------
MPI2-IO correctness: force OpenMPI to use the MPICH romio implementation for parallel I/O
--------------------------------------------------------------------
export OMPI_MCA_io=romio321
--------------------------------------
ROMIO fail with > 2GB per node read (32 bit issue)
--------------------------------------
Use later MPICH
https://github.com/paboyle/Grid/issues/381
https://github.com/pmodels/mpich/commit/3a479ab0
--------------------------------------------------------------------
MPI2-IO performance: introduced All to all based aggregation no changes required.
--------------------------------------------------------------------
https://github.com/pmodels/mpich/issues/7935
MPI_File_write_all / read_all collective buffering does not scale for 4D block-decomposed subarray views (Lustre): 6x slower than a user-level MPI_Alltoallv + POSIX transposition
/* phase 1: redistribute within the row subcommunicator.
sendcounts/senddispls derived analytically from the plan; iodata untouched. */
MPI_Alltoallv(iodata, sendcounts, senddispls, MPI_BYTE,
aggregated, recvcounts, recvdispls, MPI_BYTE, rowcomm);
/* phase 2: a handful of large contiguous writes per rank, no MPI-IO involved */
for (e = 0; e < nextent; e++) {
lseek(fd, offset + extentGsite[e]*sizeof(fobj), SEEK_SET);
write(fd, aggregated + extentLocal[e], extentSites[e]*sizeof(fobj));
}
************************
* 1. INTERCONNECT + MPI
************************
============================================================================
Piz Daint
============================================================================
https://github.com/ofiwg/libfabric/issues/11451
SEGFAULT in first MPI_Comm_dup on more than one node, after MPI_Init
Cause: symptom (ALPS/CSCS, aarch64 Grace/H200, cray-mpich 8.1.32, libfabric 1.22): the
memhooks monitor intercepts munmap by PATCHING THE PLT at MPI_Init
(ofi_memhooks_start -> ofi_write_patch with a garbage data_size); the write overruns
munmap@plt into the NEXT PLT entry (MPI_Comm_dup in Grid's binary)
leaving
`br x15` where an adrp belongs ->
Diagnosed with a hardware watchpoint on the PLT entry during MPI_Init; see the issue for the
gdb transcript. Reproduce on one node with MPICH_SINGLE_HOST_ENABLED=0.
Fix (either):
export FI_MR_CACHE_MONITOR=kdreg2
OR
export FI_MR_CACHE_MONITOR=disabled
export FI_MR_CACHE_MAX_COUNT=0
On ALPS this is a CORRECTNESS requirement for MPI on Slingshot, not a tuning.
============================================================================
Frontier
============================================================================
a) FI_MR_CACHE_MONITOR=kdreg2 (site default) leads to Runtime MPI errors with "NO_TRANSLATION"
userfaultfd and memhooks also fail.
https://github.com/ofiwg/libfabric/issues/12775
Symptom (Frontier, x86, Cray MPICH 8.1.x, ROCm 7.2): device-buffer MPI fails with
MPI_Waitall ... MPIDI_OFI_handle_cq_error: OFI poll failed
(ofi_events.c:MPIDI_OFI_handle_cq_error: Input/output error - NO_TRANSLATION)
on a LIVE, never-freed hipMalloc buffer
b) export FI_MR_CACHE_MONITOR=disabled
export FI_MR_CACHE_MAX_COUNT=0
Leads to INCORRECT RESULTS, see
https://github.com/ofiwg/libfabric/issues/12773
Fix:
export FI_HMEM_ROCR_USE_DMABUF=0
and use site default
FI_MR_CACHE_MONITOR=kdreg2
Every systems/Frontier job and sourceme now sets kdreg2 on OLCF's recommendation;
whether it resolves the Frontier NO_TRANSLATION is the pending test above.
--------------------------------------------------------------------
Slingshot: Frontier and Perlmutter libfabric slow down
and physical memory fragmentation
--------------------------------------------------------------------
export FI_MR_CACHE_MONITOR=disabled
or
export FI_MR_CACHE_MONITOR=kdreg2
--------------------------------------------------------------------
Perlmutter
--------------------------------------------------------------------
export MPICH_RDMA_ENABLED_CUDA=1
export MPICH_GPU_IPC_ENABLED=1
export MPICH_GPU_EAGER_REGISTER_HOST_MEM=0
export MPICH_GPU_NO_ASYNC_MEMCPY=0
--------------------------------------------------------------------
Frontier/LumiG
--------------------------------------------------------------------
Hiding ROCR_VISIBLE_DEVICES triggers SDMA engines to be used for GPU-GPU
cat << EOF > select_gpu
#!/bin/bash
export MPICH_GPU_SUPPORT_ENABLED=1
export MPICH_SMP_SINGLE_COPY_MODE=XPMEM
export GPU_MAP=(0 1 2 3 7 6 5 4)
export NUMA_MAP=(3 3 1 1 2 2 0 0)
export GPU=\${GPU_MAP[\$SLURM_LOCALID]}
export NUMA=\${NUMA_MAP[\$SLURM_LOCALID]}
export HIP_VISIBLE_DEVICES=\$GPU
unset ROCR_VISIBLE_DEVICES
echo RANK \$SLURM_LOCALID using GPU \$GPU
exec numactl -m \$NUMA -N \$NUMA \$*
EOF
chmod +x ./select_gpu
srun ./select_gpu BINARY
--------------------------------------------------------------------
Mellanox performance with A100 GPU (Tursa, Booster, Leonardo)
--------------------------------------------------------------------
export OMPI_MCA_btl=^uct,openib
export UCX_TLS=gdr_copy,rc,rc_x,sm,cuda_copy,cuda_ipc
export UCX_RNDV_SCHEME=put_zcopy
export UCX_RNDV_THRESH=16384
export UCX_IB_GPU_DIRECT_RDMA=yes
--------------------------------------------------------------------
Mellanox + A100 correctness (Tursa, Booster, Leonardo)
--------------------------------------------------------------------
export UCX_MEMTYPE_CACHE=n
--------------------------------------------------------------------
MPICH/Aurora/PVC correctness and performance
--------------------------------------------------------------------
https://github.com/pmodels/mpich/issues/7302
--enable-cuda-aware-mpi=no
--enable-unified=no
Grid's internal D-H-H-D pipeline mode, avoid device memory in MPI
Do not use SVM
Ideally use MPICH with fix to issue 7302:
https://github.com/pmodels/mpich/pull/7312
Ideally:
MPIR_CVAR_CH4_IPC_GPU_HANDLE_CACHE=generic
Alternatives:
export MPIR_CVAR_NOLOCAL=1
export MPIR_CVAR_CH4_IPC_GPU_P2P_THRESHOLD=1000000000
--------------------------------------------------------------------
MPICH/Aurora/PVC correctness and performance
--------------------------------------------------------------------
Broken:
export MPIR_CVAR_CH4_OFI_ENABLE_GPU_PIPELINE=1
This gives good peformance without requiring
--enable-cuda-aware-mpi=no
But is an open issue reported by James Osborn
https://github.com/pmodels/mpich/issues/7139
Possibly resolved but unclear if in the installed software yet.
************************
* 2. COMPILATION
************************
--------------------------------------------------------------------
G++ compiler breakage / graveyard
--------------------------------------------------------------------
9.3.0, 10.3.1,
https://github.com/paboyle/Grid/issues/290
https://github.com/paboyle/Grid/issues/264
Working (-) Broken (X):
4.9.0 -
4.9.1 -
5.1.0 X
5.2.0 X
5.3.0 X
5.4.0 X
6.1.0 X
6.2.0 X
6.3.0 -
7.1.0 -
8.0.0 (HEAD) -
https://github.com/paboyle/Grid/issues/100
--------------------------------------------------------------------
AMD GPU nodes :
--------------------------------------------------------------------
multiple ROCM versions broken; use 5.3.0
manifests itself as wrong results in fp32
https://github.com/paboyle/Grid/issues/464
--------------------------------------------------------------------
Aurora/PVC
--------------------------------------------------------------------
SYCL ahead of time compilation (fixes rare runtime JIT errors and faster runtime, PB)
SYCL slow link and relocatable code issues (Christoph Lehner)
Opt large register file required for good performance in fp64
export SYCL_PROGRAM_COMPILE_OPTIONS="-ze-opt-large-register-file"
export LDFLAGS="-fiopenmp -fsycl -fsycl-device-code-split=per_kernel -fsycl-targets=spir64_gen -Xs -device -Xs pvc -fsycl-device-lib=all -lze_loader -L${MKLROOT}/lib -qmkl=parallel -fsycl -lsycl -fPIC -fsycl-max-parallel-link-jobs=16 -fno-sycl-rdc"
export CXXFLAGS="-O3 -fiopenmp -fsycl-unnamed-lambda -fsycl -Wno-tautological-compare -qmkl=parallel -fsycl -fno-exceptions -fPIC"
--------------------------------------------------------------------
Aurora/PVC useful extra options
--------------------------------------------------------------------
Host only sanitizer:
-Xarch_host -fsanitize=leak
-Xarch_host -fsanitize=address
Deterministic MPI reduction:
export MPIR_CVAR_ALLREDUCE_DEVICE_COLLECTIVE=0
export MPIR_CVAR_REDUCE_DEVICE_COLLECTIVE=0
export MPIR_CVAR_ALLREDUCE_INTRA_ALGORITHM=recursive_doubling
unset MPIR_CVAR_CH4_COLL_SELECTION_TUNING_JSON_FILE
unset MPIR_CVAR_COLL_SELECTION_TUNING_JSON_FILE
unset MPIR_CVAR_CH4_POSIX_COLL_SELECTION_TUNING_JSON_FILE
************************
* 3. Visual profile tools
************************
--------------------------------------------------------------------
Frontier/rocprof
--------------------------------------------------------------------
cat << EOF > select_gpu
#!/bin/bash
export GPU_MAP=(0 1 2 3 7 6 5 4)
export NUMA_MAP=(3 3 1 1 2 2 0 0)
export GPU=\${GPU_MAP[\$SLURM_LOCALID]}
export NUMA=\${NUMA_MAP[\$SLURM_LOCALID]}
export HIP_VISIBLE_DEVICES=\$GPU
unset ROCR_VISIBLE_DEVICES
if [ \$SLURM_PROCID = "0" ]; then echo \$*; fi
if [ "1" = "1" ]
then
if [ \$SLURM_PROCID = "0" ]
then
exec numactl -m \$NUMA -N \$NUMA rocprofv3 -d $TRACE -o mgtrace.\$SLURM_JOBID \
-r -f csv pftrace --stats -D \
-P 140:20:1 \
-- \$*
else
exec numactl -m \$NUMA -N \$NUMA \$*
fi
else
exec numactl -m \$NUMA -N \$NUMA \$*
fi
EOF
chmod +x ./select_gpu
--------------------------------------------------------------------
Aurora/unitrace
--------------------------------------------------------------------
select_gpu:
#!/bin/bash
export NUMA_PMAP=(0 0 0 1 1 1 0 0 0 1 1 1 );
export NUMA_HMAP=(2 2 2 3 3 3 3 2 2 2 2 3 3 3 );
export GPU_MAP=(0.0 1.0 2.0 3.0 4.0 5.0 0.1 1.1 2.1 3.1 4.1 5.1 )
export NUMAP=${NUMA_PMAP[$PALS_LOCAL_RANKID]}
export NUMAH=${NUMA_HMAP[$PALS_LOCAL_RANKID]}
export gpu_id=${GPU_MAP[$PALS_LOCAL_RANKID]}
unset EnableWalkerPartition
export EnableImplicitScaling=0
export ZE_AFFINITY_MASK=$gpu_id
export ONEAPI_DEVICE_FILTER=gpu,level_zero
export SYCL_PI_LEVEL_ZERO_DEVICE_SCOPE_EVENTS=0
export SYCL_PI_LEVEL_ZERO_USE_IMMEDIATE_COMMANDLISTS=1
export SYCL_PI_LEVEL_ZERO_USE_COPY_ENGINE=0:4
export SYCL_PI_LEVEL_ZERO_USE_COPY_ENGINE_FOR_D2D_COPY=1
echo "rank $PALS_RANKID ; local rank $PALS_LOCAL_RANKID ; ZE_AFFINITY_MASK=$ZE_AFFINITY_MASK ; NUMA $NUMA "
if [ $PALS_RANKID = "0" ]
then
numactl -p $NUMAP -N $NUMAP unitrace --chrome-kernel-logging --chrome-mpi-logging --chrome-sycl-logging --demangle "$@"
else
numactl -p $NUMAP -N $NUMAP "$@"
fi
--------------------------------------------------------------------
Tursa/nsight-sys
--------------------------------------------------------------------
cat << EOF > mpiwrapper.sh
#!/bin/bash
lrank=\$OMPI_COMM_WORLD_LOCAL_RANK
numa1=\$(( 2 * \$lrank ))
numa2=\$(( 2 * \$lrank + 1 ))
export CUDA_VISIBLE_DEVICES=\$lrank
export UCX_NET_DEVICES=mlx5_\${lrank}:1
if [ \$OMPI_COMM_WORLD_RANK = "0" ]; then echo \$*; fi
if [ "1" = "1" ]
then
if [ \$OMPI_COMM_WORLD_RANK = "0" ]
then
exec numactl --interleave=\$numa1,\$numa2 nsys profile \
-o $TRACE/nsys.\$SLURM_JOB_ID --force-overwrite=true \
--trace=cuda,nvtx,mpi,osrt --mpi-impl=op
--sample=none --cpuctxsw=none \
--delay=140 --duration=20 --stats=true \
\$*
else
exec numactl --interleave=\$numa1,\$numa2 \$*
fi
else
exec numactl --interleave=\$numa1,\$numa2 \
fi
EOF
chmod +x ./mpiwrapper.sh
mpirun -np $SLURM_NTASKS -x LD_LIBRARY_PATH --bind-to none ./mpiwrapper.sh <binary> <args>