Barrel shift -> all to all (x2) and distributed FFT work fully load balanced without redundant work.
There is little more I can do now on FFT. Comms dominated and running distributed work dividing bandwidth optimal RingAllToAll
/ccs/home/paboyle/ParallelIO/systems/Frontier/tests/core/Test_fft_prop --mpi 3.6.4.4 --grid 48.48.48.96 --accelerator-threads 8 --shm 4096 --shm-mpi 1 --device-mem
32000 --log Error,Warning,Message,Performance
*************************************************
Benchmarking FFT of LatticeFermionD on plane wave
*************************************************
Grid : Performance : 0.501524 s : FFT took 0.001311 s (transpose P=3)
Grid : Performance : 0.501531 s : FFT pack 5.9e-05 s
Grid : Performance : 0.501533 s : FFT alltoall 0.000828 s
Grid : Performance : 0.501534 s : FFT reorder 0.000204 s
Grid : Performance : 0.501535 s : FFT kernels 1e-05 s
Grid : Performance : 0.501536 s : FFT unpack 5e-05 s
Grid : Performance : 0.509992 s : FFT took 0.001829 s (transpose P=6)
Grid : Performance : 0.510000 s : FFT pack 6e-05 s
Grid : Performance : 0.510002 s : FFT alltoall 0.001436 s
Grid : Performance : 0.510003 s : FFT reorder 0.000202 s
Grid : Performance : 0.510005 s : FFT kernels 9e-06 s
Grid : Performance : 0.510006 s : FFT unpack 5.3e-05 s
Grid : Performance : 0.517690 s : FFT took 0.001599 s (transpose P=4)
Grid : Performance : 0.517698 s : FFT pack 6e-05 s
Grid : Performance : 0.517700 s : FFT alltoall 0.001258 s
Grid : Performance : 0.517701 s : FFT reorder 0.0002 s
Grid : Performance : 0.517702 s : FFT kernels 9e-06 s
Grid : Performance : 0.517703 s : FFT unpack 4.9e-05 s
Grid : Performance : 0.524858 s : FFT took 0.001561 s (transpose P=4)
Grid : Performance : 0.524865 s : FFT pack 5.8e-05 s
Grid : Performance : 0.524867 s : FFT alltoall 0.001213 s
Grid : Performance : 0.524868 s : FFT reorder 0.000209 s
Grid : Performance : 0.524869 s : FFT kernels 8e-06 s
Grid : Performance : 0.524870 s : FFT unpack 4.9e-05 s
*************************************************
FFT of [48 48 48 96] LatticeFermionD took 0.030916 s
*************************************************
Per log2-size bucket: messages, GB, seconds inside SendToRecvFrom, GB/s,
% of ring time. Decomposes the 2 GB/s/rank average (probe: 11-20 GB/s
at 8 MB) into latency-bound small messages vs slow large ones vs wait.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RAhdQHrkzKmzfxpxW1whLn
DENSE_DEVICE_SUM=4 assumed each rank's rows of x were the contiguous
[me*nrows,...) block; x and the slab columns are in global-site order
(hX[myGsite*nbasis+b]), the contiguous block is the rank-major index.
Coincide on one rank only -> Frontier VERIFY 0.9965. Stage via myGsite,
scatter through the inverse of BuildRankMajorMap. 4-rank laptop VERIFY
7.38e-07, identical to the allreduce path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RAhdQHrkzKmzfxpxW1whLn
Arm complex instructions on M3/M4 NEON v8.3 and simplify A64FX code paths/broaden.
Grid_vector_types and Simd.h mainly reorg and prep for sComplexD and sComplexF alternat Nsimd=1 types