Barrel shift -> all to all (x2) and distributed FFT work fully load balanced without redundant work.
There is little more I can do now on FFT. Comms dominated and running distributed work dividing bandwidth optimal RingAllToAll
/ccs/home/paboyle/ParallelIO/systems/Frontier/tests/core/Test_fft_prop --mpi 3.6.4.4 --grid 48.48.48.96 --accelerator-threads 8 --shm 4096 --shm-mpi 1 --device-mem
32000 --log Error,Warning,Message,Performance
*************************************************
Benchmarking FFT of LatticeFermionD on plane wave
*************************************************
Grid : Performance : 0.501524 s : FFT took 0.001311 s (transpose P=3)
Grid : Performance : 0.501531 s : FFT pack 5.9e-05 s
Grid : Performance : 0.501533 s : FFT alltoall 0.000828 s
Grid : Performance : 0.501534 s : FFT reorder 0.000204 s
Grid : Performance : 0.501535 s : FFT kernels 1e-05 s
Grid : Performance : 0.501536 s : FFT unpack 5e-05 s
Grid : Performance : 0.509992 s : FFT took 0.001829 s (transpose P=6)
Grid : Performance : 0.510000 s : FFT pack 6e-05 s
Grid : Performance : 0.510002 s : FFT alltoall 0.001436 s
Grid : Performance : 0.510003 s : FFT reorder 0.000202 s
Grid : Performance : 0.510005 s : FFT kernels 9e-06 s
Grid : Performance : 0.510006 s : FFT unpack 5.3e-05 s
Grid : Performance : 0.517690 s : FFT took 0.001599 s (transpose P=4)
Grid : Performance : 0.517698 s : FFT pack 6e-05 s
Grid : Performance : 0.517700 s : FFT alltoall 0.001258 s
Grid : Performance : 0.517701 s : FFT reorder 0.0002 s
Grid : Performance : 0.517702 s : FFT kernels 9e-06 s
Grid : Performance : 0.517703 s : FFT unpack 4.9e-05 s
Grid : Performance : 0.524858 s : FFT took 0.001561 s (transpose P=4)
Grid : Performance : 0.524865 s : FFT pack 5.8e-05 s
Grid : Performance : 0.524867 s : FFT alltoall 0.001213 s
Grid : Performance : 0.524868 s : FFT reorder 0.000209 s
Grid : Performance : 0.524869 s : FFT kernels 8e-06 s
Grid : Performance : 0.524870 s : FFT unpack 4.9e-05 s
*************************************************
FFT of [48 48 48 96] LatticeFermionD took 0.030916 s
*************************************************