Commit Graph
8486 Commits
Author SHA1 Message Date
Peter Boyle 7a9cdb45bc Profiling split grid 2026-10-02 23:34:06 -04:00
Peter Boyle 44f24f395e TIming info. Will need to revert in future. 2026-10-02 19:46:15 -04:00
Peter Boyle 964f4c1271 Force propagation convenience 2026-10-02 12:06:39 -04:00
Peter Boyle af9d829336 Split testing 2026-10-02 12:06:23 -04:00
Peter Boyle b07ccf3d5e Force test reorg 2026-10-02 12:02:54 -04:00
Peter Boyle 260d7d2600 Be able to split operators for split CG 2026-10-02 12:01:38 -04:00
Peter Boyle 809249caa9 Latent bug fix, never hit in existing code 2026-10-02 12:01:01 -04:00
Peter Boyle 94f529bd63 Split operators 2026-10-02 12:00:37 -04:00
Peter Boyle e6d4daf194 Split operators 2026-10-02 12:00:12 -04:00
Peter Boyle 402fa0aace Complete the derivative 2026-10-02 11:58:53 -04:00
Peter Boyle d31061cd49 Batched split grid 2026-10-02 11:57:27 -04:00
Peter Boyle 3aa03d3b3f Overlap commms compute 2026-09-30 14:36:57 -04:00
Peter Boyle 87baaf1d37 Updated workaround 2026-09-29 14:58:18 -04:00
Peter Boyle 45c1a525b9 Useful script and keep the mixed precision job 2026-09-29 13:47:38 -04:00
Peter Boyle 3a3a20b8c9 Updates 2026-09-29 13:46:09 -04:00
Peter Boyle 3bd5e8883f Synch multigrid rewrite to repository and
broad unicode elimination effort
2026-09-29 12:36:34 -04:00
Peter Boyle 865c338dd2 Local coherence in the coarse space 2026-09-24 01:40:58 -04:00
Peter Boyle e583079ab9 This is not yet regressing to prior HDCG 2026-09-24 01:40:58 -04:00
Peter Boyle 507144c457 Clean organisation 2026-09-24 01:40:58 -04:00
Peter Boyle 251ac91b31 RocBLAS behaves different to HipBlas 2026-09-24 01:40:58 -04:00
Peter Boyle 41d458404a Deprecated functions 2026-09-24 01:39:08 -04:00
Peter Boyle eeb56e6b39 Clean up and unification of PVdagM and HDCG, mixed precision support 2026-09-24 01:38:39 -04:00
Peter Boyle 2043072d9c Merge pull request #493 from jdmaia/hip_stencil_padding
Fix alignment on StencilEntry struct for HIP
2026-09-11 17:32:30 -04:00
Julio Maia 487ec8d410 Fix alignment on StencilEntry struct for HIP
CUDA neatly aligns the StencilEntry struct to 16-bytes, which is not used on HIP.
Mapping the StencilEntry struct to 16-bytes on HIP yields around 5% better performance
on MI355X-class GPUs.
2026-09-11 10:56:02 -05:00
Peter Boyle 75bb6bdb11 Update for bug report 2026-09-09 18:12:52 -04:00
Peter Boyle a2ef8b39d3 New script for ORNL to try 2026-09-09 15:13:12 -04:00
Peter Boyle 03495bfcf8 Simplify 2026-09-09 14:51:40 -04:00
Peter Boyle 7f6b0f409c Better commenting 2026-09-09 14:39:56 -04:00
Peter Boyle ec682693c1 Rename to be more familial with PowerMethod 2026-09-09 14:20:24 -04:00
Peter Boyle f8f93c1443 Non herm case 2026-09-09 14:12:49 -04:00
Peter Boyle fbfb93af64 MG better organisation 2026-09-09 14:11:55 -04:00
Peter Boyle 0ac6b72783 Claude's reorg to clean up parameters 2026-09-08 15:53:52 -04:00
Peter Boyle e8a82fa683 Multigrid clean up phase, and fix shell variables for libfabric bug work around on Frontier 2026-09-08 15:52:15 -04:00
Peter Boyle a07545adc4 Preparing for multigrid parameter consolidation and clean up of code, rationalise the different variants. 2026-09-06 14:18:00 -04:00
Peter Boyle 482f3cbaa2 Real part comparisons 2026-09-05 08:05:03 -04:00
Peter Boyle a28f7ad531 Grids in FermionOperator 2026-09-03 23:18:06 -04:00
Peter Boyle 953a401366 Use PlannedFFT in momentum space propagator. 2026-09-03 20:59:42 -04:00
Peter Boyle 59f4a3729a FFT improvement by ~2x 2026-09-03 18:13:02 -04:00
Peter Boyle 357ede3664 Improved FFT -- 1.5-2x when there are 2-6 ranks in a given axis of the cartesian communicator.
Barrel shift -> all to all (x2) and distributed FFT work fully load balanced without redundant work.
There is little more I can do now on FFT. Comms dominated and running distributed work dividing bandwidth optimal RingAllToAll

/ccs/home/paboyle/ParallelIO/systems/Frontier/tests/core/Test_fft_prop --mpi 3.6.4.4 --grid 48.48.48.96 --accelerator-threads 8 --shm 4096 --shm-mpi 1 --device-mem
32000 --log Error,Warning,Message,Performance

*************************************************
 Benchmarking FFT of LatticeFermionD on plane wave
*************************************************
Grid : Performance : 0.501524 s :  FFT took     0.001311 s (transpose P=3)
Grid : Performance : 0.501531 s :  FFT pack     5.9e-05 s
Grid : Performance : 0.501533 s :  FFT alltoall 0.000828 s
Grid : Performance : 0.501534 s :  FFT reorder  0.000204 s
Grid : Performance : 0.501535 s :  FFT kernels  1e-05 s
Grid : Performance : 0.501536 s :  FFT unpack   5e-05 s
Grid : Performance : 0.509992 s :  FFT took     0.001829 s (transpose P=6)
Grid : Performance : 0.510000 s :  FFT pack     6e-05 s
Grid : Performance : 0.510002 s :  FFT alltoall 0.001436 s
Grid : Performance : 0.510003 s :  FFT reorder  0.000202 s
Grid : Performance : 0.510005 s :  FFT kernels  9e-06 s
Grid : Performance : 0.510006 s :  FFT unpack   5.3e-05 s
Grid : Performance : 0.517690 s :  FFT took     0.001599 s (transpose P=4)
Grid : Performance : 0.517698 s :  FFT pack     6e-05 s
Grid : Performance : 0.517700 s :  FFT alltoall 0.001258 s
Grid : Performance : 0.517701 s :  FFT reorder  0.0002 s
Grid : Performance : 0.517702 s :  FFT kernels  9e-06 s
Grid : Performance : 0.517703 s :  FFT unpack   4.9e-05 s
Grid : Performance : 0.524858 s :  FFT took     0.001561 s (transpose P=4)
Grid : Performance : 0.524865 s :  FFT pack     5.8e-05 s
Grid : Performance : 0.524867 s :  FFT alltoall 0.001213 s
Grid : Performance : 0.524868 s :  FFT reorder  0.000209 s
Grid : Performance : 0.524869 s :  FFT kernels  8e-06 s
Grid : Performance : 0.524870 s :  FFT unpack   4.9e-05 s
*************************************************
 FFT of [48 48 48 96] LatticeFermionD took 0.030916 s
*************************************************
2026-09-03 17:33:32 -04:00
Peter Boyle e767694b82 Better memory tracking 2026-08-30 20:16:31 -04:00
Peter Boyle 4c7953c2d1 Better memory logging 2026-08-30 20:16:17 -04:00
Peter Boyle 6c4634bc90 Command line arg 2026-08-30 10:42:10 -04:00
Peter Boyle f0a2c0c465 Changes to investigate lib fabric memory region cache fail on Frontier 2026-08-30 10:41:39 -04:00
Peter Boyle 01f504ca4b Updated test job 2026-08-29 09:08:15 -04:00
Peter Boyle 15e00edda2 POssible ROCM bug addressing 2026-08-29 01:15:12 -04:00
Peter Boyle 21b53c06d1 FI investigatins 2026-08-28 23:31:54 -04:00
Peter Boyle e03797e882 More FI_MR related 2026-08-28 22:01:36 -04:00
Peter Boyle 43c6573ca4 Memory manager update to drop cache 2026-08-28 17:03:41 -04:00
Peter Boyle f367bc8bce Solving memory pressure in NRHS=12 solver 2026-08-28 13:52:45 -04:00
Peter Boyle 63c2cdb712 Debug ulimit as core files driving me crazy 2026-08-28 11:20:43 -04:00