May finally clean up the poor MPI2 IO performance issue that has been persistent.
10 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
What This Is
Grid is a data-parallel C++ library for lattice QCD. It provides SIMD-vectorised lattice containers, MPI-based domain decomposition, GPU acceleration (CUDA/HIP/SYCL), and a full suite of QCD algorithms including HMC.
Build
Uses GNU Autotools. The bootstrap step only needs to run once (or after configure.ac changes).
./bootstrap.sh # downloads Eigen 3.4.0, generates configure
mkdir build && cd build
../configure [options]
make -j$(nproc)
make check # run root-level tests
make install
Key configure options:
| Option | Common values |
|---|---|
--enable-simd= |
AVX2, AVX512, KNL, A64FX, NEONv8, GPU |
--enable-comms= |
mpi-auto, mpi3-auto, none |
--enable-accelerator= |
cuda, hip, sycl |
--enable-shm= |
shmopen, hugetlbfs, nvlink |
--enable-Nc= |
3 (default), 2, 4, 5 |
--with-gmp=, --with-mpfr=, --with-fftw=, --with-lime= |
paths to libs |
--enable-hdf5, --enable-mkl, --enable-lapack |
optional features |
Platform recipes from README.md:
- KNL:
--enable-simd=KNL --enable-comms=mpi3-auto --enable-mkl - Skylake/Haswell:
--enable-simd=AVX512orAVX2+--enable-comms=mpi3-auto - AMD EPYC:
--enable-simd=AVX2 --enable-comms=mpi3 - A64FX (Fugaku):
--enable-simd=A64FX --enable-comms=mpi3 --enable-shm=shmget(seeSVE_README.txt)
Required external libs: GMP, MPFR, OpenSSL, zlib.
Use systems/ for real machines
systems/<machine>/ holds the known-good build for each production platform (Frontier, Aurora, Perlmutter, Summit, Tursa, Lumi, Booster, Crusher, SDCC-*, mac-arm, …). Each contains a config-command (the exact ../../configure invocation) and a sourceme.sh (module loads and env). Prefer copying/adapting these over hand-rolling configure flags — they encode compiler workarounds, LDFLAGS, and shared-memory settings that are easy to get wrong. systems/WorkArounds.txt records known vendor bugs.
Note the GPU builds use --enable-simd=GPU --enable-gen-simd-width=64, so Nsimd is not 1 on device (it is 64/sizeof(scalar)).
Regenerating Make.inc — required after adding or deleting source files
Make.inc files are generated, not tracked in git (.gitignored). scripts/filelist walks Grid/, tests/*, benchmarks/, examples/, and HMC/ and writes the file lists and per-test bin_PROGRAMS rules. Every new .cc/.h in Grid/, and every new Test_*.cc / Benchmark_*.cc / Example_*.cc, is invisible to the build until you run:
./scripts/filelist # from the source root, then re-run configure/make
bootstrap.sh runs it for you on the first setup.
Running Tests and Benchmarks
# From build directory
make check # runs only Test_simd, Test_cshift, Test_stencil, Test_dwf_mixedcg_prec
make -C tests/<subdir> tests # build (not run) the tests in a subdirectory
make bench # build benchmarks
./tests/core/Test_simd # run a single test binary directly
mpirun -n 4 ./tests/core/Test_cshift --grid 16.16.16.16 --mpi 1.1.1.4
make check is a thin smoke test — building a subdirectory with make -C tests/<subdir> tests and running the relevant binaries directly is the normal development loop. Test binaries take Grid's standard command-line arguments (--grid, --mpi, --accelerator-threads, --threads, --debug-signals, --log); see Grid/util/Init.cc.
Test subdirectories and their focus: core (SIMD, stencil, comms), solver (CG, GMRES, eigensolvers), hmc (MD integrators), forces (fermion forces), lanczos, IO, smearing, sp2n, debug.
Tests and benchmarks that need optional fermion representations are guarded by disable_tests_without_instantiations.h / disable_benchmarks_without_instantiations.h, so a --disable-fermion-reps --disable-gparity build silently compiles them to no-ops.
Architecture
Layer stack (bottom to top)
-
SIMD layer (
Grid/simd/) — platform-specific intrinsics wrapped intovRealF,vComplexD, etc. The SIMD width and layout are compile-time constants controlled by--enable-simd. -
Tensor layer (
Grid/tensors/) — Lorentz/colour/spin tensor algebra built on top of SIMD types.iMatrix,iVector,iScalartemplates compose into QCD types likeColourMatrix,SpinColourVector. -
Lattice layer (
Grid/lattice/) —Lattice<T>container: a site-local tensor replicated across a distributed Cartesian grid. All arithmetic is site-parallel and expression-template-fused. -
Cartesian/comms layer (
Grid/cartesian/,Grid/communicator/) —GridCartesianholds the MPI topology and local/global geometry.Grid/cshift/implements nearest-neighbour halo exchange;Grid/stencil/is the optimised multi-hop stencil used by Dirac operators. -
Algorithm layer (
Grid/algorithms/) — iterative solvers (CG, GMRES, BiCGSTAB, mixed-precision), eigensolvers (Lanczos, LAPACK), FFT, smearing. -
QCD layer (
Grid/qcd/) — gauge and fermion actions, HMC integrators, observables.
QCD subsystem (Grid/qcd/)
action/fermion/— Wilson, Clover, DWF (Mobius), Staggered, twisted-mass, G-parity variantsaction/gauge/— Wilson gauge, Symanzik, Iwasaki, DBW2, plaquette+rectrepresentations/— Fundamental, Adjoint, Two-index, Sp(2n)hmc/— Leapfrog, OMF2/OMF4 integrators; pseudofermion refreshment; Metropolis accept/rejectsmearing/— APE, Stout, HEX, gradient flowobservables/— Polyakov loop, plaquette, topological charge
GPU acceleration and the view/memory-manager discipline
GPU support is injected via macros in Grid/threads/Accelerator.h — accelerator_for(i, n, nsimd, {...}), accelerator_forNB (non-blocking, must be followed by accelerator_barrier()), accelerator_for2dNB, and accelerator_inline. On a CPU build these degrade to thread_for (OpenMP). Unified virtual memory is on by default (--enable-unified=yes); device-aware MPI (--enable-accelerator-aware-mpi) avoids device→host copies on transfers.
Lattice data is not directly addressable inside a kernel. You must open a view with the correct access mode so Grid/allocator/MemoryManager.h can move/mark the data:
autoView(out_v, out, AcceleratorWriteDiscard); // RAII; closes at end of scope
autoView(in_v, in, AcceleratorRead);
accelerator_for(ss, grid->oSites(), Nsimd, {
coalescedWrite(out_v[ss], coalescedRead(in_v[ss]));
});
Modes are AcceleratorRead/Write/WriteDiscard and CpuRead/Write/WriteDiscard. Getting the mode wrong (e.g. AcceleratorRead on a field you write) produces stale-data bugs that only appear on GPU builds. Inside kernels use coalescedRead/coalescedWrite rather than raw operator[] — they map the SIMD lane onto threadIdx.x so accesses stay coalesced.
Repo-local debugging skills (skills/)
skills/ contains hard-won, Grid-specific playbooks written as invocable skill files. Consult them before debugging in these areas rather than reasoning from first principles:
| File | Covers |
|---|---|
gpu-memory-performance.md |
acceleratorThreads(), LambdaApply thread mapping, coalescedRead idiom, fused vs staged HBM access |
gpu-runtime-correctness.md |
GPU runtime returning early from sync, silent wrong answers |
communication-overlap.md |
7-phase halo pipeline, per-packet events, host-staging vs GPU-direct RDMA |
mpi-heterogeneous.md |
MPI_Sendrecv device-buffer aliasing, deterministic reductions |
compiler-validation.md |
Isolating GPU compiler codegen bugs, minimal reproducers |
correctness-verification.md |
Double-run fingerprinting, per-packet checksums, flight recorder |
hang-diagnosis.md |
Diagnosing MPI/accelerator hangs |
Memory and I/O
Grid/allocator/— aligned/NUMA-aware allocators; caching allocator via--enable-alloc-cacheGrid/parallelIO/— distributed parallel reader/writer for ILDG (via LIME), SciDAC, and native binary formatsGrid/serialisation/— text, binary, HDF5, XML/JSON serialisation of arbitrary Grid objects
Executables
HMC/— production HMC driver programmes (e.g.Mobius2p1f.cc,DWF_plus_DSDR_nf2plus1_Shamir_Gparity.cc)benchmarks/—Benchmark_dwf,Benchmark_ITT,Benchmark_comms,Benchmark_memory_bandwidth, … used to qualify a new machineexamples/— small, readable programmes (Example_plaquette.cc,Example_Mobius_spectrum.cc) that are the best starting point for learning the API
Each of these directories auto-builds every top-level .cc as its own binary via scripts/filelist.
Every programme is wrapped in Grid_init(&argc, &argv) / Grid_finalize() (Grid/util/Init.h).
Key Conventions
- C++17 is required throughout.
- Template structure: most classes are templated on
<_FImpl>(fermion impl) or<Gimpl>(gauge impl), which encode the representation and precision. Instantiation is controlled by--enable-fermion-instantiations; the explicit instantiation.ccfiles live underGrid/qcd/action/fermion/instantiation/<Impl>/and are selected byscripts/filelist. - The
RealD/RealF/ComplexD/ComplexFtypedefs are used everywhere; avoid rawdouble/float. - Use
GRID_ASSERT(cond)(defined inGrid/GridStd.h), not bareassert— it prints a Grid-formatted message and aborts cleanly under MPI. - Logging is stream-based, not macro-based:
std::cout << GridLogMessage << ... << std::endl;. Channels declared inGrid/log/Log.hincludeGridLogError,GridLogWarning,GridLogDebug,GridLogPerformance,GridLogIterative,GridLogSolver,GridLogHMC,GridLogComms,GridLogMemory,GridLogDslash,GridLogIRL,GridLogMG. A subset is switched on at runtime with e.g.--log Error,Warning,Message,Performance,Iterative,Integrator,Debug,Colours(names given without theGridLogprefix). - Performance-critical paths use
GRID_TRACE(name)fromGrid/perfmon/Tracing.h(compiled out unless--enable-tracingselects a backend) and theGridStopWatchtimers inGrid/perfmon/Timer.h. - Reductions across MPI ranks go through
GridBase::GlobalSum/GlobalMax; never reduce with bare MPI calls inside library code. - Everything lives in
NAMESPACE_BEGIN(Grid)/NAMESPACE_END(Grid)macros; follow the surrounding file rather than writingnamespace Grid { }. - British spelling is used in identifiers and comments (
colour,neighbour,serialisation).