Peter Boyle
|
bc503b60e6
|
Offloadable gather code
|
2018-09-10 11:21:25 +01:00 |
|
Peter Boyle
|
6d0f1aabb1
|
Fix the multi-node path
|
2018-09-09 14:27:37 +01:00 |
|
Peter Boyle
|
28db0631ff
|
Hack to force 128bit accesses
|
2018-07-23 06:10:27 -04:00 |
|
Peter Boyle
|
b2b5137d28
|
Finally starting to get decent performance on Volta
|
2018-07-13 12:06:18 -04:00 |
|
Peter Boyle
|
c0e8bc9da9
|
Current version gets 250 - 320 GF/s on Volta on the target 12^4 volume.
|
2018-07-05 07:10:25 -04:00 |
|
Peter Boyle
|
b1265ae867
|
Prettify code
|
2018-07-05 07:08:06 -04:00 |
|
Peter Boyle
|
32bb85ea4c
|
Standard extractLane is fast
|
2018-07-05 07:07:30 -04:00 |
|
Peter Boyle
|
ca0607b6ef
|
Clearer kernel call meaning
|
2018-07-05 07:06:15 -04:00 |
|
paboyle
|
3a50afe7e7
|
GPU dslash updates
|
2018-06-27 22:32:21 +01:00 |
|
paboyle
|
3e947527cb
|
Move looping over "s" and "site" into kernels for GPU optimisatoin
|
2018-06-27 21:29:43 +01:00 |
|
paboyle
|
31f65beac8
|
Move site and Ls looping into the kernels
|
2018-06-27 21:28:48 +01:00 |
|
paboyle
|
38e2a32ac9
|
Single SIMD lane operations for CUDA
|
2018-06-27 21:28:06 +01:00 |
|
paboyle
|
efa84ca50a
|
Keep Cuda 9.1 happy
|
2018-06-27 21:27:32 +01:00 |
|
paboyle
|
5e96d6d04c
|
Keep CUDA happy
|
2018-06-27 21:27:11 +01:00 |
|
paboyle
|
df30bdc599
|
CUDA happy
|
2018-06-27 21:26:49 +01:00 |
|
paboyle
|
6c97a6a071
|
Coalescing version of the kernel
|
2018-06-13 20:52:29 +01:00 |
|
paboyle
|
73bb2d5128
|
Ugly hack to speed up compile on GPU; we don't use the hand kernels on GPU anyway so why compile
|
2018-06-13 20:35:28 +01:00 |
|
paboyle
|
b710fec6ea
|
Gpu code first version of specialised kernel
|
2018-06-13 20:34:39 +01:00 |
|
paboyle
|
b2a8cd60f5
|
Doubled gauge field is useful
|
2018-06-13 20:27:47 +01:00 |
|
paboyle
|
867ee364ab
|
Explicit instantiation hooks
|
2018-06-13 20:27:12 +01:00 |
|
Peter Boyle
|
eb7d34a4cc
|
GPU version
|
2018-05-14 19:41:47 -04:00 |
|
Peter Boyle
|
aab27a655a
|
Start of GPU kernels
|
2018-05-14 19:41:17 -04:00 |
|
Peter Boyle
|
13f50406e3
|
Suppress print statement
|
2018-05-12 18:00:00 -04:00 |
|
Peter Boyle
|
b15db11c60
|
Kernels -> pure static object to enable device execution
|
2018-03-24 19:35:20 -04:00 |
|
Peter Boyle
|
f6077f9d48
|
Kernels -> not instantiaed otherwise object ref on GPU
|
2018-03-24 19:33:44 -04:00 |
|
Peter Boyle
|
572954ef12
|
Kernels not an instantiated object, just static
|
2018-03-24 19:33:13 -04:00 |
|
Peter Boyle
|
cedeaae7db
|
Lebesge -> StencilView if necessary
|
2018-03-24 19:32:41 -04:00 |
|
Peter Boyle
|
e6cf0b1e17
|
View typedefs go to OperatorImpl
|
2018-03-24 19:32:11 -04:00 |
|
Peter Boyle
|
1f70cedbab
|
Have to make all kernel called routines static since object reference will be a host pointer on GPU
|
2018-03-24 19:29:26 -04:00 |
|
Peter Boyle
|
b50f37cfb4
|
Remove overlap comms flag
|
2018-03-24 19:28:53 -04:00 |
|
Peter Boyle
|
4e1272fabf
|
Kernels need to be static to work on GPU. No reference to host resident data
|
2018-03-22 18:44:53 -04:00 |
|
Peter Boyle
|
607dc2d3c6
|
Remove lebesgue order
|
2018-03-22 18:23:09 -04:00 |
|
Peter Boyle
|
23c880b009
|
Remove lebesgue order; stick in stencil if need
|
2018-03-22 18:13:41 -04:00 |
|
Peter Boyle
|
334bb6792f
|
Lebesgue order removed. Stick in the stencil view
|
2018-03-22 18:12:12 -04:00 |
|
Peter Boyle
|
8a1d303ab9
|
GPU friendly stencil improvements
|
2018-03-19 07:11:03 -04:00 |
|
Peter Boyle
|
bf0a4de919
|
GPU friendly params object
|
2018-03-19 07:10:12 -04:00 |
|
paboyle
|
4d60b92b7f
|
Update oSites
|
2018-03-08 21:00:25 +00:00 |
|
paboyle
|
c159c70c84
|
View introduced
|
2018-03-08 14:58:04 +00:00 |
|
paboyle
|
28b5572755
|
Merge branch 'feature/gpu-port' of https://github.com/paboyle/Grid into feature/gpu-port
|
2018-03-08 13:01:42 +00:00 |
|
Peter Boyle
|
4548523ecc
|
This modification eliminates what looks like a compiler bug
on Intel 2017.
|
2018-03-08 04:41:16 -08:00 |
|
paboyle
|
4154fc6f44
|
Revert a change
|
2018-03-07 16:54:11 +00:00 |
|
paboyle
|
4e3458516a
|
Reverting after fixing issue with extract merge
|
2018-03-07 16:50:13 +00:00 |
|
paboyle
|
40699221e2
|
Dont alias lhs and rhs in a where statement
|
2018-03-06 04:14:13 -08:00 |
|
paboyle
|
3cb1b545d0
|
Don't alias the variables with a where statement.
|
2018-03-06 04:13:26 -08:00 |
|
paboyle
|
e199ba7e88
|
Fix the Charge conjugate BC's
|
2018-03-05 13:59:02 +00:00 |
|
paboyle
|
44188a5c6f
|
AVX512 fix
|
2018-03-05 00:32:24 +00:00 |
|
paboyle
|
3277bda130
|
View introduction to prepare for accelerator offload.
Probably same problem exists for stencil object
|
2018-03-04 16:38:08 +00:00 |
|
paboyle
|
442b0b406c
|
View related changes
|
2018-03-04 16:34:14 +00:00 |
|
paboyle
|
8824a54269
|
View related changes
|
2018-03-04 16:33:33 +00:00 |
|
paboyle
|
c03423250f
|
Indexable changes
|
2018-03-04 16:31:35 +00:00 |
|