MPP version of LSDYNA was tested for various CPU configurations on new Broadwell nodes vs old Nehalem 8-CPU nodes for the same problem.
SLURM file for LS-Dyna 9.1.0 submission is here lsdyna.
For 9.1.0 version of LSDYNA use
module load lsdyna/971
module load lsdyna/pmpi
| Family | Nodes | CPUs | EXEC_TIME, hrs | lsdyna ver | mpi |
| Broadwell | 1 | 32 | 06:02 | 7.1.2 | HPMPI |
| Broadwell | 1 | 16 | 09:22 | 7.1.2 | HPMPI |
| Nehalem | 2 | 16=2×8 | >12 hrs (time limit reached) | 7.1.2 | HPMPI |
| Broadwell | 2 | 32=2×16 | 05:28 | 7.1.2 | HPMPI |
| Broadwell | 4 | 64=4×16 | 02:51 | 9.1.0 | PMPI |
| Broadwell | 8 | 128=8×16 | 01:46 | 9.1.0 | PMPI |
| Broadwell | 2 | 64=2×32 | 05:08 | 9.1.0 | PMPI |
| Sandybridge | 2 | 24=2×12 | 09:30 | 7.1.3 | PMPI |
So It looks like running LSDYNA on both CPUs (16 cores) of Broadwell is really does not make the problem solve faster. Instead use just 16 CPU-cores (see 64 CPU-cores case — its 2x faster!): either all on one socket or on different ones is still remains to be tested.
Alex Pedcenko
0 Comments.