New compute nodes (Broadwell) performance (measured by HPL Linpack benchmark)
Below are results of few HPL tests on all 56 new Broadwell 32-CPU nodes as well as on all 144 old Nehalems. Netlib xhpl was compiled with intel icc and ran with Bullx mpi. Here are the results:
| # of cores | CPU model | Config | Flops achieved | theoretical |
| 1792 | Broadwell | 56 nodes | 34 Tflops | 30.1 Tflops |
| 1152 | Nehalem | 144 nodes | 9.5 Tflops | |
| 320 | Broadwell | 10 nodes | 6.523 Tflops | 5.376 Tflops |
| 32 | Broadwell | 1 node | 675 Gflops | 537.6 Gflops |
| 12 | Broadwell | 1 node | 260 Gflops | 202 Gflops |
| 204 | Sandybridge | 17 nodes (from GPU queue) | 3.3 TFlops | 3.9 TFops |
| 12 | Sandybridge | 1 node (from GPU queue) | 200 Gflops | 230 Gflops |
| 8 | Sandybridge | 1 node (48 Gb, 12 CPU-cores) | 135.6 Gflops | 76.7 |
| 8 | Nehalem | 1 node | 70.49 Gflops | 76.6 |
| 8 | Nehalem | 4 nodes x 2 cores | 70.15 Gflops | 76.6 |
| 8 | Sandybridge | 4 nodes x 2 cores | 135.6 Gflops | 76.8 |
| 8 | Nehalem | 2 floors x 2 nodes x 2 cores | 70.17 Gflops | 76.8 |
| 216 | Sandybridge | 18 nodes x 12 cores | 3.5 Tflops | — |
| 576 | Nehalem | 72 nodes x 8 cores | 4.5 Tflops | — |
| 32 | Sandybridges on SMP node | 1 nodes x 32 cores | 0.5 Tflops | — |
Test ran using bullx mpi 1.2.9
#################### 2013 results #######################
Below are some first basic HPL (linpack) results of the cluster.
We have two types of CPU’s on Zeuse’s nodes (both @2.4 GHz):
- GPU queue: 18 nodes (zeus200-217) of Intel(R) Xeon(R) CPU E5-2440 (Sandybridge) — 12 cores per node & 48 Gb of RAM, & 2x Nvidia Kepler K20 GPUs
- Default “all” queue: 144 nodes of Intel(R) Xeon(R) CPU L5530 (Nehalem) — 8 cores per node & 48 Gb of RAM
- “SMP” queue: 1 node of Intel(R) Xeon(R) CPU E5-4620 (Sandybridge)– 32 cores per node & 512 Gb RAM
| # of cores | CPU model | Config | Flops achieved | theoretical |
| 8 | Sandybridge | 1 node | 135.6 Gflops | 76.7 |
| 8 | Nehalem | 1 node | 70.49 Gflops | 76.6 |
| 8 | Nehalem | 4 nodes x 2 cores | 70.15 Gflops | 76.6 |
| 8 | Sandybridge | 4 nodes x 2 cores | 135.6 Gflops | 76.8 |
| 8 | Nehalem | 2 floors x 2 nodes x 2 cores | 70.17 Gflops | 76.8 |
| 216 | Sandybridge | 18 nodes x 12 cores | 3.5 Tflops | — |
| 576 | Nehalem | 72 nodes x 8 cores | 4.5 Tflops | — |
| 32 | Sandybridges on SMP node | 1 nodes x 32 cores | 0.5 Tflops | — |
Test ran using bullx mpi 1.2.4
Alex Pedcenko
0 Comments.