Category Archives: HPC Announcements - Page 2

PLUTO HPC switched off

for the weekend 23-25 March, Pluto HPC is switching off, due to the cooling failure in the HPC comms  room to prevent hardware damage.

Alex Pedcenko

New Parallel File System testing

Dear Zeus HPC cluster users,

this is to let you know that we have new parallel file system available for testing and use on Zeus HPC. The new storage utilizes both existing file servers in parallel (zeus5 & 6) over 40Gbps Infiniband fabric, which should bring faster read/write access for jobs. As you may know, users’ home folders are served by one file server of Zeus (zeus5) over NFS network share, and this have led in some cases to slowing down of the file system and causing overall slowdown and poor responsiveness of HPC.

The new file system does not replace users’ homes, but is intended as a runtime space for jobs. The location of this new storage is at /mnt/beegfs/scratch/yourusername, where each user has a working folder. You can also access this space by the link from your home folder, i.e. /home/yourusername/scratch
If you like to launch jobs from this file system, you need to copy your job/project folder to /home/yourusername/scratch and submit your slurm job form there or from /mnt/beegfs/scratch/yourusername if you prefer. This would, in theory, reduce the load on Zeuse’s NFS server, which serves user’s homes.

After your job completes, it is advisable (although not necessary) to copy your results from /home/yourusername/scratch/… to your home folder /home/yourusername for safekeeping. Scratch folder is not backed up and during the test period, there may be a chance of some failures.

Best Regards,
Alex Pedcenko

Power cut 15/06/17 21:30

It seems there was a power cut in ECB comms rooms on 15 th June approx. 21:45 — all HPCs went off (If  you wonder why your jobs have died)

Alex

 

New GCC-7.1.0 compiler is installed on zeus

module load gcc/7.1.0

Alex

Google is giving a cluster of 1,000 Cloud TPUs to researchers for free

See details here: https://go.newsfusion.com//cloud-computing/item/935489

 

Alex

 

EC3-21 Temperature

HPC Temperature plots

HPC room was overheating again on Sunday 7 May

Chillers in the HPC room EC3-21 failed once again this Sunday. Broadwell nodes and half of Nehalem nodes (zeus[20-91,15]) were switched off until the cause of the faults will be finally found by Estates.
Compute nodes which are available : zeus[100-171, 200-217] (queues: all, long, GPU)

Regards,
Alex

Normality restored

Zeus HPC is operational. DataLake machine is still experiencing some problems.

Alex

Zeus HPC update

Update on the HPC issue:

Temperature in main HPC room stabilised, I brought login nodes and main server and file servers up. Until further update from Estates about the cooling system stability in the room, most of the compute nodes in that room will be offline (that includes new Broadwell nodes)

I brought some Nehalem (half of 8-CPU nodes) and Sandybridge (12-CPU “GPU” queue) compute nodes up in unaffected by cooling failure room (zeus[100-171], zeus[200-217]), they can be used as file servers now are operational.

Regards,
Alex

Zeus down due to Room overheating

Hi,

Zeus HPC is down due to cooling failure in the room EC3-23.

Regards,
Alex

P.S. will update you when it can come back….

css.php