for the weekend 23-25 March, Pluto HPC is switching off, due to the cooling failure in the HPC comms room to prevent hardware damage.
Alex Pedcenko
for the weekend 23-25 March, Pluto HPC is switching off, due to the cooling failure in the HPC comms room to prevent hardware damage.
Alex Pedcenko
Dear Zeus HPC cluster users,
this is to let you know that we have new parallel file system available for testing and use on Zeus HPC. The new storage utilizes both existing file servers in parallel (zeus5 & 6) over 40Gbps Infiniband fabric, which should bring faster read/write access for jobs. As you may know, users’ home folders are served by one file server of Zeus (zeus5) over NFS network share, and this have led in some cases to slowing down of the file system and causing overall slowdown and poor responsiveness of HPC.
The new file system does not replace users’ homes, but is intended as a runtime space for jobs. The location of this new storage is at /mnt/beegfs/scratch/yourusername, where each user has a working folder. You can also access this space by the link from your home folder, i.e. /home/yourusername/scratch
If you like to launch jobs from this file system, you need to copy your job/project folder to /home/yourusername/scratch and submit your slurm job form there or from /mnt/beegfs/scratch/yourusername if you prefer. This would, in theory, reduce the load on Zeuse’s NFS server, which serves user’s homes.
After your job completes, it is advisable (although not necessary) to copy your results from /home/yourusername/scratch/… to your home folder /home/yourusername for safekeeping. Scratch folder is not backed up and during the test period, there may be a chance of some failures.
Best Regards,
Alex Pedcenko
It seems there was a power cut in ECB comms rooms on 15 th June approx. 21:45 — all HPCs went off (If you wonder why your jobs have died)
Alex
module load gcc/7.1.0
Alex
Chillers in the HPC room EC3-21 failed once again this Sunday. Broadwell nodes and half of Nehalem nodes (zeus[20-91,15]) were switched off until the cause of the faults will be finally found by Estates.
Compute nodes which are available : zeus[100-171, 200-217] (queues: all, long, GPU)
Regards,
Alex
Zeus HPC is operational. DataLake machine is still experiencing some problems.
Alex
Update on the HPC issue:
Temperature in main HPC room stabilised, I brought login nodes and main server and file servers up. Until further update from Estates about the cooling system stability in the room, most of the compute nodes in that room will be offline (that includes new Broadwell nodes)
I brought some Nehalem (half of 8-CPU nodes) and Sandybridge (12-CPU “GPU” queue) compute nodes up in unaffected by cooling failure room (zeus[100-171], zeus[200-217]), they can be used as file servers now are operational.
Regards,
Alex
Hi,
Zeus HPC is down due to cooling failure in the room EC3-23.
Regards,
Alex
P.S. will update you when it can come back….
Recent Comments