Feeds

Storage glitches fell Australian supercomputers

DDN and SGI find fix after faulty firmware fingered for Fornax fail

Combat fraud and increase customer satisfaction

Supercomputers at two Australian research organisations have experienced substantial downtime after glitches hit their storage area networks.

West Australia’s iVEC experienced an outage, detailed here, that saw its Data Direct Networks (DDN) array and the 1152-core Fornax machine unavailable for around four days.

iVEC’s incident report for the outage says the machine was shut down on December 17th to allow the update of firmware in a “SGI DDN IS16000 storage unit.” The update produced symptoms including disks going missing from storage pools and loss of availability of some data. Instructions from the vendors concerned saw the arrays go through a lengthy “force verify process” followed by a RAID rebuild in some cases.

The rebuild started on December 18th and finished on the 20th, when it became possible to restart Fornax.

eResearch South Australia’s Tizard and Corvus supers, which offer 40 TFlops and 6 TFlops apiece, went down on December 27th and were coaxed back to life on January 3rd, a period covering three business days.

The organisation described the outage as follows on its (since updated) system maintenance page:

“During the holiday break our primary storage system locked up causing the Tizard and Corvus HPC systems to become unresponsive. At the moment there are a large number of orphaned jobs on the compute nodes of both HPC systems due to the outage.”

The culprit for the outage was not the Dell Compellent array that serves as shared NAS for the two supers, but a Dell server that sits at a critical juncture in the InfiniBand network linking the supers to their storage. Simon Brennan, a computer systems specialist at eResearch SA, told The Reg the server was overwhelmed by the quantity of data coursing through it, crashed and was restored once staff resumed from their Christmas break.

Brennan said the server at the heart of the issue was a beta, and will shortly be completely rebuilt. ®

Combat fraud and increase customer satisfaction

More from The Register

next story
This time it's 'Personal': new Office 365 sub covers just two devices
Redmond also brings Office into Google's back yard
Kingston DataTraveler MicroDuo: Turn your phone into a 72GB beast
USB-usiness in the front, micro-USB party in the back
Dropbox defends fantastically badly timed Condoleezza Rice appointment
'Nothing is going to change with Dr. Rice's appointment,' file sharer promises
BOFH: Oh DO tell us what you think. *CLICK*
$%%&amp Oh dear, we've been cut *CLICK* Well hello *CLICK* You're breaking up...
AMD's 'Seattle' 64-bit ARM server chips now sampling, set to launch in late 2014
But they won't appear in SeaMicro Fabric Compute Systems anytime soon
Amazon reveals its Google-killing 'R3' server instances
A mega-memory instance that never forgets
Cisco reps flog Whiptail's Invicta arrays against EMC and Pure
Storage reseller report reveals who's selling what
prev story

Whitepapers

Securing web applications made simple and scalable
In this whitepaper learn how automated security testing can provide a simple and scalable way to protect your web applications.
3 Big data security analytics techniques
Applying these Big Data security analytics techniques can help you make your business safer by detecting attacks early, before significant damage is done.
The benefits of software based PBX
Why you should break free from your proprietary PBX and how to leverage your existing server hardware.
Top three mobile application threats
Learn about three of the top mobile application security threats facing businesses today and recommendations on how to mitigate the risk.
Combat fraud and increase customer satisfaction
Based on their experience using HP ArcSight Enterprise Security Manager for IT security operations, Finansbank moved to HP ArcSight ESM for fraud management.