Feeds

Storage glitches fell Australian supercomputers

DDN and SGI find fix after faulty firmware fingered for Fornax fail

Application security programs and practises

Supercomputers at two Australian research organisations have experienced substantial downtime after glitches hit their storage area networks.

West Australia’s iVEC experienced an outage, detailed here, that saw its Data Direct Networks (DDN) array and the 1152-core Fornax machine unavailable for around four days.

iVEC’s incident report for the outage says the machine was shut down on December 17th to allow the update of firmware in a “SGI DDN IS16000 storage unit.” The update produced symptoms including disks going missing from storage pools and loss of availability of some data. Instructions from the vendors concerned saw the arrays go through a lengthy “force verify process” followed by a RAID rebuild in some cases.

The rebuild started on December 18th and finished on the 20th, when it became possible to restart Fornax.

eResearch South Australia’s Tizard and Corvus supers, which offer 40 TFlops and 6 TFlops apiece, went down on December 27th and were coaxed back to life on January 3rd, a period covering three business days.

The organisation described the outage as follows on its (since updated) system maintenance page:

“During the holiday break our primary storage system locked up causing the Tizard and Corvus HPC systems to become unresponsive. At the moment there are a large number of orphaned jobs on the compute nodes of both HPC systems due to the outage.”

The culprit for the outage was not the Dell Compellent array that serves as shared NAS for the two supers, but a Dell server that sits at a critical juncture in the InfiniBand network linking the supers to their storage. Simon Brennan, a computer systems specialist at eResearch SA, told The Reg the server was overwhelmed by the quantity of data coursing through it, crashed and was restored once staff resumed from their Christmas break.

Brennan said the server at the heart of the issue was a beta, and will shortly be completely rebuilt. ®

Eight steps to building an HP BladeSystem

More from The Register

next story
Sysadmin Day 2014: Quick, there's still time to get the beers in
He walked over the broken glass, killed the thugs... and er... reconnected the cables*
SHOCK and AWS: The fall of Amazon's deflationary cloud
Just as Jeff Bezos did to books and CDs, Amazon's rivals are now doing to it
Amazon Reveals One Weird Trick: A Loss On Almost $20bn In Sales
Investors really hate it: Share price plunge as growth SLOWS in key AWS division
US judge: YES, cops or feds so can slurp an ENTIRE Gmail account
Crooks don't have folders labelled 'drug records', opines NY beak
Auntie remains MYSTIFIED by that weekend BBC iPlayer and website outage
Still doing 'forensics' on the caching layer – Beeb digi wonk
Manic malware Mayhem spreads through Linux, FreeBSD web servers
And how Google could cripple infection rate in a second
BlackBerry: Toss the server, mate... BES is in the CLOUD now
BlackBerry Enterprise Services takes aim at SMEs - but there's a catch
The triumph of VVOL: Everyone's jumping into bed with VMware
'Bandwagon'? Yes, we're on it and so what, say big dogs
prev story

Whitepapers

Top three mobile application threats
Prevent sensitive data leakage over insecure channels or stolen mobile devices.
Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Boost IT visibility and business value
How building a great service catalog relieves pressure points and demonstrates the value of IT service management.
Designing a Defense for Mobile Applications
Learn about the various considerations for defending mobile applications - from the application architecture itself to the myriad testing technologies.
Build a business case: developing custom apps
Learn how to maximize the value of custom applications by accelerating and simplifying their development.