Feeds

EMC on XtremIO SSD brickup ballsup: Its LIFETIME downtime is under 3 minutes

Replace a failed X-Brick SSD once every 5 years

Designing a Defense for Mobile Applications

No DOAs here: EMC’s XTremIO arrays are expected to have less than three minutes downtime in their rated life, with X-Brick component SSDs failing once every five years or so.

We have reported that the DOA rate was too high and the co-founder and general manager of EMC's acquired XtremIO business, Ehud Rokach, has written blog which is partly a counter to that, writing that our “article presented speculations that may lead to false conclusions.”

He says that, using data obtained by monitoring “hundreds of XtremIO X-Bricks at customer sites globally and across all major verticals”:

  • XtremIO delivers world class 99.9999 per cent (six nines) field-proven availability (less than 32 Seconds of unavailability in a year, and less than 3 minutes of unavailability over the lifetime of the product.)
  • Our SSD Mean Time Between Part Replacement (MTBPR) was field-measured to be 922,240 hours, or 105 years.
  • Our Annual Replacement Rate (ARR) for SSDs was field-measured to be 0.009. For an entire X-Brick (holding 25 SSDs), the probability of encountering SSD failure at any time during a 1-year period equals (1-0.991^25), or 0.2.
  • A 0.2 ARR means that on average, based on our actual field data, you’ll need to replace a failed SSD (Due to a non-endurance related failure. …) in an X-Brick roughly once every 5 years.

Pretty darn convincing. He hammers away: “Our actual measured field performance demonstrates exceptional SSD and array-level reliability. Since initiating XtremIO’s Directed Availability program we have seen a grand total of single-digit SSD failures out of thousands of deployed SSDs.”

To refresh your memory, in the comment we reproduced from Xtremio chief techie Robin Ren, he said: “I am not too happy about our field hardware failure rate for many reasons. However, the vast majority of failures – we have seen over 150 X-Bricks so far – [pauses] in real customer environments … [and] another 200 systems internally. I think we have seen a lot of DOAs in terms of drives.”

Ren spoke in October by the way, several months after Directed Availability started.

Rokach says our story, based on Ren’s comments, “referenced (unknowingly) a couple of early DOA SSD failure events during Beta, prior to product being released for Directed Availability. The two failures during pre-release Beta were analysed, and corrective action applied (firmware update). Not surprisingly, ever since we started Directed Availability and to this very day, we have seen no excess SSD failures of any kind (in the field or DOA). This is indeed a non-issue.”

Happy to hear it. ®

The Power of One eBook: Top reasons to choose HP BladeSystem

More from The Register

next story
Apple fanbois SCREAM as update BRICKS their Macbook Airs
Ragegasm spills over as firmware upgrade kills machines
Auntie remains MYSTIFIED by that weekend BBC iPlayer and website outage
Still doing 'forensics' on the caching layer – Beeb digi wonk
Attack of the clones: Oracle's latest Red Hat Linux lookalike arrives
Oracle's Linux boss says Larry's Linux isn't just for Oracle apps anymore
THUD! WD plonks down SIX TERABYTE 'consumer NAS' fatboy
Now that's a LOT of porn or pirated movies. Or, you know, other consumer stuff
EU's top data cops to meet Google, Microsoft et al over 'right to be forgotten'
Plan to hammer out 'coherent' guidelines. Good luck chaps!
US judge: YES, cops or feds so can slurp an ENTIRE Gmail account
Crooks don't have folders labelled 'drug records', opines NY beak
Manic malware Mayhem spreads through Linux, FreeBSD web servers
And how Google could cripple infection rate in a second
prev story

Whitepapers

Designing a Defense for Mobile Applications
Learn about the various considerations for defending mobile applications - from the application architecture itself to the myriad testing technologies.
How modern custom applications can spur business growth
Learn how to create, deploy and manage custom applications without consuming or expanding the need for scarce, expensive IT resources.
Reducing security risks from open source software
Follow a few strategies and your organization can gain the full benefits of open source and the cloud without compromising the security of your applications.
Boost IT visibility and business value
How building a great service catalog relieves pressure points and demonstrates the value of IT service management.
Consolidation: the foundation for IT and business transformation
In this whitepaper learn how effective consolidation of IT and business resources can enable multiple, meaningful business benefits.