Feeds

Amazon's weekend cloud outage highlights EBS problems

The red-headed stepchild of Bezos & Co's cloud just can't keep up

Secure remote control for conventional and virtual desktops

Problems in the Amazon cloud over the weekend crushed apps like Vine, websites like Airbnb, and numerous other services that depend on Bezos & Co's hulking cloud, and the problems were due to a familiar culprit – Elastic Block Store (EBS).

EBS is a network-attached block level storage service for Amazon EC2 instances. Amazon says it is "suited for applications that require a database, file system, or access to raw block level storage," – in other words, everything.

Sunday's failure marked the third significant outage in two years to come about from EBS failures, and brought to mind the characterization of EBS as "a barrel of laughs in terms of performance and reliability" by a former Reddit sysadmin after a major outage in April 2011.

The problems on Sunday were acknowledged by Amazon in a post to the company's status dashboard at 1:22pm Pacific Time, when the company said it was "investigating degraded performance for some volumes in a single [Availability Zone] in the US-EAST-1 Region."

Amazon found that the problem was a network issue that led to elevated EBS-related API error rates in a single region. "The networking device was removed from service and we are performing a forensic investigation to understand how it failed," the company wrote.

Besides the 2011 incident, EBS also went down in December 2012. In the wake of that outage, one EBS-reliant company named Awe.sm wrote that "to maintain high uptime, we have stopped trusting EBS." Awe.sm added that in its experience, input-output rates on EBS volumes were poor, that when it fails it tends to fail across an entire data center cluster, and that if it goes down when connected to an image when running Ubuntu it fails severely.

Given the outage during the weekend just gone, cloud-first businesses might want to start looking at EBS and working out how to design their systems around potential failures in Amazon's data center hubs. ®

Top 5 reasons to deploy VMware with Tegile

More from The Register

next story
NSA SOURCE CODE LEAK: Information slurp tools to appear online
Now you can run your own intelligence agency
Azure TITSUP caused by INFINITE LOOP
Fat fingered geo-block kept Aussies in the dark
Yahoo! blames! MONSTER! email! OUTAGE! on! CUT! CABLE! bungle!
Weekend woe for BT as telco struggles to restore service
Cloud unicorns are extinct so DiData cloud mess was YOUR fault
Applications need to be built to handle TITSUP incidents
BOFH: WHERE did this 'fax-enabled' printer UPGRADE come from?
Don't worry about that cable, it's part of the config
Stop the IoT revolution! We need to figure out packet sizes first
Researchers test 802.15.4 and find we know nuh-think! about large scale sensor network ops
Turnbull should spare us all airline-magazine-grade cloud hype
Box-hugger is not a dirty word, Minister. Box-huggers make the cloud WORK
SanDisk vows: We'll have a 16TB SSD WHOPPER by 2016
Flash WORM has a serious use for archived photos and videos
Astro-boffins start opening universe simulation data
Got a supercomputer? Want to simulate a universe? Here you go
prev story

Whitepapers

Driving business with continuous operational intelligence
Introducing an innovative approach offered by ExtraHop for producing continuous operational intelligence.
Why CIOs should rethink endpoint data protection in the age of mobility
Assessing trends in data protection, specifically with respect to mobile devices, BYOD, and remote employees.
Getting started with customer-focused identity management
Learn why identity is a fundamental requirement to digital growth, and how without it there is no way to identify and engage customers in a meaningful way.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?
Simplify SSL certificate management across the enterprise
Simple steps to take control of SSL across the enterprise, and recommendations for a management platform for full visibility and single-point of control for these Certificates.