Feeds

Amazon parks human genome on cloud

Taps Exploits world of boffins

Seven Steps to Software Security

In 1993, meat space bookseller Barnes & Noble started offering Starbucks coffee to augment customers' shopping experience. Not to be outdone, the Internet's largest bookseller has finally answered.

Amazon announced yesterday that it will allow free, easy access to hard-to-find datasets like the human genome - and a few other compilations that aren't nearly as exciting. With its new Public Data Sets initiative, Amazon hopes to commoditize access to this kind of data, while positioning its EC2 system as the preferred computing platform for researchers.

For users, this service is likely to speed up the science without digging too far into the meager stack of grant money. Amazon is currently hosting data sets from biology, chemistry, and economics, having hunted down files in the public domain. To expand the service, they are soliciting people to provide more data, provided that it's not proprietary.

They boast free access to the data. "Free" is about as relative a term as “pregnant,” but Amazon is still taking some liberty with it. While technically you don't have to pay for the data sets, the only way you can access them is by spinning up an EC2 virtual machine instance and mounting the data as an Elastic Block Store drive. Amazon charges for EC2 instances by the hour, and Elastic Block Store drives by the I/O request. Conveniently for Amazon, these data sets are equally as processable as they are large.

There are already several online directories of public domain data sets, so what's the value-add? When you run an EC2 virtual instance to access the data, you can choose a machine image that contains specialized processing tools. If parallel processing is an concern, you can spin up as many instances as you need to process the data and spin them down when you're done, only paying for the resources you use.

This is still a bit dubious, because Amazon isn't really doing any work here. The data sets and processing tools come from third parties. Bezos just uses them to sell Amazon Web Services. As far as exploiting the academic spirit of sharing and mutual betterment for profit, the Public Data Sets program is a winner.

While working with the data in Amazon's environment may be a bit faster and more convenient, the target market for this type of service is the notoriously tight-assed academic researcher. While a tenured professor may see a minor productivity increase by using Amazon Web Services, she can see an order of magnitude productivity increase by enslaving a graduate student to do the same work. Said graduate student, already living below the poverty line, is unlikely to spend money on Amazon. There is a time-honored tradition among these servile few of spending late nights in the computer lab slicing up US Census data because it's difficult to get SAS or SPSS licenses for their laptops.

Be that as it may, Amazon isn't making a big bet on this one. If the analytics tools provided on EC2 take good enough advantage of the parallelism offered by virtualization, perhaps graduate students everywhere will decide to drop a little money. Less time spent analyzing data means more time spent partying, remembering that you're still under thirty and without serious responsibility. ®

Ted Dziuba is a co-founder at Milo.com You can read his regular Reg column, Fail and You, every other Monday.

The Power of One eBook: Top reasons to choose HP BladeSystem

More from The Register

next story
Beancounters tell NASA it's too poor to fly planned mega-rocket
Space Launch System would need another $400m and a lot of time
Malaysian Airlines flight MH17 claimed lives of HIV/AIDS cure scientists
Researchers, advocates, health workers among those on shot-down plane
Bad back? Show some spine and stop popping paracetamol
Study finds common pain-killer doesn't reduce pain or shorten recovery
World Solar Challenge contender claims new speed record
One charge sees Sunswift travel 500kms at over 100 km/h
SMELL YOU LATER, LOSERS – Dumbo tells rats, dogs... humans
Junk in the trunk? That's what people have
All those new '5G standards'? Here's the science they rely on
Radio professor tells us how wireless will get faster in the real world
The Sun took a day off last week and made NO sunspots
Someone needs to get that lazy star cooking again before things get cold around here
prev story

Whitepapers

Top three mobile application threats
Prevent sensitive data leakage over insecure channels or stolen mobile devices.
Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Top 8 considerations to enable and simplify mobility
In this whitepaper learn how to successfully add mobile capabilities simply and cost effectively.
Application security programs and practises
Follow a few strategies and your organization can gain the full benefits of open source and the cloud without compromising the security of your applications.
The Essential Guide to IT Transformation
ServiceNow discusses three IT transformations that can help CIO's automate IT services to transform IT and the enterprise.