Feeds

Google File System II stalked by open-source elephant

Cloudera: Anything Google can do...

High performance access to file storage

As Google rolls out GFS2 - a major update to the custom-built file system underpinning its online infrastructure - the company's former infrastructure don sees no reason why the open source world can't follow suit.

Famously, Christophe Bisciglia taught a course at the University of Washington meant to educate rising computer scientists in Google's epic data juggling ways, and earlier this year, he brought his Big Data know-how to Cloudera, a star-studded startup that's helping to mimic Google's data prowess via open source software.

Cloudera is what you might call a Red Hat for Hadoop, the Apache-hosted open source platform based on the original Google File System (GFS), and MapReduce, Mountain View's distributed number-crunching platform. The startup offers services and support around its own Hadoop distro. Hadoop already underpins online services from Yahoo!, Facebook, Microsoft (believe it or not), and other web outfits, but it might be applied to almost any task dependent on processing unusually large amounts of data.

Details on Google's GFS2 are slim. After all, it's Google. But based on what he's read, Bisciglia calls the update "the next logical iteration" of the original GFS, and he sees Hadoop eventually following in the (rather sketchy) footsteps left by his former employer.

"A lot of the things Google is talking about are very logical directions for Hadoop to go," Bisciglia tells The Reg. "One of the things I've been very happy to see repeatedly demonstrated is that Hadoop has been able to implement [new Google GFS and MapReduce] features in approximately the same order. This shows that the fundamentals of Hadoop are solid, that the fundamentals are based on the same principles that allowed Google's systems to scale over the years.

"I don't think there's anything in Google's internal systems or GFS2 that are out of scope or out of range for Hadoop. The only issue is the reality that Google has been working on it for much longer."

GFS is celebrating its 10th anniversary, and Hadoop didn't get its start until 2005, after Google published a pair of research papers describing GFS and MapReduce. The project was founded by Nutch crawler creator Doug Cutting, who named it after his son's yellow stuffed elephant.

Google says it's been brewing GFS2 - if that's what it's called - for about the last two years. And it's now part of Caffeine, a rewrite of Google's search indexing system.

With GFS - and the Hadoop File System - a master node oversees data spread across a series of distributed chunkservers. Chunkservers store, yes, chunks of data, each about 64 megabytes. Meanwhile, GFS2 uses not only distributed slaves, but distributed masters as well.

Among other things, this improves speed. "Multiple masters is going to increase your read capacity," Bisciglia says. "It's going to allow you to have many more clients accessing the file system without inducing all of that load onto the master. From a read perspective, this is very natural. It's a reasonably understood engineering task."

It's unclear whether GFS2 also writes from multiple masters. But the new setup also provides added redundancy. If one master goes down, (many) others are there to pick up the slack.

Originally, when masters went down, GFS had no hot failover at all. But Google eventually made amends, and now, the Hadoop project is developing its own hot failover. "I watched a lot of the growing pains of GFS during my time at Google," says Bisciglia. "I remember back when there was no failover for the master and there was no [user storage] quotas on in the file systems. And these are the some of the things you're now seeing from Hadoop."

Cloudera's latest test distro - CDH2, based on Hadoop 0.20 - adds quotas. "System administrators can allocate quotas to groups and users...You can much more effectively share a cluster among a larger group of users. You don't want to cut a cluster into small, isolated independent pieces. You want to manage resource sharing on a higher level, so that when one user is not using it, someone else can. And we now allow for all that."

Anything Google can do... ®

Bootnote

Cloudera touts CDH2 here. Typically, the original CDH1 is still used for production systems, but Bisciglia expects that CDH2 - which also includes more stable APIs - will receive its official release sometime this fall.

On October 2, Cloudera is hosting a Hadoop conference in New York. Until September 21, it's offering discounted tickets for developers here.

High performance access to file storage

More from The Register

next story
Seagate brings out 6TB HDD, did not need NO STEENKIN' SHINGLES
Or helium filling either, according to reports
European Court of Justice rips up Data Retention Directive
Rules 'interfering' measure to be 'invalid'
Dropbox defends fantastically badly timed Condoleezza Rice appointment
'Nothing is going to change with Dr. Rice's appointment,' file sharer promises
Cisco reps flog Whiptail's Invicta arrays against EMC and Pure
Storage reseller report reveals who's selling what
Bored with trading oil and gold? Why not flog some CLOUD servers?
Chicago Mercantile Exchange plans cloud spot exchange
This time it's 'Personal': new Office 365 sub covers just two devices
Redmond also brings Office into Google's back yard
Just what could be inside Dropbox's new 'Home For Life'?
Biz apps, messaging, photos, email, more storage – sorry, did you think there would be cake?
IT bods: How long does it take YOU to train up on new tech?
I'll leave my arrays to do the hard work, if you don't mind
prev story

Whitepapers

Securing web applications made simple and scalable
In this whitepaper learn how automated security testing can provide a simple and scalable way to protect your web applications.
Five 3D headsets to be won!
We were so impressed by the Durovis Dive headset we’ve asked the company to give some away to Reg readers.
HP ArcSight ESM solution helps Finansbank
Based on their experience using HP ArcSight Enterprise Security Manager for IT security operations, Finansbank moved to HP ArcSight ESM for fraud management.
The benefits of software based PBX
Why you should break free from your proprietary PBX and how to leverage your existing server hardware.
Mobile application security study
Download this report to see the alarming realities regarding the sheer number of applications vulnerable to attack, as well as the most common and easily addressable vulnerability errors.