Feeds

Yahoo! girds Google's bastard grid child

Open sources Hadoop with Security

5 things you didn’t know about cloud backup

Hadoop Summit Yahoo! has released a beta version of Hadoop with built-in security, while open sourcing the latest version of its in-house workflow engine for the Google-mimicking distributed number-crunching platform.

Speaking this morning at the Yahoo!-sponsored Hadoop Summit in Santa Clara, California, company Hadoop guru Eric Baldeschwieler said that both the Hadoop with Security beta and the Oozie workflow engine have been open sourced at Apache. The security beta has been deployed on the Hadoop clusters that Yahoo! maintains for research organizations, and Baldeschwieler tells The Reg that it is now part of Yahoo!'s Hadoop distro.

Based on Google’s (proprietary) software infrastructure, Hadoop is a means of crunching epic amounts of data across a network of distributed machines. Named for the stuffed elephant belonging to the son of project founder Doug Cutting, the open source platform now underpins online services operated by everyone from Yahoo! and Facebook and Twitter to — gasp! — Microsoft.

Hadoop duplicates GFS, Google's distributed file system, and MapReduce, Google's distributed number-crunching platform. In 2004, Google published a pair of research papers on these technologies, and Cutting used the papers to build a platform that would back Nutch, his open source web crawler. Hadoop was open sourced at Apache, and Yahoo! is still its largest contributor.

Hadoop with Security, Baldeschwieler says, integrates the platform with Kerberos, the open source authentication standard, while adding enhanced logging features. According to Baldeschwieler, the release is designed to allow the sharing of sensitive data with appropriate permissions and to "ease regulatory compliance". It's designed to prevent unauthorized access to data stored on Hadoop clusters, more effectively co-locate business sensitive data, and reduce costs by consolidating clusters.

Oozie — named with the Burmese term for, yes, elephant keeper — is designed for particularly complex workflows and data pipelines. At Yahoo! it's the de facto standard for ETL (extract, transform, and load) processing. This open source workflow engine is also part of the new Hadoop distro from startup Cloudera, released today.

Baldeschwieler tells The Reg that Yahoo! has been using Oozie for about six months. The version open sourced today is essentially Oozie 2.0. A 1.0 version was previously open sourced. ®

Boost IT visibility and business value

More from The Register

next story
Why has the web gone to hell? Market chaos and HUMAN NATURE
Tim Berners-Lee isn't happy, but we should be
Linux turns 23 and Linus Torvalds celebrates as only he can
No, not with swearing, but by controlling the release cycle
Apple promises to lift Curse of the Drained iPhone 5 Battery
Have you tried turning it off and...? Never mind, here's a replacement
Sin COS to tan Windows? Chinese operating system to debut in autumn – report
Development alliance working on desktop, mobe software
Eat up Martha! Microsoft slings handwriting recog into OneNote on Android
Freehand input on non-Windows kit for the first time
Linux kernel devs made to finger their dongles before contributing code
Two-factor auth enabled for Kernel.org repositories
This is how I set about making a fortune with my own startup
Would you leave your well-paid job to chase your dream?
prev story

Whitepapers

Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Endpoint data privacy in the cloud is easier than you think
Innovations in encryption and storage resolve issues of data privacy and key requirements for companies to look for in a solution.
Scale data protection with your virtual environment
To scale at the rate of virtualization growth, data protection solutions need to adopt new capabilities and simplify current features.
Boost IT visibility and business value
How building a great service catalog relieves pressure points and demonstrates the value of IT service management.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?