Feeds

Yahoo! exposes very own stuffed elephant code

Distributed data-crunching distro

Combat fraud and increase customer satisfaction

Hadoop Summit Yahoo! has released its own Hadoop distro, an internet-scale distributed data-crunching platform based on the Apache open-source project that underpins several of the web’s highest profile sites, including Yahoo!, Facebook, and - amusingly - Microsoft’s Bing.

Inspired by Google-published research papers describing Mountain View’s proprietary software infrastructure, Hadoop is the brainchild of open-source guru Doug Cutting, the Nutch crawler founder who’s now on the Yahoo! payroll.

Yahoo! has used Hadoop code on its production infrastructure for more than a year now, and after calls from the ever-growing Hadoop community, the company is opening up its internal implementation of the project.

"We’ve put a lot of investment on our testing and deployment," Yahooligan Eric Baldeschwieler said Wednesday at the Yahoo!-sponsored Hadoop Summit in Santa Clara, California. "We’re going to take that work that we put into it and put it out on the web."

The new release - known as the Yahoo! Distribution of Hadoop - is not a commercially supported distro. "We’re not getting into a new business," Baldeschwieler explained. Yahoo! is leaving that business to Cloudera, the Silicon Valley startup that unveiled a commercial Hadoop distro this spring.

According to Baldeschwieler, Yahoo! will release code identical to that tested and deployed on the company’s internal machines. "The source code release will be exactly like we use on Yahoo! clusters," he said. And he expects Yahoo!-tweaked code will be released three to six months after general release of the Apache project code it's based on.

Yahoo! will not restrict access to the code, which will be available here from the Yahoo! developer network. It will merely require an agreement before downloading. The first release will be Hadoop version 0.20, which is now under alpha test inside the company.

Yahoo! contributes about 72 per cent of all Apache Hadoop patches. And it now uses Hadoop code to crunch data for myriad Yahoo! services, including its search index and the automated system that chooses news stories for its homepage.

Cutting cooked up Hadoop in 2004, naming the project after his son’s yellow stuffed elephant. Along with Yahoo!, one of its early users was Powerset, the semantic search engine recently acquired by Microsoft. Powerset now drives at least a small portion of Redmond’s latest Google challenge, which it insists on calling Bing. ®

3 Big data security analytics techniques

More from The Register

next story
Ubuntu 14.04 LTS: Great changes, but sssh don't mention the...
Why HELLO Amazon! You weren't here last time
This time it's 'Personal': new Office 365 sub covers just two devices
Redmond also brings Office into Google's back yard
Next Windows obsolescence panic is 450 days from … NOW!
The clock is ticking louder for Windows Server 2003 R2 users
Half of Twitter's 'active users' are SILENT STALKERS
Nearly 50% have NEVER tweeted a word
OpenBSD founder wants to bin buggy OpenSSL library, launches fork
One Heartbleed vuln was too many for Theo de Raadt
Got Windows 8.1 Update yet? Get ready for YET ANOTHER ONE – rumor
Leaker claims big release due this fall as Microsoft herds us into the CLOUD
Microsoft TIER SMEAR changes app prices whether devs ask or not
Some go up, some go down, Redmond goes silent
Batten down the hatches, Ubuntu 14.04 LTS due in TWO DAYS
Admins dab straining server brows in advance of Trusty Tahr's long-term support landing
Red Hat to ship RHEL 7 release candidate with a taste of container tech
Grab 'near-final' version of next Enterprise Linux next week
prev story

Whitepapers

Mobile application security study
Download this report to see the alarming realities regarding the sheer number of applications vulnerable to attack, as well as the most common and easily addressable vulnerability errors.
3 Big data security analytics techniques
Applying these Big Data security analytics techniques can help you make your business safer by detecting attacks early, before significant damage is done.
The benefits of software based PBX
Why you should break free from your proprietary PBX and how to leverage your existing server hardware.
Securing web applications made simple and scalable
In this whitepaper learn how automated security testing can provide a simple and scalable way to protect your web applications.
Combat fraud and increase customer satisfaction
Based on their experience using HP ArcSight Enterprise Security Manager for IT security operations, Finansbank moved to HP ArcSight ESM for fraud management.