Feeds

Yahoo! exposes very own stuffed elephant code

Distributed data-crunching distro

Secure remote control for conventional and virtual desktops

Hadoop Summit Yahoo! has released its own Hadoop distro, an internet-scale distributed data-crunching platform based on the Apache open-source project that underpins several of the web’s highest profile sites, including Yahoo!, Facebook, and - amusingly - Microsoft’s Bing.

Inspired by Google-published research papers describing Mountain View’s proprietary software infrastructure, Hadoop is the brainchild of open-source guru Doug Cutting, the Nutch crawler founder who’s now on the Yahoo! payroll.

Yahoo! has used Hadoop code on its production infrastructure for more than a year now, and after calls from the ever-growing Hadoop community, the company is opening up its internal implementation of the project.

"We’ve put a lot of investment on our testing and deployment," Yahooligan Eric Baldeschwieler said Wednesday at the Yahoo!-sponsored Hadoop Summit in Santa Clara, California. "We’re going to take that work that we put into it and put it out on the web."

The new release - known as the Yahoo! Distribution of Hadoop - is not a commercially supported distro. "We’re not getting into a new business," Baldeschwieler explained. Yahoo! is leaving that business to Cloudera, the Silicon Valley startup that unveiled a commercial Hadoop distro this spring.

According to Baldeschwieler, Yahoo! will release code identical to that tested and deployed on the company’s internal machines. "The source code release will be exactly like we use on Yahoo! clusters," he said. And he expects Yahoo!-tweaked code will be released three to six months after general release of the Apache project code it's based on.

Yahoo! will not restrict access to the code, which will be available here from the Yahoo! developer network. It will merely require an agreement before downloading. The first release will be Hadoop version 0.20, which is now under alpha test inside the company.

Yahoo! contributes about 72 per cent of all Apache Hadoop patches. And it now uses Hadoop code to crunch data for myriad Yahoo! services, including its search index and the automated system that chooses news stories for its homepage.

Cutting cooked up Hadoop in 2004, naming the project after his son’s yellow stuffed elephant. Along with Yahoo!, one of its early users was Powerset, the semantic search engine recently acquired by Microsoft. Powerset now drives at least a small portion of Redmond’s latest Google challenge, which it insists on calling Bing. ®

Top 5 reasons to deploy VMware with Tegile

More from The Register

next story
Google+ goes TITSUP. But WHO knew? How long? Anyone ... Hello ...
Wobbly Gmail, Contacts, Calendar on the other hand ...
Preview redux: Microsoft ships new Windows 10 build with 7,000 changes
Latest bleeding-edge bits borrow Action Center from Windows Phone
UNIX greybeards threaten Debian fork over systemd plan
'Veteran Unix Admins' fear desktop emphasis is betraying open source
Microsoft promises Windows 10 will mean two-factor auth for all
Sneak peek at security features Redmond's baking into new OS
Google opens Inbox – email for people too stupid to use email
Print this article out and give it to someone techy if you get stuck
DEATH by PowerPoint: Microsoft warns of 0-day attack hidden in slides
Might put out patch in update, might chuck it out sooner
Redmond top man Satya Nadella: 'Microsoft LOVES Linux'
Open-source 'love' fairly runneth over at cloud event
prev story

Whitepapers

Cloud and hybrid-cloud data protection for VMware
Learn how quick and easy it is to configure backups and perform restores for VMware environments.
A strategic approach to identity relationship management
ForgeRock commissioned Forrester to evaluate companies’ IAM practices and requirements when it comes to customer-facing scenarios versus employee-facing ones.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?
Three 1TB solid state scorchers up for grabs
Big SSDs can be expensive but think big and think free because you could be the lucky winner of one of three 1TB Samsung SSD 840 EVO drives that we’re giving away worth over £300 apiece.
Security for virtualized datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.