Feeds

Yahoo! exposes very own stuffed elephant code

Distributed data-crunching distro

Providing a secure and efficient Helpdesk

Hadoop Summit Yahoo! has released its own Hadoop distro, an internet-scale distributed data-crunching platform based on the Apache open-source project that underpins several of the web’s highest profile sites, including Yahoo!, Facebook, and - amusingly - Microsoft’s Bing.

Inspired by Google-published research papers describing Mountain View’s proprietary software infrastructure, Hadoop is the brainchild of open-source guru Doug Cutting, the Nutch crawler founder who’s now on the Yahoo! payroll.

Yahoo! has used Hadoop code on its production infrastructure for more than a year now, and after calls from the ever-growing Hadoop community, the company is opening up its internal implementation of the project.

"We’ve put a lot of investment on our testing and deployment," Yahooligan Eric Baldeschwieler said Wednesday at the Yahoo!-sponsored Hadoop Summit in Santa Clara, California. "We’re going to take that work that we put into it and put it out on the web."

The new release - known as the Yahoo! Distribution of Hadoop - is not a commercially supported distro. "We’re not getting into a new business," Baldeschwieler explained. Yahoo! is leaving that business to Cloudera, the Silicon Valley startup that unveiled a commercial Hadoop distro this spring.

According to Baldeschwieler, Yahoo! will release code identical to that tested and deployed on the company’s internal machines. "The source code release will be exactly like we use on Yahoo! clusters," he said. And he expects Yahoo!-tweaked code will be released three to six months after general release of the Apache project code it's based on.

Yahoo! will not restrict access to the code, which will be available here from the Yahoo! developer network. It will merely require an agreement before downloading. The first release will be Hadoop version 0.20, which is now under alpha test inside the company.

Yahoo! contributes about 72 per cent of all Apache Hadoop patches. And it now uses Hadoop code to crunch data for myriad Yahoo! services, including its search index and the automated system that chooses news stories for its homepage.

Cutting cooked up Hadoop in 2004, naming the project after his son’s yellow stuffed elephant. Along with Yahoo!, one of its early users was Powerset, the semantic search engine recently acquired by Microsoft. Powerset now drives at least a small portion of Redmond’s latest Google challenge, which it insists on calling Bing. ®

Secure remote control for conventional and virtual desktops

More from The Register

next story
Microsoft WINDOWS 10: Seven ATE Nine. Or Eight did really
Windows NEIN skipped, tech preview due out on Wednesday
Business is back, baby! Hasta la VISTA, Win 8... Oh, yeah, Windows 9
Forget touchscreen millennials, Microsoft goes for mouse crowd
Apple: SO sorry for the iOS 8.0.1 UPDATE BUNGLE HORROR
Apple kills 'upgrade'. Hey, Microsoft. You sure you want to be like these guys?
ARM gives Internet of Things a piece of its mind – the Cortex-M7
32-bit core packs some DSP for VIP IoT CPU LOL
Microsoft on the Threshold of a new name for Windows next week
Rebranded OS reportedly set to be flung open by Redmond
Lotus Notes inventor Ozzie invents app to talk to people on your phone
Imagine that. Startup floats with voice collab app for Win iPhone
'Google is NOT the gatekeeper to the web, as some claim'
Plus: 'Pretty sure iOS 8.0.2 will just turn the iPhone into a fax machine'
prev story

Whitepapers

Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Storage capacity and performance optimization at Mizuno USA
Mizuno USA turn to Tegile storage technology to solve both their SAN and backup issues.
The next step in data security
With recent increased privacy concerns and computers becoming more powerful, the chance of hackers being able to crack smaller-sized RSA keys increases.
Security for virtualized datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.
A strategic approach to identity relationship management
ForgeRock commissioned Forrester to evaluate companies’ IAM practices and requirements when it comes to customer-facing scenarios versus employee-facing ones.