Feeds

Doug Cutting: Hadoop dodged a Microsoft-Oracle stomping

Elephant daddy on breaking into mainstream IT

Boost IT visibility and business value

Name change, game changer

The last big change went to the very heart of Hadoop’s identity: the MapReduce engine was rewritten so there is a global resource manager and per-application master that manages different applications. The rewrite was called MapReduce 2.0 (MRv2) or YARN, and has been folded into Apache Hadoop 2.0 and Cloudera’s CDH 4.

MRv2 is designed to bring the power of Hadoop to applications outside of large datasets, such as graph-processing algorithms. "Graph" is the name given to software maps constructed by social networks such as Facebook and LinkedIn to connect people, to find out who knows whom and their relationships to each other.

'Right now we are getting the low-hanging fruit of the companies that are sophisticated users of technology and that have the most glaring big data problems,' – Doug Cutting

“Graph will be moved to shared node,” Cutting said. “So you can have a node for a time doing graph process and then a minute later running a task from a MapReduce job – there’d be time-sharing on the cluster at a finer level. It’s going tog give people better resource utilisation.”

Cutting wrote Hadoop while working on Nutch, a web search project. Nutch had used a lot of manual steps and lacked a framework that would automate large-scale data crunching. MapReduce provided the automation framework and two years later Cutting was hired by Yahoo!.

Despite improvements, Hadoop still remains relatively difficult for newcomers. More work is needed on vertical-specific apps, nice UIs and tools to integrate Hadoop with existing platforms, Cutting says.

Cutting, who today spends just a third of his time working on Hadoop, is focused on the need for broader language support. He spends the rest of his time working on Apache's Avro project, designed to make Hadoop “less of a Java shop”.

“[Java] gets you 90 per cent of the performance,” Cutting told us. “Over time, the last 10 per cent of performance might become more critical and as things stabilise, we might see more things move into C/C++. In Hadoop, several critical elements have been moved into C/C++. It would be nice if wasn’t just a Java world and we worked better with other languages, but it [Java] has been a good platform for the technology.”

Cutting is bullish on Hadoop's potential to become established in the mainstream in the next five to 10 years. MRv2 should help broaden Hadoop’s adoption among different sizes of user with different types of challenges, he says. "Making Hadoop easier" should accompany such changes, he added.

The old order liveth

“Right now we are getting the low-hanging fruit of the companies that are sophisticated users of technology and that have the most glaring big data problems. Over time we can take new classes of problems and support people more easily and do it more efficiently,” Cutting says.

Cutting said he is relieved Microsoft and Oracle are on board. In fact, he says, the will actually help Hadoop grow by adding more contributors. “We want to build a large community to develop the common software... the larger the community the more refined the tools will get. Users turn into contributors over time; the more users, the more contributors over time.”

As long as Cloudera and Hortonworks don’t become Cold War proxies to the giants, maybe this IS manifest destiny. Maybe more makers of proprietary software in the field of data and analytics should be worried.

By allying with Hadoop, though, Microsoft and Oracle have done more than reassure Cutting. They have ensured the survival of their RDBMS. ®

Build a business case: developing custom apps

More from The Register

next story
KDE releases ice-cream coloured Plasma 5 just in time for summer
Melty but refreshing - popular rival to Mint's Cinnamon's still a work in progress
Leaked Windows Phone 8.1 Update specs tease details of Nokia's next mobes
New screen sizes, dual SIMs, voice over LTE, and more
Mozilla keeps its Beard, hopes anti-gay marriage troubles are now over
Plenty on new CEO's todo list – starting with Firefox's slipping grasp
Apple: We'll unleash OS X Yosemite beta on the MASSES on 24 July
Starting today, regular fanbois will be guinea pigs, it tells Reg
Another day, another Firefox: Version 31 is upon us ALREADY
Web devs, Mozilla really wants you to like this one
Secure microkernel that uses maths to be 'bug free' goes open source
Hacker-repelling, drone-protecting code will soon be yours to tweak as you see fit
Cloudy CoreOS Linux distro declares itself production-ready
Lightweight, container-happy Linux gets first Stable release
prev story

Whitepapers

Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Boost IT visibility and business value
How building a great service catalog relieves pressure points and demonstrates the value of IT service management.
Why and how to choose the right cloud vendor
The benefits of cloud-based storage in your processes. Eliminate onsite, disk-based backup and archiving in favor of cloud-based data protection.
The Essential Guide to IT Transformation
ServiceNow discusses three IT transformations that can help CIO's automate IT services to transform IT and the enterprise.
Maximize storage efficiency across the enterprise
The HP StoreOnce backup solution offers highly flexible, centrally managed, and highly efficient data protection for any enterprise.