Feeds

Doug Cutting: Hadoop dodged a Microsoft-Oracle stomping

Elephant daddy on breaking into mainstream IT

Top 5 reasons to deploy VMware with Tegile

Name change, game changer

The last big change went to the very heart of Hadoop’s identity: the MapReduce engine was rewritten so there is a global resource manager and per-application master that manages different applications. The rewrite was called MapReduce 2.0 (MRv2) or YARN, and has been folded into Apache Hadoop 2.0 and Cloudera’s CDH 4.

MRv2 is designed to bring the power of Hadoop to applications outside of large datasets, such as graph-processing algorithms. "Graph" is the name given to software maps constructed by social networks such as Facebook and LinkedIn to connect people, to find out who knows whom and their relationships to each other.

'Right now we are getting the low-hanging fruit of the companies that are sophisticated users of technology and that have the most glaring big data problems,' – Doug Cutting

“Graph will be moved to shared node,” Cutting said. “So you can have a node for a time doing graph process and then a minute later running a task from a MapReduce job – there’d be time-sharing on the cluster at a finer level. It’s going tog give people better resource utilisation.”

Cutting wrote Hadoop while working on Nutch, a web search project. Nutch had used a lot of manual steps and lacked a framework that would automate large-scale data crunching. MapReduce provided the automation framework and two years later Cutting was hired by Yahoo!.

Despite improvements, Hadoop still remains relatively difficult for newcomers. More work is needed on vertical-specific apps, nice UIs and tools to integrate Hadoop with existing platforms, Cutting says.

Cutting, who today spends just a third of his time working on Hadoop, is focused on the need for broader language support. He spends the rest of his time working on Apache's Avro project, designed to make Hadoop “less of a Java shop”.

“[Java] gets you 90 per cent of the performance,” Cutting told us. “Over time, the last 10 per cent of performance might become more critical and as things stabilise, we might see more things move into C/C++. In Hadoop, several critical elements have been moved into C/C++. It would be nice if wasn’t just a Java world and we worked better with other languages, but it [Java] has been a good platform for the technology.”

Cutting is bullish on Hadoop's potential to become established in the mainstream in the next five to 10 years. MRv2 should help broaden Hadoop’s adoption among different sizes of user with different types of challenges, he says. "Making Hadoop easier" should accompany such changes, he added.

The old order liveth

“Right now we are getting the low-hanging fruit of the companies that are sophisticated users of technology and that have the most glaring big data problems. Over time we can take new classes of problems and support people more easily and do it more efficiently,” Cutting says.

Cutting said he is relieved Microsoft and Oracle are on board. In fact, he says, the will actually help Hadoop grow by adding more contributors. “We want to build a large community to develop the common software... the larger the community the more refined the tools will get. Users turn into contributors over time; the more users, the more contributors over time.”

As long as Cloudera and Hortonworks don’t become Cold War proxies to the giants, maybe this IS manifest destiny. Maybe more makers of proprietary software in the field of data and analytics should be worried.

By allying with Hadoop, though, Microsoft and Oracle have done more than reassure Cutting. They have ensured the survival of their RDBMS. ®

Beginner's guide to SSL certificates

More from The Register

next story
Download alert: Nearly ALL top 100 Android, iOS paid apps hacked
Attack of the Clones? Yeah, but much, much scarier – report
NSA SOURCE CODE LEAK: Information slurp tools to appear online
Now you can run your own intelligence agency
Microsoft: Your Linux Docker containers are now OURS to command
New tool lets admins wrangle Linux apps from Windows
Facebook, working on Facebook at Work, works on Facebook. At Work
You don't want your cat or drunk pics at the office
Soz, web devs: Google snatches its Wallet off the table
Killing off web service in 3 months... but app-happy bonkers are fine
First in line to order a Nexus 6? AT&T has a BRICK for you
Black Screen of Death plagues early Google-mobe batch
prev story

Whitepapers

Why and how to choose the right cloud vendor
The benefits of cloud-based storage in your processes. Eliminate onsite, disk-based backup and archiving in favor of cloud-based data protection.
Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Designing and building an open ITOA architecture
Learn about a new IT data taxonomy defined by the four data sources of IT visibility: wire, machine, agent, and synthetic data sets.
10 threats to successful enterprise endpoint backup
10 threats to a successful backup including issues with BYOD, slow backups and ineffective security.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.