Feeds

Cloudera cranks out fresh Hadoop distro

All stuffed elephants in one

Build a Business Case: Developing Custom Apps

Cloudera has released a third version of its open source Hadoop distribution, boasting that the distro tightly integrates several other Apache-licensed projects designed to run in tandem with the distributed number-crunching platform.

Announced on Tuesday, the Cloudera Distribution of Apache Hadoop version 3 (CDH3) includes not only the Hadoop distributed file system (HDFS) and Hadoop MapReduce, the number-crunching platform that runs atop HDFS, but also Hive (a SQL-like query language developed at Facebook), Pig (a lower-level language developed by Yahoo!), HBase (a distributed database), Sqoop (a MySQL connector built by Cloudera), Flume (a data-loading infrastructure developed by Cloudera), Oozie (the Hadoop workflow system), Hue (a graphical user interface), and Zookeeper (a means of juggling distributed services from a central location).

"We've really evolved the distribution from being one where it packaged-up core MapReduce and HDFS to a distribution that tries to provide a complete solution for running Hadoop within an organization," Cloudera vice president of product Charles Zedlewski tells The Register. "If you look at large organizations with major Hadoop deployments, you'll find that they they use a superset of components that includes a variety of tools that lets you ingest data, author jobs, review results, etc.

"With version 3, you can get those technologies, all integrated, all combined together, and all 100 per cent open source."

In particular, Cloudera says, it has extended the Hadoop authentication and security model throughout this stack of Hadoop and Hadoop-friendly platforms.

Cloudera also claims that "small" MapReduce jobs run up to three times faster with the new distro, and that file system I/O is up to 20 per cent faster, with a 2X performance boost for HBase query input. It has added a new ODBS (open database connectivity) driver to integrate business intelligence (BI) clients such as Microstrategy. And via a new adapter framework, you can build connectors for importing from and exporting to specific relational databases.

The Silicon Valley outfit also offers a for-pay Hadoop product, which includes services and support and augments the Cloudera open source distro with proprietary management, monitoring, and administration tools.

Based on Google’s proprietary software infrastructure, Hadoop is a means of crunching epic amounts of data across a network of distributed machines. Named for a yellow stuffed elephant that belongs to the son of project founder Doug Cutting, the platform now underpins online services operated by everyone from Yahoo! to Facebook and Twitter, but Cloudera is promoting its use in the enterprise as well.

Hadoop mimics GFS, Google's distributed file system, and MapReduce, the company's distributed number-crunching platform. In 2004, Google published a pair of research papers on these infrastructure technologies, and Doug Cutting seized on these to build a platform that would back Nutch, his open source web crawler. Hadoop was open sourced at Apache, and it was bootstrapped by Yahoo!, which hired Cutting in 2006. He now works for Cloudera. ®

Using blade systems to cut costs and sharpen efficiencies

More from The Register

next story
Amazon sues former employee who took Google cloud job
Alleges breach of non-compete clause in contract
New research: Flash is DEAD. Yet resistance isn't futile - it's key
Electro-boffin may have SAVED the storage WORLD
THE GERMANS ARE CLOUDING: New AWS cloud region spotted
eu-central-1.amazonaws.com, aka, your new Amazon Frankfurt bitbarn
Airbus to send 1,200 TFlops of HPC goodness down the runway
HP scores deal to provide plane-maker with new fleet of data-crunching 'PODs'
Dimension Data cloud goes TITSUP down under... after EMC storage fail
Replacement hardware needed as Australian cloud flops for 48-plus hours
IDC busts out new converged systems charts, crowns Oracle as Platform King
Nutanix/Simplivity not shown - but they're there. Oh yes
10Gbps over crumbling COPPER: Boffins cram bits down telco wire
XG-FAST tech could finesse fiber connections
China's world-beating Tianhe-2 super has brawn, not brains
Report says need for bespoke software means low utilisation rates
Los Alamos National Laboratory likes it, puts Scality's RING on it
Improved protection as firm updates its object storage software
WANdisco plunges into the Hadoop foam party, shakes its replication booty
'This 6ft-4 Indian guy in hotpants is a friggin GENIUS'
prev story

Whitepapers

How modern custom applications can spur business growth.
In this whitepaper learn how to create, deploy and manage custom applications without consuming or expanding the need for scarce, expensive IT resources.
The Power of One eBook: Top reasons to choose HP BladeSystem
Only the Power of One delivers leading infrastructure convergence, availability and scalability with federation, and agility through data center automation.
The Essential Guide to IT Transformation
ServiceNow discusses three IT transformations that can help CIO's automate IT services to transform IT and the enterprise.
Maximizing your infrastructure through virtualization
Virtualization continues to be one of the most effective ways to consolidate, reduce cost, and make data centers more efficient.
Build a Business Case: Developing Custom Apps
In this whitepaper learn how to maximize the value of custom applications by accelerating and simplifying their development.