Feeds

Twitter to open source MySQL-to-Hadoop tool

Data Crane

Secure remote control for conventional and virtual desktops

Hadoop Summit Twitter intends to open source an additional piece of the Hadoop-happy infrastructure it uses for internal data analysis. Known as Crane, this is a tool for moving data from MySQL into Hadoop, the open source data-crunching platform based on Google's proprietary infrastructure.

Twitter uses Hadoop for ad hoc analysis of data collected from its famous microblogging service, but the platform also crunches data for use by live tools on the site, including Twitter's name-search function.

Speaking today at the Yahoo!-sponsored Hadoop Summit in Santa Clara, California, Twitter analytics man Kevin Wiel explained that the company handles Hadoop data input in essentially two ways. It does log collection with the open source Scribe developed at Facebook, logging seven terabytes of data into the Hadoop File System (HDFS) each day, and it handles tabular data with Crane.

Most of Twitter's tabular data is stored in MySQL, though "a little" is stored in the Cassandra open source distributed database and Twitter's open source "social graph" data store, Flock. "Other than that," Wiel said. "Everything you do on Twitter ends up in a MySQL table somewhere."

Crane was developed to move data from MySQL to the HDFS or to the Hadoop-friendly distributed database known as HBase, but also to other MySQL databases. "We needed to have a flexible data-moving tool, so we built Crane, which is a configuration-driven ETL [extract, transform, and load] tool," Wiel says.

The tool moves data not only into MySQL, HDFS, and HBase, but also into Flock, Google Analytics, and Facebook Insights.

Like Yahoo! — and unlike Facebook — Twitter does its Hadoop programming in Pig. Developed by Yahoo!, the open source Pig is a lower-level language than the Facebook-developed Hive. But it operates at a significantly higher level than raw Hadoop MapReduce code.

According to Wiel, Pig requires five per cent of the coding and five percent of the code compared to Hadoop MapReduce, and it comes within 30 per cent of the execution time.

Twitter employees access Hadoop via dashboard known as BirdBrain, much like Facebookers use a Hive GUI known as HiPal.

A more general Hadoop interface was just open sourced by all-star startup Cloudera. Formerly known as the Cloudera Desktop, HUE — short for Hadoop User Interface — provides a web-based graphical user interface for creating and submitting jobs on a Hadoop cluster, monitoring the cluster's health, and browsing stored data. ®

Remote control for virtualized desktops

More from The Register

next story
Download alert: Nearly ALL top 100 Android, iOS paid apps hacked
Attack of the Clones? Yeah, but much, much scarier – report
You stupid BRICK! PCs running Avast AV can't handle Windows fixes
Fix issued, fingers pointed, forums in flames
Microsoft: Your Linux Docker containers are now OURS to command
New tool lets admins wrangle Linux apps from Windows
Facebook, working on Facebook at Work, works on Facebook. At Work
You don't want your cat or drunk pics at the office
Soz, web devs: Google snatches its Wallet off the table
Killing off web service in 3 months... but app-happy bonkers are fine
First in line to order a Nexus 6? AT&T has a BRICK for you
Black Screen of Death plagues early Google-mobe batch
Whistling Google: PLEASE! Brussels can only hurt Europe, not us
And Commish is VERY pro-Google. Why should we worry?
prev story

Whitepapers

Why cloud backup?
Combining the latest advancements in disk-based backup with secure, integrated, cloud technologies offer organizations fast and assured recovery of their critical enterprise data.
A strategic approach to identity relationship management
ForgeRock commissioned Forrester to evaluate companies’ IAM practices and requirements when it comes to customer-facing scenarios versus employee-facing ones.
10 threats to successful enterprise endpoint backup
10 threats to a successful backup including issues with BYOD, slow backups and ineffective security.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?
The next step in data security
With recent increased privacy concerns and computers becoming more powerful, the chance of hackers being able to crack smaller-sized RSA keys increases.