Feeds

Steve Jobs embraces Google's bastard grid child

Cult plants ad empire on Hadoop

Next gen security for virtualised datacentres

Apple has embraced Hadoop, the open source distributed-computing platform based on Google's famously proprietary backend infrastructure.

According to a recent Apple job listing entitled "Senior Software Engineer - Hadoop", the company is using or planning to use the entire Hadoop stack, from the HDFS file system and the Hadoop MapReduce distributed number-crunching platform to the HBase distributed database, the Hive query language, and the Oozie workflow system. The stack is used or will be used for the company's iAds mobile-advertising platform.

"Apple advertising provides an opportunity to redefine the advertising on mobile devices," the job listing reads. "It's an exciting environment and a fast-paced development organization. We are looking for senior Hadoop engineer to be part of a dynamic team building highly performant and scalable applications."

Just because we felt like being ignored, we asked Apple about the job listing. It has, well, not responded.

The iAds platform has already launched — it went live this summer — but the job listing could indicate that Apple has not yet moved the platform to Hadoop. "Candidate will be responsible for the following: design and build scalable Hadoop based ETL infrastructure, build modular components in the large volume data movement and management, mentor junior engineers in Hadoop based technologies, work with architects in implementing the data pipeline."

The job requirements include not only experience with high-throughput and scalable applications and with Java programming, but also "extensive experience" with MapReduce, Hive, and HBase or Cassandra, the open source database platform originally developed by Facebook. Cassandra is separate from the Hadoop stack. The listing also says Apple would be particularly pleased with candidates who have experience with Oozie and Flume, the data-loading platform originally built by Hadoop-happy startup Cloudera.

Named for the yellow stuffed elephant belonging to the son of project founder Doug Cutting, Hadoop also underpins online services operated by everyone from Yahoo! and Facebook and Twitter to, believe it or not, Microsoft. The original open source project mimicked GFS, Google's distributed file system, and MapReduce, Mountain View's distributed number-crunching platform.

In 2004, Google published a pair of research papers on these infrastructure technologies, and Doug Cutting used the papers to build a platform that would back Nutch, his open source web crawler. Hadoop was open sourced at Apache, and was bootstrapped by Yahoo!, which hired Cutting in 2006. He has since left Yahoo! for Cloudera.

Powerset, the semantic search outfit later purchased by Microsoft, originally built HBase, which mimics Google's BigTable database. And Facebook started Hive, a SQL-like query language for Hadoop.

Google still uses MapReduce and GFS, but with its new "Caffeine" search infrastructure, it has moved on to a new version of GFS, referred to by some within the company as GFS2. ®

5 things you didn’t know about cloud backup

More from The Register

next story
The Return of BSOD: Does ANYONE trust Microsoft patches?
Sysadmins, you're either fighting fires or seen as incompetents now
Oracle reveals 32-core, 10 BEEELLION-transistor SPARC M7
New chip scales to 1024 cores, 8192 threads 64 TB RAM, at speeds over 3.6GHz
Docker kicks KVM's butt in IBM tests
Big Blue finds containers are speedy, but may not have much room to improve
US regulators OK sale of IBM's x86 server biz to Lenovo
Now all that remains is for gov't offices to ban the boxes
Gartner's Special Report: Should you believe the hype?
Enough hot air to carry a balloon to the Moon
Flash could be CHEAPER than SAS DISK? Come off it, NetApp
Stats analysis reckons we'll hit that point in just three years
Dell The Man shrieks: 'We've got a Bitcoin order, we've got a Bitcoin order'
$50k of PowerEdge servers? That'll be 85 coins in digi-dosh
prev story

Whitepapers

Endpoint data privacy in the cloud is easier than you think
Innovations in encryption and storage resolve issues of data privacy and key requirements for companies to look for in a solution.
Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Top 8 considerations to enable and simplify mobility
In this whitepaper learn how to successfully add mobile capabilities simply and cost effectively.
Solving today's distributed Big Data backup challenges
Enable IT efficiency and allow a firm to access and reuse corporate information for competitive advantage, ultimately changing business outcomes.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.