Feeds

Steve Jobs embraces Google's bastard grid child

Cult plants ad empire on Hadoop

Internet Security Threat Report 2014

Apple has embraced Hadoop, the open source distributed-computing platform based on Google's famously proprietary backend infrastructure.

According to a recent Apple job listing entitled "Senior Software Engineer - Hadoop", the company is using or planning to use the entire Hadoop stack, from the HDFS file system and the Hadoop MapReduce distributed number-crunching platform to the HBase distributed database, the Hive query language, and the Oozie workflow system. The stack is used or will be used for the company's iAds mobile-advertising platform.

"Apple advertising provides an opportunity to redefine the advertising on mobile devices," the job listing reads. "It's an exciting environment and a fast-paced development organization. We are looking for senior Hadoop engineer to be part of a dynamic team building highly performant and scalable applications."

Just because we felt like being ignored, we asked Apple about the job listing. It has, well, not responded.

The iAds platform has already launched — it went live this summer — but the job listing could indicate that Apple has not yet moved the platform to Hadoop. "Candidate will be responsible for the following: design and build scalable Hadoop based ETL infrastructure, build modular components in the large volume data movement and management, mentor junior engineers in Hadoop based technologies, work with architects in implementing the data pipeline."

The job requirements include not only experience with high-throughput and scalable applications and with Java programming, but also "extensive experience" with MapReduce, Hive, and HBase or Cassandra, the open source database platform originally developed by Facebook. Cassandra is separate from the Hadoop stack. The listing also says Apple would be particularly pleased with candidates who have experience with Oozie and Flume, the data-loading platform originally built by Hadoop-happy startup Cloudera.

Named for the yellow stuffed elephant belonging to the son of project founder Doug Cutting, Hadoop also underpins online services operated by everyone from Yahoo! and Facebook and Twitter to, believe it or not, Microsoft. The original open source project mimicked GFS, Google's distributed file system, and MapReduce, Mountain View's distributed number-crunching platform.

In 2004, Google published a pair of research papers on these infrastructure technologies, and Doug Cutting used the papers to build a platform that would back Nutch, his open source web crawler. Hadoop was open sourced at Apache, and was bootstrapped by Yahoo!, which hired Cutting in 2006. He has since left Yahoo! for Cloudera.

Powerset, the semantic search outfit later purchased by Microsoft, originally built HBase, which mimics Google's BigTable database. And Facebook started Hive, a SQL-like query language for Hadoop.

Google still uses MapReduce and GFS, but with its new "Caffeine" search infrastructure, it has moved on to a new version of GFS, referred to by some within the company as GFS2. ®

Internet Security Threat Report 2014

More from The Register

next story
Docker's app containers are coming to Windows Server, says Microsoft
MS chases app deployment speeds already enjoyed by Linux devs
IBM storage revenues sink: 'We are disappointed,' says CEO
Time to put the storage biz up for sale?
'Hmm, why CAN'T I run a water pipe through that rack of media servers?'
Leaving Las Vegas for Armenia kludging and Dubai dune bashing
Facebook slurps 'paste sites' for STOLEN passwords, sprinkles on hash and salt
Zuck's ad empire DOESN'T see details in plain text. Phew!
SDI wars: WTF is software defined infrastructure?
This time we play for ALL the marbles
Windows 10: Forget Cloudobile, put Security and Privacy First
But - dammit - It would be insane to say 'don't collect, because NSA'
Oracle hires former SAP exec for cloudy push
'We know Larry said cloud was gibberish, and insane, and idiotic, but...'
prev story

Whitepapers

Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Cloud and hybrid-cloud data protection for VMware
Learn how quick and easy it is to configure backups and perform restores for VMware environments.
Three 1TB solid state scorchers up for grabs
Big SSDs can be expensive but think big and think free because you could be the lucky winner of one of three 1TB Samsung SSD 840 EVO drives that we’re giving away worth over £300 apiece.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.
Security for virtualized datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.