Feeds

Presto: Facebook reveals exabyte-scale query engine

'Fast queries over a 250 PETABYTE data warehouse? That's nothing!'

3 Big data security analytics techniques

Facebook has revealed a query engine for data warehouses that blows the doors off Hive, and plans to publish it as open source this year.

The "Presto" technology is a query execution engine built for Facebook's vast data warehouse. It was announced on Thursday at a data analytics conference hosted at Facebook's HQ in Menlo Park, California. Presto gets rid of some of the failings of Hive – the Hadoop data-warehouse tool – and highlights how the Hadoop ecosystem is maturing.

"We built Presto from the ground up to deal with FB scale," Facebook engineer Matin Traverso, says. "It can handle all the 250PB of data we have in our data warehouse – thousands of machines across multiple global regions."

Presto has demonstrated a four-to-seven times improvement over Hadoop Hive for CPU efficiency, and is eight to 10 times faster than Hive in returning the results of queries.

"The problem with Hive is it's designed for batch processing," Traverso said. "We built Presto from the ground up to deal with FB scale."

Presto is Facebook's attempt to speed up queries on Hadoop, and the project is functionally similar to Cloudera's Impala and Hortonworks' Stinger technologies.

Like these, Presto gets rid of the job part of Hadoop – MapReduce – and instead uses a special-purpose query engine, which is SQL-like ANSI-SQL compatible, with some additional features that Facebook will reveal in the next few months.

This is for both ease of use by Facebook developers, and to supercharge the performance of queries over very, very big datasets.

"One of the things Presto can do that MapReduce can't – Presto can start all the stages at once and can stream all the data through the stages," Traverso says.

Before our dear commentards point out that most of Facebook is pictures of cats, updates about bodily functions, nihilistic ramblings, and the pingings of Zynga games feeding e-stims to folk, it bears noting that none of this really matters for designing massive data systems – when you abstract away from the content, you have a set of different things that are deluging your system in data, and you need to deal with them.

And Facebook has more data than most. The company's existing data warehouse is 250PB in size, and growing rapidly: 600TB is added to the warehouse every day.

"As we project our growth, it's quite clear that at some point soon we will reach one exabyte," Ravi Murthy, a Facebook engineering manager, says. "We have to rethink a lot of different things. Not just the software pieces of it, but literally the entire stack."

Most of Facebook's data ends up being stored in the Hadoop Distributed File System, so although some may question why Facebook doesn't just use a SQL DB engine for its queries, the reason is that it needs to have as few layers of abstraction between it and the underlying HDFS data. For that reason, creating add-ons that inteface directly with HDFS, such as Presto, is better for performance than abstracting away.

Since launching at the end of last year, Presto has grown to have 850 internal users per day performing 27,000 queries and fiddling with 320TB of data. Scale aside, the adoption is impressive given Facebook's penchant for a flat organizational stucture that means engineers are not forced to use any particular software package – they either do or they don't, and an app's fortunes are tied closely to its adoption. Presto seems to have been given the thumbs up.

The software should be available as open source by the end of this year. ®

SANS - Survey on application security programs

More from The Register

next story
This time it's 'Personal': new Office 365 sub covers just two devices
Redmond also brings Office into Google's back yard
Kingston DataTraveler MicroDuo: Turn your phone into a 72GB beast
USB-usiness in the front, micro-USB party in the back
Dropbox defends fantastically badly timed Condoleezza Rice appointment
'Nothing is going to change with Dr. Rice's appointment,' file sharer promises
Inside the Hekaton: SQL Server 2014's database engine deconstructed
Nadella's database sqares the circle of cheap memory vs speed
BOFH: Oh DO tell us what you think. *CLICK*
$%%&amp Oh dear, we've been cut *CLICK* Well hello *CLICK* You're breaking up...
Just what could be inside Dropbox's new 'Home For Life'?
Biz apps, messaging, photos, email, more storage – sorry, did you think there would be cake?
IT bods: How long does it take YOU to train up on new tech?
I'll leave my arrays to do the hard work, if you don't mind
Amazon reveals its Google-killing 'R3' server instances
A mega-memory instance that never forgets
prev story

Whitepapers

Designing a defence for mobile apps
In this whitepaper learn the various considerations for defending mobile applications; from the mobile application architecture itself to the myriad testing technologies needed to properly assess mobile applications risk.
3 Big data security analytics techniques
Applying these Big Data security analytics techniques can help you make your business safer by detecting attacks early, before significant damage is done.
Five 3D headsets to be won!
We were so impressed by the Durovis Dive headset we’ve asked the company to give some away to Reg readers.
The benefits of software based PBX
Why you should break free from your proprietary PBX and how to leverage your existing server hardware.
Securing web applications made simple and scalable
In this whitepaper learn how automated security testing can provide a simple and scalable way to protect your web applications.