Facebook query software crawls data mountains more quickly

Facebook has updated its Presto SQL query engine for faster performance

Facebook engineers have revamped the company's open source Presto query engine to run up to four times faster, making it even more feasible for very large scale data warehouse work.

"We have a large number of internal users at Facebook who use Presto on a continuous basis for data analysis. Improving query performance directly improves their productivity, so we thought through ways to make Presto even faster," wrote Dain Sundstrom, in a blog post introducing the improvements.

Facebook created Presto to run SQL commands against large amounts of unstructured data, notably data the company kept in Hadoop file systems. Understood by countless database administrators, SQL (Structured Query Language) has long been the cornerstone of relational database systems and data analysis systems.

Last November, Facebook open sourced Presto. Like with many of the programs the company has developed in-house, Facebook is hoping others will deploy the code and submit bug fixes and improvements. Organizations could use the software to run data analysis on sets of data that would be too large, or too expensive to fit into a commercial data warehouse.

To improve performance, Facebook engineers revamped Presto's software for reading data. The software can now directly ingest data from Hadoop systems in columns, rather than by reading them in by rows, which required additional time for restructuring the data.

The software now takes note of the minimum and maximum values of each column, which allows it to more quickly find a set of data with a small range of values. It also evaluates queries more intelligently, employing the user's filter terms to minimize the data set being inspected, saving processing time.

With these improvements, Facebook engineers found that Presto was able to execute single single column reads up to four times faster, as measured against a 600-million-row data set residing on 14 servers, with each server running 16 cores and 64GB of memory.

Presto seems to be gaining popularity outside of Facebook. Argyle uses the software as the basis of its own real-time fraud detection software for telecommunications companies. Presto "gives us a lot of flexibility," said Arshak Navruzyan, Argyle vice president of product management.

Argyle paired Presto with open source data store Accumolo. Argyle's customers keep terabytes, and even petabytes worth of user call data, which is far too large to hold in traditional data warehouses.

Presto queries serve as the basis of Argyle's alert messaging system, which flags unusual behavior for its customers. For Argyle, Presto provides an interface for user-facing business intelligence software such as Tableau.

Airbnb also uses Presto, in conjunction with its own software, to allow employees to access the company's operational data.

Presto is one of a number of SQL query engines developed for interrogating massive amounts of data across multiple servers in parallel. Cloudera developed Impala for similar reasons. Pivotal also created HAWQ for the same purpose, borrowing technology from its Greenplum database technology.

Join the newsletter!

Or

Sign up to gain exclusive access to email subscriptions, event invitations, competitions, giveaways, and much more.

Membership is free, and your security and privacy remain protected. View our privacy policy before signing up.

Error: Please check your email address.

Tags Facebooksoftware

Keep up with the latest tech news, reviews and previews by subscribing to the Good Gear Guide newsletter.

Joab Jackson

IDG News Service
Show Comments

Cool Tech

Toys for Boys

Family Friendly

Stocking Stuffer

SmartLens - Clip on Phone Camera Lens Set of 3

Learn more >

Christmas Gift Guide

Click for more ›

Brand Post

Most Popular Reviews

Latest Articles

Resources

PCW Evaluation Team

Aysha Strobbe

Microsoft Office 365/HP Spectre x360

Microsoft Office continues to make a student’s life that little bit easier by offering reliable, easy to use, time-saving functionality, while continuing to develop new features that further enhance what is already a formidable collection of applications

Michael Hargreaves

Microsoft Office 365/Dell XPS 15 2-in-1

I’d recommend a Dell XPS 15 2-in-1 and the new Windows 10 to anyone who needs to get serious work done (before you kick back on your couch with your favourite Netflix show.)

Maryellen Rose George

Brother PT-P750W

It’s useful for office tasks as well as pragmatic labelling of equipment and storage – just don’t get too excited and label everything in sight!

Cathy Giles

Brother MFC-L8900CDW

The Brother MFC-L8900CDW is an absolute stand out. I struggle to fault it.

Luke Hill

MSI GT75 TITAN

I need power and lots of it. As a Front End Web developer anything less just won’t cut it which is why the MSI GT75 is an outstanding laptop for me. It’s a sleek and futuristic looking, high quality, beast that has a touch of sci-fi flare about it.

Emily Tyson

MSI GE63 Raider

If you’re looking to invest in your next work horse laptop for work or home use, you can’t go wrong with the MSI GE63.

Featured Content

Product Launch Showcase

Don’t have an account? Sign up here

Don't have an account? Sign up now

Forgot password?