Amazon Elastic MapReduce
Based on Hadoop, MapReduce equips users with potent distributed data-processing tools
- Doesn't take long to get the hang of
- Currently available in the US region only
You'll want to be familiar with the Apache Hadoop framework before you jump into Elastic MapReduce. It doesn't take long to get the hang of it, though. Most developers can have a MapReduce application running within a few hours.
Have you got a few hundred gigabytes of data that need processing? Perhaps a dump of radio telescope data that could use some combing through by a squad of processors running Fourier transforms? Or maybe you're convinced some statistical analysis will reveal a pattern hidden in several years of stock market information? Unfortunately, you don't happen to have a grid of distributed processors to run your application, much less the time to construct a parallel processing infrastructure.
Well, cheer up: Amazon has added Elastic MapReduce to its growing list of cloud-based Web services. Currently in beta, Elastic MapReduce uses Amazon's Elastic Compute Cloud (EC2) and Simple Storage Service (S3) to implement a virtualized distributed processing system based on Apache Hadoop.
Hadoop's internal architecture is the MapReduce framework. The mechanics of MapReduce are well documented in a paper by J. Dean and S. Ghemawat [PDF], and a full treatment is beyond the scope of this article. Instead, I'll illustrate by example.
Suppose you have a set of 10 words and you want to count the number of times those words appear in a collection of e-books. Your input data is a set of key/value pairs, the value being a line of text from one of the books and the key being the concatenation of the book's name and the line's number. This set might comprise a few megabytes big -- or gigabytes. MapReduce doesn't much care about size.
You write a routine that reads this input, a pair at a time, and produces another key/value pair as output. The output key is a word (from the original set of 10) and the associated value is the number of times that word appears in the line. (Zero values are not emitted.) This routine is the map part of map/reduce. Its output is referred to as the intermediate key/value pairs.
The intermediate key/value pairs are fed to another function (another "step" in the parlance of MapReduce). For this step, you write a routine that iterates through the intermediate data, sums up the values, and returns a single pair whose key is the word and whose value is the grand total. You don't have to worry about grouping the results of like keys (i.e., gathering all the intermediate key/values for a given word), because Hadoop does that grouping for you in the background.
Join the newsletter!
Samsung QLED 8K TV
Apple iMac Pro
Ballistix Tactical Tracer RGB 3000
Bang and Olufsen Beoplay A9 Speaker
Cartier Calibre de Cartier Diver Watch
Ballistix Sport AT
Toys for Boys
Little Bits DROID Inventor Kit
Oregon Pro WMR500 Weather Station
Nix Pro Colour Sensor
Tivoli PAL BT
ESET Cyber Security Pro for Mac
ESET Internet Security
Osmo Coding Awbie Game
ESET Smart Security Premium
Ikea RIGGAD work lamp with wireless charging
TimeFlip Magnet Simple Time Tracking Device
SmartLens - Clip on Phone Camera Lens Set of 3
Ultimate Ears Wonderboom Bluetooth Speaker
Naztech Xtra Drive Mini + 256GB microSD Card
Need to buy a gift for somebody who loves technology but you can’t afford the big ticket items?
Most Popular Reviews
- 1 Oppo R17 Pro review: Oppo's thriftiest flagship yet drives a hard bargain
- 2 Nokia 7.1 review: A modest and modern mid-tier option
- 3 Tenda Nova MW6 review: A gateway drug for mesh Wi-Fi
- 4 Huawei Mate 20 Pro review: Expensive, but probably the best phone you can buy right now
- 5 Apple iPhone XS review: Astonishment at a price
Latest News Articles
- Brother pitch themselves at SMBs with new 'Inkvestment' options
- Ted’s World of Imaging opening in Sydney
- McAfee QTR sees cryptocurrency mining surge continue in second quarter
- RMIT Online introduces two new Australian University courses for blockchain skills
- Telstra announces new IoT products to help locate things that matter most
PCW Evaluation Team
Microsoft Office continues to make a student’s life that little bit easier by offering reliable, easy to use, time-saving functionality, while continuing to develop new features that further enhance what is already a formidable collection of applications
I’d recommend a Dell XPS 15 2-in-1 and the new Windows 10 to anyone who needs to get serious work done (before you kick back on your couch with your favourite Netflix show.)
It’s useful for office tasks as well as pragmatic labelling of equipment and storage – just don’t get too excited and label everything in sight!
The Brother MFC-L8900CDW is an absolute stand out. I struggle to fault it.
I need power and lots of it. As a Front End Web developer anything less just won’t cut it which is why the MSI GT75 is an outstanding laptop for me. It’s a sleek and futuristic looking, high quality, beast that has a touch of sci-fi flare about it.
If you’re looking to invest in your next work horse laptop for work or home use, you can’t go wrong with the MSI GE63.
- PC World 2018 Editor's Choice Awards
- Huawei Mate 20 Pro review: Full, in-depth, Australian review
- Razer Phone 2 review: One for the fans
- Everything you need to know about Smart TVs
- What's the difference between an Intel Core i3, i5 and i7?
- Laser vs. inkjet printers: which is better?