Amazon Elastic MapReduce
Based on Hadoop, MapReduce equips users with potent distributed data-processing tools
- Doesn't take long to get the hang of
- Currently available in the US region only
You'll want to be familiar with the Apache Hadoop framework before you jump into Elastic MapReduce. It doesn't take long to get the hang of it, though. Most developers can have a MapReduce application running within a few hours.
Have you got a few hundred gigabytes of data that need processing? Perhaps a dump of radio telescope data that could use some combing through by a squad of processors running Fourier transforms? Or maybe you're convinced some statistical analysis will reveal a pattern hidden in several years of stock market information? Unfortunately, you don't happen to have a grid of distributed processors to run your application, much less the time to construct a parallel processing infrastructure.
Well, cheer up: Amazon has added Elastic MapReduce to its growing list of cloud-based Web services. Currently in beta, Elastic MapReduce uses Amazon's Elastic Compute Cloud (EC2) and Simple Storage Service (S3) to implement a virtualized distributed processing system based on Apache Hadoop.
Hadoop's internal architecture is the MapReduce framework. The mechanics of MapReduce are well documented in a paper by J. Dean and S. Ghemawat [PDF], and a full treatment is beyond the scope of this article. Instead, I'll illustrate by example.
Suppose you have a set of 10 words and you want to count the number of times those words appear in a collection of e-books. Your input data is a set of key/value pairs, the value being a line of text from one of the books and the key being the concatenation of the book's name and the line's number. This set might comprise a few megabytes big -- or gigabytes. MapReduce doesn't much care about size.
You write a routine that reads this input, a pair at a time, and produces another key/value pair as output. The output key is a word (from the original set of 10) and the associated value is the number of times that word appears in the line. (Zero values are not emitted.) This routine is the map part of map/reduce. Its output is referred to as the intermediate key/value pairs.
The intermediate key/value pairs are fed to another function (another "step" in the parlance of MapReduce). For this step, you write a routine that iterates through the intermediate data, sums up the values, and returns a single pair whose key is the word and whose value is the grand total. You don't have to worry about grouping the results of like keys (i.e., gathering all the intermediate key/values for a given word), because Hadoop does that grouping for you in the background.
Join the newsletter!
Panasonic OLED 4K Ultra HD TV - TH-55EZ950U
Dyson Supersonic™ Hair Dryer Fuchsia/Iron
Apple iPhone X
Bang and Olufsen BeoVision 14
Panasonic OLED 4K Ultra HD TV - TH-77EZ1000U
SanDisk MicroSDXC™ for Nintendo® Switch™
Breitling Superocean Heritage Chronographe 44
WD MY PASSPORT™ X Gaming Storage
WD MY PASSPORT™ Gaming Storage
cloudandco Smart Cane
Toys for Boys
Ubiquiti Network’s Front Row Camera
Propel Star Wars T-65 X-Wing Drone
Lego Mindstorms EV3
Bose SoundLink Micro
UBTech First Order Stormtrooper Robot
Onyx Smart Walkie Talkie
LaCie Rugged USB-C Portable Hard Drive
Google Daydream View VR Headset
Leica M10 Digital Rangefinder Camera
Nest Protect Smart Smoke Alarm
Dearear Endear In-ear Wireless Earphones
WD MY CLOUD™ HOME Personal Cloud Storage
Amazon Echo Bluetooth Speaker
iRobot Roomba 980 Vaccum Cleaning Robot
Xbox One X
Belkin Pocket Power 10,000mAh
PETKIG Go Smart Dog Leash
Panasonic Hi-Fi - SC-UA7GS-K
Panasonic 4K UHD Blu-Ray Player and Full HD Recorder with Netflix - UBT1GL-K
Toffee Bags Commuter Satchel
Panasonic Portable Splashproof Fun - RF-D20U
Tile Pro Bluetooth Tracker
Logitech Doodle Collection Wireless Mouse
Kogan Bluetooth Soundbar
Fallout Geeki Tikis
Urbanworx Full HD Action Camera
3SIXT 3-in-1 Smartphone Lens Kit
Ikea NORDMÄRKE Wireless Charging Pad
Lexon Flip Alarm Clock
Razer DeathAdder Expert Ergonomic Gaming Mouse
Raspberry Pi Starter Kit
Most Popular Reviews
- 1 Fitbit Ionic review: Impressive but not quite iconic
- 2 Acer Spin 5 review: Value for money but conditions apply
- 3 Huawei Mate 10 Pro Review: A solid winter flagship that cribs from the best
- 4 Sony LF-S50G review: Google Assistant and then some
- 5 Google Pixel 2 review: not quite 'pixel perfect' but damn close
Latest News Articles
- ASUS Announces Two New Entries into the VivoBook Range with the VivoBook 14 and VivoBook 15
- US says laptop ban may expand to more airports
- Epson launches new high-speed Enterprise inkjet printer
- HP's Spectre x360 puts Kaby Lake and Thunderbolt into a thinner, faster package
- HP upgrades the Envy 13 laptop with Kaby Lake, debuts the 4K Envy 27 display
PCW Evaluation Team
I would recommend this device for families and small businesses who want one safe place to store all their important digital content and a way to easily share it with friends, family, business partners, or customers.
It’s easy to set up, it’s compact and quiet when printing and to top if off, the print quality is excellent. This is hands down the best printer I’ve used for printing labels.
Brainstorming, innovation, problem solving, and negotiation have all become much more productive and valuable if people can easily collaborate in real time with minimal friction.
The print quality also does not disappoint, it’s clear, bold, doesn’t smudge and the text is perfectly sized.
The Huddle Board’s built in program; Sharp Touch Viewing software allows us to easily manipulate and edit our documents (jpegs and PDFs) all at the same time on the dashboard.
The biggest perks for me would be that it comes with easy to use and comprehensive programs that make the collaboration process a whole lot more intuitive and organic
- CES 2018: Belkin go big on wearables accessories
- CES 2018: Alcatel Embrace 18:9 Aspect Ratio In 2018
- OPPO Load Up A73 Smartphone With Flagship Features
- What's the difference between an Intel Core i3, i5 and i7?
- Laser vs. inkjet printers: which is better?
Product Launch Showcase
- FTNetwork Technical Specialist L3 x 2 ? Large Telco ? 6 month contract initiallyNSW
- CCTest Analyst - Port Macquarie NSWNSW
- FTSenior Java DeveloperNSW
- TP3 x Change and Adoption Managers | Health | 12 month contractsQLD
- FTIT Field TechnicianNSW
- FTNetwork Engineering Team Lead/Network ManagerACT
- TPSystem AnalystACT
- FTIT Contracting & Licensing AnalystOther
- CCNetwork Technical Specialist L3 x 2 ? Large Telco ? 6 month contract initiallyNSW
- CCCloud EngineerNSW
- CCNetwork Engineer (Juniper)NSW
- CCJunior to Mid-level DeveloperQLD
- FTCRM Solution ArchitectACT
- FTiOS DeveloperWA
- CCMachine Learning SpecialistNSW
- FTSolution Architect- SAP S/4 HANA - Orange NSWOther
- CCSenior Project ManagerNSW
- FTFront-End DeveloperSA
- FTPega DeveloperACT
- FTBusiness Analyst - Operational Performance ReportingOther
- CCApplications Development Delivery LeadVIC
- FTSenior Solution Architect - Data CentreACT
- CCSecurity SpecialistQLD
- CCIteration Lead - Telco - Melbourne CBDVIC
- FTUX Design Manager (Urgent!!)Other