Cloudera positions Hadoop as an enterprise data hub

Hadoop can now serve as the single source of data for an enterprise, Cloudera says

Taking note of how customers have been working with its Hadoop distribution, Cloudera has expanded the scope of its software so that it can serve as a hub for all of an organization's data, not just data undergoing Hadoop MapReduce analysis.

Some of Cloudera's enterprise customers have "started to use our platform in a new way, as the center of their data centers," said Mike Olson, Cloudera's chairman and chief strategy officer.

"We think this is a very big deal. It will change the way the industry thinks about data," Olson said.

Cloudera has released a new beta of its commercial distribution, Cloudera Enterprise, that provides tools for managing an organization's data, as well as tools from Cloudera and third parties for data analysis.

Olson announced the beta of Cloudera Enterprise 5 at the O'Reilly Strata-Hadoop World conference, being held this week in New York.

"It used to be that an organization had lots of balkanized data silos," Olson said. "The stuff that you used to run on a data warehouse because you had no choice, now you can run on the hub."

Putting the data in a Hadoop-based storage repository has many advantages, Olson argued. You can run different types of analytical workloads against the data in the hub. It can easily feed data to other systems, such as content management systems. It can work as an archiving system.

An enterprise data hub, Olson said, can store data as it is generated, even if the organization isn't sure how the data will be needed. Such data may be valuable later for machine learning analysis or other uses not considered.

An enterprise hub also puts security and governance mechanisms in place to safeguard the data. Cloudera has been working on these tools for several releases, Olson said.

"Our ambition is to draw more workloads in and make the hub more valuable over time," he said.

Part of Hadoop's newfound ability to act as a data hub comes from software additions in the latest version of the open-source software, Apache Hadoop 2, on which Cloudera Enterprise is built.

The inclusion of YARN (Yet Another Resource Manager), for instance, allows Hadoop to handle multiple analysis applications, not just those that run on the batch process-oriented MapReduce.

To facilitate the hub, Cloudera has also set up a management framework that third-party analysis applications can plug into. SAS, Revolution Analytics, Syncsort and other organizations have ported some of their software to the platform. Porting analysis software requires that the operations be executed in parallel, as data in Hadoop is typically distributed across multiple nodes, Olson said.

Cloudera Enterprise 5 also adds the ability to cache HDFS (Hadoop Distributed File System) contents in the working memory of a server, which can boost query response and data processing times.

The company's Navigator auditor tool now allows analysts and data modelers to search, explore, define and tag datasets. Users can add customized queries to Cloudera's Impala SQL engine. And Cloudera Enterprise 5 can work with the NFS (Network File System) nodes, which should make the process of injecting data into HDFS much easier, Olson said.

The software also now can take snapshots of the data, providing a backup if the original data is lost or destroyed.

Joab Jackson covers enterprise software and general technology breaking news for The IDG News Service. Follow Joab on Twitter at @Joab_Jackson. Joab's e-mail address is Joab_Jackson@idg.com

Join the newsletter!

Or

Sign up to gain exclusive access to email subscriptions, event invitations, competitions, giveaways, and much more.

Membership is free, and your security and privacy remain protected. View our privacy policy before signing up.

Error: Please check your email address.

Tags softwarecloudera

Keep up with the latest tech news, reviews and previews by subscribing to the Good Gear Guide newsletter.

Joab Jackson

IDG News Service
Show Comments

Brand Post

Most Popular Reviews

Latest Articles

Resources

PCW Evaluation Team

Luke Hill

MSI GT75 TITAN

I need power and lots of it. As a Front End Web developer anything less just won’t cut it which is why the MSI GT75 is an outstanding laptop for me. It’s a sleek and futuristic looking, high quality, beast that has a touch of sci-fi flare about it.

Emily Tyson

MSI GE63 Raider

If you’re looking to invest in your next work horse laptop for work or home use, you can’t go wrong with the MSI GE63.

Laura Johnston

MSI GS65 Stealth Thin

If you can afford the price tag, it is well worth the money. It out performs any other laptop I have tried for gaming, and the transportable design and incredible display also make it ideal for work.

Andrew Teoh

Brother MFC-L9570CDW Multifunction Printer

Touch screen visibility and operation was great and easy to navigate. Each menu and sub-menu was in an understandable order and category

Louise Coady

Brother MFC-L9570CDW Multifunction Printer

The printer was convenient, produced clear and vibrant images and was very easy to use

Edwina Hargreaves

WD My Cloud Home

I would recommend this device for families and small businesses who want one safe place to store all their important digital content and a way to easily share it with friends, family, business partners, or customers.

Featured Content

Product Launch Showcase

Don’t have an account? Sign up here

Don't have an account? Sign up now

Forgot password?