Social context for data analysis
How data analysis could work productively on the Web
Jon Udell (InfoWorld) 14/12/2006 13:36:03

I'm a huge fan of the CAPStat (formerly DCStat) program, but despite my cheerleading, the hoped-for citizen-led mashups haven't yet materialized in a big way.

In principle, the data is there for the taking, and there's an open invitation for anyone to scoop it up and do useful analysis. In practice, only half the battle is won -- thanks to the immediate availability of data represented as RSS, Atom, and the district's own, richer flavor of XML. It's great to lay your hands on the data, but as Bob Glushko rightly insists on reminding me, XML only seems to be a self-describing format. What do tags or field names really mean? Which elements or fields are or are not comparable? We can only answer these questions by pointing to instances of data (records, documents), discussing them, and coming to agreements.

Lately, I'm seeing some intriguing glimpses of how that process could work productively on the Web. One stunning example Dabble DB , which enables you to pluck data right from the surface of a Web page and inject it into a shareable Web database. Once it's there, the whole panoply of Web-2.0-style techniques -- linking, tagging, blogging -- can support a loosely coupled conversation about the provenance and the semantics of the data.

Today I found another piece of the puzzle -- a new site called Swivel . It's done in the standard Web 2.0 style, complete with regulation Flickr-blue search buttons and Ruby on Rails URL syntax. To tell you the truth, I'm not sure how useful it'll turn out to be. But the idea at the core of Swivel -- inviting people to publish, annotate, and share datasets -- is spot on.

As a first experiment, I grabbed the CAPStat reported-crime feed for November, sucked it into Excel 2003, consolidated incidents by day, pivoted them on type of offense (homicide, burglary), and exported them back out as a CSV (comma-separated value) file that Swivel could import. The service immediately produced a chart for each of the nine crime types in my data set. Eventually the site will "swivel" my data, a process of further analysis that it assures me will be "worth the wait." I dunno, maybe -- I'm not holding my breath. Poking around, I haven't found any breathtaking examples of mechanical insight.

But there's something a lot simpler, yet I think also a lot more useful, going on here. The charts are fun to look at, but it's the data (and the source attributions) that really matter. When it's parked in the cloud, other people can find it by way of search terms ('washington,' 'burglary,' 'arson,' 'dcstat'). And whether Dabble DB massages it online or Excel does so locally, they can gather around a common URL to discuss how to use and interpret the data.

Data analysis is an inherently social act. Until now it has lacked an appropriate social context. But that's going to change -- and soon, I hope.

More about ACT
Recommend this article?
Yes0 votes
No0 votes

Comments

Post new comment

The content of this field is kept private and will not be shown publicly.
  • Web page addresses and e-mail addresses turn into links automatically.
  • Allowed HTML tags: <a> <em> <strong> <cite> <code> <ul> <ol> <li> <dl> <dt> <dd>
  • Lines and paragraphs break automatically.

More information about formatting options

Enter the fully qualified URL, eg. http://www.example.com/
Users posting comments agree to the PC World comments policy.
Login or register to link comments to your user profile, or you may also post a comment without being logged in.
Syndicate content
 
Gift Guide
MWave
Samsung

CXO Latest

LED Advisor
 

Colour your world with Samsung

A chance to win with every
Samsung Consumable purchase*