
For the past few weeks I've been interested in how large data sets are being "visualized". Not only is this a fun (and easy) way to digest large issues, but the data must come from somewhere.
Some blogs I've been looking at lately include Information Aesthetics, Information is Beautiful, and Ben Fry's Projects, especially a visualization of changes in editions of Darwin's Origin of the Species.
I explored these sites to see where the authors are getting their data. Information is Beautiful provides links to his sources, which open as Google document spreadsheets. David McCandless, the author, also links to The Guardian's data (since the newspaper's DataBlog is also linked) and this also opens to spreadsheets. Information Aesthetics' Andrew Vande Moere doesn't provide his data sources, but does offer links to other data/information sites for hours of wasting/exploring. Ben Fry's The Preservation of Favoured Traces project states the "text for each edition was sourced from their careful transcription of Darwin's books" although more information about how the text was actually transcribed. The "Data" link on his site, however, is grayed out, so apparently not available. Fry does explain all his projects in detail, however, just maybe not all the information we would want.
An interesting article on the Visual Journalism blog talks about how some data is just boring and not everything needs to be graphically interpreted. The author, Gert K. Nielsen, is discussing a NY Times interactive visual about how people spend their day and mentions that the original data of the graphic, "American Time Use Survey, is in danger of being shut down due to serious cuts in the budget of the Bureau of Labor Statistics". That would hamper the continuation of this visualization, but the Bureau of Labor Statistics will surely have other great data to troll.
The popularity of information visualization will probably only increase (until the novelty runs out?). Ben Fry has written a book about it, Visualizing Data, published by O'Reilly. I wonder if the graphs and pretty colors will distract people from questioning the source of the data, or highlight the nature of statistics and data even more, creating a demand for source knowledge. I'm hoping the latter, especially with easy sources like data.gov, there is no reason not to provide the raw data. I also like the possibility of increased data sharing this might facilitate. Even if it starts out small, such as the GoogleDocs on Information is Beautiful, this creates a society of openness that can be emulated.
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.