Wednesday, October 28, 2009

The Information Bottleneck


In his article Institutional Repositories and Research Data Curation in a Distributed Environment, Michael Witt discusses the lack of a framework for organizing information digitally and the effect that has had on datasets.
Witt begins by talking about scientific research of the past and how carefully lab notebooks were kept by scientists and later preserved in archives as part of the scientific record. While I have trouble believing this was always the case, I would agree that changes in technology have changed the way the records are being kept by scientists.
The bulk of his article is spent talking about the existing repository infrastructure of the Purdue Libraries, but I found his ideas about an information bottleneck more interesting. Witt believes that the original data is narrowed for use within the scope of a particular article. This new data is all that most people ever see of the original data. Assuming that the long tail holds true, then that data is valuable for its many different future uses. Unfortunately, the nature of the information bottleneck has left the data stripped of much of its original information and perhaps also of its future value.
Witt insists that datasets need to be presented in context to remain meaningful and useful. And this seems to me a fairly obvious idea. So why aren't we getting the raw data into repositories along with the narrow-focus versions of the data? Based on some of the other readings that I have done, I can only suggest it is because we are still having trouble getting even the narrowly focused data into repositories.
Witt does not address how this is to be done, he just suggests that "...at some point in the future, the process and units of scholarly communication may be reconsidered to fully recognize and include research datasets. In some cases, such as the Human Genome Project, the value of a genome dataset itself is generally recognized to be greater than any single, published finding resulting from its analysis."
I agree with his sentiment, and can imagine many uses for such a collection of datasets. But while we are still struggling to have truly useful repositories, it seems like an idea that will have to remain in the future.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.