Wednesday, September 23, 2009

A Library Where No One Came

Issue 461 of Nature magazine is devoted to issues surrounding open access to data, with a feature on what happens when researchers don’t take advantage of data repositories and keep their data to themselves. This article is a good reminder that despite all the hard work that information professionals often put into building tools, it is all for naught if the tools aren’t used. It is important here to distinguish sharing of data from sharing of publications in online open access. While the latter is becoming much more common to the point of being the norm in some fields, the former lags behind.


In this article Bryn Nelson explains that different groups of researchers have differing cultures that affect their views on sharing data. Some disciplines may have rigidly structured data, such as genome sequences, while others may have data that doesn’t fit any predefined schema, such as climate data. In addition, researchers are often concerned about misuse of their data or lack of attribution and losing the original rights they have to the data. Some data can’t be understood without looking at a whole other group of related data: an example given in the article is a group who is trying to track down weak gravitational fields and must gather a huge amount of background data such as ocean currents and atmospheric conditions to use as filters on their data. Releasing such primary data to the public runs the risk of people taking it out of context and thinking they have found spurious results.


Nelson suggests that three authorities in these fields exercise their control to help corral the data: publishers grant funders, and scientific societies. Grant agencies can make it a requirement that the projects they fund will have their data collected in a repository, and publishers can ask that authors share their data, although they must be careful not to scare away authors who would turn elsewhere. In addition, Nelson says that large government efforts are required to build the kinds of data standards that will work across scientific communities: it’s not enough to think that ad hoc solutions will work in the long run.


On the one hand, the problems of the researchers are all ones we have seen before: they worry about losing the rights to their work, misattribution, and people drawing unwarranted conclusions from the data. While this is nothing new and can occur with any work put out into the public, there is an expectation of objectivity to data that makes releasing it especially dangerous. The traditional gatekeepers of journals should be able to hold at bay the worst examples of pseudoscience that draw on public data, but perhaps it would be beneficial to build in dependencies to data streams that link data that can’t be separated. Data could be provided as a package, not independent streams.


Nelson goes against much of the current thinking in online collaboration when he states that a government group must facilitate data sharing. This approach is anathema to the rabid supporters of open source software and the like who hold it as doctrine that dividing work up among a group will provide better results than top-down control. I believe Nelson is right however, because data standards, no matter if they are eventually superseded by better standards, must have some sort of common ground to start from in the beginning, and this relies upon the work of an authority. Competing standards could actually serve to slow down progress, as work must be done to translate between standards and generate tools to do this.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.