Wednesday, September 16, 2009

Creating a culture of data sharing

This week I was also interested in data sharing. I read Bob O’Hara’s blog post on why he wishes there was more data sharing http://network.nature.com/people/boboh/blog/2009/09/14/data-sharing-some-ramblings

The complexities of a data sharing system is something that I kept thinking about while reading Borgman as well, and something that has come up briefly in class a few times (is someone writing a paper on this?) To find out more about why researchers do and don’t share data, I looked at the Nature issue addressing this. http://www.nature.com/news/specials/datasharing/index.html

The article, Data Sharing: Empty Archives http://www.nature.com/news/2009/090909/full/461160a.html discusses why researchers have disincentives to share, and how a community of sharing could be fostered.
In Data Sharing: Empty Archives the author, Nelson, writes that while data sharing is always touted as a good idea, and researchers seem to support it in theory, in practice projects to create data repositories remain empty. The example he gives is a University of Rochester project that cost $200,000 in 2003 and still does not have much in it.

The reasons he gives for supporting data in theory aren’t much different from Borgman’s or from other familiar pro-data-sharing sources, “It opens up observations to independent scrutiny, fosters new collaborations and encourages further discoveries in old data sets.” But they also suggest several questions that complicate the matter, “What will keep work from being scooped, poached or misused? What rights will the scientists have to relinquish? Where will they get the hours and money to find and format everything?”

So why do some open-data projects succeed? The article uses the example of Cornell’s arXiv among others, which is for physicists, mathematicians and computer-scientists. Open-source programs are very successful online today as well, so there are some sharing projects that work, but why can’t most projects fill their databases while others can?

Researchers may be worried about the safety of their data, or may feel that they have intellectual rights to their data, or may simply not have the time to format anything for a database or project. Even when projects “wrangle data” as Nelson puts it, often there is no clear way to categorize it or place it in an infrastructure.

Perhaps our inability to plan how to create standards for data is the obstacle to open-data that we should focus on first. Successful projects such as the human genome project are projects where creating standards is perhaps more straightforward. It does seem that open-source programs may be shared to avoid what would clearly be duplication of effort. Many times, a javascript code to create an alert pop-up would be written very similarly by two completely different programmers. With some research data, it may be more difficult to see how sharing would circumvent that duplication of effort (though I still believe it is often the case.) In fact, many of the more creative computer programming solutions (though not all) are proprietary.


Secondly, as Borgman suggests, Nelson also writes that funds and rewards should be based on data sharing for researchers to see the benefit of it. Nelson writes that sharing is more common when the expectations are clear, and I believe this is true for both standards and rewards.

We do have some new tools that probably will encourage data sharing such as Creative Commons, but we will need a cultural shift to continue to go in that direction.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.