This week we read about all the reasons why scientists do not share data, and why they should do so. I agree that data sharing is important. Unlike the depressing article we read by Sterling, T., & Weinkam, J., I even think that it may be possible to encourage more data sharing.
In response to the articles we read this week, I read this article, in which the authors make the point that in sustained knowledge sharing, "a distinction between knowledge-contribution and knowledge-seeking behaviors and an adequate emphasis on their variance in terms of user belief is needed." The authors argue that a Knowledge Management System should support both rolls, so that people will trust it as a place to put their knowledge. This will encourage knowledge sharing because people will seek prestige. I think that is true, infrastructure makes things easier, but I still have lingering questions about why scientists would share their data freely, and if they should have to.
There are reasons for not sharing data that I sympathize with, and reasons that I fundamentally think are wrong. For example, a scientist who will not share their data because they do not want anyone to disprove their finding is wrong. I think that disproving a theory is a reason to force data to be open, not the opposite. Authors are not protected from criticism, and I think it is equally important that scientists should not be protected from assessment.
The privacy of data can be protected. Anonymize the data, and make available what you can. Yes, it takes more work, but we have a responsibility to say, this is part of being a good scientist. You must share your data, so if some information needs to be protected, that should be planned for from the start.
The reasons that I can understand are more difficult to address. For example, if a scientist works long and hard to acquire data, they should have the right to not only use that data first and for whatever they can think of to do with it, but also to get credit when someone else uses the gathered data. However, data is a slippery slope of definitions. And getting credit for data seems to be a worse tangle.
According to the Stanford Fair Use Overview, "There are some things that copyright law will not protect. Copyright will not protect the titles of a book or movie, nor will it protect short phrases such as "Make my day." Copyright protection also doesn't cover facts, ideas or theories. These things are free for all to use without authorization."
There are others laws, like patent laws, that may cover some of these, but patent law can be a bit misunderstood as well, take the famous patented peanut butter and jelly sandwich. In addition, a phrase could be trademarked, but the likelihood is that scientists are not using anything in their data set as an advertising tool (I hope...)
But facts, ideas, and theories are not protected. So if data is a fact, then that data is not meant to be protected by copyright law. The way that fact is presented may be, but wouldn't that be the paper written about the data, not the data itself?
According to one article we read this week, Data at Work, there are both experimentalists, and theoretical modelers. If we need both (and I hope we do, because I consider myself an experimentalist) then we need to be certain that the experimentalist is encouraged to contribute. Regathering data seems awfully inefficient.
If I collect data, and share it, is there a way that I can be repaid for the time and effort I put into that? Or do I just need to hoard my data and keep publishing using it?
Is collecting data a public good, like utilities, and therefore should it be funded, alleviating the need for sustained reward as a motivation for researchers?
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.