While poking around in Michael Nielsen's bookmarks on delicious, I found this one for The Open Dinosaur Project, and I thought it was a good example both of a digital curation project and of the use of crowd-sourcing. The goals of The Open Dinosaur Project are " to involve scientists and the public alike in developing a comprehensive database of dinosaur limb bone measurements, to investigate questions of dinosaur function and evolution." Phase I of the project will involve using data submitted by the public to discover patterns of limb bone evolution in ornithischian dinosaurs, and how this relates to the evolution of locomotion in this group. They accept submissions of measurements harvested from scholarly literature or measurements obtained first hand (though scholarly literature is their primary source). And, they accept submissions from anyone, regardless of age, education, or qualifications. Preliminary results will be blogged on the website, and the final paper will be submitted to a scholarly journal for peer review. All of those who contribute to data collection will be listed as junior authors. The projected dates for completion are as follows: completion of data collection by 1 February 2010, completion of data analysis by 1 March 2010, and submission of final paper by 1 April 2010. The process for submitting measurements involves several levels of verification. First, investigator 1 fills out and emails a Data Entry Sheet and relevant bibliographic references to the project lead (the "curator" of the database, Andrew Farke). The project lead posts the full dataset, excluding the measurements, to the verification list. At this point, investigator 2 downloads the verification list, sees that species 1 needs to be verified, and downloads the PDFs of the relevant papers. Investigator 2 then enters the data for species 1 into the verification list, and emails it to the project lead. The project lead then compares the data sets from investigators 1 and 2, identifies and corrects any discrepancies, and posts the data to the public spreadsheet. The verification lists and public spreadsheet are both available to the investigators through Google docs.
As you can see, the process of submission basically involves culling information from the existing scholarly literature on paleontology. The role of the investigator is to filter out this information and submit it to be included in the ODP database. The participation of two investigators in the verification of each data set insures that mistakes are kept to a minimum. While the acceptable sources of data do not seem to be restricted, the project lead (Andrew Farke) presumably has the ultimate say in whether or not a data set is included in the ODP database. It is not clear whether or not the reliability of the source is considered when data sets are submitted for verification. The ODP does have a number of links to openly accessible journals, but it does not say that these are the only acceptable data sources.
The fact that every data contributor will be listed as a junior author on the published scholarly paper is one of the most interesting aspects about this project. We have talked a lot in class about the current system of tenure, and the changes that might allow scholars to get credit for data sharing. In the ODP system of data contribution, scholars can get credit for data sharing because submitting their own data will allow them to become authors of a scholarly work. Even if people don't contribute original data, they still get credit for their contribution in a form that is accepted in the current paradigm, namely publication. This kind of crowd-sourcing might be a viable way to integrate data sharing into the current system of tenure, and thereby change the idea that the only tenure-worthy contribution to academia is scholarly publication.
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.