Wednesday, September 16, 2009

Data Sharing: Empty archives (if you build it... will they come?)

Bryn Nelson, Nature, 461, 160-163 (2009) doi:10.1038/461160a (Published online September 9, 2009)

I'm very interested in the traditions, notions, and expectations of data sharing that we've been discussing the past couple weeks. Even with the best intentions and adequate incentive, the complexities appear to outweigh the ultimate benefit (at least at this point in time).

This article opens with a brief history of the University of Rochester, NY digital archive. Launched in 2003, the university had spent six months of research and marketing to determine that a university repository would be supported and used - both internally and by the public. Today, the repository is mostly empty. Researchers were very supportive of the idea, but once available, couldn't find their data, didn't know how to use the archive, or complained of lack of time and resources to properly transfer their data.

Nelson observes that like institutional repositories, "if you build it, they will come" also does not yet apply to other data sharing efforts. The concept is widely embraced, but the advantages often do not outweigh the concerns of the researchers. Repositories such as arXiv.org, the Protein Data Bank and the International Virtual Observatory Alliance should be considered exceptions, rather than the norm; their respective fields have established traditions of open access and data requirements.

Mark Parsons, an open-data advocate, manages a global program aiming to preserve and organize data from the International Polar Year project. Data needed, and was mandated, to be made available quickly. "Part of what is driving that is the rapidness of change in the poles," says Parsons. "If we're going to wait five years for data to be released, the Arctic is going to be a completely different place." Unfortunately, he found that the infrastructure needed to be created, and this was not available through his program's resources. His team was able to delegate the collection work to national coordinators, "data-wranglers", who contact the investigators and prepare the data for the databank. Sweden has become one of the most successful at data-wrangling - it formed a subcommittee to correct the lag in data collecting and is now housing data from smaller projects that may not reach the international databanks. However, unlike many countries, there is no practice that requires funded projects to submit project data to data centers.

The complexities to open data vary wildly. Parsons points out that even if the data is available, it is not always clear where the data belongs. Furthermore, funded data centers may be forced to reject data from projects without a related funding source. Other fields grapple with the concerns over quantity and quality. Do data centers keep only the information most likely to be of value (as understood by current science), or keep the vast quantities of data accumulated? Should they be concerned with naive users making premature discoveries, or allow for a fresh perspective? Szabolcs Márka, the lead scientist for LIGO (Laser Interferometer Gravitational-Wave Observatory)
asserts that data centers are not in the business of data analysis, but are charged with the task of provide access to accurate information. Other issues involve proper citation of data and credit given to researchers.

Infrastructure, standards, and culture all need to be created or adapted in order for open data to be as effective as the ideals imply. It is clear that it will take global initiatives, financial support, and likely many test cases to create adequate middleware and workflow techniques to push the data to the centers.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.