Salo, Dorothea. "Costs and Service Models for Data Curation." The Book of Trogool (Blog). Posted September 19, 2009. Available at http://scienceblogs.com/bookoftrogool/2009/09/cost_and_service_models_for_da.php (Accessed 10/06/2009).
This week I want to talk about a blog post (rather than an article) that caught my attention: Dorothea Salo's "Costs and Service Models for Data Curation." In her post, Salo discusses her concerns about the differences between how "Big Science" and "small science" will fare in matters of data curation and preservation. Perhaps because it is a blog post and not an article, Salo never quite defines "Big Science" and "small science," but from context and some reading around on the internet, it seems that "Big Science" generally refers to large multiple investigator projects, with big budgets, that tend to involve advanced (and expensive) technologies with large centralized labs (such as CERN), whereas "small science" involves a single investigator or a small team, usually working at a single lab on an independent research program whose funding is dependent on smaller limited-term grants. Salo sees the differences between big and small science as resulting in important disparities that affect data curation.
The problem as Salo, an academic librarian invested in e-research, sets it out is such: Big Science produces large, generally homogeneous data in huge quantities. However, because it's all part of one project, data standards evolve quickly and procedures can be institutionalized and set for the entire project. Thus once you get past the problem of storing huge amounts of data (and Salo is confident that they are coming up with solutions on this front), the human resources cost for curation is relatively small per terabyte of data. Salo characterizes small science, on the other hand, as tending to have ad hoc procedures without real data standards (since there are not enough people working on similar data who are willing to share that data and work together to generate standards). This is a problem because, without standards, each of small science's highly heterogeneous pieces of data will require "individual attention if it is to be adequately described and future-proofed." Accordingly, small science has the potential to have a frighteningly high human resources cost per terabyte.
This is significant because Salo intuits that small science, as a whole, has the potential to generate more research data than Big Science. It is also significant because, as she argues, all research data is potentially equally important since you can't know in advance where the groundbreaking, paradigm-altering insights will come from (especially since those breakthroughs may come from subsequent mash-ups of data). So, small science, which is least equipped to pay for long-term data curation, is likely to have the highest costs. This puts its "equally important" data at risk of getting lost in the data deluge (Salo explains in the subsequent comment posts that she has already seen instances of this type of loss). Salo is skeptical about the model of institutional cost-recovery cyberinfrastructure since she thinks Big Science will opt out (as she also explains in a subsequent comment) in favor of doing their own data curation with an embedded librarian whereas whatever money small science has to contribute may not be enough to adequately fund cost-recovery operations. The question that emerges for Salo is what type of business model will help us to address in a more equitable fashion the data curation disparities between Big and small science.
The comments that follow Salo's post also raise important questions. One of the respondents, Sayeed Choudhury, Director of digital curation at JHU (see Megan's email shout-out from earlier today), questions whether Big Science data really is as homogeneous as Salo presumes it to be. He also takes her to task in general (though, in the nicest way possible) for working from assumptions rather than concrete facts. He notes, for example that medical science is presumably "small science" but that the NIH has much more funding to distribute than the NSF. Choudhury suggests that libraries should become more involved in scientific data curation so that they can actually discover (rather than speculate about) what it entails in order to better build their infrastructures.
Choudhury's criticism that we need to actually compile and generate data on these issues is an important one. Salo's questions are provocative and, I think, important, which is why I've chosen to write about them here, but it is also unclear whether data curation in Big and small science really does work in the manner in which she supposes it does. Someone want to do a fact finding study?
In Atkins et al's NSF report "Revolutionizing Science and Engineering Through Cyberinfrastructure", the authors recognize that there is a "significant need [...] in many disciplines for long-term, distributed, and stable data and metadata repositories that institutionalize community data holdings" (42). In their imagining, these would provide tutorials on data formatting and quality control as well as offer tools for data preparation. This does seem like a potentially useful tactic for coping with data curation costs, but the question remains whether these tutorials and tools will really be equipped to cope with the heterogeneity of data. It also seems that these tools and standards would need to be engaged during the planning stages of researchers' projects if they wanted to avoid prohibitive human-resources costs further downstream. These are potentially important propositions, but we won't know how well such resources will serve small science until the tools are up and running and being widely used.
Tuesday, October 6, 2009
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.