Wednesday, September 9, 2009

So who is doing data curation?

Reading the articles for this week, I kept wondering if any university libraries are actually doing data curation. I even had the opportunity to ask a couple of librarians associated with the UT Digital Repository about it. They kinda smirked and looked at each other and said, "sure, we'd love to do that." As in, maybe some other lifetime. So I set about trying to find institutions that are starting to actually preserve and provide access to scientific data.

One article I found, [G. Sayeed Choudhury. "Case Study in Data Curation at Johns Hopkins University."Library Trends 57.2 (2008): 211-220. Project MUSE. 11 Jun. 2009]. discusses data curation at Johns Hopkins. In the beginning, Johns Hopkins wanted to use its DSpace repository to house scientific data, specificially data from observatories. They quickly found this wasn't an ideal solution because the repository is designed for documents, not data. Also, the repository is branded for JHU and the data comes from a collaboration of many institutions. I found it interesting that they wanted to use their repository for the data. Is that considered cyberinfrastructure? It turns out that JHU doesn't actually curate data, but they do propose a method for linking online datasets to electronic publications. It's a good idea, but not really a case study in data curation.

University of Minnesota is also interested in data curation. In 2006 three librarians were hired to spend at least some of their time researching and documenting emerging needs for e-science. Before that, UMN conducted a survey of their researchers asking them questions related to the library and e-science. According to the article [Leslie M. Delserone. "At the Watershed: Preparing for Research Data Management and Stewardship at the University of Minnesota Libraries." Library Trends 57.2 (2008): 202-210. Project MUSE.], the scientists agreed that the libraries could play a role in helping researchers organize and manipulate their data. They also expressed frustration at the lack of standards for storing, securing, and sharing data. But regarding the preservation of data the researchers had some interesting responses. One was, "Am I worried it [data] won’t be there in 20 years? No. Am I worried it won’t be there in 100? It doesn’t matter. By that point, data become irrelevant except as historical curiosity" and another: "It’s important to maintain data for two or three years—saved on disks—but after that the field moves so quickly that it’s no longer relevant…I hadn’t really thought much about [researchers who might be interested in the work in 10, 20 or 30 years]. But it wouldn’t be good if they couldn’t find the data, would it?" Should we chalk that up to an ignorant (i.e. non-LIS) response, or is that attitude a legitimate one? That aside, UMN, like JHU, considered using their institutional repository as the place to store their data,but ultimately decided it wouldn't work. The UMN Libraries and the university itself each created working groups to explore data curation at UMN, but to date, there is no actual curation taking place. At least they're talking about it.

Purdue has formed D2C2 (Distributed Data Curation Center) to assist researchers in preserving and providing access to their data collections. "In its first eighteen months, the D2C2 tracked the submission of over forty grant proposals that included more than twenty different librarians as named collaborators. The center attained its goal of procuring over $1,000,000 of research support its first year, the majority of which supported research into data curation." [Michael Witt. "Institutional Repositories and Research Data Curation in a Distributed Environment." Library Trends 57.2 (2008): 191-201. Project MUSE. 11 Jun. 2009 ]. So, they are doing research into data curation, but don't appear to have started yet.

These articles were found in the most recent issue of
Library Trends. So, to sum up, it appears that data curation by libraries is crazy new and digital repositories are not a good solution for housing data. I guess the cyberinfrastructure isn’t quite there yet.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.