Tuesday, September 8, 2009

Exploring Research Data Hosting at the HKUST Institutional Repository

An article by Gabrielle K.W. Wong

This article is linked on the front page of the Distributed Data Curation Center site.

Wong makes it clear that the problem of collecting a large amount of data has changed from being an issue with drive space to an issue of management of the data. Wong claims that although it is difficult to manage data, the value of data increases when it is accessible. She also claims that universities are well suited for the responsibility of managing data because there are existing repositories at so many universities. This suggestion is congruent with the advice from other experts, too. For instance, in the article we read this week from Inge Angevaare, she points out that universities are ensnared in the access model in terms of scholarly journals, but universities remain in control of original data collected. Angevaare even suggests that curating this data has the potential to "revive the library’s unique position at the very heart of the university’s information network".

Wong is from the Hong Kong University of Science and Technology (HKUST) Library and this year they took on the task of creating a DSpace repository. Wong is the institutional repository coordinator so she used OpenDOAR to find 53 repositories and then analyzed how they curate datasets. Not surprisingly, Wong found many different practices throughout all of these repositories. In terms of metadata, DSpace uses Dublin Core, which allows each repository flexibility in their metadata. Of course, this is not neccesarily a good thing for catalogers or users.

The functional problems with existing digital curation does not end with metadata! There's a number of different file formats found throughout the repositories and this brings with it issues of software versions for end-users. Also, there is no standard way to cite datasets. Actually, the very practice of linking data to the research papers written about them isn't standardized. This lack of standardization throughout repositories not only keeps the repositories from being fully accessible, it also creates a significant amount of work for the curators of the collections. Instead of having unified practices, individual projects have to reinvent standards for their datasets.

In addition to her study of DSpace repositories, Wong summarizes the many reasons why researchers do not make their data public, which has been examined by the RIN and explained in the article To share or not to share. This is a brief account of their reasons:


  • lack of career rewards as a major disincentive
  • wish to retain exclusive use of data until all the publication value is extracted
  • lack of time, resources and expertise to handle the data management
  • legal and ethical constraints, such as data ownership issue and confidentiality issue when personal data is involved
  • lack of appropriate archive service
  • fear of exploitation or inappropriate use of the data.


To be honest, I don't see any of these barriers going away, especially in private industry where innovations result in profit. Even if digital curators developed standard practices, there would be a number of researchers reluctant to participate.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.