Tuesday, October 6, 2009

Data-driven Scholarship in the Sciences and Humanities


In 2007, a piece entitled "The Virtual Observatory and the Roman de la Rose: Unexpected Relationships and the Collaborative Imperative" was posted on Academic Commons. It was authored by Sayeed Choudhury, the Director for Library Digital Programs and Hodson Director of the Digital Knowledge Center for the Sheridan Libraries at Johns Hopkins University, and Timothy Stinson, a post-doc fellow at Johns Hopkins with dual appointment in the Digital Research Curation Center and the Department of English. The article is concerned about the impact of cyberinfrastructure on the way that scholarship is conducted. The authors cite the Virtual Observatory and the Roman de la Rose Digital Library as two examples from different fields that exhibit the benefits afforded to scholarship thanks to cyberinfrastructure. The Virtual Observatory is a resource that brings together data from ground- and space-based telescopes, allowing scientists to find, retrieve, and analyze this astronomical data. The Roman de la Rose Digital Library, a collaborative effort between the Sheridan Libraries of Johns Hopkins University and the Bibliotheque national de France, is a digital library the goal of which is to create and house digital copies of all extant manuscripts of the Roman de la Rose.

Most would agree that the hard sciences are moving toward a data-driven model of scholarship, as evidenced by the Virtual Observatory project. However, the authors argue that the humanities, too, are beginning to be able to exploit (thanks to cyberinfrastructure) "data" for research in much that same way as the sciences. The parallel might at first seem not to fit, until one considers that data for the sciences and data for the humanities are two different things. For the sciences, data consist of sensor readings, scanner output, measurements of various kinds, and other things that are largely numerical in nature. However, humanities scholarship is based on a different, less straightforward type of data. As the authors of the Academic Commons article say, humanities materials can be considered data-rich in a variety of ways. For example, a single manuscript of a medieval text like the Roman de la Rose can contain many types of data in the form of illuminations, "artwork"/images, marginalia, annotations, as well as the semantic and linguistic data contained in the text itself. The creation of large digital libraries of this type is, for the humanities, analogous the creation of huge, complex datasets that make up digital curation projects like the Virtual Observatory, even though they seem comparatively data-poor.

The authors of this article further state that by exploiting resources such as the Roman de la Rose Digital Library, the humanities are poised to enter into the kind of collaborative scholarship that characterizes the hard sciences. "Digital tools," they assert, "are allowing us to capture, manipulate, and examine books and their data in ways that are revolutionizing the humanities." Digital versions of entire libraries can be created, and the contents of libraries that have been distributed over time can be virtually reassembled, creating new possibilities for scholarship. The authors recommend that humanities scholars work to define their own needs in the realm of cyberinfrastructure, and to create the specifications for and build the tools that they require for research.

The authors conclude by recognizing the imperative role that libraries and librarians will play in the future of these new data sets. While the role of the library in preserving the physical objects on which something like the Roman de la Rose Digital Library is based will remain vital, the library's role in preserving and making accessible the digital data that is created is one that is even more vital, all the more so because "the datasets are often not as highly regarded by libraries."

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.