Showing posts with label humanities. Show all posts
Showing posts with label humanities. Show all posts

Wednesday, September 9, 2009

Geospatial Data for the Humanities

Elliott, T. & and Gillies, S. (2009). Digital geography and the classics. Digital Humanities Quarterly, 3(1).


The authors here discuss the brief history of digital geographic data in use by humanities studies (specifically classics) through pivotal scholarly projects, ending in a discussion of future possibilities for geographic or spatial data as well as their own Pleiades project.


Certainly spatial information has lagged behind HTML documents in their general findability, but as the authors point out the gap is closing: Googlebot began indexing web documents encoded in the Keyhole Markup Language (KML) along with feeds containing GeoRSS markup in late 2006; Microsoft has added support to its Local Search and Virtual Earth services for similar data in late 2007. In general KML and GeoRSS resources, if properly marked up in an indexable web page, will be discoverable.


Typically the actual coordinates of some spatial resource is extracted from an external dataset (examples: Getty Thesaurus of Geographic Names, GeoNames database, Alexandria Digital Library Gazetteer) based on the name. This can naturally lead to erroneous information on account of name ambiguity (names changing through time, coordinates changing through time, etc.). The authors believe that over time such misjudgments will be minimized through sophisticated name-variant databases as well as automated best-guesses algorithms, like examining other geographic points in an article to determine the likely coordinates of a single ambiguous place name. This is mass auto-extraction of metadata is similar to our discussion of Google Books metadata problems, with a similar refinement-over-time solution.


The authors examine the beginnings of maps and spatial data on the web, charting a path from the earliest server-side applications to heterogenous, continuous-panning, browser-side interfaces that pull data and overlays from multiple resources. I cannot help but look at the battle maps in my group’s chosen digital archive, The Valley of the Shadow, and see that these geographic resources fall into the earlier paradigm. While the interface allows the user to apply a set number of overlays to the unfolding battle animations, the application itself is closed off from other web resources: there are no coordinates, no historical geographic or topological data, etc., available, and there is no way to see such information except by requesting the programmers to add it.


The authors also distinguish between large top-down geo-historical projects like the National Geospatial Digital Archive (US) and Global Spatial Data Infrastructure Association and smaller, “hand-crafted” datasets created by dispersed volunteered geographic information (VGI). The latter, they argue, is quickly eclipsing the achievements of the former. VGI datasets have become more prominent on account of a rise in web applications that can easily publish such information. These applications are increasingly adopting metadata standards, and collaborative projects are arising that address interoperable issues in geo-spatial data like the Register of Geographic Entities (RAGE).


Last the authors discuss their own Pleiades Project. This attempts to establish a standard reference dataset for Classical geography. Most interesting here is their rejection of coordinates and toponyms as the primary organizing theme for geo-historical data, using instead the concepts of place, “understood as a bundle of associations between attested names and measured (or estimated) locations (including areas).” I think this is a really great innovation: it moves away from a more scientific, exacting sense of place (coordinates) to a more humanly and historically meaningful sense of place, one that has the advantage of being more easily discussed and contested in a collaborative environment. Certainly the computer will need to understand something like the Roman Empire in terms of coordinates, but a layer of abstraction (locations and areas) is extremely useful when humans need to play with such data.

Tuesday, September 8, 2009

Keeping up with the Sciences: Humanities and Cyberinfrastructure

Crane, Gregory, Alison Babeu, and David Bamman. "eScience and the Humanities." International Journal on Digital Libraries 7, no. 1 (2007): 117-122.

In this article, Crane, Babeu, and Bamman suggest that though those in the humanities are increasingly working with large digital datasets, they are lagging behind their scientific peers in developing the cyberinfrastructures necessary to support and maintain digital resources. As with the sciences, the humanities need cyberinfrastructure to help address both the massive scale of digital data and the fact that managing this data requires specialized knowledge beyond the capacity of any single researcher. Unfortunately, the humanities have been slow to develop such infrastructures due, in part, to disparities in funding between the sciences and the humanities. The authors note, for example, that the budget of the National Science Foundation (NSF) is 39 times larger than that of the National Endowment for the Humanities (NEH). However funding is not the only culprit. Crane et al. also suggest that a failure of imagination has impeded the humanities: that is, they have been slow to recognize the types of intellectual activity that emerging cyberinfrastructures might support.

In order to grow cyberinfrastructure, the article recommends that the humanities systematically develop alliances with the sciences and other better-funded disciplines, as well as collaborate with them on shared technological interests. International collaboration must also be encouraged. To this end, the authors enumerate five core services that they assert reflect "a convergence of interests that extends beyond the humanities" (120). These services include: 1) Conversion of page images to digital text (including handwritten documents that pre-date the advent of printing), 2) conversion from raw text to structured data (including semantic classification, indentifications, and morphological and syntactic analysis) 3)support of multiple languages (including cross-language information retrieval), 4)customization and personalization of data retrieval, and 5) the support of continuous user contributions (such as corrections of OCR errors). Crane et al. close by recommending that the humanities strive to develop larger, more stable organizational structures (as opposed to constantly re-inventing the wheel in numerous small-scale projects) in order to ensure the maintenance of digital data services. They also hold out hope for the emergence of disciplinary centers--such as one for classicists--to attend to the specialized needs of their constituents.

Though one imagines that the article is meant to rally the humanities and provide real-world strategies for confronting difficult funding situations, the effect is still somewhat depressing. Identifying overlapping interests makes good financial sense, but it seems important too to recognize and investigate significant areas of divergence between the sciences and the humanities. For example, how do humanities "datasets" differ from those in the field of science and social science and what type of functionality would humanities data benefit from most? What constitutes pre-publication "raw data" in the humanities? How does the data life cycle for humanities materials compare with that which Anna Gold outlines for the sciences in "Cyberinfrastructure, Data, and Libraries, Part 1?" Where significant differences exist, what are the cyberinfrastructure implications of these differences? As this week's article by Inge Angevaare notes, a vital part of any digital curation plan is identifying and attending to the needs of one's designated community.

Collaboration with the sciences is undoubtedly key to developing a more robust cyberinfrastructure for the humanities, but, at the same time, humanities agendas should not be unduly shaped by the interests of the sciences. Sharing resources only works in so far as the end product serves the needs of both parties.