Showing posts with label digital curation. Show all posts
Showing posts with label digital curation. Show all posts

Monday, November 23, 2009

Digital Curation & User Testing

Marchionni, Paola. “Why Are Users So Useful?: User Engagement and the Experience of the JISC Digitisation Programme.” Ariadne, no. 61 (October 2009). http://www.ariadne.ac.uk/issue61/marchionni/.


This recent article by JISC's Paola Marchionni refocuses attention on the purpose of curated digital collections: that is, their use by different user groups. Marchionni begins by noting that many digitization projects are still not paying enough attention to their users that their users' needs. They become so caught up in trying to make their content accessible online, that they don't adequately research their key users. As a result many publicly-funded projects are going un- (or under-) used. Marchionni illustrates the insight users can provide by presenting two case studies of projects that incorporated users into their development process: the British Library's Archival Sound Recordings 2 (ASCR2) project (a collection consisting of over 25,000 recordings) and Oxford University's First World War Poetry Digital Archive (WW1PDA) (a collection that contains over 7000 items pertaining to WWI poets, including digitized images of materials held at UT's own Harry Ransom Center).

Though Marchionni's article helpfully reminds digitization (and digital curation) projects to keep their electronic eyes on the prize and really take their users into account, I'm not sure that much of what Marchionni presents in her list of suggestions for user engagement is particularly surprising. She recommends first recognizing the importance of interacting with users and even having an "Engagement Officer" position as the ASR2 project did. She also advises establishing an early and on-going relationship with users. The WW1PDA project, for example, developed a typology of users, with a steering committee of scholars in the field of WWI literature advising which materials should be digitized and participating in quality control, and a separate group of secondary school and higher education instructors helping to develop and offer feedback on the education section of the project. Marchionni also emphasizes the importance of knowing what to do with user feedback. When users expressed anxieties about the integration of Web 2.0 tools out of fear that they might undermine the authority of the WW1PDA archive, the project decided to integrate such functionality in a way that made the lines between the archivists' and the users contributions more clear.

Some of the more interesting lessons regarding users came from the WW1PDA's approach to educational resources. The project held workshops for teachers in order to discover what functionality this group would like to see on the site. In a rather ballsy move, the project then asked the workshop members to help author a number of learning resources for the website. Though this did result in the creation of some resources, ultimately the project realized that perhaps it had overreached in what it was asking their busy users to produce. (Frankly, I would be a little annoyed if I agreed to participate in a workshop on a new resource and then came away having been assigned the time-consuming "homework" of creating a bunch of resources for that project).

Though asking users to create lessons plans and other teaching materials was not as successful as the WW1PDA project might have hoped, users were willing (and excited!) to contribute materials from their own familial archives to the project. In fact, the project received such a high level of response to their requests that they held extra workshops to help the public digitize their items.

Some of Marchionni's suggestions seem to blend user engagement and marketing. For example, the WW1PDA's teachers' workshops seemed to have functioned in part as a source of user feedback, but also as a forum for promoting and publicizing the resource. Teachers were seen as the key to two user groups: teachers and students. Similarly, Marchionni also suggests targeting any information dissemination activities at specific user groups. The ASR project, for example, publicized its Holocaust collection by contacting networks for Historians and those in the field of Jewish and Theological Studies. Though his may seem more like advertising than user engagement, it's nevertheless important to remember that we sometimes need to market our resources if we want them to be used. Finally, after highlighting the importance of user engagement, Marchionni ends with a reminder not to lose sight of the project's mission: though it's important to listen to user feedback, we shouldn't be bullied by it. Focus on the the needs of one's primary users and keep in mind that you can't satisfy everyone.

One thing that I wish Marchionni had addressed in greater detail is the expense involved in maintaining a high level of user engagement. Obviously it's more expensive in the long run to pour money into a resource that doesn't get used than to devote some money to engaging users, but nonetheless creating sustained relationships with users can be a drain on already strained budgets and staff schedules. I'd love to hear more about how small projects or ones meager financial resources might effectively develop ongoing relationships with users.

Tuesday, October 20, 2009

Remediation & Early English Books Online

Kichuk, Diana. "Metamorphosis: Remediation in Early English Books Online (EEBO)." Literary and Linguistic Computing. (published June 18, 2007). Available at http://llc.oxfordjournals.org.ezproxy.lib.utexas.edu/cgi/content/full/fqm018v1 (Accessed October 13, 2009).

Out of morbid curiosity I had wanted to read an article written about a project that I had worked on, Early English Books Online (EEBO), and this article by Diana Kichuk explores the layers of remediation in the project's production of its digital archives of Early English Books. As Kichuk defines it, remediation is the "re-presentation of one medium in another" (2) as well as the "appropriation or re-purposing of old media in new media" (a definition she attributes to R. Grusin and J.D. Bolter's Remediation: Understanding New Media [2000]). EEBO provides an interesting case study because its associated projects involve several levels of remediation.

EEBO consists of an archive of images of over 100, 000 books printed in English between 1473-1700, and EEBO-Text Creation Partnership (EEBO-TCP)--the part of the project that I worked for--now offers hand-keyed searchable, reading texts for 25,000 of these titles. Access in both cases is through individual or institutional subscription. The project entails several levels of remediation. The majority of the images of the Early English books come from the Early English Books microfilm project, which was began in the mid-1930s to preserve copies of Britain's important early books, and the project continued in the years following the Second World War. Beginning in 1998, these microfilms were then digitized to produce EEBO's digital archive of images. The clear reading texts produced by EEBO-TCP are drawn from human transcriptions the digitized microfilm images. Accordingly, Kichuk asserts that EEBO is a surrogate of a surrogate, with remediation acting like a "distorting lens or opaque veil through which the scholar 'sees' the mediated Early English book" (6). Though Kichuk states that the collection is of "formidable scholarly valuable," her article emphasizes the need to recognize that that EEBO does not offer an exact copy of the original and that scholars must be better aware of the limitations entailed in remediation.

Kichuk sees ProQuest's decision to digitize the microfilm facsimiles rather than print copies as ultimately wise: digitizing print copies that it did not own would have been prohibitively expensive and slow. Moreover, the technological limitations present at the time the project was initiated would likely have meant that the original books would have to be dis-bound to be scanned, thus endangering the original artifacts.

There are, however, important sacrifices associated with this decision, resulting in content amputation and page distortion, among other things. The EEB microfilms generally do not include the endpapers or bindings of the books, and the pages were often cropped, resulting both in the loss of some marginalia and the misrepresentation of the physicality of the original book. The page curvature resulting from open book photography generates distortions in font appearance and darkens the gutters, affecting both legibility and accuracy of representation. Also, the resolution used by the EEB microfilms was too low for grayscale capture, so the images are bi-tonal black and white, which makes capturing any traces of color in the original works near impossible. Also, in the transition from microfilm to digital image, EEBO again sacrificed detail by further lowering its image resolution in order to ensure acceptable download transmission rates (the images are 440 PPI [pixels per inch], too low for a preservation quality facsimile).


(This figure, taken from the EEBO homepage (http://eebo.chadwyck.com/marketing/eebo_demo11_tcp.htm) offers a sample of what the EEBO interface and text images look like.)

As a digital codex, Kichuk notes that the project does not accurately reflect both the text and the physicality of the original--to achieve this would require applying the latest imaging technologies to the original books. It would be a different project carried out at a different time. In terms of improvements though, Kichuk would like EEBO (and the scholars who use it) to be more conscious of and clear about the loss remediation entails. She argues for the vendor guard to against claims of authenticity and identical-ness. She would also like to see EEBO include duplicate copies, which would be of use to bibliographers (this seems to me a slightly odd request given that EEBO fulfills so little of a bibliographer's needs).

Though I think the issues of remediation are fascinating, the questions that EEBO and Kichuk's article leave me with are ones about what it is we want from our digital surrogates. EEBO is a hugely important source of access to the content of rare early printed books that most users would never have the opportunity to access otherwise. EEBO-TCP adds further value by allowing readers to engage in full-text searches across over 25,000 titles. This said, it is not as useful a resource if you are interested in thinking about the book as object. Its images are not high-quality representations and it has only limited functionality to support this type of investigation. (For example, I had a friend who was interested in using EEBO to study investigate the significance of when blackletter was used as opposed to roman or italic fonts in early English books. Since EEBO-TCP doesn't specify font-type in its encoding and only records emphasis by noting the distinction between the normative font for the book and any "highlighted" font by using tags, there's simply no way of searching for this short of examining each text.) I would argue, however, that in many cases consulting EEBO could help someone in determining which books to examine in person.

Ultimately, reading about EEBO foregrounds important questions about how we should prioritize financial feasibility, technical limitations, functionality, and speed of collection development in generating digital libraries and large curated collections. Though EEBO is more about access than preservation, the issues it raises overlap at least partially with the discussion of essential elements in this week's reading (Harvey 16). Though Ross Harvey's context is somewhat different, his question: "Is the value tied to the way the material looks? (Would it be lost or significantly degraded if the material looked different?)" remains instructive in thinking about EEBO. In developing such projects, we must decide which features of the original are most important to capture and at what cost, while keeping in mind that remediation entails both loss and gain.

Tuesday, September 8, 2009

Keeping up with the Sciences: Humanities and Cyberinfrastructure

Crane, Gregory, Alison Babeu, and David Bamman. "eScience and the Humanities." International Journal on Digital Libraries 7, no. 1 (2007): 117-122.

In this article, Crane, Babeu, and Bamman suggest that though those in the humanities are increasingly working with large digital datasets, they are lagging behind their scientific peers in developing the cyberinfrastructures necessary to support and maintain digital resources. As with the sciences, the humanities need cyberinfrastructure to help address both the massive scale of digital data and the fact that managing this data requires specialized knowledge beyond the capacity of any single researcher. Unfortunately, the humanities have been slow to develop such infrastructures due, in part, to disparities in funding between the sciences and the humanities. The authors note, for example, that the budget of the National Science Foundation (NSF) is 39 times larger than that of the National Endowment for the Humanities (NEH). However funding is not the only culprit. Crane et al. also suggest that a failure of imagination has impeded the humanities: that is, they have been slow to recognize the types of intellectual activity that emerging cyberinfrastructures might support.

In order to grow cyberinfrastructure, the article recommends that the humanities systematically develop alliances with the sciences and other better-funded disciplines, as well as collaborate with them on shared technological interests. International collaboration must also be encouraged. To this end, the authors enumerate five core services that they assert reflect "a convergence of interests that extends beyond the humanities" (120). These services include: 1) Conversion of page images to digital text (including handwritten documents that pre-date the advent of printing), 2) conversion from raw text to structured data (including semantic classification, indentifications, and morphological and syntactic analysis) 3)support of multiple languages (including cross-language information retrieval), 4)customization and personalization of data retrieval, and 5) the support of continuous user contributions (such as corrections of OCR errors). Crane et al. close by recommending that the humanities strive to develop larger, more stable organizational structures (as opposed to constantly re-inventing the wheel in numerous small-scale projects) in order to ensure the maintenance of digital data services. They also hold out hope for the emergence of disciplinary centers--such as one for classicists--to attend to the specialized needs of their constituents.

Though one imagines that the article is meant to rally the humanities and provide real-world strategies for confronting difficult funding situations, the effect is still somewhat depressing. Identifying overlapping interests makes good financial sense, but it seems important too to recognize and investigate significant areas of divergence between the sciences and the humanities. For example, how do humanities "datasets" differ from those in the field of science and social science and what type of functionality would humanities data benefit from most? What constitutes pre-publication "raw data" in the humanities? How does the data life cycle for humanities materials compare with that which Anna Gold outlines for the sciences in "Cyberinfrastructure, Data, and Libraries, Part 1?" Where significant differences exist, what are the cyberinfrastructure implications of these differences? As this week's article by Inge Angevaare notes, a vital part of any digital curation plan is identifying and attending to the needs of one's designated community.

Collaboration with the sciences is undoubtedly key to developing a more robust cyberinfrastructure for the humanities, but, at the same time, humanities agendas should not be unduly shaped by the interests of the sciences. Sharing resources only works in so far as the end product serves the needs of both parties.