Showing posts with label digital humanities. Show all posts
Showing posts with label digital humanities. Show all posts

Monday, November 23, 2009

Digital Curation & User Testing

Marchionni, Paola. “Why Are Users So Useful?: User Engagement and the Experience of the JISC Digitisation Programme.” Ariadne, no. 61 (October 2009). http://www.ariadne.ac.uk/issue61/marchionni/.


This recent article by JISC's Paola Marchionni refocuses attention on the purpose of curated digital collections: that is, their use by different user groups. Marchionni begins by noting that many digitization projects are still not paying enough attention to their users that their users' needs. They become so caught up in trying to make their content accessible online, that they don't adequately research their key users. As a result many publicly-funded projects are going un- (or under-) used. Marchionni illustrates the insight users can provide by presenting two case studies of projects that incorporated users into their development process: the British Library's Archival Sound Recordings 2 (ASCR2) project (a collection consisting of over 25,000 recordings) and Oxford University's First World War Poetry Digital Archive (WW1PDA) (a collection that contains over 7000 items pertaining to WWI poets, including digitized images of materials held at UT's own Harry Ransom Center).

Though Marchionni's article helpfully reminds digitization (and digital curation) projects to keep their electronic eyes on the prize and really take their users into account, I'm not sure that much of what Marchionni presents in her list of suggestions for user engagement is particularly surprising. She recommends first recognizing the importance of interacting with users and even having an "Engagement Officer" position as the ASR2 project did. She also advises establishing an early and on-going relationship with users. The WW1PDA project, for example, developed a typology of users, with a steering committee of scholars in the field of WWI literature advising which materials should be digitized and participating in quality control, and a separate group of secondary school and higher education instructors helping to develop and offer feedback on the education section of the project. Marchionni also emphasizes the importance of knowing what to do with user feedback. When users expressed anxieties about the integration of Web 2.0 tools out of fear that they might undermine the authority of the WW1PDA archive, the project decided to integrate such functionality in a way that made the lines between the archivists' and the users contributions more clear.

Some of the more interesting lessons regarding users came from the WW1PDA's approach to educational resources. The project held workshops for teachers in order to discover what functionality this group would like to see on the site. In a rather ballsy move, the project then asked the workshop members to help author a number of learning resources for the website. Though this did result in the creation of some resources, ultimately the project realized that perhaps it had overreached in what it was asking their busy users to produce. (Frankly, I would be a little annoyed if I agreed to participate in a workshop on a new resource and then came away having been assigned the time-consuming "homework" of creating a bunch of resources for that project).

Though asking users to create lessons plans and other teaching materials was not as successful as the WW1PDA project might have hoped, users were willing (and excited!) to contribute materials from their own familial archives to the project. In fact, the project received such a high level of response to their requests that they held extra workshops to help the public digitize their items.

Some of Marchionni's suggestions seem to blend user engagement and marketing. For example, the WW1PDA's teachers' workshops seemed to have functioned in part as a source of user feedback, but also as a forum for promoting and publicizing the resource. Teachers were seen as the key to two user groups: teachers and students. Similarly, Marchionni also suggests targeting any information dissemination activities at specific user groups. The ASR project, for example, publicized its Holocaust collection by contacting networks for Historians and those in the field of Jewish and Theological Studies. Though his may seem more like advertising than user engagement, it's nevertheless important to remember that we sometimes need to market our resources if we want them to be used. Finally, after highlighting the importance of user engagement, Marchionni ends with a reminder not to lose sight of the project's mission: though it's important to listen to user feedback, we shouldn't be bullied by it. Focus on the the needs of one's primary users and keep in mind that you can't satisfy everyone.

One thing that I wish Marchionni had addressed in greater detail is the expense involved in maintaining a high level of user engagement. Obviously it's more expensive in the long run to pour money into a resource that doesn't get used than to devote some money to engaging users, but nonetheless creating sustained relationships with users can be a drain on already strained budgets and staff schedules. I'd love to hear more about how small projects or ones meager financial resources might effectively develop ongoing relationships with users.

Sunday, November 8, 2009

Who's Who in Digital Humanities Collaboration

Spiro, Lisa. "Examples of Collaborative Digital Humanities Projects." Digital Scholarship in the Humanities (blog). Posted June 1, 2009. Available at: http://digitalscholarship.wordpress.com/2009/06/01/examples-of-collaborative-digital-humanities-projects/. (Accessed November 8, 2009).

In a long and wonderfully detailed blog post (yes, it even has footnotes!), Lisa Spiro provides a descriptive overview of collaboration in digital humanities projects. She begins by noting that historically collaboration has not been a part of the publication model in the humanities. As evidence, she cites her own finding that between 2004 and 2008 only 2% or the articles published in American Literary History were co-authored. Similarly, she notes that Cronin et al found in their longer-ranging survey that only 2% of articles published between 1900 and 2000 in the philosophy journal Mind had more than one author. Spiro, however, observes that the humanities do have an extensive tradition of circulating and providing feedback on one another's work, and that new digital technologies such as CommentPress and Zotero are helping to facilitate the exchange of ideas. For Spiro (and John Unsworth, whom she cites), collaboration in the digital humanities holds much potential: "Through online collaboration, scholars can divide labor (whether in making a translation, developing software, or building a digital collection), exchange and refine ideas (via blogs, wikis, listservs, virtual worlds, etc.), engage multiple perspectives, and work together to solve complex problems." Indeed she suggests that the incidence of humanities collaboration in the digital environment is higher than in the paper and ink world.

Spiro's discussion of collaboration only sometimes overlaps with the notion of data sharing in the sciences. I wonder if the differences between what counts as raw "data" in the humanities (i.e. primary sources--books, artworks, historical documents) vs. in sciences (i.e. observed and experimental data) means that collaboration is a more useful concept for the humanities than is data sharing.

Spiro spends the lion's share of her blog post providing detailed examples of different types of collaboration in the humanities. She explains that she was having difficulty articulating how collaboration functions in humanities research until she began exploring concrete examples. She divides the types of collaboration into three main categories: "facilitating communication and knowledge building," "sharing and aggregating content," and "collaborative annotation, transcription, and knowledge production." Classified under each heading are more discrete project types. Here is a skeleton of how she schematizes collaboration in the digital humanities (though, as she notes, there is inevitably some overlap among the types of collaboration:

Facilitating communication and knowledge building:
  • Online communities/virtual organizations (e.g. listservs, online forums, online communities, advanced video conferencing)
  • Collaboratories (which are virtual research environments that use advanced networking, remote instrumentation, databases, and digital libraries to foster "communication, collaboration, resource sharing, and research regardless of physical distance.")
Sharing and aggregating content:
  • Digital memory banks/user-contributed content (i.e. various projects to which users can contribute their own content, sometimes with the help of flickr and youtube, such as The Hurricane Digital Memory Bank and the Oxford-sponsored Great War Archive.)
  • Content aggregation and integration (i.e. federated digital collections which draw from a variety of archives to overcome the "silo" effect that can plague individual institutional collections. Two examples Spiro gives are the Walt Whitman Archive’s Finding Aids for Poetry Manuscripts and the the Quilt Index.)
  • Data sharing (e.g.Open Context , an archaeology project that permits researchers to upload, tag, analyze and share data sets).
Collaborative annotation, transcription, and knowledge production:
  • Crowdsourcing transcription (these include efforts that attempt to crowdsource transcription, just as Project Gutenberg and Project Madurai are crowdsourcing the proofreading of OCR texts).
  • Collaborative translation (e.g. Suda Online (SOL), which "brings together classicists to collaborate in translating into English the Suda, a tenth century encyclopedia of ancient learning written by a committee of Byzantine scholars.")
  • Collaborative editing (wherein collaborative online editions of texts are made)
  • Social bibliographies, collaborative filtering, and annotation (including platforms like zotero and eComma, which enable sharing bibliographies and collaborative annotation respectively)
  • Collaborative writing (e.g. subject wikis such as the Pynchon Wiki.)
  • Gaming: "collaborative play" and games as research (wherein "games provide motivation and a structure for collaboration" and "teamwork enables puzzles to be solved more rapidly.")
  • Publishing (Spiro's examples include posting materials online for peer-to-peer reviews prior to print publication)
  • Social learning (wherein participating in digital projects is a form of apprenticeship for undergraduates and graduate students, as they digitize materials, provide metadata, do programming, or contribute to a wiki).
I'd actually like to pause on this last item for a minute since it raises some of the same questions regarding the changing nature of authorship in the digital environment that we've been discussing throughout the course of the semester. Spiro's entry on "social learning" makes me a little uneasy since she frames the students' contributions as an interactive mode of "learning" rather than "authoring"--sure, they're learning, but they are also generating content as well. One thing I'd really like to see is a discussion of how digital collaboration is transforming notions of authorship in the humanities. Lisa Spiro recently blogged on the topic, but it wasn't quite as down and dirty as I wanted it to be. It is, however, an area in which she's conducting ongoing research.

In general, Spiro's blog entry on humanities collaboration is more descriptive than analytic. She doesn't really delve into the logic of her classifications or the problems associated with "social scholarship" (though she does link to an earlier blog of hers addressing this second topic). Accordingly, her piece is useful for familiarizing oneself with the types of collaboration going in the humanities, rather than thinking through the meaty issues associated with collaboration and sharing. Perhaps it might be worthwhile to consider such matters from a humanities perspective in class.

Tuesday, October 20, 2009

Remediation & Early English Books Online

Kichuk, Diana. "Metamorphosis: Remediation in Early English Books Online (EEBO)." Literary and Linguistic Computing. (published June 18, 2007). Available at http://llc.oxfordjournals.org.ezproxy.lib.utexas.edu/cgi/content/full/fqm018v1 (Accessed October 13, 2009).

Out of morbid curiosity I had wanted to read an article written about a project that I had worked on, Early English Books Online (EEBO), and this article by Diana Kichuk explores the layers of remediation in the project's production of its digital archives of Early English Books. As Kichuk defines it, remediation is the "re-presentation of one medium in another" (2) as well as the "appropriation or re-purposing of old media in new media" (a definition she attributes to R. Grusin and J.D. Bolter's Remediation: Understanding New Media [2000]). EEBO provides an interesting case study because its associated projects involve several levels of remediation.

EEBO consists of an archive of images of over 100, 000 books printed in English between 1473-1700, and EEBO-Text Creation Partnership (EEBO-TCP)--the part of the project that I worked for--now offers hand-keyed searchable, reading texts for 25,000 of these titles. Access in both cases is through individual or institutional subscription. The project entails several levels of remediation. The majority of the images of the Early English books come from the Early English Books microfilm project, which was began in the mid-1930s to preserve copies of Britain's important early books, and the project continued in the years following the Second World War. Beginning in 1998, these microfilms were then digitized to produce EEBO's digital archive of images. The clear reading texts produced by EEBO-TCP are drawn from human transcriptions the digitized microfilm images. Accordingly, Kichuk asserts that EEBO is a surrogate of a surrogate, with remediation acting like a "distorting lens or opaque veil through which the scholar 'sees' the mediated Early English book" (6). Though Kichuk states that the collection is of "formidable scholarly valuable," her article emphasizes the need to recognize that that EEBO does not offer an exact copy of the original and that scholars must be better aware of the limitations entailed in remediation.

Kichuk sees ProQuest's decision to digitize the microfilm facsimiles rather than print copies as ultimately wise: digitizing print copies that it did not own would have been prohibitively expensive and slow. Moreover, the technological limitations present at the time the project was initiated would likely have meant that the original books would have to be dis-bound to be scanned, thus endangering the original artifacts.

There are, however, important sacrifices associated with this decision, resulting in content amputation and page distortion, among other things. The EEB microfilms generally do not include the endpapers or bindings of the books, and the pages were often cropped, resulting both in the loss of some marginalia and the misrepresentation of the physicality of the original book. The page curvature resulting from open book photography generates distortions in font appearance and darkens the gutters, affecting both legibility and accuracy of representation. Also, the resolution used by the EEB microfilms was too low for grayscale capture, so the images are bi-tonal black and white, which makes capturing any traces of color in the original works near impossible. Also, in the transition from microfilm to digital image, EEBO again sacrificed detail by further lowering its image resolution in order to ensure acceptable download transmission rates (the images are 440 PPI [pixels per inch], too low for a preservation quality facsimile).


(This figure, taken from the EEBO homepage (http://eebo.chadwyck.com/marketing/eebo_demo11_tcp.htm) offers a sample of what the EEBO interface and text images look like.)

As a digital codex, Kichuk notes that the project does not accurately reflect both the text and the physicality of the original--to achieve this would require applying the latest imaging technologies to the original books. It would be a different project carried out at a different time. In terms of improvements though, Kichuk would like EEBO (and the scholars who use it) to be more conscious of and clear about the loss remediation entails. She argues for the vendor guard to against claims of authenticity and identical-ness. She would also like to see EEBO include duplicate copies, which would be of use to bibliographers (this seems to me a slightly odd request given that EEBO fulfills so little of a bibliographer's needs).

Though I think the issues of remediation are fascinating, the questions that EEBO and Kichuk's article leave me with are ones about what it is we want from our digital surrogates. EEBO is a hugely important source of access to the content of rare early printed books that most users would never have the opportunity to access otherwise. EEBO-TCP adds further value by allowing readers to engage in full-text searches across over 25,000 titles. This said, it is not as useful a resource if you are interested in thinking about the book as object. Its images are not high-quality representations and it has only limited functionality to support this type of investigation. (For example, I had a friend who was interested in using EEBO to study investigate the significance of when blackletter was used as opposed to roman or italic fonts in early English books. Since EEBO-TCP doesn't specify font-type in its encoding and only records emphasis by noting the distinction between the normative font for the book and any "highlighted" font by using tags, there's simply no way of searching for this short of examining each text.) I would argue, however, that in many cases consulting EEBO could help someone in determining which books to examine in person.

Ultimately, reading about EEBO foregrounds important questions about how we should prioritize financial feasibility, technical limitations, functionality, and speed of collection development in generating digital libraries and large curated collections. Though EEBO is more about access than preservation, the issues it raises overlap at least partially with the discussion of essential elements in this week's reading (Harvey 16). Though Ross Harvey's context is somewhat different, his question: "Is the value tied to the way the material looks? (Would it be lost or significantly degraded if the material looked different?)" remains instructive in thinking about EEBO. In developing such projects, we must decide which features of the original are most important to capture and at what cost, while keeping in mind that remediation entails both loss and gain.