Tuesday, September 22, 2009

Making Data Mining Sexy (Even to Literary Scholars)

C. Plaisant, J. Rose, B. Yu, L. Auvil, M. G. Kirschenbaum, M. N. Smith, T. Clement, and G. Lord, "Exploring Erotics in Emily Dickinson's Correspondence with Text Mining and Visual Interfaces," Proceedings of the 6th ACM/IEEE-CS Joint Conference on Digital Libraries (JCDL '06), 2006: 141-150.

It is no accident that this article by Catherine Plaisant et al. bears the provocative title "Exploring Erotics in Emily Dickinson's Correspondence with Text Mining and Visual Interfaces." Its authors, a multidisciplinary group of researchers associated with the Nora Project (www.noraproject.org), are intentionally trying to make the text mining of digital collections both sexier to and more accessible for literary scholars. Their project is rooted in the premise that to impact research in the humanities, computational methods must facilitate and participate in the process of interpretation. Accordingly, the Nora Project seeks to develop a digital architecture that will enable literary scholars with no special IT expertise to perform text-mining on an electronic corpus of 18th and 19th century literature.

Plaisant et al.'s specific case study reports on their use of computational tools to classify individual Emily Dickinson letters as "hot" or "not hot" (that is, erotically-charged or non-erotic). The user first manually rates a representative set of sample documents on a five point color-coded scale of "hot-ness." These ratings then serve as "training" that permits a multinomial naive Bayes (NB) algorithm (executed by a D2K data mining tool) to rate the remainder of the letters in the collection on the hot-ness scale. The user then can examine some of the automatically-rated letters and accept, reject, or modify ratings before re-running the process in order to increase the accuracy of the rating predictions. The better the training set, the better the classification results.


As you can see in the above figure, which depicts the interface, color is used to help visualize the classification. The tool also suggests a list of words (see right side of figure) that it posits may be indicators of hotness or not-hotness. To help the literary scholars "read" the results, the interface offers several pop-up FAQs (E.g. "What do the purple squares mean"). The tool also enables the user to create scatter plots to visualize the rated documents with regards to variables such as date (no correlation between letter date and eroticism was found). A video demonstration of the interface is available through the Nora Project website at http://noraproject.org/nora_ol_video/. Though the study does not appear to have been large, the authors report that an Emily Dickinson expert who served as a test-user found that the results of the data-mining genuinely shed new interpretive light on familiar texts. Additional user feedback led to ideas for improving the interface, such as allowing users to assign extra weight to certain words.

Though Plaisant et al concede that the merit of the project will depend on how successful literary scholars are at publishing papers that draw on their tools, they assert that their initial findings are positive: their architecture and interfaces are usable by those without experience in data-mining, and their tools appear to yield provocative and inspiring insights for literary scholars.

Both this article and Blackwell and Crane's "Conclusion: Cyberinfrastructure, the Scaife Digital Library and Classics in a Digital Age" constructively provide clear examples of how computational and analytic tools applied to curated digital collections might advance scholarship in the humanities. However, features such as citation identification, syntactic and metrical analysis, or interpretive classification all rely on unfettered access to digital collections. Does this mean that literary scholarship will advance unevenly across periodizing lines? Since much of modernist and all of post-modernist literature still falls under copyright, there's no available digital corpus on which to apply these new computational tools. Will Sappho and Emily Dickinson scholarship benefit from added digital functionality while Denise Levertov scholarship remains mired in analog methodologies? One solution might be to make data-mining and other textual analysis services more portable. This could allow a motivated scholar to scan his or her own personal library of relevant books and, using OCR, create his or her own digital research collection. Data-mining could then be performed on this mirco-collection. It's not an ideal solution, but neither is excepting most post-1923 literature from the methodological innovations of digital scholarship!

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.