Wednesday, September 9, 2009

Geospatial Data for the Humanities

Elliott, T. & and Gillies, S. (2009). Digital geography and the classics. Digital Humanities Quarterly, 3(1).


The authors here discuss the brief history of digital geographic data in use by humanities studies (specifically classics) through pivotal scholarly projects, ending in a discussion of future possibilities for geographic or spatial data as well as their own Pleiades project.


Certainly spatial information has lagged behind HTML documents in their general findability, but as the authors point out the gap is closing: Googlebot began indexing web documents encoded in the Keyhole Markup Language (KML) along with feeds containing GeoRSS markup in late 2006; Microsoft has added support to its Local Search and Virtual Earth services for similar data in late 2007. In general KML and GeoRSS resources, if properly marked up in an indexable web page, will be discoverable.


Typically the actual coordinates of some spatial resource is extracted from an external dataset (examples: Getty Thesaurus of Geographic Names, GeoNames database, Alexandria Digital Library Gazetteer) based on the name. This can naturally lead to erroneous information on account of name ambiguity (names changing through time, coordinates changing through time, etc.). The authors believe that over time such misjudgments will be minimized through sophisticated name-variant databases as well as automated best-guesses algorithms, like examining other geographic points in an article to determine the likely coordinates of a single ambiguous place name. This is mass auto-extraction of metadata is similar to our discussion of Google Books metadata problems, with a similar refinement-over-time solution.


The authors examine the beginnings of maps and spatial data on the web, charting a path from the earliest server-side applications to heterogenous, continuous-panning, browser-side interfaces that pull data and overlays from multiple resources. I cannot help but look at the battle maps in my group’s chosen digital archive, The Valley of the Shadow, and see that these geographic resources fall into the earlier paradigm. While the interface allows the user to apply a set number of overlays to the unfolding battle animations, the application itself is closed off from other web resources: there are no coordinates, no historical geographic or topological data, etc., available, and there is no way to see such information except by requesting the programmers to add it.


The authors also distinguish between large top-down geo-historical projects like the National Geospatial Digital Archive (US) and Global Spatial Data Infrastructure Association and smaller, “hand-crafted” datasets created by dispersed volunteered geographic information (VGI). The latter, they argue, is quickly eclipsing the achievements of the former. VGI datasets have become more prominent on account of a rise in web applications that can easily publish such information. These applications are increasingly adopting metadata standards, and collaborative projects are arising that address interoperable issues in geo-spatial data like the Register of Geographic Entities (RAGE).


Last the authors discuss their own Pleiades Project. This attempts to establish a standard reference dataset for Classical geography. Most interesting here is their rejection of coordinates and toponyms as the primary organizing theme for geo-historical data, using instead the concepts of place, “understood as a bundle of associations between attested names and measured (or estimated) locations (including areas).” I think this is a really great innovation: it moves away from a more scientific, exacting sense of place (coordinates) to a more humanly and historically meaningful sense of place, one that has the advantage of being more easily discussed and contested in a collaborative environment. Certainly the computer will need to understand something like the Roman Empire in terms of coordinates, but a layer of abstraction (locations and areas) is extremely useful when humans need to play with such data.

Tuesday, September 8, 2009

Keeping up with the Sciences: Humanities and Cyberinfrastructure

Crane, Gregory, Alison Babeu, and David Bamman. "eScience and the Humanities." International Journal on Digital Libraries 7, no. 1 (2007): 117-122.

In this article, Crane, Babeu, and Bamman suggest that though those in the humanities are increasingly working with large digital datasets, they are lagging behind their scientific peers in developing the cyberinfrastructures necessary to support and maintain digital resources. As with the sciences, the humanities need cyberinfrastructure to help address both the massive scale of digital data and the fact that managing this data requires specialized knowledge beyond the capacity of any single researcher. Unfortunately, the humanities have been slow to develop such infrastructures due, in part, to disparities in funding between the sciences and the humanities. The authors note, for example, that the budget of the National Science Foundation (NSF) is 39 times larger than that of the National Endowment for the Humanities (NEH). However funding is not the only culprit. Crane et al. also suggest that a failure of imagination has impeded the humanities: that is, they have been slow to recognize the types of intellectual activity that emerging cyberinfrastructures might support.

In order to grow cyberinfrastructure, the article recommends that the humanities systematically develop alliances with the sciences and other better-funded disciplines, as well as collaborate with them on shared technological interests. International collaboration must also be encouraged. To this end, the authors enumerate five core services that they assert reflect "a convergence of interests that extends beyond the humanities" (120). These services include: 1) Conversion of page images to digital text (including handwritten documents that pre-date the advent of printing), 2) conversion from raw text to structured data (including semantic classification, indentifications, and morphological and syntactic analysis) 3)support of multiple languages (including cross-language information retrieval), 4)customization and personalization of data retrieval, and 5) the support of continuous user contributions (such as corrections of OCR errors). Crane et al. close by recommending that the humanities strive to develop larger, more stable organizational structures (as opposed to constantly re-inventing the wheel in numerous small-scale projects) in order to ensure the maintenance of digital data services. They also hold out hope for the emergence of disciplinary centers--such as one for classicists--to attend to the specialized needs of their constituents.

Though one imagines that the article is meant to rally the humanities and provide real-world strategies for confronting difficult funding situations, the effect is still somewhat depressing. Identifying overlapping interests makes good financial sense, but it seems important too to recognize and investigate significant areas of divergence between the sciences and the humanities. For example, how do humanities "datasets" differ from those in the field of science and social science and what type of functionality would humanities data benefit from most? What constitutes pre-publication "raw data" in the humanities? How does the data life cycle for humanities materials compare with that which Anna Gold outlines for the sciences in "Cyberinfrastructure, Data, and Libraries, Part 1?" Where significant differences exist, what are the cyberinfrastructure implications of these differences? As this week's article by Inge Angevaare notes, a vital part of any digital curation plan is identifying and attending to the needs of one's designated community.

Collaboration with the sciences is undoubtedly key to developing a more robust cyberinfrastructure for the humanities, but, at the same time, humanities agendas should not be unduly shaped by the interests of the sciences. Sharing resources only works in so far as the end product serves the needs of both parties.

Image socialbookmarking sites as Digital Curation?

Recently I have been intrigued by sites such as http://imgfave.com that aggregate photos/images through socialbookmarking and personal uploading. Imgfave allows the user to 'favorite' what image appeals to them, and then displays the collections of others who have fave'ed that particular image.

One may add those images to particular 'collections', and one may become a 'fan' of (or you may 'follow') other users. There are different "feeds": "Popular" (most fave'ed), "Everyone," "Friends," and "My Profile" that one may browse and work from. One also may link images or one's profile to Facebook.

The only major drawback with the site is that it lacks metadata. That is, until I discovered its rival: http://vi.sualize.us/.

Visualize uses metadata heirarchies - It has a section of Popular tags with general categories such as: photography, illustration, design, nature, art. Each contain sub-categories of qualities from: "vintage" and "humor" to elements or examples such as "typography" or "landscape."

There is also a section "Most used tags" to guide one. These are listed in alphabetical order, with most used ones in bolder, larger font. You can view (and RSS subscribe to) "popular pictures feed" and "recent pictures feed."

Each image displays a file name or title, when it was added, how often it's been "liked," comments on it, and its tags. And it appears that not only the 'creator' or person who uploaded the image, but everyone can tag the image.

This blog article gives a list of 10 social bookmarking visual sites: including also Picocool, We love typography, and others.

At a glance - the browsability seems more enjoyable and sublime with imgfave.com than with vi.sualize.us/. I am not sure if this is because the thumbnail images are smaller with vi.sualize.us/ or because the random, metadata-lacking feed in imgfave makes the experience more prone to serendipity and thus feel more direct and personal than one mediated by a metadata schema.

I think in regards to the size of the image, what imgfave does best is feature larger images that one needs to scroll down to view. So in scrolling down, one is discovering the images, experiencing them each individually. In vi.sualize.us/, on the other hand, one is presented with rows of smaller images, displayed three across, and typically 2 1/2 rows will fill one's screen. This in combination with all of the text lends itself to a visually overwhelming experience, whereas imgfave feels far more like a poetic and artistic discovery. In imgfave one does not know what one is "looking for" but favorites those visual items when one sees them. Yet, the only real drawback is then, once having found that image (or others like it), the lack of metadata makes it difficult to further seek out these kinds of pictures.

Another feature in imgfave that I enjoy is its bookmarklet widget: Similar to the del.icio.us widget in the toolbar, when one sees an image online one clicks the <3imgfave text in the toolbar and all the images on that page become surrounded with a red border with a button in the center prompting one to 'add to imgfave."

An interesting argument in this comparative discussion might be: "What about Flickr?" Why was Flickr not mentioned as a "visual social bookmarking site?" Flickr has marvelous social networking features, from the ability to add contacts and create communities and image pools. It also has a tagging dimension (although one far less standardized than vi.sualize.us's fine art/design metadata schema (although vi.sualize.us's IS customizable). There is a component in Flickr whereby one can 'favorite' other people's images. However, when one clicks on one of the images that one has favorited, one does not see recommendations of images from others who have favorited that same image. This is one of the things that intrigues me - following this trail. Does it yield reliable "data?" Sometimes, there is still a signal to noise ratio - but it provides a slightly more relevant result than searching randomly or by ranking.

The most important question however - "Is all this Digital Curation?" This is something I am still investigating. I believe personally that digital curation in these examples would necessitate a combination of the associative discovery factor in imgfave with the aesthetic metadata elements of vi.sualize.us, along with the social networking/collaborative potential in Flickr - and ultimately, the ability in using these features to observe and discover emergent visual trends and patterns, and to not only come up with innovative stylistic metadata to describe these things, but to have that information emerge via tag ranking as an influence to how we communicate about contemporary visual culture as a whole.

Gov 2.0: Suitable for viewing in school

In Paul Miller's blog The Cloud of Data, a post this week focuses on Gov 2.0 called "If Government is a Platform, what are people building?" Miller muses on a recent Tim O'Reilly's blog entry (also a good read about the subject, as he is a co-chair on the upcoming gov2.0summit) and discusses the beneficial ways an open platform for government information and data will impact aspects of technology growth, community activism, and government accountability.

Basically Miller and O'Reilly are enthusiastic about the idea of the government using Web 2.o as a platform to make connections, not just focusing on the 'social media' part that is most associated with the term. O'Reilly cites several great examples of ways in which government data has been re-purposed to provide more meaningful information for the public, such as everyblock.com, that originally started out as a site that took crime data and put it with a Google map. Users can find police reports, restaurant inspections, home sales, and news stories by neighborhood/zip code/etc in select cities.

Data.gov is mentioned by O'Reilly, and Miller picks up on the site's use to further technology. Because of its large data sets, researchers are "turning to resources like Data.gov in search of interesting technological problems and large pools of data upon which to test new techniques and ideas". This is an great experiment with a free resource that might not have too many other uses. I'm assuming other large data sets are in the sciences and are proprietary or just not meant to be played with.

While I don't know much about creating or using applications, both O'Reilly and Miller discuss the benefits of user-generated apps, and how the platforms that Apple or the government put in place can creatively impact new uses and designs for sometimes boring or information that is not applicable to much in its current form.

I think that libraries, archives, and museums should think about this approach to Web 2.0, or Library 2.0, instead of just trying to engage Generation (Insert Letter Here) through boring tweets or attempts at blogging. While it may be fun and dorky to be a fan of a film archive on Facebook, what is it really doing to help push a library/archive into a larger public technological role? These institutions are a bit behind in this Web 2.0 (please just get out of Second Life) and could benefit by letting users find new appropriations for the information and services offered. Some JSTOR app on an iPhone sounds amazing! Or an ILL Book Finder app! Maybe there already is one. I don't have any specific solutions, but libraries should be involved with tech goings-on such as the gov2.0summit for inspiration and ideas about their own Foray into the Future.

Exploring Research Data Hosting at the HKUST Institutional Repository

An article by Gabrielle K.W. Wong

This article is linked on the front page of the Distributed Data Curation Center site.

Wong makes it clear that the problem of collecting a large amount of data has changed from being an issue with drive space to an issue of management of the data. Wong claims that although it is difficult to manage data, the value of data increases when it is accessible. She also claims that universities are well suited for the responsibility of managing data because there are existing repositories at so many universities. This suggestion is congruent with the advice from other experts, too. For instance, in the article we read this week from Inge Angevaare, she points out that universities are ensnared in the access model in terms of scholarly journals, but universities remain in control of original data collected. Angevaare even suggests that curating this data has the potential to "revive the library’s unique position at the very heart of the university’s information network".

Wong is from the Hong Kong University of Science and Technology (HKUST) Library and this year they took on the task of creating a DSpace repository. Wong is the institutional repository coordinator so she used OpenDOAR to find 53 repositories and then analyzed how they curate datasets. Not surprisingly, Wong found many different practices throughout all of these repositories. In terms of metadata, DSpace uses Dublin Core, which allows each repository flexibility in their metadata. Of course, this is not neccesarily a good thing for catalogers or users.

The functional problems with existing digital curation does not end with metadata! There's a number of different file formats found throughout the repositories and this brings with it issues of software versions for end-users. Also, there is no standard way to cite datasets. Actually, the very practice of linking data to the research papers written about them isn't standardized. This lack of standardization throughout repositories not only keeps the repositories from being fully accessible, it also creates a significant amount of work for the curators of the collections. Instead of having unified practices, individual projects have to reinvent standards for their datasets.

In addition to her study of DSpace repositories, Wong summarizes the many reasons why researchers do not make their data public, which has been examined by the RIN and explained in the article To share or not to share. This is a brief account of their reasons:


  • lack of career rewards as a major disincentive
  • wish to retain exclusive use of data until all the publication value is extracted
  • lack of time, resources and expertise to handle the data management
  • legal and ethical constraints, such as data ownership issue and confidentiality issue when personal data is involved
  • lack of appropriate archive service
  • fear of exploitation or inappropriate use of the data.


To be honest, I don't see any of these barriers going away, especially in private industry where innovations result in profit. Even if digital curators developed standard practices, there would be a number of researchers reluctant to participate.

Beyond Digital Incunabula

Like Gregory Crane mentioned in the article assigned as part of last week's reading entitled "What Do You Do with a Million Books?," digital libraries are limited by preconceived notions based on the paradigm of print libraries. Our digital libraries, he says, "remain filled with digital incunabula." In the article "Beyond Digital Incunabula: Modeling the Next Generation of Digital Libraries," Crane and his colleagues explore these "incunabular assumptions," and use the Perseus Digital Library as an example of the ways in which digital libraries can move beyond such limitations.

In the first section of the article, the authors identify three notions that have carried over from paper libraries and have limited the design of digital libraries. The first is that scholarly writing is based on well defined forms in which the units of information (the journal article, the book chapter, etc.) are relatively coarse. This has translated into the digital world as PDF or HTML files that mimic their print predecessors in format. Secondly, metadata in a digital library replicates the card catalog of the print library, and thus has all the same constraints. Finally, the static nature of print libraries has carried over to digital libraries -- just as no new knowledge is created by interplay of books in a print library, digital libraries are not being exploited to create new knowledge "by learning from their collections and from their users."

The authors next identify three characteristics that they believe distinguish "post-incunabular" digital libraries from their antecedents. One characteristic is finer granularity. More advanced digital libraries are beginning to break down digital objects into smaller and smaller chunks, making it easier for users to get to the information they are looking for without having to wade through a whole lot of content in which they have no interest. Another characteristic is autonomous learning. The documents in a digital library should constantly be looking for resources with which to enrich themselves, as, for example, by scanning for new secondary sources and new editions of primary source material. The final characteristic of a post-incunabular digital library is decentralized, real-time community contributions, as exemplified by Wikipedia. The article continues, using elements of the Perseus DL to exemplify each of these three characteristics.

The greater consequence of these incunabular assumptions is the continued "hegemony of library, author, and publisher" over the needs of the user. In order to counteract this, digital libraries should offer the user the ability to customize their environment. The user should not just be able to change the page layout and appearance, but should benefit from more substantive forms of customization. For example, in the Perseus DL, the user is able to select the textbook with which he or she learned Latin or Greek, and thereby be shown which words in a given passage are likely to be new or unfamiliar. The user should also benefit from the digital library's ability to analyze his or her behavior and background and on the basis of this offer automatically generated content. On the basis of four initial questions, the Perseus DL is able to predict, based on past question patterns, which of the words that user is most likely to query.

This article is another instance of the conviction that we saw expressed in the readings from last week, namely that the next generation of digital libraries should be greater than the sum of their parts. The idea that digital libraries should not just provide content, but also add value to that content, pervades the literature and the discourse on the subject of the future of digital libraries. We need to find practical ways to make the content of digital libraries useful in a way that print resources are not, and to do this we need to think outside of the paradigm of the print library and to view digital objects not as copies of their print counterparts but as distinct entities with unique properties. The example of the Perseus DL provides concrete examples of the way in which this is actually being accomplished.

Expiration dates on digital information?

Clicking around on the DCC-inspired Digital Curation blog led me to an interesting article from the Times Online called “Google Must Let Us Forget” which talks about life caching and a book by Viktor Mayer-Schönberger titled Delete: The Virtue of Forgetting in the Digital Age. This article discusses a downside to putting your digital information online – the fact that sometimes it is difficult to get rid of the information once it has been released into the wilds of the internet. Mayer-Schönberger offers an interesting solution to this potential problem by suggesting that there be a set expiration date on digital information, after which the information would be purged by storage devices.

Google, the Internet Archive, and other places on the internet, manage to capture so much information and keep it around, maybe much longer than those who first posted the information intended. The Neil Beagrie reading from this week mentioned life caching (the capturing and sharing of life memories online) and the problems individuals are having with protecting privacy and organizing the ever-increasing amount of digital information. And we’ve been talking about digital preservation and reading about all the issues that come with trying to keep digital data usable and preserved as technology changes. With all of that in mind, the notion of setting any digital information that we put out or take in with an expiration date intrigues me. The article mostly talks about setting expiration dates for personal, social networking-type information, but it could apply in other areas as well.

I’ve been trying to wrap my mind around it, but I’m still not sure how I feel about it. For organizing personal digital documents this might be great. But if this could be done with any bit of digital information that is put online, how would that work? Would this be helpful or harmful in when thinking about digital curation? After all, the point of archiving something is to keep a lasting record of it. How much harder would it be to curate information with expiration dates?

Yes, on one hand it would be nice to have digital data available for a set period of time (and it would give those a little too fond of life caching a clean slate after awhile… which would be good for getting those embarrassing undergrad party photos that just about everyone has posted at one time or another off the internet). At the same time, what type of data expiration standards would need to be developed; what types of digital information would need longer or shorter expiration times? How would this work?? Obviously, this article has raised some questions... Anyway, it was an interesting idea that I wanted to share.

Interedition

Interedition (an interoperable supranational infrastructure for digital editions) is an European concerted research action. The action is made up of a group of researchers looking for international scholarly cooperation and communication that facilitates the exchange of literary research and information technology. The web site states that their “aim is to promote the interoperability of the tools and methodology we use in the field of digital scholarly editing and research.” The group holds meetings to encourage the discussion of computer tools created by individual researchers in the hopes of building a networked infrastructure for digital scholarly editing and analysis. One main focus is to become an international body that organizes, creates, and perhaps govern such an initiative.

Interedition’s current primary objective is to create a manual or roadmap for developing an infrastructure specifically for literary materials. Researchers from thirteen countries are listed as participating (as of 2008): Belgium, Denmark, Finland, Former Yugoslav Republic of Macedonia, France, Germany, Ireland, Israel, Italy, Netherlands, Norway, Poland, and the United Kingdom. One of the goals of the group is to expand the involvement to other European countries.

A Memorandum of Understanding, available as a PDF, outlines the group’s intention as a four-year process of research and construction of a roadmap for implementing an infrastructure. In the first year, members would inventory of related projects on a European and global level. Year two, disseminate results of that inventory to members at conferences. The most recent meeting was held on May 6, 2009 at the Royal Irish Academy in Dublin. In the minutes from that recent meeting, members of one of the working groups acknowledged that they had not fully accomplished the goals of year one and must continue to inventory relevant projects.

This initiative is interesting in light of our discussion and reading. So far, most of the information available on the project consists of their intentions and organizational goals. The group is putting a lot of effort into coordinating interactions of a large number of individuals representing many countries and so far there not an indication of what the ultimate deliverable will look like.

From Babel to Knowledge

In "From Babel to Knowledge: Data Mining Large Digital Collections" Daniel J. Cohen answers the question from last week's discussion, what do you do with a million books? Specifically, Cohen explores ways in which algorithms developed by computer scientists might be applied to humanities-themed collections to extract useful information from the chaos caused by an overabundance of sources. (Cohen artistically begins his article with a reference to a Borges short story involving a "library of babel.")

In the article, Cohen discusses two major cs techniques he has used to harvest historical information. (Cohen is squarely in the history-is-a-humanity camp.) The first method involves an API he designed to seek out and gather course syllabi on the web. His API uses a "dictionary of notions" to identify keywords and sequences of keywords common to course syllabi, and can be narrowed by keywords to look for specific topics. Cohen reports about a ninety percent success rate for his Syllabus Finder. (Interestingly, the concept of a "dictionary of notions" originated with one Hans Luhn who built the groundwork for the idea while creating an inverted index for cocktail recipes by ingredient in what may have been the most every-day practical application of information science to date.)

The second tool discussed by Cohen is H-Bot, a powerful search tool designed to answer specific historical fact-based questions (thereby freeing historians for what Cohen describes as "higher levels of history.) H-Bot, in addition to being designed to understand complex questions, is able to gather information on two levels: a quick mode which mines data from (more or less) trusted encyclopedias and dictionaries and a "pure" mode which extracts info from pages across the entire web. The program than evaluates the most recurring terms and determines which terms makes sense syntactically. Cohen illustrates the process with the example of the query "When did Charles Lindbergh fly to Paris?" The article did not mention how H-Bot deals with more complicated questions, but I took the liberty of experimenting with it a bit. H-bot sincerely apologized to me for not being able to tell me what caused the civil war, told me that Robert Anton Wilson shot Kennedy, and that the Beatles landed in New York on February 7th, 1964. (H-Bot can be found here: H-Bot.) Cohen did admit that answering questions was a lot more complicated than finding syllabi.

Beyond answering the question on how to deal with a million books (or attempting to, at any rate), Cohen's article directly touched on our last class discussion through its conclusions. Cohen's first conclusion is that more non-profit organizations need to come out with open API's. (Currently, H-Bot gives you the choice of using Google or Yahoo.) This sentiment seems to echo a lot of the mistrust leveled at Google. Second, Cohen asserts that free tools, no matter how imperfect, are more beneficial than pay services, which also casts misgivings onto Google Books. In his final conclusion, however, Cohen alleges that quantity is more important than quality in digital collections and objects, as the more individual sources that can be mined, than the greater chance an API has of providing the most relevant information. In this regard, Cohen seems to bolster at least partially Google's position.

Overall, I found Cohen's article to be a fairly interesting exploration of the juntion of information science and the humanities.

Google Books follow up: > or < the sum of its parts?


I have not yet received a response to my request for more specifics on metadata import -- "Deear Googlez, plz giv more metadata info. Lov, Ramona." Keep your fingers crossed everyone.


An answer in About Google Books states, "We use automated methods to analyze the book, and in some cases we use content from third-party sources; as a result, we're currently unable to accept any manual edits to information that's displayed in the 'Related books,' 'Contents,' 'Key terms,' 'References from books,' 'References from scholarly works' or 'Selected pages' sections."

I can't think why automating import of metadata would preclude future changes to it, except that "manual edits" as they put it, would be very time consuming if not properly crowd-sourced. It is interesting that they state this, since according to the Nunberg article that Meg wrote about last week, "Dan Clancy also suggests that users could fix the errors one by one, in the way they fix errors on Wikipedia."

Google does offer a connection to some library metadata via a link to WorldCat ("Find in a Library"). However, if the metadata for the book is incorrect, this link would, no doubt, be incorrect as well, so that the right information in WorldCat would be linked with the wrong book.

As for the Google Books API, Dan Cohen wrote this blog post in 2008. He said then that although Google had just released their API, it was aptly and opaquely named the "
Google Book Search Book Viewability API." Cohen writes that the API does not allow enough access for scholars to mine the full text even of pre-1923 books -- public domain books that are out of copyright -- nor does it allow for the creation of enough tools besides simple viewing and search tools.
Related to our discussion of the API last week, he lists what can be done with it. It is a start:

  • Link to Books in Google Book Search using ISBNs, LCCNs, and OCLC numbers
  • Know whether Google Book Search has a specific title and what the viewability of that title is
  • Generate links to a thumbnail of the cover of a book
  • Generate links to an informational page about a book
  • Generate links to a preview of a book
If you want to find out more about the Google Books API, you can visit the page here.

The recent interest in the Google Books settlement has come because Tuesday was the deadline for filing court briefs for the case. EFF and Microsoft, two organizations I think rarely filing briefs on the same side of a case, both participated in the opposition. Read more here.

And finally, as to what universities or libraries might get out of a Google Books agreement, U-M posted this, which I noted says, "The agreement also calls for Google to contribute millions of dollars to establish up to two new research centers." Some of the contacts are publicly available here. According to the Google page, "... we can say that all of them are non-exclusive." I like Stanford sites, and found their page on the agreement with Google Books to be the most informative.

I think that like it or not, Google Books has it's own (cyber)infrastructure, and is a curation project. Although Google, and other digital curation projects, may not employ anyone that fits our mental image of a curator, digital projects are large and complex and probably necessitate the addition of structures that support the addition of sense-making data (or metadata?) to collections rather than individual curation efforts that may have been traditional in the past.
Below: The first result from Google images for the search term, "curator." -- Right after it were images of an enemy(?) from World of Warcraft.

Post 1: Retro Reprint Requests

In the September 2009 issue of the Scientist Steven Wiley has a short commentary on the disappearing phenomenon of the reprint request in academic publishing. Before online journals with open access were common, researchers commonly sent request cards directly to the authors of articles they were interested in. Authors would then mail the articles directly to the researchers, who may not have been able to find the full-text articles and only found a citation in an index.

Wiley laments the passing of the reprint request for several reasons, but mostly because they provided a valuable feedback mechanism. Authors could know which articles provoked the most interest and know who was working on similar topics, as well as even sometimes read short notes from the researchers involved. In addition, they provide a valuable reality check in the case of articles that the author may have thought to have more importance than the academic community acknowledges. Requests often provided detailed information on the researcher, their department, affiliations and subject interests.

It is interesting that Wiley bemoans the loss of reprint requests as a "rapid feedback mechanism" when we have a similar function in the online sphere: e-mail. However, his point in general is definitely valid. Authors submitting to institutional repositories and similar online collections don't have any direct way of knowing who exactly is reading their journals and what their level of involvement is. We seem to be bound by standard web practice, which encourages simple anonymous statistics (page counts, clickthroughs) that are of little importance to the academic author interested in the community that surrounds their work. Simple page counts do not reflect more subtle details such as whether someone was simply browsing and scanning articles, or whether they took the time to read one in full.

We need to consider the community building aspects of academic cyberinfrastructure in addition to their technological aspects. Systems such as validated logins which would identify users to authors of articles might be useful in future attempts to bridge the gap between readers and authors. We can easily imagine an institutional repository with social-networking-style features that tracks affiliations, connections and research interests, as well as articles read, tagged and commented on.

Post2: Twitter, meet Science: The World of Scientwists



“532 Scientific Twitter Friends” is a July 9, 2009 article that appeared on Sciencebase, an online magazine written by David Bradley, who describes it as the “place to be for reactive science communication on the web, covering all areas of scientific, technical, engineering, and medical topics in a timely and informative manner” (faq). This article marks the passage of the 500 marks for Bradley’s science related friends ("friends," in this case, seems to mean that there is a reciprocal relationship of "following") on Twitter. At the beginning of 2009 Bradley assessed his Twitter friends (through a process of downloading his friends list, saving the CSV files and opening a spreadsheet in order to sort), and discovered that 100 of his current contacts were members of the scientific community, he listed these “Scientwists” on his blog and so began a list outside of Twitter of Twitter science tweeters. Similar to our discussion of Twitter as a means of gaining tenure, he states that list “formed the basis of a nice trade – you retweet or comment on the list and I’ll add you if you’re aren’t already on it and if you are I’d activate your twitter name and so it grows.” This has proved true for his own Twitter account, as well as others – some tweeters listed on Bradley’s list gained hundreds of new followers.
As evidence for Bradley’s argument the graph illustrates how over the course of a few months his followers grew by hundreds due to retweets and articles related to twitter and science writers.

To date, he is up to 5,220 followers and follows 1,856 tweeters.


Inspired by Bradley’s use of Twitter in the scientific community’s communication, Andrew Maynard of the blog 2020 Science wanted to “get a feel for how science information is beginning to flow between different communities and users on the web” in his article "As Twitter users skyrocket, how are the science tweeps doing?" Thus he created a bubble chart representing science tweeters and their number of followers, including some tweeters followers in the tens of thousands.


Though he recognizes that his measures are crude and the study incomplete he believes that it “at least suggest that scientists and science writers are beginning to embrace new social media.”

Wednesday, September 2, 2009

The good, The bad, and Google Books.

The article by the Open Book Alliance sets out some of the key arguments against what could end up being one of the defining moments in how books are accessed by the next generation of students and users. The Key issue is Google's continuing push to digitize large collections of books (some ten million) and the settlement with the Association of American Publishers (AAP) which clears a path for them to do so. The key issue that the authors raise is access and I for one see their point on many of the issues that they raise but I have some reservations with their arguments.


The nature of the settlement with Google or rather I should more correctly say the nature of the settlement as I have had it explained to me by various news sources many with biases that should be acknowledged, is that this settlement hands Google a great amount of control over the collections Google a monopoly on the dataset. This monopoly, it is feared, will make competition with Google in the realm of book searching impossible. Google will have the ability to set the price for access to the dataset and this pricing structure could be prohibitive to small institutions widening the gap between the haves and the have-nots.


This argument is very interesting to me. I for one understand the fear of handing over vast sums of information to a corporate entity whose primary concern is the fiduciary well being of its share holders, however I still use Gmail. I used it to get the account I am posting this on. I still use a large number of Google programs, documents, reader, search , etc. Every time I use these I know that somewhere out there Google is keeping track of the information that I give it and using it to do.... something. I guess my point here is that I am not convinced it is in the fiduciary well being of Google to act in a restrictive way with the digital collection that they are building. (And yes I do think that this is a digital collection rather than a digital library despite the fact that books are being digitized).


However, I do believe in balance and I support what the Open Book Alliance is doing by promoting a “fair and flexible” solution. This project has been in the offing for quite some time (2006 by my count) why was this coalition so long in the building?


I do want to say that there are two main issues here which are pros and cons in my mind concerning this settlement.


Pro: What really excites me about Google acting as the driving force behind the push for book digitization is the possibilities of interaction that with the written word that this project promises. Google has extensive API libraries and is very friendly towards processes that allow data to be manipulated in a variety of ways. This could potentially be a great boon to those who have been looking for interactive libraries for years. This could be the catalyst that sort of interoperability.


Con: The way this deal was reached. I really think that this settlement set a difficult precedent for the way issues of copyright and use to be decided. I am not at all comfortable with the settlement.


Google might not be evil but that does not make this sort of legal wrangling beneficial to the culture as a whole.

DCC&U: An Extended Digital Curation Lifecycle Model



By Panos Constantopoulos, et. al., Athena Research Center Digital Curation Unit

The International Journal of Digital Curation, Issue 1, Vol. 4, 2009

This article was written by the Digital Curation Unit (DCU) at the Athena Research Center in Athens, Greece. The goal of the paper was to compare digital curation models introduced by the Digital Curation Center (DCC) and the Digital Curation Unit (DCU). Their proposal is to combine the two models, in order to highlight the “the need to extend the digital curation lifecycle by adding provisions for the registration of usage experience, a stage for knowledge enhancement, and controlled vocabularies used by the convention to denote concepts, properties and relations.”

The authors describe some of the primary questions that are arising out of the deluge of digital information as follows: how do we ensure authenticity and integrity of information; what information do we preserve and how do we preserve it; and how do we ensure usability and accessibility, especially as its context and uses are continually changing? The DCU proposes that digital curation works to answer these questions, and that its fundamental principle is that ensuring “future fitness for purpose” of digital information requires active management and appraisal over the entire lifecycle of the digital assets.

The DCC model approaches digital information management from a lifecycle perspective and addresses authenticity, integrity, and support for predictable preservation (the original DCC model is above, left). The DCU model has grown from cultural heritage informatics and has greater concerns regarding ontologies, metadata, natural language processing and dynamic representation of data. Both models focus on the trustworthiness and preservation of information, as well as value-added services. However, the DCU model is additionally concerned with the production and communication of knowledge as well as information about the users of this knowledge. Here, the “fit for purpose” depends on “adequate and appropriate” representation of the data.

I appreciate the DCU’s consideration that knowledge enhancement as an essential component of the curation lifecycle. It takes into account the increasing use of Web 2.0 and Semantic Web technologies, and helps to differentiate the concept of data curation from the slightly-more-straightforward digital archiving or digital libraries. In order to understand the goals of digital curation, it is important to understand not only the lifecycle needs but also the expectations of any collection of digital information and knowledge.

The enhanced model, above right, incorporates knowledge enhancement (into the “Curate and Preserve” ring), authority’s information – in the form of controlled vocabularies, etc. (into the “Description and Representation Information” ring), and user experience (following “Access, Use and Reuse” in the outer ring). This can be compared to the original, seen below in Ramona’s blog about the Edinburgh Mouse Atlas Project case study conducted by the DCC.

I am interested in following this and future proposals regarding the original DCC model to watch how it evolves as refinements are suggested. Is a “unified theory” feasible, or will models vary based on the function, use and expectation of the digital information.

Google Books: A Metadata Train Wreck?

A recent post on Language Log titled “Google Books: A Metadata Train Wreck” by Geoff Nunberg argues that the errors in Google Books’s metadata are egregious and unforgivable. He wrote the blog entry after presenting his conclusions at the Google Books Settlement Conference. He cites many examples of poor metadata in Google Book’s collection – dates that are so wrong they’re funny, searches that return impossible results, and ridiculous classification errors. Nunburg says that Google should take responsibility for fixing the metadata problems, and he questions whether Google’s engineers should be trusted with the “Last Library,” as he calls it.

Nunberg reports that Dan Clancy of Google responds to his charges by saying the errors are the fault of the libraries (and metadata providers) supplying Google with the content. According to Nunberg, Dan Clancy also suggests that users could fix the errors one by one, in the way they fix errors on Wikipedia.

Finally, Nunberg suggests that the Antitrust Division of the Justice Department should consider Google’s poor metadata design and implementation resulting from “no contractual obligation, and only limited commercial incentives, to get it right.” Check out the comments to the post - they're pretty interesting as well.

We’re all probably familiar with the quality vs. quantity question: do you use your resources to digitize 1,000 items with great metadata or do you use your resources to digitize 1,000,000 with not-so-good metadata? Crane touches on this when he talks about noise. Lynch seems to argue that large collections – even with not-so-good metadata – have the potential to be more useful because they provide a larger “data set” on which to aggregate and recombine information. Are Google’s errors unforgivable, or are they an unfortunate but inevitable consequence of creating such a large collection? Who should take responsibility for them? Is Google the “curator” of Google Books?

Digital Curation and Privacy

Last week in class, we talked about the difference between a digital library and a digital curation project. While the difference between the two hasn't been entirely fleshed out, the essential divide seems to be that digital libraries have a specific and narrow focus in terms of the scope of their curation whereas digital curation projects (e.g., Google Books) are more broadly focused and their curation efforts are less constrained. While there are certainly many positive aspects to a broadly focused curation project, the unwieldy nature of such projects can give rise to a number of issues, one of which is privacy.

In his talk at the Web Wise Conference in 2002, Clifford Lynch delineates some of the positive and negative aspects of digital curation projects. While his talk predates the implementation of most of the major projects which are currently ongoing, Lynch presciently predicts that books will begin to "talk to each other," meaning that books (or artifacts in general) in a digital collection will create linkages between each other effectively making relevant information more accessible. Lynch predicts (presciently again) that the downside to all of this sharing is that there will come a time when books do more than talk to each other, they will also talk with "external programs and organizations and people." Lynch states that this will create a situation that is "potentially annoying and invasive and raises some privacy issues."

This morning, I noticed that the most e-mailed article on the website of the New York Times was the article "Facebook Exodus." The author cites a number of reasons why people seem to be leaving Facebook, but the most widespread reason seems to be Facebook's lack of attention to the privacy of its users1 personal information. "It's not 'your' Facebook profile. It is Facebook's profile about you," one former user declares. Another laments that the site seemed to be "crawling with mercenaries trying to sell books and movies."

While not a digital repository in the traditional sense, I think that it would be fair to characterize Facebook as a digital curation project. Like Twitter, it curates information about its users and creates linkages between the various bits of information. As digital curation projects, Facebook and Twitter raise some interesting and difficult issues regarding the ownership of the information contained in their sites because unlike a digitial library where there is a clear distinction between the 'object of study' and the 'user,' users of Facebook and Twitter effectively function as both 'user' and 'object.' In this instance, privacy becomes a tricky subject and it's a subject that has received, in my opinion, far too little consideration especially because both Facebook and Twitter are being used by millions of people most of whom must have no idea what is being done with their personal information and whether they have any rights to protect that information.

In October 2007, at the OECD Participative Web Forum, Richard Ackerman of the blog Science Library Pad asked a panel the following question regarding Facebook Applications: "How do we deal with privacy when we expect that sites will want to interlink like this, that people will want to connect their information..." The response Mr. Ackerman received from Mozelle Thompson of Facebook boiled down to "...if you don't want to share your information with that application, you should not download that application." Surely, there must be a better way to restrict access to personal information.

In his blog post, Ackerman points to a press release from the Office of the Privacy Commissioner of Canada dated August 27, 2009. First of all - Canada has a "Privacy Commissioner?" Wow. The Canadians seem to be at the forefront of reigning in social networking sites so their users have more power to restrict third-party access to their personal information. Up to this point, application developers have had "virtually unrestricted access to Facebook users' personal information." However now, Facebook has agreed to "retrofit its application platform in a way that will prevent any application from accessing information until it obtains express consent for each category of personal information it wishes to access." Check out the press release for a full listing of the privacy restrictions that Facebook has agreed to.

One of the themes I think would be interesting to discuss over the semester is - what happens to privacy considerations when the distinction between 'user' and 'object' becomes blurred? Do the benefits of creating linkages between information trump the user's desire for privacy? Is there a happy medium, and could that medium be provided through government regulation? Or is the best option for an individual concerned about privacy to simply opt-out? Incidentally, this is the option that I have decided to take. I am, admittedly, one of the few non-Facebook holdouts, but it seems like there will be more people joining my ranks in the near future.

Academic Evolution

Scholarly Communications must Transform

I got this series of articles from the blog "Academic Evolution"; Gideon Burton, the author of the blog itself, has written a series of articles discussing how and why scholarly communication needs to undergo a radical transformation in order to remain useful and "alive" as we enter the digital age. The series focuses on ways that scholarly communication must change in order to "keep up with the Joneses," so to say. Burton picks openness, standard compliance and syndication as three major areas in need of development.

Burton started out with outlining the traditional method of scholarly communication and discussing how academia is eager to join the digital environment as its potential for elevating the level of scholarship is so possible and evident; however the openness of the digital culture and the exclusiveness of traditional scholarly communication are at odds with one another. Burton maintains that the online/digital environment is fast becoming the new and accepted method of dispensing information in the future and that if academia wants to maintain its current exalted position in the information world, then some major changes in scholarly communication will have to take place.

The first area that Burton tackles is the idea of openess; he lists areas where openness is needed
  • Open Access
  • Open Review
  • Open Dialogue
  • Open Process
  • Open Formats
  • Open Data
These are all areas where traditional scholarly communication are known for their exclusivity and, according to Burton, greater transparency and flexibility needs to become an accepted characteristic within the academic publishing world if they are going to join the digital environment.

The second area that Burton focuses on standards compliabilty. He maintains that, due to the rise of the semantic web and increased demands of interoperability, information will have to become "format agnostic," in other words, be independent of format but be formattable for the semantic web.

The third area of change that scholarly communication needs to undergo is syndication. The presence of Really Simple Syndication (RSS) ensures that information can be easily disseminated quickly and effortlessly to a large group of people. Burton suggests that scholarly communication needs to be syndicated in order to gain its audience and create dialogue which in turn, inspires new areas of thought, discoveries, new insight.

The way that Burton's series of articles on transforming scholarly communication directly applies to our in-class materials is that it, like most of the other articles, explores the difference between print sources and digital sources and examines the long term ramifications of the changing format of information.

OpenWetWare.org

Open Source Science, Open Notebook Science, Science 2.0 – call it what you want, but the notion of the scientific community sharing their data openly online is a heavily written about topic in magazines and blogs. In January 2008 Scientific American published an article called "Science 2.0: Great New Tool or Great Risk?" which talked about the future of science, how the scientific community could be using the web more effectively and why many are reluctant to do so. One of the success stories mentioned in the article that I found particularly interesting was OpenWetWare.org

In April 2005 graduate students at MIT created OpenWetWare, which was initially a private lab wiki for two science laboratories at MIT. However, in June of 2005 the site opened itself up for any lab to join. According to their information page, the site is now currently editing over 13,404 pages by 6,571 users. These users come from more than 100 research laboratories in over 40 institutions including Caltech, Harvard, MIT, Tufts, Stanford, Johns Hopkins and here at the University of Texas . Labs involved in OpenWetWare come from around the world - in Asia, Europe, Australia, South America and the United States.

OpenWetWare was created in an “effort to promote the sharing of information, know-how, and wisdom among researchers and groups who are working in biology & biological engineering.” The wiki site provides a place for people and laboratories to “organize their own information and collaborate with others easily and efficiently.” It offers links to blog posts, a place to share research protocols, materials, resources, and other lab techniques. The wiki page even allows for users hosts courses where they can post lab results, ask each other questions, discuss the answers and collaborate on papers.

The site runs on customized MediaWiki software and uses Linux servers. According to the Wikipedia entry about OpenWetWare, all of the content “is available under free content licenses, specifically the GNU free documentation license (GFDL) and the Creative Commons Attribution ShareAlike license.”

While I'm not certain exactly how much raw research data gets shared there, OpenWetWare is still interesting to me for two reasons. The first reason is that it started out simply and small as a wiki intended for two laboratories to be able to better communicate and share information, and it has grown and evolved into so much more than that. It even earned a grant from the National Science Foundation in May 2007 in an effort to continue evolving into a self-sustaining community. Secondly, the success of OpenWetWare, and new early success stories like the already mentioned PLoS Currents: Influenza, gives me hope that more in the scientific community will see the benefit of open sharing and online collaboration, and will overcome their reluctance to join in the open source science revolution.

Article: A new website for the rapid sharing of influenza research

The Public Library of Science has launched a beta site to facilitate the rapid sharing of research being done on influenza. PLoS Currents: Influenza is encouraging researchers to submit their findings in order to facilitate rapid progress in influenza research. All of the content is available under the Creative Commons Attribution License.

Clearly, the development on H1N1 as a global concern has created a need for an organized network of information sharing. Since there is so much research being done on influenza and there is such an urgency in understanding and containing the virus, PLoS has a great opportunity to demonstrate the benefits of open communication between scientists.

The curation, organization and presentation of PLoS Currents: Influenza generally relies on Google knol which means that anybody with a Google account is able to contribute an article or comment on an article. However, Google knol has recently allowed communities of knols to be moderated for content, which allows PLoS to rapidly review all articles that are submitted to the knol and maintain moderation throughout the life of the community. The PLoS Currents: Influenza beta has 25 moderators and they are described as an "expert group of influenza researchers". It would be interesting to know if 25 moderators will be necessary to successfully oversee everything within the knol or if the knol could function with fewer moderators while requiring some level of credentials to contribute to the site.

While PLoS Currents: Influenza is available and in development through Google knol, the collection is simultaneously being archived permanently at the Rapid Research Notes, hosted by the National Center for Biotechnology Information. While all of the articles are available in full-text at the RRN site, they also include a link to the Google knol version. The process for adding articles to the RRN archive involves many more restrictions than adding an article to the PLoS knol, the first criteria being that only publishers may submit articles. Additionally, the publishers must submit the articles already marked up in XML according to the NLM's Journal Publishing DTD. I personally would be interested in seeing other repositories that require documents to come already marked up.

Right now there is a lot of buzz about what open access to research should look like. Actually, there appears to be a full-blown fervor over the speculation of how researchers and information organizations should work together to present findings. It seems that everybody has an opinion on how information needs to be free. PLoS is in a unique situation because they already have experience in open access to science, they are an authority on the subject with real-life experience in applying these practices to see how well they work. The success of PLoS Currents: Influenza will provide further insight as to how scientists do their work including when and why scientists reveal their work to the rest of the world. Since PLoS is asking for very preliminary work from researchers, there really is no time that is too early to submit findings.
The article I looked at is "Reinventing academic publishing online" at the First Monday journal. Brian Whitworth and Rob Friedman argue for a change in academic publishing, describing the present journal publication system as feudal in regard to knowledge exchange. Their concern is specifically with the computing and information systems (IS) disciplines, and they examine how the traditional journal publication practice has squandered the fields' academic and theoretical potential.

Problems with the traditional journal publishing route are: small readership, slow article gestation, small dissemination, conservatism, subject-specialized, and geared toward career advancement and business over pure knowledge dissemination. The authors also point out the conflation of "rigor" with "quality," arguing that rigor alone stifles theories and new ideas, eventually signaling the end of the journal itself. The conservatism charge is especially potent; the authors highlight how a seemingly dated idea like media richness theory (MRT), "which links 'rich' media to rich interactions," which has been challenged by the successes of email and text messages over predicted successes like video chat, is still prominent.

They argue that IS, as a relatively new and cross-disciplinary field, has been hit especially hard. The sentiment that theory is worthless while practice alone is sufficient has its origins, they argue, in this dated publication model that cannot afford to nurture or entertain young theories or new researchers. New ideas are then frequently manifested by practitioners in the commercial realm. While this isn't bad, it could be better; the authors believe theory and practice have a synergistic relationship.

As others have noted stagnation and hardening of ideas is not a new problem, nor is institutional self-interest, but digital publication and digital curation of articles and research at least promise a faster cycle, or more Kuhn-type revolutions per decade, or some speeding up of ideas and counter ideas. The authors will likely get into this in future parts (this is Part I), but it is interesting that they associate so much of what is problematic with current scholarly publishing with the old analog model. It's a lot of weight to put on digital journals alone, and I bet some of the solution will be allocated to an open-access policy along with digital publication.

Since the authors believe knowledge flowers at the crossroads of various disciplines, it seems especially important that data sharing and curation be robust there. This makes me concerned for metadata interoperability; it seems significantly harder to administrate a repository that contains an article with datasets on historical music composition, music composition algorithms, and network technology, or some other convergence of disciplines and data, and then facilitate sharing of this data with data from other fields. Add to this the authors' dissatisfaction with specialized cross-disciplinary journals ("intersection journals") like International Journal of Computational Models and Algorithms in Medicine (IJCMAM). For the authors a journal like this "satisfies the need of the many to publish, but retains the tradition of dividing knowledge into artificial and disconnected fiefdoms." This seems to put a lot of weight on the shoulders of digital publication, collection and curation since these practices will be partially responsible for breaking down these "artificial fiefdoms" with interoperable, open metadata and highlighting how one study informs another.