Wednesday, September 30, 2009

A Real Solution to High Textbook Prices

This week I read about a piece of legislation sure to be a hit with my fellow students: Dick Durbin (D-Ill.) has introduced the Open College Textbook Act, which would provide competitive grants to create open college textbooks. Although Durbin calls for the textbooks to be "freely" available, it seems like there may be a bit of confusion over whether these textbooks are open in the sense of open source software (free to modify and re-distribute) or whether they are truly free.

Durbin's plan avoids the pitfalls of Wikipedia and other similar open sites (which many college students use anyways) by providing for a review process of the material. In addition, Durbin is seeking to bypass the inflated prices we have seen recently by textbook publishers who bundle software and online portions to their textbooks. Personally, these addons often seemed like nothing more than an additional investment in time that students would have to take, instead of enrichment. Often, addons such as CD-ROMs are just ways for publishers to save money, by including PDF answer keys and similar content that could just as easily live online.

There are already free textbooks for basic subject such as physics and calculus online, freely accessible to students and the public alike. Since these subjects are well-established and not likely to change any time soon (you could just as easily teach yourself how to find derivatives with a textbook from the 1960s as from 2009) this is the right approach for these topics. Also, Durbin's approach follows other recent projects such as MIT's Open Courseware that make class materials available to anyone online. I can even think of one class here at the iSchool that has no textbook but instead pointers to free online tutorials that one can use to study HTML and PHP.

The question with these free content projects often falls down to "who's going to pay for it?" Since these are grant-funded textbooks, that problem has been solved, as long as the grants don't dry up. Universities seem to be of two minds on this topic: either they gladly make their course material available (as in the MIT case) or are worried about losing their intellectual property by putting the materials online. In addition, professors who have developed custom coursework often are in a bind because the university has rights to their own material, and they don't want to see it stolen and reused at other universities. Anyone whose professor has put passwords on downloadable course lecture notes will be familiar with this problem. This open solution could bypass the problem for professors: it's owned by no one (or everyone?) and is freely reusable.

The open access approach seems to be a nice step away from the consumer mentality that seems to plague current college students.Students "shop" for courses and professors and evaluate courses as if they are consumer products. See this article by Leah Schweitzer for a better analysis of this subject. In a way, education is a high-priced consumer good, but it is the work that students put in, in collaboration with professors, TAs and fellow students that can raise it above a passive experience. We want to believe that the halls of academia are more noble and lofty than the ITT Tech outlet in the local strip mall. Now, students with enough determination could in fact bypass the university and teach themselves with open course material.

I think this bill is a step in the right direction. Universities would have the chance to distinguish themselves in the level of personal service and creative projects they provide to students while still using the same materials as other schools. If academic papers and basic encyclopedias are available through database subscriptions at almost every college, there's no reason why basic textbooks shouldn't be available to everyone online. It would definitely make looking up the definition of Simple Harmonic Motion at 3AM easier right before that exam, without giving blind trust to Wikipedia.

PLoS and the new Metrics

This week I am looking at the Article Level Metrics recently implemented at PloS and what the might mean for the journal publishing world. One of the sublime pleasures of closely watching a growing field is how quickly the terrain can change.

In the tradition model of academic publishing the key metric to a “successful” article was the number of times it was cited in the literature. It should be noted here what a very poor metric this is to judge the influence of an article, the information that is contained in the publication could be well, suspect but if it provokes conversation or is used as a counter example then it will still be cited and still be seen as influential when counting citations. While perhaps there is no system that is perfect it is obvious that this traditional model could use a bit of up-dating.

The folks over at PloS have implemented a new set of article level metrics that go beyond the tradition of counting only citations and have begun counting other ways that articles are accessed and used.

The new metrics count the number of times the article is access, book marks, comments, ratings blog links as well as citation. The ability to collect all of this information is one of the exiting changes that online journals can offer. This collection of way that the journal is used can tell more about the quality of the work than the one dimensional view provided by examining citations alone.

Of course there are some know issues already when it comes to using these new metrics. For the most part PloS is up front about the issues inherent in these metrics and lists the steps that they are taking to address them. We should remember that each of these metrics has strengths and weakness but taken as a whole they reinforce each other and potentially have the power to change the way journals are used.

Now if we can just change the tenure paradigm in tandem.

Government Data

Government data is hugely useful in lots of applications. Three examples are social science research, transparency in government, and an individual's decision about where to live. It appears government data is making its way to the web. Data.gov hosts datasets from many different federal agencies, such as the Environmental Protection Agency, the Bureau of Transportation Statistics, the Centers for Disease Control and Prevention, and the Department of Education. You can get the raw data - XML, CSV or text - or you can use (more or less) people-friendly tools from their tool catalog. The tool catalog is a collection of widgets and extraction tools that are supposed to make the data on the site useful. For example, the MyEnvironment Widget allows a user to enter her zip code and get all kinds of related environmental data on a single page, complete with pie charts, maps, and graphs. I discovered a landfill in my neighborhood that I had no idea about. Crap. Related to data.gov is DataMasher, which we looked at in class earlier in the semester. It takes government data and allows registered users to create "mashups" of that data. For example, there are mashups of Hate Crimes vs Population, State Education Spending vs SAT scores, and Federal Spending per Congressional Representative. There's definitely a push out there to go beyond making the data available, but also making it useful to developers and end users.

And it's not just the federal government that is putting its data online. San Francisco recently launched datasf.org, which has datasets related to elections, environment, geography, health, housing, public safety, public works, and transportation. The site also features an App Showcase, which at this point mostly contains interactive maps and phone apps related to crime and the BART trains. I think it's a good idea to put the applications in one place; it makes it plainly obvious how important open data is. It occurs to me that if I were an app builder in Austin, I'd start pushing Austin to open up its data. Economic incentives might make the open data/transparent government pill a little easier to swallow.

The W3C, the definitive authority on all things web, published a working draft of Publishing Open Government Data in early September of 2009 that provides "straightforward steps to publishing government data." They recommend making the data available in well-known machine-readable formats such as XML, RDF, and CSV, while also making the data available in human-readable formats by putting it in HTML and using "permanent" and findable URIs. They provide guidance on documenting the data, preserving the data, not creating interfaces that prohibit access by other applications, and choosing what data to publish. The last point is an interesting one. How do you choose what data to publish? One of the people behind datasf.org, Jay Nath, writes about the 80-20 rule of government data. He discovered by analyzing downloads on the Washington DC Data Catalog that 80% of dataset downloads came from 20% of the datasets, mostly crime- and 311-related data. Nath and his group prioritized the data they made available on datasf.org based on these kinds of download statistics. One has to wonder, though, if the most unpopular datasets will ever make it online. If the only motivation for opening up government data is economic, then we might not ever see true transparency.

Cyberinfrastructure = Hardware + Software + Bandwidth + People

Cyberinfrastructure = Hardware + Software + Bandwidth + People

This week I was really not sure what I could write about, but then I came across Michael Roy's article in the Academic Commons that discussed the idea of cyberinfrastructure and building a computer research center at Northeastern University. I'm still working out my own ideas of what a cyberinfrastructure exactly is and this article featured a graspable array of issues and ideas that make up the necessity of establishing an infrastructure for research.

Roy begins his article by stating that computers were once thought to be for research, but then they became known to be teaching aids, but now the pendulum is swinging back to the idea that computers are used for research. He divides his article up into three parts that each address particular areas of establishing a IT research center: the practical issues, such as equipment, servers, staffing, the politics and the process and then the academic research process.

The first section elaborates on how Northeastern University decided to establish a Computer Research center. A committee was created and they interviewed the faculty to see what kind of research projects were currently being done and also asked about future projects to gain a picture of research needs. They decided at the end to build a cluster with features that focused on job management, size and interactvity. Then they had to set up a proposal to find the vendors that would be able to provide Northeastern with the technological equipment needed.

Needless to say, the IT Computer Research Center unleashed a series of critiques and skepticism amongst the faculty and within the academic world (mainly from the liberal arts world, predictably); however Leslie Hitch, Director of Academic Technology at Northeastern University, pointed out that Northeastern had a made an organizational goal to increase the quanity and quality of faculty research and the establishment of the Computing Research Center fell within the parameters of the university's mission. Then Roy began to discuss some of the challenges that they've had: staffing--how to staff the computing center, how much control to allow for, what defines productivity and so on.

Essentially, this article was useful for me to read because it provided some kind of framework of an individual institution's efforts to bank on digital research and clearly outlined some of the challenges we continue to have with defining the parameters of computing research and exactly what to do with it.

Twitter, Curation, and Horizontality

I'd like to start off by quoting The Boston Globe as reported by Tim Leberecht on The PopTech Blog: "Just as technology is giving us the ability to amplify every word we utter we have nothing really meaningful to say." While we could simply label this injunction as yet another irascible Bostonian diatribe, I think there is a good point here and it's not (of course) that no one is saying anything interesting anymore. The point is that people are saying SO MUCH that maybe we need some sort of filter or curator to let us know what is interesting and worthwhile and separate it from what is not. Does Twitter need a curator to make it meaningful?

We are already seeing some web-sites and applications which are doing exactly that. Breaking Tweets lists links to people are are tweeting about global "news worthy" events. Alltop lets you know about the celebrities and bloggers whom you "should" be following on Twitter. Applications like TweetDeck allow you to create your own filters so that you can organize your tweets and Facebook updates into categories like "news" and "friends" and "ridiculous" (there's a lot of that). From this, one can extrapolate ways in which other Twitter users could use websites to access tweets about various subjects: botanists could more easily access tweets about plant research, and journalists could access tweets about global events. Right now, the only way to filter tweets is by keyword searching for hash tags, but that can be a time-consuming and frustrating process.

While I can definitely appreciate the increased accessibility that such websites afford to Twitter users who have a specific interest, I also think there is something to be said for the horizontality of the current state of Twitter. All Twitter users have only 140 characters in which to express themselves, and no one is given more or less space based on their social status or another other factor. There's something egalitarian about the current state of Twitter and I think we're losing something important if we start imposing a hierarchical system of selection.

While I stated in last week's blog post that I like the idea of individual Twitter users curating web-content and posting it on Twitter, the idea of placing filters on Twitter itself seems to miss the point of Twitter altogether. In an interview for the Los Angeles Times, the co-creator of Twitter, Jack Dorsey, talks about the original concept of Twitter (this is a lengthy quote, but I feel like it's worth reproducing in its entirety): "The concept is so simple and so open-ended that people can make of it whatever they wish. They seek value and they add value. I've always said that Twitter is whatever you make of it. Because the first complaint we hear from everyone is: Why would I want to join this stupid useless thing and know what my brother's eating for lunch? But that really misses the point because Twitter is fundamentally recipient-controlled -- you choose to listen and you choose to leave. But you also choose what to put down and what to share. So if you decide to hook your plants up to Twitter and have it report when it needs to be watered, then that's a valid usage, or if you just decide to report what you're eating for lunch, that's a valid usage too." I think this gets at what I think is the uniqueness of Twitter - it's a giant behemoth of little snippets of information which the user must wade through to determine what he or she does or does not find interesting or useful. I don't know that we've really seen anything like this before, and I'm worried that if we start using websites that filter tweets for us, Twitter is going to become just another news website.

Twitter recently announced that it is, in fact, saving everyone's tweets. I like the idea of having one long, historical feed of 140 character blurbs with no form of organization apart from a date stamp. I think it will afford a unique, unfiltered way of looking at history.

Digitizing Historical Documents at the Mint

This week I found a rather old web article that describes a relatively early effort at creating cyber-infrastructure in the realm of the social sciences, namely the digitizing of historical records and photos at the United States Mint.

The article, posted in 2001, briefly describes a digitization effort at the Mint that involved 225k pages, 5k photographs, 2.6k microfiche reels, and over 400 monographs being scanned and stored on servers as tif files and pdf files for display. Two separate databases were used to keep track of the resulting digital output, one for accessioning and item records, and a second for increased cataloging and description. The databases each possessed both internal and external interfaces. Finally, finding aids were put on the web via EAD.

What struck me about this article is that it lists specific difficulties encountered in digitizing historical records, many of which still plague us nearly a decade later. For example, Rothfeld bemoans the lack of funding for large projects such as at the Mint. While electronic storage space has become remarkably cheaper, scanning unique items by hand remains a labor intensive, costly process. Furthermore, software costs and limitations often determine what precisely a collection can do. Rothfeld also found it problematic to be attempting to instill "tech people" such as the database designers with the full idea of how historical archives get used. Conversely, her archivists did not program. Finally, Rothfeld noted that the big challenge that remained for her project was achieving interoperability with other, outside valuable sources and incorporating outside materials into the Mint's online collection.

The answer to those problems, of course, is to further the cyber-infrastructure, not just in terms of the goods (i.e. digital replications of historical or other social scientific data) but also in terms of channels of information. Of particular importance, I think, is the education aspect of building cyber-infrastructure. Many of the problems Rothfeld encountered would have been mitigated if the programming and historical/archival planning skills had been vested in the same individuals. With a higher level of 21st century information literacy, cross-trained individuals would also be able to better take advantage of some of the low-cost open source software that can be customized to be at least somewhat field specific.

In conclusion, while it is encouraging that much historical (and other social scientific) data is finding its way into a digital format, further effort does need to be spent on developing cyber-infrastructure within the field.

UK asks data developers for feedback

If you've ever tried to access public data from a government website, you've probably noticed how difficult it is. It seems like the government should be better at organizing data, and certainly it is their responsibility to make that public data accessible. The problem is that the expensive websites the government builds are rarely usable.
The UK may change that. The UK's Cabinet Office website government website has posted an open call of sorts to data developers. While the details of the project are vague, the site states that there are "over 1000 existing data sets, from 7 departments (brought together in re-useable form for the first time) and community resources".
They are asking developers to offer feedback on their current site design and ideas about how to transform their data into a usable, single point of access for government-held public data. They are interested in any features that should be added or changes that should be made to the site or even any data that should be there but isn't. To participate, developers need only to join the Google Group.
The post states: "From today we are inviting developers to show government how to get the future public data site right - how to find and use public sector information."
I am having trouble envisioning how this kind of crowdsourcing could be bad, but I have no trouble imagining the good that could come from it. The added input may help anticipate broader uses of the data in addition to providing design changes that will make the public data usable. I am excited to see this kind of openness in the site creation process and look forward to seeing whether or not it proves to be successful. All the digitally stored data in the world isn’t going to do us any good if it isn’t usable.

How Last.fm inspired a scientific breakthrough

http://www.guardian.co.uk/technology/blog/2009/sep/16/last-fm-mendeley-victor-keegan

Mendeley has launched a research paper organizing service with software that is similar to the existing PDF organizing tools (Zotero, CiteULike, Connotea). Not only does Mendeley work really well to help organize citations, it also keeps statistics on what people are reading and from what field of study those readers come from. This way, users can look at the statistics for what is most popular in their field.

The article from the Guardian is particularly enthusiastic about Mendeley being the revolutionizing thing that the internet/researchers/science have been waiting for. They say Mendeley "shows the second phase of the dotcom boom is throwing up great, practical ideas". This is from a technology blog. After reading a few statements like this I downloaded Mendeley and tried it out. It is pretty interesting to log in and look at the statistics they provide, but I'm not sure if it will go much further than being interesting to look at. The have a lot of people joining very quickly, so it might be possible that Mendeley will change the face of science (this is what Dr. Werner Vogels of Amazon predicts).

In Mendeley authors can upload PDFs of their work and include it in the Mendeley database. Mendeley has a very interesting take on the legality of this in their Mendeley FAQ. They say that most journals are usually okay with authors "self-archiving", which is not defined. Mendeley's simple answer to self-archiving is "Simply self-archive your preprint as well as your postprint, and wait to see whether the publisher ever requests removal." After downloading and using Mendeley a little bit I am still not clear on how much access is provided to the articles that authors upload.

There's one thing that the author of this column wrote that I do not understand:

"Mendeley says that instead of waiting for papers to be published after a lengthy procedure of acquiring citations, they could move to a regime of "real-time" citations, thereby greatly reducing the time taken for research to be applied in the real world and actually boost economic growth."

What is this lengthy procedure of acquiring citations? I've asked a couple of people about this and nobody knows what he's talking about. Is he somehow confusing this with the lengthy procedure of peer review? He doesn't actually mention peer review in the blog at all. Maybe the technology column of the Guardian isn't the most reliable source on how scholarly publishing works.

Blogs (hangingtogether.org, Open Access News) have mentioned that the metadata in Mendeley is horrible. This looks like another case of why won't fill in organization start fill in the practice we like. The metadata might not be that great, but providing stellar metadata is probably outside the scope of Mendeley's goals. One component of Mendely is to crowdsource articles, which points the argument back to folksonomies vs. cataloging. While it would be great to have a service that uses both tagging and quality metadata, I'm not sure what that kind of service would look like or who would actually do the work to create it.

A data model for interoperability between scholarly repositories

This 2007 paper (actually submitted October 2006) from the International Journal of Digital Libraries proposes a model for improving interoperability of metadata and data between scholarly journals. Regardless the proposed model presented here, the authors point the most pressing shortcomings in the current state of the art. But as a side note, since this article was written about three years ago, and since I haven't done much scholarly research or any publishing, let me know if a critique is no longer true.

They start out by noting, as we have in class, that the traditional unit in paper publishing, the article (a single "unit of communication"), need not prevail in digital publishing where units of communication can be tracked and displayed. Repositories are already housing complex digital objects that do service to these units of communication (datasets, citations) but interoperability is limited to the OAI-PMH and its mandatory stipulation of a Dublin Core metadata format for sharing.

The authors establish three "high-level requirements" in their Pathways project that would allow much greater interoperability:
  • a shared data model for the digital objects
  • a surrogate format that serializes the digital object according to the digital model
  • support for a universal way of handling repository materials as surrogates: obtain, harvest and put.
As you might imagine, the authors proceed to explain their terms and concepts. I will do that briefly and then relate an example that illustrates how their model would work.

Their data model consists of an entity with leaves of datastreams attached. Entities are recursive: one entity could be the ACM Digital Library holding multiple entities of conference proceedings, which in turn contain entities of talks and speakers. An entity has any number of datastreams attached to it, these point to actual bits stored somewhere that constitute an instance of that entity or components of that entity. This allows for both abstract (entity) and concrete (datastreams) representations of the digital object. Entities and datastreams have various values attached to them like providerInfo, format, and crucially hasLineage. This lets entities unambiguously point to the repositories of origin. Since entities are recursive this allows evidential citation chain that preserves the complex interrelationships of scholars and their work. For example, translations would natively point to the original article, the original article would natively point to preprint articles it references, and so on.

The surrogate format would serialize the digital object so it could be represented by reference. This is like creating a skeleton of the digital object where each bone points to the actual corresponding bone in a repository. That way, unless a user or repository wants to ingest the digital object, work can be done without shipping the entire digital object over the network. The authors found RDF useful for modeling surrogates.

The last piece is three proposed essential services for repositories: obtain, harvest, put. Obtain requests a surrogate of an identified object, harvest reads the surrogate, and put would submit the surrogate or parts of the surrogate to a repository.

Their example details a user selecting articles from three different repository types (DSpace, arXiv and aDORe) and putting them, to various degrees, in his own Fedora repository so they can be issued in a single journal. The enhanced usability described mainly entails the user ingest the surrogates of these articles. This lets the new journal point to the articles by reference and point to all the complex data about the objects as well (any citations, data and so on). The custodial chain of the articles is preserved as well.

A central registry of repositories is the key to their model. Although this registry is not necessarily complex (a single line for each participating repository) it is essential for the success of their model because every entity would have a providerInfo set that would let a harvester application (they argue OAI-PMH would work with some adjustments) look up the provider in a registry, contact it, etc.

The main drift of their work is to move interoperability away from simply discovering resources with a federated search via stripped-down, minimal metadata markup (Dublin Core) and allow repositories their more custom, nuanced handling. They do this by providing some very (they argue) low-barrier principles to which participating repositories can subscribe. The two most fundamental of these are the shared data model of entities and datastreams and support for the three services of obtain, harvest and put.














Tuesday, September 29, 2009

Mendeley - not Manderley

Having some difficulty finding an article to blog about this week, I chose to investigate further a document management tool called Mendeley that was brought to my attention by a professor a few days ago. I found several blog entries about it, coming from varying disciplines with varying levels of investigation.

One entry, found on Savage Minds, an anthropology blog, very logically lays out the impediments of several different bibliographic softwares, coming to the conclusion that Mendeley works best for the author, probably a student or scholar, since it is "both a web application and stand alone desktop software". This means the tool combines the best aspects of Zotero, Sente, Evernote, and others. The blogger realizes Mendeley is still evolving but has high hopes.

Another post on Hanging Together, a blog written by OCLC/RLG staff, quotes an article from The Guardian and explains Mendeley's sharing techniques, comparing it to last.fm's profile building and recommendation system. The blog rightly assumes librarians will wonder how a "web application based on an entertainment model should have proved so much more attractive than the painstakingly built repositories we have been holding under the noses of our academic authors over the last several years?". This is answered by the fact that Mendeley runs statistics about who is looking at what paper, offers cross referencing and recommendation systems, and will integrate with programs like Zotero. "Mendeley has grabbed the attention of users because it understands what they like. They like simplicity. They like instantaneous results". Mendeley apparently very lightly touches upon copyright and metadata (using OpenSource ethics and 'change it if it's wrong' metadata concept), which means "the requirements libraries often put up front are almost dismissed as non-issues" by Mendeley.

Its real time editing abilities are what grabbed my attention. I'm not sure if users can actually comment on others' papers, but by posting pre-published papers, feedback on use will be available, at least. This is just another tool that is changing the academic scholarship traditions. I'm interested to see where this bibliographic tool goes in the near future, and if it delivers on the so-far limited yet excited hype. Maybe someone will upload Rebecca, and we can all go to Manderley. Easily.

PLoS Currents

Once again this week I've found myself drawn to the topic of data sharing. In this case, I've discovered PLoS Currents, a beta version of a service that is described as "a moderated collection for rapid and open sharing of useful new scientific data, analyses, and ideas." This particular PLoS Currents forum is dedicated to the H1N1 influenza virus (aka swine flu). The purpose of PLoS currents is to allow for expeditious communication between experts on pertinent issues to enable rapid advancements in these areas. A forum such as this is a novel way to approach time-sensitive issues such as the swine flu epidemic by presenting a setting in which scientists can share their data, research, and theories without the high level of peer review required by scholarly journals.

Anyone can submit to the PLoS currents colloquium. Publication and review of submissions are done via the Google Knol software. First, an author assembles their content at the Google Knol site. There it can be shared with other users at the author's discretion. Once ready, the author submits the content to PLoS currents, and it is reviewed by the moderators. This panel of experts in the field decide whether or not a given submission is "intelligent, relevant, and suitable." All accepted submissions are made available on the PLoS Currents: Influenza website, and publicly archived at the National Center for Biotechnology Information.

One key to achieving the goals put forth by the creators of PLoS currents is the idea of openness. All content on PLoS currents is openly accessible, and is made available under a Creative Commons attribution license. Another key is speed. Work submitted to PLoS currents is quickly screened (though the exact time frame is not specified) by a panel of expert moderators, but is not subjected to the intense level of peer review required for acceptance by a scholarly journal. The contents of PLoS currents are not considered to be formal publications, but are citable. Every accepted PLoS currents submission receives a permanent identifier that can be linked to and cited. And, as mentioned above, submissions are archived at the National Center for Biotechnology Information. Since acceptance by PLoS currents is not considered formal publication, the same content can be published in a peer-reviewed journal.

The idea of PLoS currents is interesting because it takes to heart the ideal of open science , and attempts to exploit it in pursuit of something that we mentioned in class last week, namely the greater good. Indeed, the greater good seems to be what the creators of PLoS currents had in mind when founding the project. In response to the FAQ "Why should researchers submit content to PLoS Currents: Influenza?" the creators write that "researchers submitting results to PLoS Currents: Influenza will also be contributing to a global effort to respond to the recent emergence of H1N1 influenza and to the urgent and ever-present public health threat posed by influenza." The greater good also seems to be the chief driving force behind submission. Since publication on PLoS currents is not considered formal publication, it cannot count toward tenure. And acceptance by PLoS currents is no guarantee of acceptance by any of the journals published by PLoS. The only advantage (apart from benefitting the greater good) of publishing on PLoS currents that I can discern is that one's work will be made available to a larger audience, and that its preservation will be ensured due to being archived at the National Center for Biotechnology Information. Of course, the same benefits can be achieved through publishing, so I suppose that only those authors who, for want of acceptance, have not published in an academic journal would find these benefits attractive.

Shifting roles, shifting space

Libraries of the Future
September 24, 2009

In light of our discussion last week about similarities and differences between traditional physical libraries and digital libraries, I found “Libraries of the Future” from September 23, 2009 to be an interesting read. The article summarizes statements made by Daniel Greenstein, vice provost for academic planning and programs at the University of California System, to a group of academic librarians meeting to discuss sustainable scholarship. Steve Kolowich, the author of the article, states that Greenstein’s vision is one in which, “the university library of the future will be sparsely staffed, highly decentralized, and have a physical plant consisting of little more than special collections and study areas.” In effect, the digital repository would become the defining entity for library, while storage and maintenance of physical materials would be shared between institutions, and if possible, outsourced.

The primary motivation for this line of thinking seems to be economic, with budget shifts requiring university libraries to restructure and reallocate resources. Greenstein points to a decrease in individual libraries staff, collections, and services, but an increase in collaboration with other repositories. The goal would be for libraries to become drastically different over the next decade, which some librarians present at the talk, as well as those who posted comments, thought was an unreasonable amount of time for such a shift in the library world.

Most comments made seemed primarily concerned with the changing role of the librarian who, perhaps in the eyes of Greenstein, has just lost their job. I think the questions that arise are about what the roles of librarians might be in a system that outsources most of the services they performed. Might librarians move to other positions in which the take on the roles within the companies performing the tasks needed? Who remains a part of the small staff at the university library? Is it really cheaper to outsource cataloging, conservation, etc?

I am not sure if the answer to the recent economic problems is to dismantle the system and depend on outsourcing to provide services. I think the changes that will happen to the university library system will involve many shifts in the various roles and potential employment opportunities, of librarians. Still, I find the physical shift suggested unsettling. Where do all the books go? How does the space shift?

Disappearing Digital Archives

Digital Archives That Disappear” (April 22, 2009) is a reaction article to the sudden disappearance of the online newspaper archive “Paper of Record.” Began in 1999, Paper of Record is a digital collection of newspapers (consisting of 21 million images), with a large portion of the collection consisting of Mexican newspapers.


Paper of Record was secretly bought from its former owner in 2006, however, this purchase was not publicly revealed until 2008, when the archive suddenly disappeared. Amidst an outcry from historians who relied on the materials available through paperofrecord.com, Google agreed to let the site resume services under the previous owner, however, its future is still uncertain. For example, there is concern over the possibility of Google charging prohibitive institutional subscription fees for access.


This case raises many interesting questions about the future of digital archives and scholarly research. First, in the case of paperofrecord.com, the Mexican government as well as Mexican universities paid for many of the Mexican newspapers to be digitized. This has upset many members of Mexico’s historian community. Finally, this raises questions about the continued accessibility of online archival collections.

But no really, be practical, HOW will we do all that?

This week I also was unsure what to blog about; I considered my most frequently scribbled marginalia. The question is quite simply, how? It may be a mundane question, but I am a person who can't plan without understanding it.

Here's one example from this week: "Data at both the individual and firm level, as well as national and international levels, must be integrated. Complex qualitative and quantitative data from a variety of modes must be combined." (p. 20 of http://ucdata.berkeley.edu/pubs/CyberInfrastructure_FINAL.pdf)

While I am aware that technology changes quickly, I thought it might help to find some examples of how current digital library projects are structured. I began to hunt for articles about social science databases. Megan's comment about persistent identifiers was timely; 404 not found and I got to know each other quite well.

Finally, I found an article about the Exploratorium, an online science museum. http://saturn.exploratorium.edu/partner/nsdl/pubs.html
The article, From Playful Exhibits to LOM: Lessons from Building an Exploratorium Digital Library, can be found here. The authors describe the metadata standards adopted to create the interactive Exploratorium.

Specifically, they designed their metadata for use with "Learning Objects." They knew in advance what their mission was, to provide educational material, so they knew the metadata should be structured to support that. The standard they chose (IEEE LOM) also maps to Dublin Core, so I think it is extensible in the way that Berman and Brady were suggesting from my earlier quote.

They assessed what they had to put in the collection before starting, and then used an assets management solution called Canto, which supports scalability and APIs.

Should they have tried to combine more projects and possibly be more extensible? Do they have enough tools, enough scope, or enough added information to be classified as a curation project? I'm not sure. I know I found their website much easier to understand than some others. An image of one view of their structure and metadata is below.

The question of "how?" deserves more consideration, but I believe this site is a good example of knowing what you want in the collection, setting goals and standards, and adapting when necessary. The authors acknowledge that their use of URLs as stable locations may need to change, but currently they do not have a better option.

(I did not receive a single 404 error while exploring their site.)

Who's Afraid of Google Scholar?

Jasco, Peter. "Newswire Analysis: Google Scholar's Ghost Authors, Lost Authors, and Other Problems." Library Journal (9/24/2009), available at http://www.libraryjournal.com/article/CA6698580.html (accessed 9/28/2009).

The latest critique of Google's metadata comes from Peter Jasco writing on Google Scholar for Library Journal. In his analysis, Jasco illustrates how ill-trained parsers and web crawlers are responsible for creating a variety of metadata errors (particularly as pertains to authorship) which in turn produce incorrect publication and citation counts.

Jasco begins by noting that whereas Google Books blamed the embarrassing metadata errors exposed by Geoffrey Nunberg on wonky data from libraries and publishers, they do not have this excuse for errors in Google Scholar since Google chose not to use the metadata offered them by publishers and instead rely on its crawlers and parsers. The result, Jasco asserts, is that Google's algorithms replace real authors with ersatz ones like "P Login" (here Google has misconstrued a login prompt on a publisher's page, "Please Login," as an article's author). Alternately, Google Scholar also inflates publication and citation rates by counting citations as records in author searches and by generating multiple master records for the same article. Jasco notes that a search for his own name as author (e.g. author:jasco) yielded 578 records, but 403 (or 79%) of these records were for articles that cite his papers, not articles he authored. Such problems are compounded when Google Scholar's figures are naively adopted by tools such as the Google Scholar Citation Count gadget or Publish or Perish (PoP) software.

Jasco's figure depicting the now-corrected "Password" as author problem.


I'm torn as to how to respond to Jasco's article. On the one hand, it is clear that Google Scholar's metadata is error-riddled. One can't help but wince at some of Jasco's examples: the Google parser identified as author names things such as headers, section titles, and publication information, including Methods (42,700 author records) Contents (25,200), Limited (234,000) and Ltd (452,000). On the other hand, any university or department that's willing to make tenure decisions based on a cursory search of Google Scholar has such deep-rooted problems that inaccuracies on the part of Google Scholar can only be the least of these. It is to be expected that an unwitting high school or college student might be led astray by mis-attributed authorship or incorrect publication dates in Google Books, but it's less understandable for professionals in the field to outsource their committee work to Google Scholar or (even more perplexing) for an individual to assume that Google Scholar knows more about how many articles s/he has published than s/he does. It is part of being a professional in the field to keep an up-to-date CV and be cognizant of the publications one has authored or co-authored. Citation indexes are admittedly a trickier issue since such information is not as easily compiled by individuals, however shouldn't we expect (demand?) that academics and administrators would vet any results Google Scholar yields?

Jasco is not optimistic about Google Scholar's future. He sees their refusal of publisher/indexer metadata in favor of their own under-trained crawlers and parsers as a "lethal mix of ignorance and arrogance." Though Google Scholar has corrected some widely publicized errors, Jasco asserts "[t]he parsers [themselves] have not improved much in the past five years despite much criticism."

In addition to the actively "malevolent" threats to cyberinfrasctructure posed by hackers, phishers, and other cybertransgressors that are discussed in the Workshop on Cyberinfrastructure for the Social and Behavioral Sciences Final Report, cyberinfrastructure is at risk from the perhaps even more damaging effects of incompetence, laziness, and insouciance--factors which undermine our confidence in cyberinfrastructures and the collections they support. There is plenty of blame to go around. Google needs to be more concerned about inaccuracies in its products. Perhaps the responsible thing would be for Google Scholar either to preface certain functionalities, such as citation counts, with clear "user beware" warnings or to suspend them until it can ensure that its records are more accurate (it would still yield record results for the user to vet and tabulate). However, it is also incumbent upon users to educate themselves better regarding the tools they are adopting for important tasks. After all, Google Scholar announces itself as "Beta" in its logo. Hiring, promotion, and funding decisions should not be based on numbers drawn from tools that are still very much works-in-progress simply because it's easier to get your numbers from Google Scholar than to crunch them yourselves.

Monday, September 28, 2009

NSF awards $7 Million to Texas Advanced Computing Center

Article: NSF Awards $7 Million To TACC for Remote Visualization and Data Analysis

I wasn’t sure what to blog about this week, so on a whim I went to Google news and searched for the words “science” and “cyberinfrastructure”; one of the articles I found was actually pretty exciting. The news broke today – according to the University of Texas public affairs news site, the National Science Foundation awarded the Texas Advanced Computing Center here at UT a 7 MILLION dollar grant to fund a three-year project that aims to "provide a new computing resource and the largest, most comprehensive suite of visualization and data analysis (VDA) services to the open science community."

According to the article, this new computing resource, unsurprisingly named "Longhorn", will help science communities both nationally and internationally. Longhorn will enable scientists to "interactively visualize and analyze datasets of near petabyte scale." I didn't know this before now, but a petabyte is apparently equal to one quadrillion bytes or 1,000 terabytes... which is quite a large number.

The article goes on to give many more technical system capabilities and components that Longhorn will have, which all sound really impressive, but I have no idea what any of it really means because I am not a deeply technical person. This project is being touted as “unprecedented” though, so the technology must be pretty spectacular. Check out the article for the specifics. To sum up the important part that I got out of this article: basically, this three-year project, which is slated to “enter full production” on January 4, 2010, will “provide the framework for leading researchers to collaborate and analyze their terascale datasets while still offering low barriers to entry and usability.” The 7 million dollar grant aims to "stimulate and support" new research in the area of visualization and data analysis. The grant will expire July 2012.

John Mullen, the vice president and general manager of Education, State and Local Government at Dell, praises this “long-term strategic investment” and says that Longhorn’s new advancements in technologies (particularly in CPU, GPU and networking technologies) will allow scientists to take on new, complex challenges. What TACC is developing could also have an impact on other fields. Kelly Gaither, who is the principal investigator and director of data and information analysis at TACC, says that Longhorn’s developments in interactive visualization and data analysis can help solve problems not only in science but also in engineering, medicine, and in the areas of national security and safety. The humanities aren't mentioned in this article, but I imagine it and other similar fields could probably benefit as well.

Basically the entire article mentions points we keep discussing in class: science and technology are both evolving rapidly; the internet and high performance computing are creating more and more data than people can handle at once; new technologies must be created to handle such large datasets and help link together those who are working on complex projects; and, finally, the fact that strategic investing is important in building parts of a cyberinfrastructure.

This seems like an exciting opportunity for the university, and I’m interested in finding out more about the project. As I mentioned, this was the first article that I found on the subject, and it came out only this morning. UT's much quoted motto is "what starts here changes the world"... we will have to see during the next three years if Longhorn will live up to this lofty goal.

Wednesday, September 23, 2009

Visible Archive Series Browser

Two days ago Vimeo demonstrated the "Series Browser": a visualisation of approximately 65,000 series from the National Archives of Australia's collection. Mitchell Whitelaw created this as a prototype and part of his Visible Archive - "a research project on the interactive visualisation of archival datasets."

I am very excited about the prospects these kinds of visualization technologies can have on digital scholarship, particularly in the areas of the arts and history. The browser is interactive, with visualized archival series (represented by squares) arranged both chronologically and in terms of relations to other series. When one scrolls over a series a bubble appears with general scope and content information. The size of the series is illustrated by how big the square is. An inner square represents the volume and the outer square represents the physical shelf space that the series occupies. By selecting and zooming in on a series - lines demonstrating relationships with other series appear (different kind of link relationships are categorized by different colors).

One can also browse and research by agencies outside of series to gain broader context. All this is handled through interaction with the visualized data.

Three hours ago Mitchell launched a video demo describing the A1 Explorer, a similarly interactive visualization of the A1 series in the National Archives of Australia. In this we are looking at what appears to be a kind of tag cloud - as a word-frequency visualization of the contents based upon the item titles. When you hover over one of the words, lines linking it to related words appear. The thickness of the connecting lines is also an additional layer of information. Selecting a text item displays a list of items belonging to that category.

Another aspect of this example is the histogram which displays visually the number of items appearing in the collection per year. This demonstrate the distribution over time but also works as another port of entry into exploring the archive. The text cloud and histogram are also linked in that hovering over a term will display how many of such items appear in the collection per year.

An additionally cool functionality of Mitchell's visualization is that if there is a term that is dominating the text cloud - if you select and click on it you can remove it from your view and the cloud will adjust accordingly, allowing you to see and explore the remaining items in the collection. If you select an item from the list it will load a digital image of that item from the archive.

I couldn't help but think about how cool this application would be for navigating my links in delicious.com, but really - the applicability of this for all manner of archives is immense. I hope to see this technology catch on!

Not To Share

For this week I am examining the findings reported in a paper by Caroline Savage and Andrew Vickers about data sharing published on PLos. The authors of the study devised an experiment to request that researchers whom have published a paper on PLos carry through with the data requirements stipulated by the publishing agreement and actually make the data referenced in their study available.

It should probably come as no surprise that out of the ten researchers contacted only one was willing or able to actually produce the dataset, and then only after talking with the requester about what would be done with the data.

This raises so many of the key issues concerning scientific data sharing that it could almost act as a summary of why instituting a data sharing methodology has been so problematic. Some researchers were “too Busy” to fulfill the data requests, even though “we pre-specified that any request must not create undue work for the investigators”. Others claimed that they were “forbidden” to release the dataset. Some simply did not respond to the request at all.

The reasons given for not sharing data are easy to sympathize with, gathering a data set, annotating it so that others will find it useful takes an investment in time. It is difficult to persuade anyone to take a few hours out of their busy day to give you something they feel attached to, when they have no idea of your intentions or how important the data is or even who you are. It is illuminating that the one researcher who complied with the request did so after an interview process; almost like a vetting procedure to see if the requesters intentions were true.

It must be pointed out that however reasonable the denials are they come from a group of people whom have published in a journal which explicitly states the requirement that datasets need to be accessible. This active participation in the journal should nullify any claims of being “too busy”. It does not, and this demonstrates one of the key faults with data sharing as it is: journals lack the teeth and methodology to ensure that datasets be made available.

The main reason for this is that journals are not and should not be concerned with enforcement, journals should not be the police ensuring that everyone play nice with data. The best a journal could do is to deny further publication. The real and effective enforcement should come from the community in general. There needs to be a ground up call for data sharing. This sort of change in the way science works will only take place slowly and if there is a proper system in place to reward data sharing.

A Library Where No One Came

Issue 461 of Nature magazine is devoted to issues surrounding open access to data, with a feature on what happens when researchers don’t take advantage of data repositories and keep their data to themselves. This article is a good reminder that despite all the hard work that information professionals often put into building tools, it is all for naught if the tools aren’t used. It is important here to distinguish sharing of data from sharing of publications in online open access. While the latter is becoming much more common to the point of being the norm in some fields, the former lags behind.


In this article Bryn Nelson explains that different groups of researchers have differing cultures that affect their views on sharing data. Some disciplines may have rigidly structured data, such as genome sequences, while others may have data that doesn’t fit any predefined schema, such as climate data. In addition, researchers are often concerned about misuse of their data or lack of attribution and losing the original rights they have to the data. Some data can’t be understood without looking at a whole other group of related data: an example given in the article is a group who is trying to track down weak gravitational fields and must gather a huge amount of background data such as ocean currents and atmospheric conditions to use as filters on their data. Releasing such primary data to the public runs the risk of people taking it out of context and thinking they have found spurious results.


Nelson suggests that three authorities in these fields exercise their control to help corral the data: publishers grant funders, and scientific societies. Grant agencies can make it a requirement that the projects they fund will have their data collected in a repository, and publishers can ask that authors share their data, although they must be careful not to scare away authors who would turn elsewhere. In addition, Nelson says that large government efforts are required to build the kinds of data standards that will work across scientific communities: it’s not enough to think that ad hoc solutions will work in the long run.


On the one hand, the problems of the researchers are all ones we have seen before: they worry about losing the rights to their work, misattribution, and people drawing unwarranted conclusions from the data. While this is nothing new and can occur with any work put out into the public, there is an expectation of objectivity to data that makes releasing it especially dangerous. The traditional gatekeepers of journals should be able to hold at bay the worst examples of pseudoscience that draw on public data, but perhaps it would be beneficial to build in dependencies to data streams that link data that can’t be separated. Data could be provided as a package, not independent streams.


Nelson goes against much of the current thinking in online collaboration when he states that a government group must facilitate data sharing. This approach is anathema to the rabid supporters of open source software and the like who hold it as doctrine that dividing work up among a group will provide better results than top-down control. I believe Nelson is right however, because data standards, no matter if they are eventually superseded by better standards, must have some sort of common ground to start from in the beginning, and this relies upon the work of an authority. Competing standards could actually serve to slow down progress, as work must be done to translate between standards and generate tools to do this.

Listening In

This ACM article is a write-up of a study conducted on the social aspects of limited music sharing via I-tunes technology that manages to circumvent the complicated legal landscape of ip rights management and music sharing. (The music can only be shared via subnet and the file is streamed instead of actually copied.) A link to the full article can be found here, though you probably have to be on UT's campus to use it: http://portal.acm.org/citation.cfm?id=1054972.1054999&coll=ACM&dl=ACM&CFID=53223505&CFTOKEN=15055051.

What intrigued me about this article is that it describes both practices of popular curation and shared curation of digital objects. In terms of popular curation, I-tunes technology allows users to share their playlists with co-workers. However, you can choose which specific playlists you want to share, and you retain control over the naming of the list as well as the metadata tags of the individual songs (i-tunes retrieves song info via the net but the user has final authority to change the info on his or her individual machine.) What is really interesting about this, is that the study found that people used i-tunes' sharing features as a way of projecting identity, and thus the choice of which playlists to share (exhibit?) became tantamount. For instance one participant worried about some of the songs he purchased for his wife being spotted by co-workers:

I mean if people are looking at my playlist to get a picture of
the kind of music I like and don’t like, you know. Or to get a
little insight into what I’m about, it’d be kind of inaccurate
‘cuz there’s, you know, there’s Justin Timberlake and
there’s another couple of artists on here that…Michael
McDonald, you know. Some of this stuff I would not, you
know, want to be like kind of associated with it.…I guess
part of it is it wouldn’t be bad if, you know, people thought I
was kind of hip and current with my music instead of like an
old fuddy duddy with music. I mean I sort of like to
experiment a little bit with stuff. I mean I’m not like totally
wild but I like to experiment with, you know, some newer
stuff. So I guess it would be okay if people thought that I
had good taste. It wouldn’t be so good if they said “God! He
likes Justin Timberlake? That sucks!” (P1).

The whole concept reminded me of our discussion of the popular curation project at the Brooklyn museum but on a much more informal and individualistic scale. The fact that informal, individualistic curation regularly occurs also highlights the importance of creating infrastructure in the terms of knowledge, skill sets, and education on digital technology. Or perhaps, such knowledge is already ingrained but not formally articulated? I'm not sure.

However, the other thing that struck me here is that the curation responsibility is shared by i-tunes. As i mentioned earlier, the metadata tags associated with songs are initially included with their purchase or are downloaded via song recognition software if being imported from cd. Also, the default for sharing is to share nothing, so i-tunes helps enforce privacy. Finally, the availability of certain fields only as tags in someway informs the curation through software constraints.

All in all, though, I think popular curation as an expression of identity is a trend that will grow to encompass ever more fields in the future.

Data Sharing

Continuing on last week’s discussion about data, I read a blog post by Cameron Neylon about the reluctance of authors to share their data. The post was in response to a paper titled Empirical Study of Data Sharing by Authors Publishing in PLoS Journals, published last week in PloS ONE.

The paper details a study in which 10 authors of PLoS articles were asked to share their data for use in a Master educational project. PLoS authors were chosen because of the journal’s policy that “publication is conditional upon the agreement of authors to make freely available any materials and information associated with their publication that are reasonably requested by others for the purpose of academic, non-commercial research.”

It stands to reason then, that the authors of articles published by PLoS would be willing to share their data. That was not the case. Only 1 of the 10 authors provided their data. Although the study was small, the results suggest a greater problem. Why are authors so reluctant to share their data even when the publication requires it?

Neylon suggests that the findings of this study are not new, and the reasons the authors gave for not sharing their data are not new either. Some authors did not reply, others were unable to be contacted, some said it was “too hard” or their institution “forbids it.” She goes on to suggest that none of these excuses are acceptable.

Neylon insists that it is time for the publications to step up and take responsibility for ensuring their data sharing policies are followed. She believes that the most effective way to go about it is to pull the papers of authors who refuse to share their data. By making an example of a few authors, the data sharing problems for a particular journal will quickly become a thing of the past as authors choose to either share their information or publish elsewhere. Moreover, the publication will benefit from being seen as a source of quality data, further differentiating it from other journals in the field.

And for those who say sharing their data is “too hard” Neylon states that if an author cannot maintain their data so that it can be used again, they shouldn’t be publishing anyway. Neylon does allow that there are occasions when data must be kept confidential, but she argues that these should be the exceptions- exceptions made only after careful consideration of the individual circumstances.

In theory, requiring publications to follow up on the data sharing practices of their authors sounds like a good idea. But I can’t help but wonder if we are putting too much weight on the publishing industry. Is it really their job to make sure their authors share their data? It seems that the authors should step up and make data sharing happen, but it is also clear that they are not currently doing it. What kinds of fundamental changes need to occur to ensure data sharing?

The Virtual Observatory and the Roman de la Rose: Unexpected Relationships and the Collaborative Imperative | Academic Commons

The Virtual Observatory and the Roman de la Rose: Unexpected Relationships and the Collaborative Imperative | Academic Commons

I found this article in the December 2007 Academic Commons issue. It intrigued me because it took two digital projects, seemingly as opposite on the academia spectrum as they possibly could be and drew some strong parallels to one another, and by doing so, was able to pinpoint cyberinfrastructure within a humanist point of view. This was very helpful me, as an old fashioned English major used to stuffy texts and houndstooth jackets, to be able to apply the idea of cyberinfrastructure into the study of literature and other humanities areas.

The article, by Sayeed Choudhury and Timothy Stinson, both from the John Hopkins University, starts off by picking two digital libraries: The Virtual Observatory (VO) and the Roman De La Rose Digital Library. The VO is a typical area, structure, application that one would expect from a digital project: large complex datasets shared, visualized and analyzed by a group of astronomers. The Roman de la Rose Digital Library is a collection of digitized medieval manuscripts chronicling the many different editions, write-ups, illuminations, citations and so on and so forth of the Roman de la Rose.

The traditional view is that these two areas of academic discipline have nothing in common: one is numbers and the other is a qualitative analysis of a medieval work of literature, but Choudhury and Stinston draw some strong parallels in order to establish the shape that cyberinfrastructure can take place in the humanities by suggesting that we expand our defintion of the term "data-rich." According to Choudhury and Stinson, "data-rich" does not have to refer to technology, but rather the ease and difficulty that the practitioners are able to generate, acquire, and process data. For example, in 1627, the Roman de la Rose, had been written, illuminated and studied many times over and had a wealth of data that could be transmitted and used to create new data, and this was way before computers.

Then article goes on to talk about the data deluge: it is not just happening in the more quantitative disciplines, but also in the humanities because manuscripts, being "data-rich," still retain their original values and meanings, but also acquire new interpretations and values with each new generation of academia. This also includes metadata generated. Therefore, the humanities field has just as much presence and need for a cyberinfrastructre to keep it all from losing the progress of information through the ages.

In a sense, the scientific datasets are a modern equivalent of medieval manuscripts. Choudhury and Stinston point out that the many different versions of the Roman de la Rose that build off of one another mirror the intertextuality of World Wide Web and digital media provides an opportunitu reflect more accurately on forms of medieval textuality and transmission of information that may have been lost during the print era.

I felt that Stinston and Choudhury's article heavily drew on ideas and theories of new forms of scholarship and how digital curation can make that possible. It also exposes the need for humanities computing and gives it an adequate place within the information spectrum of the humanities discipline by doing something inherently "liberal arts," comparing two distinct opposites and drawing parallels to show that its not all that different!

Online archive opens access to research

I was interested in this article in response to our class discussion regarding faculty submitting their work into an institutional repository, and the restrictions that must be placed on the work (imposed by the publishing entity). It was also here that I learned of the September 14 announcement of the "Compact for Open Access Publishing Equity" initiated by MIT, Cornell, Dartmouth, Harvard, and UC Berkeley. The Compact pledges to develop systems to pay for open access journals for the articles published by the institutions' scholars.

The University of Western Ontario has launched a new initiative, Scholarship@Western, meant to provide open access to articles and other scholarly materials produced by faculty and students. Like many institutions world-wide, Canadian funding agencies are beginning to require open access to published findings. For example, the Canadian Institute of Health Research will withdraw funding if open access to peer-reviewed articles is not provided within six months of publication. This is not only an effort to distribute knowledge, but also to promote transparency and public accountability. Launched ten months ago, over 460 papers have been added to Scholarship@Western.

Individuals create a homepage using a Western template which may include a profile, c.v., and contact information. The page may provide abstracts and citations, links to published and non-published articles, contributions, books, conference proceedings, presentations, etc., and PDFs may be uploaded. Site visitors are invited to join a mailing list to receive notifications of additions to the page. All information is maintained by the individual, and they are "encouraged" to review journals' copyright agreements before posting reprints to the repository. If the individual leaves Western, the page is not removed, but the template will be reverted to a generic version - it will continue to be accessible.

Scholarship@Western is intended to showcase Western achievements, increase the researcher and university profile, connect researchers and create collaborations, prevent duplicate research efforts, and advance new research developments. Additionally, it serves to collect, disseminate, archive, and preserve materials created or sponsored by Western. It is an institutional repository that provides a simple user interface allowing browsing and searching, and provides access to the institutions scholarly output via a single portal.

Western is admittedly behind the curve with providing open access to materials and does not currently require faculty to submit materials to an open access repository. Scholarship@Western was created as a response to the growing need for open access. The Compact for Open Access Publishing Equity is yet another response to this need, and has provided a framework for institutions to support open access journals as an alternative to traditional subscription journals. The agreement commits institutions to “the timely establishment of durable mechanisms for underwriting reasonable publication charges for articles written by its faculty and published in fee-based open-access journals and for which other institutions would not be expected to provide funds.”

Open access is increasingly necessary and increasingly prohibited by the traditional subscription journal system. Institutional repositories are one method of preservation of the achievements of an institution and it's researchers, but do not adequately provide open access. Wide-spread support of the Compact, in addition to open access mandates by funding and other sponsoring agencies, should serve to create and encourage a new environment for open access scholarly communication and publishing.

Tuesday, September 22, 2009

Flickr Commons and the Smithsonian

The situation covered in this article was pretty intriguing: based on the Library of Congress' success in putting some of their photographs on Flickr Commons, the Smithsonian decides to do the same. The authors of the article were at the head of the project, and here they recount the success of their efforts.

They provide some background information on their own digitization efforts. Digitization itself is headed by the different museums and research centers comprising the Smithsonian. The Smithsonian Photographic Services handles the requests for the majority of physical photographs, but the authors state they aspire to virtually collect all the photographs dispersed through various parts of the Smithsonian in a single digital repository under a comprehensive, unified metadata scheme. The idea is that such a repository will dramatically enhance the use of their holdings.

The Flickr Commons project was seen as a first step in this direction, in part because it would help indicate how the Smithsonian could present a single face to the public in spite of its highly distinctive, individuated collecting institutions (presenting a unified, impressive front to the public seems to be a recurring perceived benefit of repositories), and in part because the Smithsonian would begin to learn how to normalize its metadata when it began pulling images from its various departments.

Without quoting their statistics, the Smithsonian finds the project a success. Traffic is thorough and is boosted with each new submission of photographs. The institute is researching ways to employ the tags and comments into their own site's search functions and into the object catalog records.

An approach like this to digital curation partly hinges on the nature of the digital objects, which I would describe as humanities objects. This is a simplification (their photographs are data in the technical lineage of photographs, and many depict social and natural phenomenon such that they're "data" too, among other exceptions), but the subject matter lends itself well to the folksonomy Flickr provides. This is similar to the creative texts marked up in the IVANHOE project Jerome McGann discussed. In both cases the objects being described don't suffer from overlapping or contradictory descriptions; in fact it seems the only appropriate way to mark them. Considerably uncontrolled community engagement with the objects is seen as the best way to their digital aspect. Certainly scientific data can support layered metadata that describes different levels of a measurement, but these terms and layers are carefully circumscribed. In the case of more open or ambiguous digital objects markup approaches like IVANHOE and Flickr seem to fit well (although granted that IVANHOE does not appear to receive much traffic these days).

Interesting too is that the Smithsonian essentially gets to have it both ways with Flickr. They have applied machine tags to a group of photos, the Belize Larval Fish Group (it doesn't look like they're displayed to users though). These tags have a simple machine-readable format ("taxonomy:genus = Sphyraena") that other services can harvest. This achieves the kind of controlled specificity typical in scientific datasets, but with the same tool (Flickr).

What's challenging about digital humanities objects is that while traditional tools like XML or Flickr's machine tags will work well for storage and access, actual value-adding markup depends upon a different kind of tool that relies on a lot more human agency. I think these tools are still being developed as it's difficult to apprehend what's desired. Mass computation on Big Science datasets one could argue is a difference in degree; digital tools for humanities might be a difference in kind.

Journalism as Digital Curation

Who says newspapers are dead? Well, maybe they are financially, but that isn't stopping some papers from attempting to adapt their reporting procedures to the digital age. In an article for the online newsletter of the Poynter Institute, a journalism school in St. Petersburg, Florida, Patrick Thornton writes about a recent trend in newspaper journalism that upends the traditional reportage hierarchy. Normally, the major news blogs "report" by responding to stories that are published by the "mainstream media," and if the major newspapers have any actual presence in the blogosphere, it's usually in the form of an automated RSS feed that generates streams of headlines for Twitter or Google Reader or the like.

The Chicago Tribune, however, is upending the traditional reportage hierarchy by engaging in what is being referred to in the journalism world as a "curated approach" to their Twitter feed. The Tribune actually has two Twitter feeds: @ChicagoTribune, which is a standard RSS feed of headlines from the online edition of the newspaper, and @ColonelTribune which is something entirely different. The idea for ColonelTribune came about because headlines would often be cut-off due to Twitter's 140 character limit. The Tribune decided to create another feed which didn't simply parrot the headline but created something more succinct and Twitter-friendly. Instead of the headline, they might lift an interesting quote or statistic from the article. Then the Tribune started doing something really revolutionary with ColonelTribune - they started tweeting about blog posts that were relevant to news stories in the Chicago metropolitan area. They would still post articles from their own website, but they would also post articles from blogs, websites of local Chicago television stations, and even the website of the Chicago Sun-Times. Readers were also allowed to respond to tweets and reader tweets are included in the ColonelTribune stream. I've even seen tweets where ColonelTribune responds to reader tweets. The interest in the "curated" vs. "automated" models is indicated by the relative number of followers of the two feeds. While 19,520 people follow @ChicagoTribune, 584,779 follow @ColonelTribune.

Obviously, there are downsides to the curated model. While the automated model requires very little effort, the curated model has less of a return on investment because the Tribune has to pay someone to run the feed (that being the case - I have to say that tweeting about interesting things going on in Chicago sounds like the best job in the world!) But the implications of the curated model are much more interesting. The Tribune is essentially curating a daily history of relevant digital articles relating to the Chicago metropolitan area from a number of sources outside of their own newspaper. That seems like a project that is worth the cost. If only they could find a way to make it profitable. Any ideas?