Tuesday, October 20, 2009

Specifics in digital collection policies

Last year, in this post on the DCC blog, Chris Rusbridge wrote about an Australian national code that all universities signed regarding research data and primary sources. The code is quite specific, I was surprised that the term recommended for keeping data is only 5 years, however, as Rusbridge writes, since it is specific, it is stronger than no policy at all.

The policy also states where and how to keep the data (not just in the researcher's own computer) and outlines researcher responsibilities as well as repository requirements. We've talked about the importance of being specific both to the field that a digital curation project is targeted toward, and in policy language for data sharing. This strikes me as an important step in Australia to mandate data preservation, and to some extent sharing.

Rusbridge writes that he does not know of any other policies exactly like the Australian one, which mandates that, "Each institution must have a policy on the retention of materials and research data." Although it is commonly known that collection policies are important, but at the time of this post the author mentions only one other similar policy, from the UK, the UK Research Integrity Office's Code of Conduct and Policy.

The UK code is certainly a good idea, but in my view it loses ground to the Australian code on specifics. For example, while the Australian code specifies that institutions must hold the data as well as the researcher, the UK code vaguely writes, "This is a shared responsibility between researcher and the research organisation, but individual researchers should always ensure that primary material is available to be checked..." Of course they *should but sometimes they don't, as we have discussed.

Another example of the language waffling in the UK document, "In addition, it is always a good idea, even if it is not required, to seek advice from an institutional official." (if unsure about revealing a conflict of interest.) Compare this to the Australian code, "Researchers must foster and maintain a research environment of intellectual honesty and integrity, and scholarly and scientific rigour." followed by 7 bullet points including, "...cite awards, degrees conferred and research publications accurately, including the status of any publication, such as under review or in press."

Since the Rusbridge post is a year old, I retrieved some recent collection policies from institutions to look at. Successful projects (meaning projects with active links and not empty) often have their collection policies available, and the policies are specific but still flexible.

For example, the Internet Children's Digital Library Collection Policy (not research data but still a curation project I believe, and given as an example on NSF's website) states as a goal, "to create a collection of more than 10,000 books in at least 100 languages that is freely available to children, teachers, librarians, parents, and scholars throughout the world via the Internet;" This specifies how many and what kind of materials without limiting the project.

The UK Data Archive Collections policy provides criteria for selecting and acquiring collection sets as well, the summary of which is, "The UKDA acquires data for
four central purposes:
• archival preservation;
• secondary use and analysis for research;
• teaching and learning use;
• replication and validation of research."

I then wanted to look a little closer to home so I tried the DSpace instance at UT libraries. The introduction states only, "The Repository's purpose is to collect, record, provide access to, and archive the scholarly and research works of the University of Texas at Austin, as well as works that reflect the intellectual and service environment of the campus." The collection policy provides several guidelines on the structure and metadata, but each collection must write their own description of what materials should be included.
Here is the description from the UT Faculty/Researcher Works, "Peer-reviewed pre-print articles, published articles, technical reports, white papers, presentations, collections of digitized data, field notes, etc…"

Compare this to the California Digital Library collection development framework. The CDL is another federation of several institutions and collections, much as the UT library attempts to network many diverse collections. In their overall collection development document, they have dictated a balance between disciplines, and the policy also states that,
"Conventional collection development criteria are paramount and will be applied consistently across formats, including digital resources.
  • Establishing a coherent rationale for the acquisition of each resource, meeting faculty and student information needs
  • Providing orderly access and guidance to the digital resources, and integrating them into library service programs
  • Ensuring that the advantages of the digital resource are significant enough to justify its selection in digital format."
Although not a national code, the CDL code is here doing what Rusbridge praised the Australian Code for doing.

Any policy is better than none, but one that states specific goals without limiting future extensibility and, when necessary, mandates further collection policies from participating institutions, will probably mean a more stable project infrastructure.

Remediation & Early English Books Online

Kichuk, Diana. "Metamorphosis: Remediation in Early English Books Online (EEBO)." Literary and Linguistic Computing. (published June 18, 2007). Available at http://llc.oxfordjournals.org.ezproxy.lib.utexas.edu/cgi/content/full/fqm018v1 (Accessed October 13, 2009).

Out of morbid curiosity I had wanted to read an article written about a project that I had worked on, Early English Books Online (EEBO), and this article by Diana Kichuk explores the layers of remediation in the project's production of its digital archives of Early English Books. As Kichuk defines it, remediation is the "re-presentation of one medium in another" (2) as well as the "appropriation or re-purposing of old media in new media" (a definition she attributes to R. Grusin and J.D. Bolter's Remediation: Understanding New Media [2000]). EEBO provides an interesting case study because its associated projects involve several levels of remediation.

EEBO consists of an archive of images of over 100, 000 books printed in English between 1473-1700, and EEBO-Text Creation Partnership (EEBO-TCP)--the part of the project that I worked for--now offers hand-keyed searchable, reading texts for 25,000 of these titles. Access in both cases is through individual or institutional subscription. The project entails several levels of remediation. The majority of the images of the Early English books come from the Early English Books microfilm project, which was began in the mid-1930s to preserve copies of Britain's important early books, and the project continued in the years following the Second World War. Beginning in 1998, these microfilms were then digitized to produce EEBO's digital archive of images. The clear reading texts produced by EEBO-TCP are drawn from human transcriptions the digitized microfilm images. Accordingly, Kichuk asserts that EEBO is a surrogate of a surrogate, with remediation acting like a "distorting lens or opaque veil through which the scholar 'sees' the mediated Early English book" (6). Though Kichuk states that the collection is of "formidable scholarly valuable," her article emphasizes the need to recognize that that EEBO does not offer an exact copy of the original and that scholars must be better aware of the limitations entailed in remediation.

Kichuk sees ProQuest's decision to digitize the microfilm facsimiles rather than print copies as ultimately wise: digitizing print copies that it did not own would have been prohibitively expensive and slow. Moreover, the technological limitations present at the time the project was initiated would likely have meant that the original books would have to be dis-bound to be scanned, thus endangering the original artifacts.

There are, however, important sacrifices associated with this decision, resulting in content amputation and page distortion, among other things. The EEB microfilms generally do not include the endpapers or bindings of the books, and the pages were often cropped, resulting both in the loss of some marginalia and the misrepresentation of the physicality of the original book. The page curvature resulting from open book photography generates distortions in font appearance and darkens the gutters, affecting both legibility and accuracy of representation. Also, the resolution used by the EEB microfilms was too low for grayscale capture, so the images are bi-tonal black and white, which makes capturing any traces of color in the original works near impossible. Also, in the transition from microfilm to digital image, EEBO again sacrificed detail by further lowering its image resolution in order to ensure acceptable download transmission rates (the images are 440 PPI [pixels per inch], too low for a preservation quality facsimile).


(This figure, taken from the EEBO homepage (http://eebo.chadwyck.com/marketing/eebo_demo11_tcp.htm) offers a sample of what the EEBO interface and text images look like.)

As a digital codex, Kichuk notes that the project does not accurately reflect both the text and the physicality of the original--to achieve this would require applying the latest imaging technologies to the original books. It would be a different project carried out at a different time. In terms of improvements though, Kichuk would like EEBO (and the scholars who use it) to be more conscious of and clear about the loss remediation entails. She argues for the vendor guard to against claims of authenticity and identical-ness. She would also like to see EEBO include duplicate copies, which would be of use to bibliographers (this seems to me a slightly odd request given that EEBO fulfills so little of a bibliographer's needs).

Though I think the issues of remediation are fascinating, the questions that EEBO and Kichuk's article leave me with are ones about what it is we want from our digital surrogates. EEBO is a hugely important source of access to the content of rare early printed books that most users would never have the opportunity to access otherwise. EEBO-TCP adds further value by allowing readers to engage in full-text searches across over 25,000 titles. This said, it is not as useful a resource if you are interested in thinking about the book as object. Its images are not high-quality representations and it has only limited functionality to support this type of investigation. (For example, I had a friend who was interested in using EEBO to study investigate the significance of when blackletter was used as opposed to roman or italic fonts in early English books. Since EEBO-TCP doesn't specify font-type in its encoding and only records emphasis by noting the distinction between the normative font for the book and any "highlighted" font by using tags, there's simply no way of searching for this short of examining each text.) I would argue, however, that in many cases consulting EEBO could help someone in determining which books to examine in person.

Ultimately, reading about EEBO foregrounds important questions about how we should prioritize financial feasibility, technical limitations, functionality, and speed of collection development in generating digital libraries and large curated collections. Though EEBO is more about access than preservation, the issues it raises overlap at least partially with the discussion of essential elements in this week's reading (Harvey 16). Though Ross Harvey's context is somewhat different, his question: "Is the value tied to the way the material looks? (Would it be lost or significantly degraded if the material looked different?)" remains instructive in thinking about EEBO. In developing such projects, we must decide which features of the original are most important to capture and at what cost, while keeping in mind that remediation entails both loss and gain.

Monday, October 19, 2009

A Tale of Compromise

Jessica Branco Colati, Robin Dean, and Keith Maull, “Describing Digital Objects: A Tale of Compromise,” Cataloging & Classification Quarterly 47, no. 3 (4, 2009): 326-369.

This article reports the efforts Alliance Digital Repository (ADR), a consortial digital repository service for the Colorado Alliance of Research Libraries (Alliance), to create a standard descriptive metadata policy for their repository records. Members in the Alliance included twelve separate libraries at nine different academies and institutions, public and private. The goal was to support each library's community standards while also ensuring interoperability. This would facilitate a central repository.

Needless to say the scope for a project like this, which is charged with accommodating a dozen libraries each with their own practices and metadata (not to mention unique data both from within and without the library), expands at a rate equal to thought itself.

The ADR eventually pared down the libraries' metadata they had to consider to:
  • MARC
  • MARC XML
  • MODS
  • DC
  • Various extensions to these standard schemas
  • ProQuest's Digital Dissertations XML-based metadata schema
ADR chose Fedora as the main repository software for its flexibility, conformance to OAIS, and automatic versioning of data and metadata. This is a really useful consideration, since metadata is a digital object too that needs to have its own records of when it was made, changed, by whom, etc., creating an easy audit trail of what has happened in the repository. Of course, one could go on and on with this, pointing out that that metadata too ought to have some metadata associated with it. As a quick solution to this problem I would offer that the point at which all metadata is automatically generated is the point where you can stop adding metadata to metadata, since automatically generated metadata doesn't need to document its own automatic generation (right?).

As they note Fedora does not have a user-friendly interface for submission and searching built-in (like DSpace) so they chose Fez as a configurable front end, adding that other interfaces could be attached to Fedora as needed.

MODS was chosen as the "normalizing schema," the one all records would be converted to at a minimum no matter what metadata scheme they had upon ingestion. The report gives a really helpful list of reasons why, and I think it's interesting that #1 is the simple fact that Fez and MODS "are proven partners: The Fez system was already using MODS widely throughout the system as the primary means of describing objects." They give other factors: the right granularity, crosswalks to the popular data-sharing DC, and plain familiarity, but it's notable that the tool in this case (Fez) partly determined yet another tool.

ADR determined minimal metadata field requirements with DLF Aquifier, a Digital Library Federation initiative to help distributed library networks and content.

A very significant technical problem for ADR was validation of incoming records to make sure the minimum MODS metadata was there and was correct. They decided that XSD could both the valid structure for XML encoded MODS records and the correct data types as they had agreed upon. For instance a record would have not only the correct structure of metadata that a DTD might confirm, but also the correct data inside the metadata fields (for example the correct time, date, price, or URI format).

The problem derived from the fact that the packaged XSDs as used by Fez were not applicable for ADR's purposes, which meant the creation of new XSDs. Integration of new custom XSDs with Fez was poorly documented and ADR encountered functional bugs with Fez. A survey of Fez use revealed that no one else was trying to significantly modify Fez's XSD templates, and ADR's number of digital objects, each with unique metadata structures, was very large. Added to this was the need to crosswalk MODS to DC that required a whole new set XSDs for the DC metadata records.

ADR ended up authoring their own "document types" (MODS and DC XSD-validated metadata records) for use. That process is long and complex and ultimately did not provide the best treatment for heterogeneous objects (objects containing multiple genres of document types, like a web page) since Fez allowed such minimal description for a single document type. This forced them to store these objects as separate records.

What interests me with this report is not so much the specific technical problems and solutions they encountered (which are kind of torturous to read about) but that such huge obstacles can be created when the technical tools are insufficient to the task. On the one hand, MODS was easily selected and implemented for a number reasons, not least of which is that Fez worked smoothly with it. There are just a few paragraphs on it. On the other hand, entire collections of XSDs and the creation of a new conceptual entity (document types) is created largely because Fez has a poor capacity to deal with new or modified XSDs. Four pages detail this struggle, and it's unresolved at the conclusion of the article. I conclude that while the biggest struggles for digital curation and interoperability may be social, political, cultural and so on, the extent to which technical can grease the wheels is tremendous.

Saturday, October 17, 2009

Cooliris

Cooliris, is a 3-D image wall that allows you to visually browse images, whether they be in Google Image, YouTube, Flickr, Facebook, Picassa, images in one's harddrive and recent versions of the photo database program Adobe Lightroom.

I was sad to discover that I could not download this program to my own Mac Mini as it features Mac OX 10.4, Power PC, and apparently I cannot upgrade to the necessary Leopard version that supports Cooliris. Update: Previous versions (tho unsupported) are available.

You can share your 3-D image wall through a url, as well as bookmark and save it. It allows one to "jump freely" through Flickr image pools and sets - though not sure how that differs from the current way we maneuver from image pools to sets - does "jump freely" = "move fluidly?" Not sure.

It claims usability with hundreds of sites due to its being built around Media RSS format. It advertises its ability to be used to view numerous television and movie episodes on Hulu, etc. Its slideshow feature allows one to double click and launch slide shows that one can pause and rewind. It claims to have the fastest way online to search images with its' style of "zipping" through the 3-D image wall.

Again this makes me think back to the Visible Archive from Australia - and its archival metadata visualization. If there was a way to integrate these technologies and allow one to browse visualizations of data - the future could be very exciting.

Wednesday, October 14, 2009

Dude, Your in the NYT

Happy off week everybody, here is a light piece but one which should still be interesting to us.

Congratulations, now when people ask you about what exactly it is that you are doing in school you don't have to give them vocab heavy monologues about metadata, data migration and petabytes, you can point them to this article in the New York Times.

The heart of this article is what we have been discussing through much of the semester, the sciences are going through a major change where the shear amount of data that is being gathered threatens to overwhelm researchers who don't have the training to deal with such datasets. The article goes then discusses how current students in the sciences are unable to access the vast scope of the information sets which are a part of their field s due to the limited nature of the computational facilities available to students. This causes the young minds to “imprint on these small systems, that becomes their frame of reference and what they’re always thinking about...”. I have doubts on the validity of this but still Google and IBM are going to solve these problems for scientist by providing them with access to better computers.

There is of course the unstated belief that shadows this article that all of these issues are inherently technical, and the application of computational power will solve the petabyte problem. Indeed the article quotes on researcher saying “Science these days has basically turned into a data-management problem”. A brief Google search does find that the scientist quoted here is a biology/ genetics researcher so in that field this might be a perfectly valid point, but I wonder if all scientist feel this way. And if they do why, and if they don't why not?

The question that I suppose I have from this article is where do we fit in? Do we work with the scientist to help them understand their data or are we only concerned with where the data is kept, and how to ensure that it remains available for use now and in the future?
I don't have good answers I have been looking at these problems for the better part of a year now and I am still not sure how we as a profession fit with this area, or if we do.

google scholar. not there yet.


Given how much class time has been devoted to discussing the many aspects of Google, I thought I would find this week’s article using Google Scholar. A quick search for the key term “data sharing” (all articles) revealed a top hit for a 1994 article titled “The Bayou architecture: Support for data sharing among mobile users” (http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=512726). Cited by 329 other sources, this article is a good 50 citations ahead of its closest competitor. Clicking through to the list of articles that cited the Bayou article in order of the article’s citation instances might be more useful if there were options as to how this list were arranged (say chronologically). Going back to the original search page and changing the search parameters to limit articles by date (say “since 2008”) still yields some 72,000 hits, however, in this function one can no longer see the returns based on number of citations.

In an article titled "Newswire Analysis: Google Scholar’s Ghost Authors, Lost Authors, and Other Problems," which appeared in Library Journal (9/24/2009), the author criticizes Google Scholar for other reasons: namely their poor metadata standards, and how this will effect citation numbers. He notes that because Google Scholar relies on their own algorithms rather than utilizing the metadata available through libraries. This, he argues “creates phantom authors for millions of papers. They derive false names from options listed on the search menu, such as P Login (for Please Login).”

It seems clear, that Google Scholar is a long way from having any sort of reliable measure of citation count, and also is not yet the powerful search tool it purports to be.

Blog Post About Nothing


So, I ran across this hypothetical Seinfeld scene mocking the twitter phenomenon, and couldn't resist turning it into a blog post. Basically, in the scene, Elaine and Jerry are giving George a hard time because he has become obsessed with tweeting and is engaging in said activity compulsively. George is staunchly defending his right to tweet (declaring "No one stops George Costanza from tweeting!") When Jerry points out that George, who has no job or girlfriend, does not in fact have anything of interest to tweet about, George emphatically asserts, "I got plenty to tweet about, baby!" After Kramer wanders in and mentions that he had a friend who tweeted himself to death, the scene ends with Jerry asking rhetorically whether typing in increments greater than 140 characters would break the internet.

The scene, though obviously farcical, illustrates several points relevant to the consideration of digital curation. First, it brings to mind the problem with noise potentially drowning out information of value. Digital publication is so cheap and easy that anyone, even a consummate loser like Costanza, can flood cyberspace with information about nothing. Because of the potential to lose relevant information in a sea of trivia, I think that it is extremely important that we emphasize the filter function involved in digital curation. For example, can any single scientist really humanly go through ALL data pertaining to a study? Particularly in fields like climatology it seems essential to perform some sort of data triage. (Many of those fields benefit from shared data, but usually also along the lines of collaborative efforts instead of straight data dumps.) It seems to me that digital curators could at least perform a cursory, primary triage of data.

A second point raised by the Seinfeld skit is that no matter how great and convenient we the information pros find new information technology, there will always be the likes of ludites, tenured liberal arts professors emeriti, and half-deranged hipsters who will hate, fear, or refuse to accept new technological methods. In expanding access via digital means, we must be careful not to limit access to other portions of the population, which may be more a problem of infrastructure than straight-up digital curation. At the very least, education programs in new i.t. methods seem necessary.

Finally, Jerry's joke about 141 characters breaking the internet drives home the fact that we really are dealing with new media. Twitter has changed the way people communicate by rigidly enforcing a strict brevity. This may change people's reading/learning habits. I'm not precisely suggesting we limit our descriptive elements to 140 characters per digital item, but it is something to keep in mind.

So, basically, the sketch brought to the surface of my mind some of the digital curation issues that have been spinning around in there such as filtering noise, adapting to new user habits, yada, yada, yada...

Curating Architectural 3D CAD Models

MacKenzie Smith's article presents MIT’s FACADE (Future-Proofing Architectural Computer-Aided DEsign) research project as an attempt to determine how to preserve and provide access to 3D CAD digital records. The digital curation project includes working with three architectural firms to determine their workflow, usage, and expectations about archival preservation of their digital design records. Traditionally, architectural archives have included records about the design of architectural projects, as well as firm records that support the creation of buildings. Numerous documents are generated through the process of architectural design and these records are now largely digital born – email correspondence, digital photography, 2D CAD renderings, and 3D CAD models, to name a few. Architects generally use proprietary software to create many of their design. Most of these files are easily migrated or converted to standard formats, such as PDF or JPEG. The problem arises when trying to preserve or access 3D architectural models, created using proprietary software in non-standard formats, which cannot be easily migrated or converted.

This presents a major problem for libraries, archives, and museums whose goals would include providing access to the digital records, but would not have the funding to support the ownership and maintenance of proprietary software systems, especially when the software used to create the documents rapidly changes. The FACADE project looked at strategies for exporting files and creating standard formats. This required expertise in both the CAD software and the underlying data, which many librarians or archivists at repositories would not have. The hope is that this process will become more automated over time.

One of the primary issues is how to connect the various records created within context of the design process. The 3D model is most valuable when viewed in relation to other drawings, photographs, and project correspondence. To this end, they have established a “Project Information Model” which will collect 3D model data and connect it to other project records. They have created a prototype for open source software that will assist curators in creating metadata.

The FACADE project web site includes further information, including a final report on the project. The final report outlines four versions in different formats that will be used to preserve the 3D CAD models.

1. Original (the originally submitted version of the CAD model)
2. Display (an easily viewable format to present to users, normally 3D PDF)
3. Standard (full representation in preservable standard format, normally IFC or STEP)
4. Dessicated (simple geometry in a preservable standard format, normally IGES)

This project was funded by a two-year grant from IMLS. The project predominantly set out to conduct research on the best methods and practices for curating 3D CAD models. There are still many research questions to be addressed, such as those pertaining to privacy and copyright of the data.

More on Google Books

Sergey Brin wrote an opinion article late last week for the New York Times titled "A Library to Last Forever." The same article appeared a day later as a blog post titled A Tale of 10,000,000 Books on the official Google blog. The post is essentially a defense of Google Books and the recent settlement they have reached with the Authors Guild and the Association of American Publishers.
In the post, Brin talks about a "black hole" of information printed after 1923. These books are not in the public domain, but are frequently out of print and increasingly difficult to find. Brin argues that Google Books is making important parts of the "world's collective knowledge and cultural heritage" available. I don't think any of us would argue the potential for the positive impact of Google Books on the way research is conducted. I am more concerned about what Google Books means for future mass digitization projects.
Brin states that "If Google Books is successful, others will follow." But what about the large startup costs? To Google's credit, they are the only ones with the means and the drive to even attempt such a large-scale book digitization project. But with the expense being so high, will there ever be anyone else with both the means and drive to create another digital collection? It seems to me that it would be highly unlikely, especially if Google Books is successful. Who wants to take on a giant when the cost of entrance is so high?
While I don't pretend to know everything about the settlement, it appears that it will create a registry of rights for books which helps to identify the rights holders for individual works and to make obtaining permission and assigning appropriate revenue to the rights holders easier. Again, such a registry would be very helpful, but how many entities exist with the kind of funding it would take to pay the appropriate fees and digitize the books?
And if Google Books is the only entity to attempt such a large-scale book digitization project, we find ourselves back where we started with the problems of inaccurate and misleading metadata. Brin addresses these issues briefly, only to say that Google is "working hard to address them."
In spite of all of my doubts, it is hard not to get swept up by Brin's call to save the literary works of the 20th century. I suppose we'll have to wait and see how Google Books affects the future of mass digitization projects. Meanwhile, Google would like us believe that they have found the answer.

Visualize This.


For the past few weeks I've been interested in how large data sets are being "visualized". Not only is this a fun (and easy) way to digest large issues, but the data must come from somewhere.
Some blogs I've been looking at lately include Information Aesthetics, Information is Beautiful, and Ben Fry's Projects, especially a visualization of changes in editions of Darwin's Origin of the Species.

I explored these sites to see where the authors are getting their data. Information is Beautiful provides links to his sources, which open as Google document spreadsheets. David McCandless, the author, also links to The Guardian's data (since the newspaper's DataBlog is also linked) and this also opens to spreadsheets. Information Aesthetics' Andrew Vande Moere doesn't provide his data sources, but does offer links to other data/information sites for hours of wasting/exploring. Ben Fry's The Preservation of Favoured Traces project states the "text for each edition was sourced from their careful transcription of Darwin's books" although more information about how the text was actually transcribed. The "Data" link on his site, however, is grayed out, so apparently not available. Fry does explain all his projects in detail, however, just maybe not all the information we would want.

An interesting article on the Visual Journalism blog talks about how some data is just boring and not everything needs to be graphically interpreted. The author, Gert K. Nielsen, is discussing a NY Times interactive visual about how people spend their day and mentions that the original data of the graphic, "American Time Use Survey, is in danger of being shut down due to serious cuts in the budget of the Bureau of Labor Statistics". That would hamper the continuation of this visualization, but the Bureau of Labor Statistics will surely have other great data to troll.

The popularity of information visualization will probably only increase (until the novelty runs out?). Ben Fry has written a book about it, Visualizing Data, published by O'Reilly. I wonder if the graphs and pretty colors will distract people from questioning the source of the data, or highlight the nature of statistics and data even more, creating a demand for source knowledge. I'm hoping the latter, especially with easy sources like data.gov, there is no reason not to provide the raw data. I also like the possibility of increased data sharing this might facilitate. Even if it starts out small, such as the GoogleDocs on Information is Beautiful, this creates a society of openness that can be emulated.

Galaxy Zoo 2

These past couple of weeks, I’ve been blogging about developing cyber initiatives and grants being given to fund projects that designed to generate and display data like never before. In fact, the words “revolutionary data collection” kept popping up again and again. But in most of the articles I’ve been reading about developing projects, no one ever really delves into what, exactly, will be done with the massive amount of data that will be collected; no one ever mentions how the data could be used – just that it will be utilized in “revolutionary” ways. With this in mind, this week I set out to find an example of something revolutionary – or at the very least something cool and interesting – that was being done with digital data.

My search led me to the zoo, the Galaxy Zoo, that is. Galaxy Zoo is a site that invites the public to help classify millions of galaxies using data and images from the Sloan Digital Sky Survey at the Apache Point Observatory in New Mexico. All of the images come from a digital camera which is mounted on a telescope.



I think this is really neat. People like you and me can help classify and study the universe, no prior knowledge of astronomy needed (or so claims Wikipedia). The original Galaxy Zoo first made its debut in 2007 “with a data set made up of a million galaxies imaged with the robotic telescope of the Sloan Digital Sky Survey”; the new version, Galaxy Zoo 2, launched back in February of this year. According to their website, within the first day of Galaxy Zoo’s launch the site was getting 70,000 classifications an hour; during the first year over 50 million classifications were received from nearly 150,000 people.

Galaxy Zoo 2 has focused on almost a quarter of a million of the “nearest, brightest and most beautiful galaxies” for users to view and classify. Apparently, new discoveries by the amateur astronomers at Galaxy Zoo are being made constantly about galaxy colors, shapes and patterns. Professional telescopes (and I’m assuming professional astronomers along with them) have even followed up on some of the Galaxy Zoo findings; the list of professionals includes “the Isaac Newton and William Herschel Telescopes on the island of La Palma in the Canaries, Gemini South in Chile, the WIYN telescope on Kitt Peak, Arizona, the IRAM radio telescope in Spain’s Sierra Nevada, the Swift and GALEX satellites, and the Hubble Space Telescope.”

I've been playing around with the website's classification tutorial page, and I have to say that it is pretty fun. This is a great example of the neat things that can be accomplished with crowd-sourcing and open data. Regular people with an interest in astronomy having access to data to classify galaxies, make discoveries, and help figure things out about the universe; Galaxy Zoo 2 may not be mind-blowingly revolutionary but it definitely fit into the cool and interesting category.

Wednesday, October 7, 2009

Google ambitions ignore digital ephemera of yesteryear

Wired recently had an article critiquing Googles ambitious reach into digital curation of books with a reminder of the last large digital library it undertook and all but abandoned. Usenet.

First Google rescued post-1995 Usenet (a dial-up message board system founded in 1980) from Dejanews in 2001. Google later was able to add millions of posts to this from the magtape of an old Unix guru, Marc Spencer. This gave Google a digital library from over 2 decades of 700 million articles from 35,000 newsgroups.

Wired interviews the Unix guru whose archive supplies a bulk of Google groups (and can't seem to get his name right - is it Marc or Henry Spencer?). Speaking on behalf of his community of Usenet users, Mr. Spencer is very disappointed with the poor searchibility of the material he provided Google.

I would ask: Where are the finding aids? How are things catalogued? Does Google use any of the methods that have worked for traditional archives spanning hundreds of years? No, Google can't even retrieve the legendary alt.gothic flamewars of 1993.

Well, perhaps the world is a better place for that.

But in all seriousness, if anything required professional curation, Usenet could be that. Many from my generation compiled subject knowledge in the early 90s in the form of F.A.Qs, discographies, videographies, and many other encyclopediac collections of unpublished popular (or not very popular) material. Did we save it all? Probably not. There were news-groups for great swaths of information that probably did not carry over to the world wide web post 1995, or even post 2001.

Granted, an enormous amount of Usenet one would not want to have see the light of day, for reasons legal and sundry that could make Craigslist pale in comparison. Still, there are volumes of music, film, and sub-cultural anthropology I have in my file cabinets printed on dot-matrix printers in 1992 that I have yet to see mirrored online or in zines, magazines or books. As I get older and de-clutter, and deem certain material unnecessary or not keeping in with my current interests or values, that information will most likely be lost.

I think a great idea would be for there to be a way to feed and properly code Usenet material (i.e. Usenet group "such and such" date: 1996) and feed the archive anonymously into Wikipedia.

I think the anonymous Wiki framework would be the best platform to mine and add Usenet knowledge and curate it in a central way that would be searchable. Usenet entries would be date-logged as historical data, and contestation, edits and updates would need to follow the historical Usenet information. This would give Wiki a valuable historical layer that I currently find missing. There are times when Wiki feels to me like urban Las Vegas, where the old is torn down and forgotten to make room for the new. What is Wiki kept historical layers to their entries? I think it would be interesting to see a Wiki article from 7 years ago on a given topic.

Data Requests with an Agenda

This week I encountered an interesting example showing the problems inherent in sharing data with others who may have an agenda. Data is of no use to those who are not versed in the field they are examining, as we shall see.

In 2008 Richard E. Lenski (an evolutionary biologist) submitted a paper showing the result of a 20-year experiment that showed evolution in action: after more than 30,000 generations of 12 different populations of E. coli bacteria, one variant evolved the ability to use citrate (a component of the growth medium that can't usually be used as a carbon source by E. coli) as a carbon source. Andy Schafly (a Creationist) of the website Conservapedia sent a letter to Lenski asking for his data, quoting the submission guidelines for the Proceedings of the National Academy of Science that state that data must be made available to readers. To briefly summarize the dialog, Lenski's reply to Schafly's original request and followup made clear that Schafly had not read the article in-depth and did not have a background in evolutionary biology. Lenski pointed out the relevant methods and data were all shown in the paper, and offered to post data on three minor points which were left out of the paper (none of which concerned the existence of the citrate-using bacteria).

Lenski points out that Schafly seems to not be sure what he is asking for- his request for "data" is vague. One of Schafly's "acolytes" even states that Lenski should share the actual bacterial populations, since they form the basis of the data, as a way to "keep tax-payer-funded scientists honest". Lenski states that he will share the bacteria, but only with competent scientists, which Shafly clearly is not. A considerable amount of back and forth discussion has occured over this issue on blogs. (One notable post by a biophysics graduate student dissects Conservapedia's misinformed use of statistics.)

It is important to note that Lenski correctly identifies that Schafly is not asking for the data in good faith. The PNAS guidelines state that data sharing is for the purpose of "replicating and building upon work", while Schafly is only concerned with poking holes or disproving Lenski's results. Schafly takes a tactic often employed by those trying to promote the pseudoscience of "Intelligent Design" and attempts to poke holes anywhere he can in the paper while not being able to disprove the actual results. He even resorts to casting suspicion on the "astoundingly short 14-day peer review period", as if this were a valid argument.

We can see from this issue that making data freely available can have unexpected and troubling results. As has been mentioned in our readings, non-scientists without the skills to correctly interpret the data can draw incorrect conclusions, or may not even be able to recognize the data is staring them plain in the face within the paper. Schafly would probably not be satisfied even if the experiment was performed right in front of his face, even if he had 20 years to wait, but he seems to think that access to data means he should be allowed to see every piece of information generated in the experiment, down to lab notebooks that may only be comprehensible to Lenski and his assistants. Even when the raw data is presented in a table (as in the colony counts in Lenski's paper), those without scientific training may misuse the ideal of "data sharing" to attack scientists with baseless and misinformed accusations.

A new way to search?

Recently Nature.com implemented a new tool for searching the content of the Journal Nature called, appropriately, Nature.com Opensearch.
Over the course of the semester we have examined a many solutions and plans for storing data in repositories and in digital libraries, but the storage of information in these “data silos” is just one part of the data issue. Another piece of the puzzle is how data is found in a digital repository and this is the focus of natures new search strategy. What Nature.com has done is utilized the Opensearch specifications along with CQL to allow for the data in their collection to be searched freely by anyone who would like to build a widget to do so. Not a very revealing sentence I know, believe me I know, so let’s delve.
Opensearch is a set of specifications developed by Amazon (the geniuses who brought you the mechanical Turk) which are meant to allow information to be searched and used in ways that traditional search engines cannot.
Nature.com combined this set of specifications with the CQL, Contextual Query Language, which searches based on ‘semantic’ relationships rather than tradition syntax. This combined methodology offers several dimensions of flexibility to the system at Nature.com.
First the Open side of the application means that it can be embedded in a number of places (such as blogs etc) and results returned in any method that utilizes a standard XML format. The second and in my mind much more interesting aspect of this is that it allows a user to create specific search methodologies and returns. This page has several really nice examples.
Now what does this mean for us. Well in the end in this particular implementation is kind of limited in use. This is a pay Journal, so while anyone can see the articles that are returned not everyone will have access to the content of those article. Researchers with institutional access to the Journal will no doubt find this far more useful that the lay person. In the end what I am more interested in is the idea of turning the search function over to the general population. The institution in this case has set up a framework and turned access (read only, of course) over to anyone who might have a need for it. This way of searching where the data sets conform to standards and the applications which search them are fluid and changeable point to the promise that well curated digital collections can hold.

Microsoft's MyLifeBits: TOTAL DIGITAL CURATION

In the spirit of Microsoft trying to be absolutely everything to everyone, Jim Gemmell and Gordon Bell of the Microsoft Bay Area Research Center are in the process of developing an application that will archive literally every aspect of a person's life. The application is called "MyLifeBits." Well, it's not exactly an application yet, the project is still in the research phase. I heard about it while reading an article by Richard Cox of the School of Information Sciences at the University of Pittsburgh titled "Digital Curation and the Citizen Archivist." In the article, Cox discusses the role of the archivist in schooling the general public on strategies to preserve their digital artifacts, and he mentions the MyLifeBits project as a possible method to do just that.

The MyLifeBits project was inspired by Vannevar Bush's theoretical computer system - the Memex. In 1945, Bush developed a plan for a machine which would combine the functions of storage and electronic capture of images. This "Memex" (the word is a combination of "memory" and "index") would store a library of information on microfilm and all of the images would be connected by "associative links." The Memex has been credited as a pre-cursor to everything from the development of hypertext to the search engine, but in their paper "MyLifeBits: A Personal Database for Everything" Gemmell and Bell see themselves as fulfilling Bush's vision. They state that Bush "posited Memex as 'a device in which an individual stores all his books, records, and communications, and which is mechanized so that it may be consulted with exceeding speed and flexibility. It is an enlarged intimate supplement to his memory."

Microsoft is taking the word "enlarged" seriously. Included in this application are the following tools (keep in mind that this is not an exhaustive list):
  1. TV capture tool (to capture your TV watching)
  2. SenseCam (to capture images IRL)
  3. GPS import and Map display (to track your movements)
  4. Radio capture
  5. Telephone capture
  6. IM capture
  7. Browser tool (to track your web-browsing)
  8. Legacy email client (you get the idea)
  9. Voice annotation tool
  10. Text annotation tool
Gemmell and Bell's argument for developing such an ambitious application with such wide-reaching curation tools is that since no one can predict when some item might be useful the "safest thing is to simply [sic] keep it all. Everything."

As I'm sure you can imagine, there are a whole host of issues in making such an enormous amount of information accessible, and so far it doesn't sound like Gemmell and Bell have developed a workable solution to these problems. They have been experimenting with hierarchical classification systems and have discovered (surprise!) that such systems quickly become unwieldy when dealing with large amounts of information. They haven't abandoned the idea of hierarchical folder systems, however, and are actually experimenting with "hierarchical classifications that will be developed by others to be downloaded by the user, and which contain extra information such as synonyms and descriptions to ease their use."

Classifications schemes are not the only thorny issue for the MyLifeBits team. Constructing useful metadata capture for video has also proved difficult. In developing a method to garner metadata for video they have suggested harvesting audio and converting it to searchable text so that, e.g., one could search for the name of a given person and jump to a segment of the video where the speaker mentions that person's name. Whether this would necessarily lead the searcher to a segment of the video with the person he/she is looking for is, well, who knows?

Clearly there are a number of issues that need to be ironed out before Microsoft can release this product, but Gemmell and Bell seem intent on doing so. Just this year, they released a book, Total Recall: How the E-Memory Revolution Will Change Everything, to proselytize their vision of total personal digital curation. While this project currently seems a little premature, it may be that in 30 years applications like MyLifeBits could be commonplace. If this is the direction that personal digital curation is going, what are some of the issues presented by such an all-encompassing view of "curation?" Certainly there are going to be copyright issues, classification issues, privacy issues (not to mention dystopian issues! Who wants their life to turn into some monitored, post-apocalyptic version of a Tom Cruise movie?)

Comparison of citation retrieval by search engines and the ISI Web of Knowledge

I continue to be curious about the practice of citation for promotion and tenure, and am always amazed when I haphazardly come across a print article that is so pertinent to our class discussion.

The September 2009 issue of College and Research Libraries (Vol. 70, No. 5, pp. 460-472) contains the article "A Citation Analysis of College & Research Libraries Comparing Yahoo, Google, Google Scholar, and ISI Web of Knowledge with Implications for Promotion and Tenure" by Charles Martell. (Web access is available to ALA members.)

1985 and 2005 studies determined that College & Research Libraries (C&RL) was ranked number one in "journal prestige in terms of value for tenure and promotion" by the Association of Research Libraries (ARL) Library Directors. Given this ranking, Martell determined the following: 1) frequency of citations of C&RL articles, 2) citation retrieval strengths of Yahoo, Google, Google Scholar, and ISI Web of Knowledge, 3) the relevance and applicability of these findings, and 4) classification and quantification of retrieved entries that did not qualify as citations. Every refereed C&RL article from 2000-2006 was used.

Advantages and drawbacks to each search method was discussed. Yahoo and Google are readily available and expansive, but laborious to filter. ISI WOK has a high level of quality and ease of use, but is fee-based and proprietary. Google Scholar has a broad coverage, but still has varying drawbacks: for example, for two of the top ten articles, over half the citations were from China. That said, Google Scholar returned more than twice as many citations as ISI WOK, and of course, the advantages and disadvantages should be weighed before selecting one over the other for citation counting.

It is valuable to note that this study was found to be particularly relevant to academic librarians in tenure-track positions. I found it curious that Martell discovered that deans of ALA-accredited education programs ranked C&RL 11th behind six information science journals and four library science journals. On second thought, perhaps not so curious given the trend of rebranding programs from "library science" to "information science" to incorporate the increasing importance of information technologies.

I think that this data would be very interesting to follow over the long term, in light of the growing interest in, or opposition to, open journals, and also our current discussion of broadening the definition of citations. It would be useful to note spikes/fluctuations in journal and/or citation usage and how it relates to the increasing effectiveness of search engines or initiatives such as the "Compact for Open Access Publishing Equity" (initiated by MIT, Cornell, Dartmouth, Harvard and UC Berkeley). If this data was made readily traceable, would it change the decisions and activities of the stakeholders (authors, promotion committees, publishers, etc.)?

What's the HapMap??

While clicking through the blogosphere earlier this week, I came across a project that seems to be an extension of the National Human Genome Research Institute. It's the International HapMap Project, with the intended focus to find "genes that affect health, disease, and individual responses to medications and environmental factors". The US, Canada, China, the UK, Nigeria, and Japan are collaborating on the research with public and private funds. Because it started in 2002, phase one of the project finished in 2005, with Phase II ending in 2007. I am interested in how accessible the data is and process of the project, which seems fairly documented.

The core of this project is to share information. The website says: "The International HapMap Project is not using the information in the HapMap to establish connections between particular genetic variants and diseases. Rather, the Project is designed to provide information that other researchers can use to link genetic variants to the risk for specific illnesses, which will lead to new methods of preventing, diagnosing, and treating disease."
Very commendable! It appears that this consortium is so expectant that their data will be used, that they offer tips on how to cite the project, especially how to refer to the sample populations (from Nigeria, Japan, Utah, etc). There is also a push to publish, as "researchers are encouraged to publish results based on combining HapMap data with data from other projects, particularly in efforts to find genes affecting a disease or a drug response. Researchers also are encouraged to use HapMap data to publish on the development of novel methods to analyze polymorphism, linkage disequilibrium, and association data."

Why all the good will? Is it because this is funded through the NIH? And must scholars who integrate the HapMap data with their own research make that data available? Interested in how much the data is being used, I looked at the site's publication page, but that hadn't been updated since 2007 when their last phase ended. So I went to GoogleScholar (gasp!) so I could just do a quick look up and found that the HapMap site has been cited 160 times. That could be good!

If I could even slightly understand anything about this project, I might be tempted to download the data (they provide tutorials!). Instead I just looked at the page from whence the files could come and was confused. There is a place to upload your own data, which I'm not sure I think is a particularly good idea, since their research seems quite controlled (270 participants only). I'm assuming the HapMap data and others' data is separate.

The International HapMap Project appears to be an example of a successful data-sharing hard science project, although its proximity in topic and funding agency to the Human Genome Project might also account for that. From an outsider's perspective, HapMap appears to have a socially conscious purpose (healing sick people, yay!) with an integrative approach to the data and publishing.

Challenges of Open-access Journals

An article titled "Ten Challenges for Open-access Journals" appeared in this month's issue of The Scholarly Publishing & Academic Resources Coalition (SPARC) Open Access Newsletter. In the article, Peter Suber explores what he sees as the ten greatest challenges facing open-acess journals.
Three of Suber's challenges deal with disparities between what is intended and what has been achieved, and the remaining seven have to do with doubts about open-access publications.
One of the disparities lies in the Impact Factors (IF) used to measure the impact of a particular journal. While IFs are intended to reveal the quality of the journal, a journal isn't even eligible for an IF until it has been around for 2 years. This is particularly hard on open-access journals because they are relatively new.
A second disparity is the failure of open-access journals to use a creative commons license to allow more than fair use of a journal. Many apen-access journals believe that by making their content free, they are truly open access. Suber argues that institutions are not free to exceed fair use. He suggests that until open-access journals explicitly allow for more uses (such as including the works in a database or archiving copies for preservation) they are not truly open-access.
The final disparity can be found in the gap between quality and prestige. It is impossible for a new journal to be prestigious from its start though it could be of a high quality. Subter argues that quality should matter more than prestige an until it does, open-access journals are at a disadvantage.
Subter moves on to discuss seven doubts that cause challenges for open-access publications. This include doubts about quality, preservation, honesty, publication fees, sustainability, redirection of funds, and even doubts about the strategies for addressing the challenges of open-access journals. These doubts are fairly self-explaintory and Subter proposes ways to deal with each of them. Subter suggests that an awareness and willingness to address these issues is absolutely necessary for the success of open-access journals.
We have previously discussed the need for changes in the publication process and even in the way an article's contribution is measured, but I still feel a bit doubtful that the change will happen anytime soon.

NSF's DataNet Partners so far: Data Conservancy and DataONE

Last spring, the Office of Cyberinfrastructure of the NSF announced its plan to support data preservation and access through the Sustainable Digital Data Preservation and Access Network Partners (DataNet) grant program. The program aims to address "one of the major challenges of this scientific generation: how to develop the new methods, management structures and technologies to manage the diversity, size, and complexity of current and future data sets and data streams." Through funds from the grant, "exemplar national and global data research infrastructure organizations," called DataNet Partners, will do just that. These organizations will combine library and archival sciences, cyberinfrastructure, computer and information sciences, and domain science expertise to provide preservation, access, integration and analysis capabilities over a decades-long timeline. As if that wasn't enough, these organizations should also be able to anticipate and adapt to changes in technologies and user expectations, stay cutting edge, and drive research and development. Oh, and one more thing: "potential applicants should note that this program is not intended to support narrowly-defined, discipline-specific repositories." The NSF is holding out for a hero!

Since the solicitation for proposals, two organizations have been awarded grants - both in August. (It seems that TeraGrid is categorized under the program but was awarded its $32 mil way back in '05.) $3.7 mil was given to Sayeed Choudhury of Johns Hopkins University in order to create the Data Conservancy (DC). Along with a long list of partners from universities and scientific organizations, Johns Hopkins will create "a new model in which libraries regard digital data as a special collection that must be maintained and served like their other collections." The Data Conservancy is concerned with the global carbon cycle and the carbon-climate-human system. The second award was given to William Michener of University of New Mexico and a ton of other partners including the CDL, Amazon, UIC, and Intel. That $12.3 million award is for the creation of DataONE (Observation Network for Earth). DataONE will address four challenges: preserving data, providing access to dispersed data, integrating data, and creating best practices for managing data. Its focus is huge: it will make "biological data available from the genome to the ecosystem; make environmental data available from atmospheric, ecological, hydrological, and oceanographic sources; provide secure and long-term preservation and access; and engage scientists, land-managers, policy makers, students, educators, and the public through logical access and intuitive visualizations."

It's interesting to me that both these DataNet Partners are doing pretty similar things, except the scope of DataONE seems to be much larger than that of the Data Conservancy. DataONE was presented at the 4th International Conference on Open Repositories at Georgia Tech in May. The presentation outlines some of the problems we've discussed with putting multi-disciplinary data in one place, but it purports that DataONE will "provide one-stop shopping for data" and will create "investigator toolkit" that will include nifty things like data visualization and Kepler workflow diagrams (remember those?) The presentation also gives an outline of how DataONE will be technologically implemented using the micro-services approach - or an "unbundled alternative to monolithic systems." Both Coudhury and Michener will speak at the next Educause conference in November, along with our pal Clifford Lynch, in a session called "Initiatives from the NSF's DataNet Program: DataONE and the Data Conservancy."

Sum of Humankind’s Knowledge Available Online 

Great title for an article, huh? It appears that scientific research in Europe is becoming available to the public and other researchers through the DRIVER (Digital Repository Infrastructure Vision for European Research) Search Portal, which serves as a door to European Open Access research. The idea is to create a “library of libraries,” where individuals can readily search across repositories throughout Europe.

European researchers have created D-NET, software that links “information collected on diverse computer platforms, using legacy software which can still ‘talk’ or work with older systems in more than 25 European languages.” The software regularly harvests publications to provide access to information from numerous repositories. Visitors to the site are encouraged to use the software to set up other portals and are pointed toward examples of test cases where other groups are using it. The DRIVER portal site reports the following information about the collection:

• approximately 1,000,000 documents
• found in journal articles, dissertations, books, lectures, reports, etc.
• harvested regularly from more than 200 institutional or thematic repositories
• from 23 European countries
• in 25 languages.

The portal allows you to search broadly across repositories, but also limit your search by repository, language, document type, and date. At present the focus is on textual documents, but the idea is extend into other types of media. Current goals are to continue building the infrastructure and increasing the number of participating repositories. Another development of the initiative was the creation of Guidelines for Repository Managers, a document that outlines how to make content compatible with DRIVER. This type of document is an instructional tool to ease the process of increasing interoperability. The initiative is funded by the Research Infrastructure priority of the EU’s Sixth Framework Programme for research.

As a “library of libraries,” the DRIVER portal seems like a super repository that allows you to access a larger network of published materials through a single search, but I do not believe it goes beyond published works to data sets. However, the possibility for expansion into broader scholarly communication is there and the openness to using the software may allow for other uses within the scientific community.