Monday, October 19, 2009

A Tale of Compromise

Jessica Branco Colati, Robin Dean, and Keith Maull, “Describing Digital Objects: A Tale of Compromise,” Cataloging & Classification Quarterly 47, no. 3 (4, 2009): 326-369.

This article reports the efforts Alliance Digital Repository (ADR), a consortial digital repository service for the Colorado Alliance of Research Libraries (Alliance), to create a standard descriptive metadata policy for their repository records. Members in the Alliance included twelve separate libraries at nine different academies and institutions, public and private. The goal was to support each library's community standards while also ensuring interoperability. This would facilitate a central repository.

Needless to say the scope for a project like this, which is charged with accommodating a dozen libraries each with their own practices and metadata (not to mention unique data both from within and without the library), expands at a rate equal to thought itself.

The ADR eventually pared down the libraries' metadata they had to consider to:
  • MARC
  • MARC XML
  • MODS
  • DC
  • Various extensions to these standard schemas
  • ProQuest's Digital Dissertations XML-based metadata schema
ADR chose Fedora as the main repository software for its flexibility, conformance to OAIS, and automatic versioning of data and metadata. This is a really useful consideration, since metadata is a digital object too that needs to have its own records of when it was made, changed, by whom, etc., creating an easy audit trail of what has happened in the repository. Of course, one could go on and on with this, pointing out that that metadata too ought to have some metadata associated with it. As a quick solution to this problem I would offer that the point at which all metadata is automatically generated is the point where you can stop adding metadata to metadata, since automatically generated metadata doesn't need to document its own automatic generation (right?).

As they note Fedora does not have a user-friendly interface for submission and searching built-in (like DSpace) so they chose Fez as a configurable front end, adding that other interfaces could be attached to Fedora as needed.

MODS was chosen as the "normalizing schema," the one all records would be converted to at a minimum no matter what metadata scheme they had upon ingestion. The report gives a really helpful list of reasons why, and I think it's interesting that #1 is the simple fact that Fez and MODS "are proven partners: The Fez system was already using MODS widely throughout the system as the primary means of describing objects." They give other factors: the right granularity, crosswalks to the popular data-sharing DC, and plain familiarity, but it's notable that the tool in this case (Fez) partly determined yet another tool.

ADR determined minimal metadata field requirements with DLF Aquifier, a Digital Library Federation initiative to help distributed library networks and content.

A very significant technical problem for ADR was validation of incoming records to make sure the minimum MODS metadata was there and was correct. They decided that XSD could both the valid structure for XML encoded MODS records and the correct data types as they had agreed upon. For instance a record would have not only the correct structure of metadata that a DTD might confirm, but also the correct data inside the metadata fields (for example the correct time, date, price, or URI format).

The problem derived from the fact that the packaged XSDs as used by Fez were not applicable for ADR's purposes, which meant the creation of new XSDs. Integration of new custom XSDs with Fez was poorly documented and ADR encountered functional bugs with Fez. A survey of Fez use revealed that no one else was trying to significantly modify Fez's XSD templates, and ADR's number of digital objects, each with unique metadata structures, was very large. Added to this was the need to crosswalk MODS to DC that required a whole new set XSDs for the DC metadata records.

ADR ended up authoring their own "document types" (MODS and DC XSD-validated metadata records) for use. That process is long and complex and ultimately did not provide the best treatment for heterogeneous objects (objects containing multiple genres of document types, like a web page) since Fez allowed such minimal description for a single document type. This forced them to store these objects as separate records.

What interests me with this report is not so much the specific technical problems and solutions they encountered (which are kind of torturous to read about) but that such huge obstacles can be created when the technical tools are insufficient to the task. On the one hand, MODS was easily selected and implemented for a number reasons, not least of which is that Fez worked smoothly with it. There are just a few paragraphs on it. On the other hand, entire collections of XSDs and the creation of a new conceptual entity (document types) is created largely because Fez has a poor capacity to deal with new or modified XSDs. Four pages detail this struggle, and it's unresolved at the conclusion of the article. I conclude that while the biggest struggles for digital curation and interoperability may be social, political, cultural and so on, the extent to which technical can grease the wheels is tremendous.

Saturday, October 17, 2009

Cooliris

Cooliris, is a 3-D image wall that allows you to visually browse images, whether they be in Google Image, YouTube, Flickr, Facebook, Picassa, images in one's harddrive and recent versions of the photo database program Adobe Lightroom.

I was sad to discover that I could not download this program to my own Mac Mini as it features Mac OX 10.4, Power PC, and apparently I cannot upgrade to the necessary Leopard version that supports Cooliris. Update: Previous versions (tho unsupported) are available.

You can share your 3-D image wall through a url, as well as bookmark and save it. It allows one to "jump freely" through Flickr image pools and sets - though not sure how that differs from the current way we maneuver from image pools to sets - does "jump freely" = "move fluidly?" Not sure.

It claims usability with hundreds of sites due to its being built around Media RSS format. It advertises its ability to be used to view numerous television and movie episodes on Hulu, etc. Its slideshow feature allows one to double click and launch slide shows that one can pause and rewind. It claims to have the fastest way online to search images with its' style of "zipping" through the 3-D image wall.

Again this makes me think back to the Visible Archive from Australia - and its archival metadata visualization. If there was a way to integrate these technologies and allow one to browse visualizations of data - the future could be very exciting.

Wednesday, October 14, 2009

Dude, Your in the NYT

Happy off week everybody, here is a light piece but one which should still be interesting to us.

Congratulations, now when people ask you about what exactly it is that you are doing in school you don't have to give them vocab heavy monologues about metadata, data migration and petabytes, you can point them to this article in the New York Times.

The heart of this article is what we have been discussing through much of the semester, the sciences are going through a major change where the shear amount of data that is being gathered threatens to overwhelm researchers who don't have the training to deal with such datasets. The article goes then discusses how current students in the sciences are unable to access the vast scope of the information sets which are a part of their field s due to the limited nature of the computational facilities available to students. This causes the young minds to “imprint on these small systems, that becomes their frame of reference and what they’re always thinking about...”. I have doubts on the validity of this but still Google and IBM are going to solve these problems for scientist by providing them with access to better computers.

There is of course the unstated belief that shadows this article that all of these issues are inherently technical, and the application of computational power will solve the petabyte problem. Indeed the article quotes on researcher saying “Science these days has basically turned into a data-management problem”. A brief Google search does find that the scientist quoted here is a biology/ genetics researcher so in that field this might be a perfectly valid point, but I wonder if all scientist feel this way. And if they do why, and if they don't why not?

The question that I suppose I have from this article is where do we fit in? Do we work with the scientist to help them understand their data or are we only concerned with where the data is kept, and how to ensure that it remains available for use now and in the future?
I don't have good answers I have been looking at these problems for the better part of a year now and I am still not sure how we as a profession fit with this area, or if we do.

google scholar. not there yet.


Given how much class time has been devoted to discussing the many aspects of Google, I thought I would find this week’s article using Google Scholar. A quick search for the key term “data sharing” (all articles) revealed a top hit for a 1994 article titled “The Bayou architecture: Support for data sharing among mobile users” (http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=512726). Cited by 329 other sources, this article is a good 50 citations ahead of its closest competitor. Clicking through to the list of articles that cited the Bayou article in order of the article’s citation instances might be more useful if there were options as to how this list were arranged (say chronologically). Going back to the original search page and changing the search parameters to limit articles by date (say “since 2008”) still yields some 72,000 hits, however, in this function one can no longer see the returns based on number of citations.

In an article titled "Newswire Analysis: Google Scholar’s Ghost Authors, Lost Authors, and Other Problems," which appeared in Library Journal (9/24/2009), the author criticizes Google Scholar for other reasons: namely their poor metadata standards, and how this will effect citation numbers. He notes that because Google Scholar relies on their own algorithms rather than utilizing the metadata available through libraries. This, he argues “creates phantom authors for millions of papers. They derive false names from options listed on the search menu, such as P Login (for Please Login).”

It seems clear, that Google Scholar is a long way from having any sort of reliable measure of citation count, and also is not yet the powerful search tool it purports to be.

Blog Post About Nothing


So, I ran across this hypothetical Seinfeld scene mocking the twitter phenomenon, and couldn't resist turning it into a blog post. Basically, in the scene, Elaine and Jerry are giving George a hard time because he has become obsessed with tweeting and is engaging in said activity compulsively. George is staunchly defending his right to tweet (declaring "No one stops George Costanza from tweeting!") When Jerry points out that George, who has no job or girlfriend, does not in fact have anything of interest to tweet about, George emphatically asserts, "I got plenty to tweet about, baby!" After Kramer wanders in and mentions that he had a friend who tweeted himself to death, the scene ends with Jerry asking rhetorically whether typing in increments greater than 140 characters would break the internet.

The scene, though obviously farcical, illustrates several points relevant to the consideration of digital curation. First, it brings to mind the problem with noise potentially drowning out information of value. Digital publication is so cheap and easy that anyone, even a consummate loser like Costanza, can flood cyberspace with information about nothing. Because of the potential to lose relevant information in a sea of trivia, I think that it is extremely important that we emphasize the filter function involved in digital curation. For example, can any single scientist really humanly go through ALL data pertaining to a study? Particularly in fields like climatology it seems essential to perform some sort of data triage. (Many of those fields benefit from shared data, but usually also along the lines of collaborative efforts instead of straight data dumps.) It seems to me that digital curators could at least perform a cursory, primary triage of data.

A second point raised by the Seinfeld skit is that no matter how great and convenient we the information pros find new information technology, there will always be the likes of ludites, tenured liberal arts professors emeriti, and half-deranged hipsters who will hate, fear, or refuse to accept new technological methods. In expanding access via digital means, we must be careful not to limit access to other portions of the population, which may be more a problem of infrastructure than straight-up digital curation. At the very least, education programs in new i.t. methods seem necessary.

Finally, Jerry's joke about 141 characters breaking the internet drives home the fact that we really are dealing with new media. Twitter has changed the way people communicate by rigidly enforcing a strict brevity. This may change people's reading/learning habits. I'm not precisely suggesting we limit our descriptive elements to 140 characters per digital item, but it is something to keep in mind.

So, basically, the sketch brought to the surface of my mind some of the digital curation issues that have been spinning around in there such as filtering noise, adapting to new user habits, yada, yada, yada...

Curating Architectural 3D CAD Models

MacKenzie Smith's article presents MIT’s FACADE (Future-Proofing Architectural Computer-Aided DEsign) research project as an attempt to determine how to preserve and provide access to 3D CAD digital records. The digital curation project includes working with three architectural firms to determine their workflow, usage, and expectations about archival preservation of their digital design records. Traditionally, architectural archives have included records about the design of architectural projects, as well as firm records that support the creation of buildings. Numerous documents are generated through the process of architectural design and these records are now largely digital born – email correspondence, digital photography, 2D CAD renderings, and 3D CAD models, to name a few. Architects generally use proprietary software to create many of their design. Most of these files are easily migrated or converted to standard formats, such as PDF or JPEG. The problem arises when trying to preserve or access 3D architectural models, created using proprietary software in non-standard formats, which cannot be easily migrated or converted.

This presents a major problem for libraries, archives, and museums whose goals would include providing access to the digital records, but would not have the funding to support the ownership and maintenance of proprietary software systems, especially when the software used to create the documents rapidly changes. The FACADE project looked at strategies for exporting files and creating standard formats. This required expertise in both the CAD software and the underlying data, which many librarians or archivists at repositories would not have. The hope is that this process will become more automated over time.

One of the primary issues is how to connect the various records created within context of the design process. The 3D model is most valuable when viewed in relation to other drawings, photographs, and project correspondence. To this end, they have established a “Project Information Model” which will collect 3D model data and connect it to other project records. They have created a prototype for open source software that will assist curators in creating metadata.

The FACADE project web site includes further information, including a final report on the project. The final report outlines four versions in different formats that will be used to preserve the 3D CAD models.

1. Original (the originally submitted version of the CAD model)
2. Display (an easily viewable format to present to users, normally 3D PDF)
3. Standard (full representation in preservable standard format, normally IFC or STEP)
4. Dessicated (simple geometry in a preservable standard format, normally IGES)

This project was funded by a two-year grant from IMLS. The project predominantly set out to conduct research on the best methods and practices for curating 3D CAD models. There are still many research questions to be addressed, such as those pertaining to privacy and copyright of the data.

More on Google Books

Sergey Brin wrote an opinion article late last week for the New York Times titled "A Library to Last Forever." The same article appeared a day later as a blog post titled A Tale of 10,000,000 Books on the official Google blog. The post is essentially a defense of Google Books and the recent settlement they have reached with the Authors Guild and the Association of American Publishers.
In the post, Brin talks about a "black hole" of information printed after 1923. These books are not in the public domain, but are frequently out of print and increasingly difficult to find. Brin argues that Google Books is making important parts of the "world's collective knowledge and cultural heritage" available. I don't think any of us would argue the potential for the positive impact of Google Books on the way research is conducted. I am more concerned about what Google Books means for future mass digitization projects.
Brin states that "If Google Books is successful, others will follow." But what about the large startup costs? To Google's credit, they are the only ones with the means and the drive to even attempt such a large-scale book digitization project. But with the expense being so high, will there ever be anyone else with both the means and drive to create another digital collection? It seems to me that it would be highly unlikely, especially if Google Books is successful. Who wants to take on a giant when the cost of entrance is so high?
While I don't pretend to know everything about the settlement, it appears that it will create a registry of rights for books which helps to identify the rights holders for individual works and to make obtaining permission and assigning appropriate revenue to the rights holders easier. Again, such a registry would be very helpful, but how many entities exist with the kind of funding it would take to pay the appropriate fees and digitize the books?
And if Google Books is the only entity to attempt such a large-scale book digitization project, we find ourselves back where we started with the problems of inaccurate and misleading metadata. Brin addresses these issues briefly, only to say that Google is "working hard to address them."
In spite of all of my doubts, it is hard not to get swept up by Brin's call to save the literary works of the 20th century. I suppose we'll have to wait and see how Google Books affects the future of mass digitization projects. Meanwhile, Google would like us believe that they have found the answer.

Visualize This.


For the past few weeks I've been interested in how large data sets are being "visualized". Not only is this a fun (and easy) way to digest large issues, but the data must come from somewhere.
Some blogs I've been looking at lately include Information Aesthetics, Information is Beautiful, and Ben Fry's Projects, especially a visualization of changes in editions of Darwin's Origin of the Species.

I explored these sites to see where the authors are getting their data. Information is Beautiful provides links to his sources, which open as Google document spreadsheets. David McCandless, the author, also links to The Guardian's data (since the newspaper's DataBlog is also linked) and this also opens to spreadsheets. Information Aesthetics' Andrew Vande Moere doesn't provide his data sources, but does offer links to other data/information sites for hours of wasting/exploring. Ben Fry's The Preservation of Favoured Traces project states the "text for each edition was sourced from their careful transcription of Darwin's books" although more information about how the text was actually transcribed. The "Data" link on his site, however, is grayed out, so apparently not available. Fry does explain all his projects in detail, however, just maybe not all the information we would want.

An interesting article on the Visual Journalism blog talks about how some data is just boring and not everything needs to be graphically interpreted. The author, Gert K. Nielsen, is discussing a NY Times interactive visual about how people spend their day and mentions that the original data of the graphic, "American Time Use Survey, is in danger of being shut down due to serious cuts in the budget of the Bureau of Labor Statistics". That would hamper the continuation of this visualization, but the Bureau of Labor Statistics will surely have other great data to troll.

The popularity of information visualization will probably only increase (until the novelty runs out?). Ben Fry has written a book about it, Visualizing Data, published by O'Reilly. I wonder if the graphs and pretty colors will distract people from questioning the source of the data, or highlight the nature of statistics and data even more, creating a demand for source knowledge. I'm hoping the latter, especially with easy sources like data.gov, there is no reason not to provide the raw data. I also like the possibility of increased data sharing this might facilitate. Even if it starts out small, such as the GoogleDocs on Information is Beautiful, this creates a society of openness that can be emulated.

Galaxy Zoo 2

These past couple of weeks, I’ve been blogging about developing cyber initiatives and grants being given to fund projects that designed to generate and display data like never before. In fact, the words “revolutionary data collection” kept popping up again and again. But in most of the articles I’ve been reading about developing projects, no one ever really delves into what, exactly, will be done with the massive amount of data that will be collected; no one ever mentions how the data could be used – just that it will be utilized in “revolutionary” ways. With this in mind, this week I set out to find an example of something revolutionary – or at the very least something cool and interesting – that was being done with digital data.

My search led me to the zoo, the Galaxy Zoo, that is. Galaxy Zoo is a site that invites the public to help classify millions of galaxies using data and images from the Sloan Digital Sky Survey at the Apache Point Observatory in New Mexico. All of the images come from a digital camera which is mounted on a telescope.



I think this is really neat. People like you and me can help classify and study the universe, no prior knowledge of astronomy needed (or so claims Wikipedia). The original Galaxy Zoo first made its debut in 2007 “with a data set made up of a million galaxies imaged with the robotic telescope of the Sloan Digital Sky Survey”; the new version, Galaxy Zoo 2, launched back in February of this year. According to their website, within the first day of Galaxy Zoo’s launch the site was getting 70,000 classifications an hour; during the first year over 50 million classifications were received from nearly 150,000 people.

Galaxy Zoo 2 has focused on almost a quarter of a million of the “nearest, brightest and most beautiful galaxies” for users to view and classify. Apparently, new discoveries by the amateur astronomers at Galaxy Zoo are being made constantly about galaxy colors, shapes and patterns. Professional telescopes (and I’m assuming professional astronomers along with them) have even followed up on some of the Galaxy Zoo findings; the list of professionals includes “the Isaac Newton and William Herschel Telescopes on the island of La Palma in the Canaries, Gemini South in Chile, the WIYN telescope on Kitt Peak, Arizona, the IRAM radio telescope in Spain’s Sierra Nevada, the Swift and GALEX satellites, and the Hubble Space Telescope.”

I've been playing around with the website's classification tutorial page, and I have to say that it is pretty fun. This is a great example of the neat things that can be accomplished with crowd-sourcing and open data. Regular people with an interest in astronomy having access to data to classify galaxies, make discoveries, and help figure things out about the universe; Galaxy Zoo 2 may not be mind-blowingly revolutionary but it definitely fit into the cool and interesting category.

Wednesday, October 7, 2009

Google ambitions ignore digital ephemera of yesteryear

Wired recently had an article critiquing Googles ambitious reach into digital curation of books with a reminder of the last large digital library it undertook and all but abandoned. Usenet.

First Google rescued post-1995 Usenet (a dial-up message board system founded in 1980) from Dejanews in 2001. Google later was able to add millions of posts to this from the magtape of an old Unix guru, Marc Spencer. This gave Google a digital library from over 2 decades of 700 million articles from 35,000 newsgroups.

Wired interviews the Unix guru whose archive supplies a bulk of Google groups (and can't seem to get his name right - is it Marc or Henry Spencer?). Speaking on behalf of his community of Usenet users, Mr. Spencer is very disappointed with the poor searchibility of the material he provided Google.

I would ask: Where are the finding aids? How are things catalogued? Does Google use any of the methods that have worked for traditional archives spanning hundreds of years? No, Google can't even retrieve the legendary alt.gothic flamewars of 1993.

Well, perhaps the world is a better place for that.

But in all seriousness, if anything required professional curation, Usenet could be that. Many from my generation compiled subject knowledge in the early 90s in the form of F.A.Qs, discographies, videographies, and many other encyclopediac collections of unpublished popular (or not very popular) material. Did we save it all? Probably not. There were news-groups for great swaths of information that probably did not carry over to the world wide web post 1995, or even post 2001.

Granted, an enormous amount of Usenet one would not want to have see the light of day, for reasons legal and sundry that could make Craigslist pale in comparison. Still, there are volumes of music, film, and sub-cultural anthropology I have in my file cabinets printed on dot-matrix printers in 1992 that I have yet to see mirrored online or in zines, magazines or books. As I get older and de-clutter, and deem certain material unnecessary or not keeping in with my current interests or values, that information will most likely be lost.

I think a great idea would be for there to be a way to feed and properly code Usenet material (i.e. Usenet group "such and such" date: 1996) and feed the archive anonymously into Wikipedia.

I think the anonymous Wiki framework would be the best platform to mine and add Usenet knowledge and curate it in a central way that would be searchable. Usenet entries would be date-logged as historical data, and contestation, edits and updates would need to follow the historical Usenet information. This would give Wiki a valuable historical layer that I currently find missing. There are times when Wiki feels to me like urban Las Vegas, where the old is torn down and forgotten to make room for the new. What is Wiki kept historical layers to their entries? I think it would be interesting to see a Wiki article from 7 years ago on a given topic.

Data Requests with an Agenda

This week I encountered an interesting example showing the problems inherent in sharing data with others who may have an agenda. Data is of no use to those who are not versed in the field they are examining, as we shall see.

In 2008 Richard E. Lenski (an evolutionary biologist) submitted a paper showing the result of a 20-year experiment that showed evolution in action: after more than 30,000 generations of 12 different populations of E. coli bacteria, one variant evolved the ability to use citrate (a component of the growth medium that can't usually be used as a carbon source by E. coli) as a carbon source. Andy Schafly (a Creationist) of the website Conservapedia sent a letter to Lenski asking for his data, quoting the submission guidelines for the Proceedings of the National Academy of Science that state that data must be made available to readers. To briefly summarize the dialog, Lenski's reply to Schafly's original request and followup made clear that Schafly had not read the article in-depth and did not have a background in evolutionary biology. Lenski pointed out the relevant methods and data were all shown in the paper, and offered to post data on three minor points which were left out of the paper (none of which concerned the existence of the citrate-using bacteria).

Lenski points out that Schafly seems to not be sure what he is asking for- his request for "data" is vague. One of Schafly's "acolytes" even states that Lenski should share the actual bacterial populations, since they form the basis of the data, as a way to "keep tax-payer-funded scientists honest". Lenski states that he will share the bacteria, but only with competent scientists, which Shafly clearly is not. A considerable amount of back and forth discussion has occured over this issue on blogs. (One notable post by a biophysics graduate student dissects Conservapedia's misinformed use of statistics.)

It is important to note that Lenski correctly identifies that Schafly is not asking for the data in good faith. The PNAS guidelines state that data sharing is for the purpose of "replicating and building upon work", while Schafly is only concerned with poking holes or disproving Lenski's results. Schafly takes a tactic often employed by those trying to promote the pseudoscience of "Intelligent Design" and attempts to poke holes anywhere he can in the paper while not being able to disprove the actual results. He even resorts to casting suspicion on the "astoundingly short 14-day peer review period", as if this were a valid argument.

We can see from this issue that making data freely available can have unexpected and troubling results. As has been mentioned in our readings, non-scientists without the skills to correctly interpret the data can draw incorrect conclusions, or may not even be able to recognize the data is staring them plain in the face within the paper. Schafly would probably not be satisfied even if the experiment was performed right in front of his face, even if he had 20 years to wait, but he seems to think that access to data means he should be allowed to see every piece of information generated in the experiment, down to lab notebooks that may only be comprehensible to Lenski and his assistants. Even when the raw data is presented in a table (as in the colony counts in Lenski's paper), those without scientific training may misuse the ideal of "data sharing" to attack scientists with baseless and misinformed accusations.

A new way to search?

Recently Nature.com implemented a new tool for searching the content of the Journal Nature called, appropriately, Nature.com Opensearch.
Over the course of the semester we have examined a many solutions and plans for storing data in repositories and in digital libraries, but the storage of information in these “data silos” is just one part of the data issue. Another piece of the puzzle is how data is found in a digital repository and this is the focus of natures new search strategy. What Nature.com has done is utilized the Opensearch specifications along with CQL to allow for the data in their collection to be searched freely by anyone who would like to build a widget to do so. Not a very revealing sentence I know, believe me I know, so let’s delve.
Opensearch is a set of specifications developed by Amazon (the geniuses who brought you the mechanical Turk) which are meant to allow information to be searched and used in ways that traditional search engines cannot.
Nature.com combined this set of specifications with the CQL, Contextual Query Language, which searches based on ‘semantic’ relationships rather than tradition syntax. This combined methodology offers several dimensions of flexibility to the system at Nature.com.
First the Open side of the application means that it can be embedded in a number of places (such as blogs etc) and results returned in any method that utilizes a standard XML format. The second and in my mind much more interesting aspect of this is that it allows a user to create specific search methodologies and returns. This page has several really nice examples.
Now what does this mean for us. Well in the end in this particular implementation is kind of limited in use. This is a pay Journal, so while anyone can see the articles that are returned not everyone will have access to the content of those article. Researchers with institutional access to the Journal will no doubt find this far more useful that the lay person. In the end what I am more interested in is the idea of turning the search function over to the general population. The institution in this case has set up a framework and turned access (read only, of course) over to anyone who might have a need for it. This way of searching where the data sets conform to standards and the applications which search them are fluid and changeable point to the promise that well curated digital collections can hold.

Microsoft's MyLifeBits: TOTAL DIGITAL CURATION

In the spirit of Microsoft trying to be absolutely everything to everyone, Jim Gemmell and Gordon Bell of the Microsoft Bay Area Research Center are in the process of developing an application that will archive literally every aspect of a person's life. The application is called "MyLifeBits." Well, it's not exactly an application yet, the project is still in the research phase. I heard about it while reading an article by Richard Cox of the School of Information Sciences at the University of Pittsburgh titled "Digital Curation and the Citizen Archivist." In the article, Cox discusses the role of the archivist in schooling the general public on strategies to preserve their digital artifacts, and he mentions the MyLifeBits project as a possible method to do just that.

The MyLifeBits project was inspired by Vannevar Bush's theoretical computer system - the Memex. In 1945, Bush developed a plan for a machine which would combine the functions of storage and electronic capture of images. This "Memex" (the word is a combination of "memory" and "index") would store a library of information on microfilm and all of the images would be connected by "associative links." The Memex has been credited as a pre-cursor to everything from the development of hypertext to the search engine, but in their paper "MyLifeBits: A Personal Database for Everything" Gemmell and Bell see themselves as fulfilling Bush's vision. They state that Bush "posited Memex as 'a device in which an individual stores all his books, records, and communications, and which is mechanized so that it may be consulted with exceeding speed and flexibility. It is an enlarged intimate supplement to his memory."

Microsoft is taking the word "enlarged" seriously. Included in this application are the following tools (keep in mind that this is not an exhaustive list):
  1. TV capture tool (to capture your TV watching)
  2. SenseCam (to capture images IRL)
  3. GPS import and Map display (to track your movements)
  4. Radio capture
  5. Telephone capture
  6. IM capture
  7. Browser tool (to track your web-browsing)
  8. Legacy email client (you get the idea)
  9. Voice annotation tool
  10. Text annotation tool
Gemmell and Bell's argument for developing such an ambitious application with such wide-reaching curation tools is that since no one can predict when some item might be useful the "safest thing is to simply [sic] keep it all. Everything."

As I'm sure you can imagine, there are a whole host of issues in making such an enormous amount of information accessible, and so far it doesn't sound like Gemmell and Bell have developed a workable solution to these problems. They have been experimenting with hierarchical classification systems and have discovered (surprise!) that such systems quickly become unwieldy when dealing with large amounts of information. They haven't abandoned the idea of hierarchical folder systems, however, and are actually experimenting with "hierarchical classifications that will be developed by others to be downloaded by the user, and which contain extra information such as synonyms and descriptions to ease their use."

Classifications schemes are not the only thorny issue for the MyLifeBits team. Constructing useful metadata capture for video has also proved difficult. In developing a method to garner metadata for video they have suggested harvesting audio and converting it to searchable text so that, e.g., one could search for the name of a given person and jump to a segment of the video where the speaker mentions that person's name. Whether this would necessarily lead the searcher to a segment of the video with the person he/she is looking for is, well, who knows?

Clearly there are a number of issues that need to be ironed out before Microsoft can release this product, but Gemmell and Bell seem intent on doing so. Just this year, they released a book, Total Recall: How the E-Memory Revolution Will Change Everything, to proselytize their vision of total personal digital curation. While this project currently seems a little premature, it may be that in 30 years applications like MyLifeBits could be commonplace. If this is the direction that personal digital curation is going, what are some of the issues presented by such an all-encompassing view of "curation?" Certainly there are going to be copyright issues, classification issues, privacy issues (not to mention dystopian issues! Who wants their life to turn into some monitored, post-apocalyptic version of a Tom Cruise movie?)

Comparison of citation retrieval by search engines and the ISI Web of Knowledge

I continue to be curious about the practice of citation for promotion and tenure, and am always amazed when I haphazardly come across a print article that is so pertinent to our class discussion.

The September 2009 issue of College and Research Libraries (Vol. 70, No. 5, pp. 460-472) contains the article "A Citation Analysis of College & Research Libraries Comparing Yahoo, Google, Google Scholar, and ISI Web of Knowledge with Implications for Promotion and Tenure" by Charles Martell. (Web access is available to ALA members.)

1985 and 2005 studies determined that College & Research Libraries (C&RL) was ranked number one in "journal prestige in terms of value for tenure and promotion" by the Association of Research Libraries (ARL) Library Directors. Given this ranking, Martell determined the following: 1) frequency of citations of C&RL articles, 2) citation retrieval strengths of Yahoo, Google, Google Scholar, and ISI Web of Knowledge, 3) the relevance and applicability of these findings, and 4) classification and quantification of retrieved entries that did not qualify as citations. Every refereed C&RL article from 2000-2006 was used.

Advantages and drawbacks to each search method was discussed. Yahoo and Google are readily available and expansive, but laborious to filter. ISI WOK has a high level of quality and ease of use, but is fee-based and proprietary. Google Scholar has a broad coverage, but still has varying drawbacks: for example, for two of the top ten articles, over half the citations were from China. That said, Google Scholar returned more than twice as many citations as ISI WOK, and of course, the advantages and disadvantages should be weighed before selecting one over the other for citation counting.

It is valuable to note that this study was found to be particularly relevant to academic librarians in tenure-track positions. I found it curious that Martell discovered that deans of ALA-accredited education programs ranked C&RL 11th behind six information science journals and four library science journals. On second thought, perhaps not so curious given the trend of rebranding programs from "library science" to "information science" to incorporate the increasing importance of information technologies.

I think that this data would be very interesting to follow over the long term, in light of the growing interest in, or opposition to, open journals, and also our current discussion of broadening the definition of citations. It would be useful to note spikes/fluctuations in journal and/or citation usage and how it relates to the increasing effectiveness of search engines or initiatives such as the "Compact for Open Access Publishing Equity" (initiated by MIT, Cornell, Dartmouth, Harvard and UC Berkeley). If this data was made readily traceable, would it change the decisions and activities of the stakeholders (authors, promotion committees, publishers, etc.)?

What's the HapMap??

While clicking through the blogosphere earlier this week, I came across a project that seems to be an extension of the National Human Genome Research Institute. It's the International HapMap Project, with the intended focus to find "genes that affect health, disease, and individual responses to medications and environmental factors". The US, Canada, China, the UK, Nigeria, and Japan are collaborating on the research with public and private funds. Because it started in 2002, phase one of the project finished in 2005, with Phase II ending in 2007. I am interested in how accessible the data is and process of the project, which seems fairly documented.

The core of this project is to share information. The website says: "The International HapMap Project is not using the information in the HapMap to establish connections between particular genetic variants and diseases. Rather, the Project is designed to provide information that other researchers can use to link genetic variants to the risk for specific illnesses, which will lead to new methods of preventing, diagnosing, and treating disease."
Very commendable! It appears that this consortium is so expectant that their data will be used, that they offer tips on how to cite the project, especially how to refer to the sample populations (from Nigeria, Japan, Utah, etc). There is also a push to publish, as "researchers are encouraged to publish results based on combining HapMap data with data from other projects, particularly in efforts to find genes affecting a disease or a drug response. Researchers also are encouraged to use HapMap data to publish on the development of novel methods to analyze polymorphism, linkage disequilibrium, and association data."

Why all the good will? Is it because this is funded through the NIH? And must scholars who integrate the HapMap data with their own research make that data available? Interested in how much the data is being used, I looked at the site's publication page, but that hadn't been updated since 2007 when their last phase ended. So I went to GoogleScholar (gasp!) so I could just do a quick look up and found that the HapMap site has been cited 160 times. That could be good!

If I could even slightly understand anything about this project, I might be tempted to download the data (they provide tutorials!). Instead I just looked at the page from whence the files could come and was confused. There is a place to upload your own data, which I'm not sure I think is a particularly good idea, since their research seems quite controlled (270 participants only). I'm assuming the HapMap data and others' data is separate.

The International HapMap Project appears to be an example of a successful data-sharing hard science project, although its proximity in topic and funding agency to the Human Genome Project might also account for that. From an outsider's perspective, HapMap appears to have a socially conscious purpose (healing sick people, yay!) with an integrative approach to the data and publishing.

Challenges of Open-access Journals

An article titled "Ten Challenges for Open-access Journals" appeared in this month's issue of The Scholarly Publishing & Academic Resources Coalition (SPARC) Open Access Newsletter. In the article, Peter Suber explores what he sees as the ten greatest challenges facing open-acess journals.
Three of Suber's challenges deal with disparities between what is intended and what has been achieved, and the remaining seven have to do with doubts about open-access publications.
One of the disparities lies in the Impact Factors (IF) used to measure the impact of a particular journal. While IFs are intended to reveal the quality of the journal, a journal isn't even eligible for an IF until it has been around for 2 years. This is particularly hard on open-access journals because they are relatively new.
A second disparity is the failure of open-access journals to use a creative commons license to allow more than fair use of a journal. Many apen-access journals believe that by making their content free, they are truly open access. Suber argues that institutions are not free to exceed fair use. He suggests that until open-access journals explicitly allow for more uses (such as including the works in a database or archiving copies for preservation) they are not truly open-access.
The final disparity can be found in the gap between quality and prestige. It is impossible for a new journal to be prestigious from its start though it could be of a high quality. Subter argues that quality should matter more than prestige an until it does, open-access journals are at a disadvantage.
Subter moves on to discuss seven doubts that cause challenges for open-access publications. This include doubts about quality, preservation, honesty, publication fees, sustainability, redirection of funds, and even doubts about the strategies for addressing the challenges of open-access journals. These doubts are fairly self-explaintory and Subter proposes ways to deal with each of them. Subter suggests that an awareness and willingness to address these issues is absolutely necessary for the success of open-access journals.
We have previously discussed the need for changes in the publication process and even in the way an article's contribution is measured, but I still feel a bit doubtful that the change will happen anytime soon.

NSF's DataNet Partners so far: Data Conservancy and DataONE

Last spring, the Office of Cyberinfrastructure of the NSF announced its plan to support data preservation and access through the Sustainable Digital Data Preservation and Access Network Partners (DataNet) grant program. The program aims to address "one of the major challenges of this scientific generation: how to develop the new methods, management structures and technologies to manage the diversity, size, and complexity of current and future data sets and data streams." Through funds from the grant, "exemplar national and global data research infrastructure organizations," called DataNet Partners, will do just that. These organizations will combine library and archival sciences, cyberinfrastructure, computer and information sciences, and domain science expertise to provide preservation, access, integration and analysis capabilities over a decades-long timeline. As if that wasn't enough, these organizations should also be able to anticipate and adapt to changes in technologies and user expectations, stay cutting edge, and drive research and development. Oh, and one more thing: "potential applicants should note that this program is not intended to support narrowly-defined, discipline-specific repositories." The NSF is holding out for a hero!

Since the solicitation for proposals, two organizations have been awarded grants - both in August. (It seems that TeraGrid is categorized under the program but was awarded its $32 mil way back in '05.) $3.7 mil was given to Sayeed Choudhury of Johns Hopkins University in order to create the Data Conservancy (DC). Along with a long list of partners from universities and scientific organizations, Johns Hopkins will create "a new model in which libraries regard digital data as a special collection that must be maintained and served like their other collections." The Data Conservancy is concerned with the global carbon cycle and the carbon-climate-human system. The second award was given to William Michener of University of New Mexico and a ton of other partners including the CDL, Amazon, UIC, and Intel. That $12.3 million award is for the creation of DataONE (Observation Network for Earth). DataONE will address four challenges: preserving data, providing access to dispersed data, integrating data, and creating best practices for managing data. Its focus is huge: it will make "biological data available from the genome to the ecosystem; make environmental data available from atmospheric, ecological, hydrological, and oceanographic sources; provide secure and long-term preservation and access; and engage scientists, land-managers, policy makers, students, educators, and the public through logical access and intuitive visualizations."

It's interesting to me that both these DataNet Partners are doing pretty similar things, except the scope of DataONE seems to be much larger than that of the Data Conservancy. DataONE was presented at the 4th International Conference on Open Repositories at Georgia Tech in May. The presentation outlines some of the problems we've discussed with putting multi-disciplinary data in one place, but it purports that DataONE will "provide one-stop shopping for data" and will create "investigator toolkit" that will include nifty things like data visualization and Kepler workflow diagrams (remember those?) The presentation also gives an outline of how DataONE will be technologically implemented using the micro-services approach - or an "unbundled alternative to monolithic systems." Both Coudhury and Michener will speak at the next Educause conference in November, along with our pal Clifford Lynch, in a session called "Initiatives from the NSF's DataNet Program: DataONE and the Data Conservancy."

Sum of Humankind’s Knowledge Available Online 

Great title for an article, huh? It appears that scientific research in Europe is becoming available to the public and other researchers through the DRIVER (Digital Repository Infrastructure Vision for European Research) Search Portal, which serves as a door to European Open Access research. The idea is to create a “library of libraries,” where individuals can readily search across repositories throughout Europe.

European researchers have created D-NET, software that links “information collected on diverse computer platforms, using legacy software which can still ‘talk’ or work with older systems in more than 25 European languages.” The software regularly harvests publications to provide access to information from numerous repositories. Visitors to the site are encouraged to use the software to set up other portals and are pointed toward examples of test cases where other groups are using it. The DRIVER portal site reports the following information about the collection:

• approximately 1,000,000 documents
• found in journal articles, dissertations, books, lectures, reports, etc.
• harvested regularly from more than 200 institutional or thematic repositories
• from 23 European countries
• in 25 languages.

The portal allows you to search broadly across repositories, but also limit your search by repository, language, document type, and date. At present the focus is on textual documents, but the idea is extend into other types of media. Current goals are to continue building the infrastructure and increasing the number of participating repositories. Another development of the initiative was the creation of Guidelines for Repository Managers, a document that outlines how to make content compatible with DRIVER. This type of document is an instructional tool to ease the process of increasing interoperability. The initiative is funded by the Research Infrastructure priority of the EU’s Sixth Framework Programme for research.

As a “library of libraries,” the DRIVER portal seems like a super repository that allows you to access a larger network of published materials through a single search, but I do not believe it goes beyond published works to data sets. However, the possibility for expansion into broader scholarly communication is there and the openness to using the software may allow for other uses within the scientific community.

Earth Simulator Center

Earth Simulator Center

In order to be able to visualize our readings for class this week, I selected one of the projects mentioned: the Japanese Earth Simulator Center and found it very interesting, innovative and a good example of using digital images to aid research.

The Earth Simulator Center (ESC) is part of the Yokohama Institute for Earth Sciences in Japan and is research center that appears to be a carbon copy of technologists and researchers working together create data. They research ocean temperatures, storm patterns, atmospheric conditions, and seem to track marine fluctuations. They take the data and construct different types of simulation diagrams. ESC states that their mission and statement for doing this is to "build a harmonius relationship between the Earth and human beings...through various areas of research and development" (Jamstec.g.jp).

The data gathered is used construct elaborate digital pictures of natural forces as they develop. There are pictures demonstrating geothermic changes with the presence of urban communities and without.

ESC also prides itself in working collaboratively, both on a domestic (in Japan) and on an international level with other research groups, mainly with marine scientists, and scientists that study clouds, weather, and other atmospheric forces, scientists that track changes on earth due to civilization or climate change. They have a feature on their website that lists all the projects that they're involved in with a link that either takes you another website or a webpage with their data on a pdf. However, most of the domestic project sites are written in Japanese and it is difficult for me ascertain what exactly the PDF reports could be.

I'm no scientist, but the pictures were easy for me to grasp the implications and meaning behind their constructions. They also exemplify a lot of qualities that are listed as progressive in building a science cyberinfrastructure community, such as collaboration, the cross disciplanary approach to research, and showing clearly how this technology can benefit scientific research.

SWORD2 and Infrastructure

This week I read JISC's SWORD2 Project Final Report [pdf] in conjunction with SWORD: Cutting Through the Red Tape to Populate Learning Materials Repositories, a February 2009 article introducing and explaining the protocol.

SWORD stands for Simple Web-service Offering Repository Deposit and is designed to "lower the barriers to deposit." It's on the default install of DSpace and is an optional interface for Fedora, EPrints, and IntraLibrary. Various repositories use the protocol and there are client-side applications that utilize it as well: Feedforward is a desktop application for personal information management, OfficeSWORD plugs into Microsoft Office so one could upload data to a repository directly from the program, and there's a widget and Facebook app for the protocol too.

The "Cutting Through the Red Tape" article briefly explains the underlying technology (it operates off of the Atom Publishing Protocol) and applications for use. It does this through four hypothetical use scenarios. The first details a university lecturer who can deposit notes, syllabi, transcripts and such simply by dragging them to an icon, and then a refinement he only needs to press a button in the application where he's creating the document (MS Office, say) and it's deposited. There's an optional form that could pop up for extra metadata he could attached (comments, etc.) but all the metadata the repository needs is already discernible (it's his work machine so they know who he is, time of deposit, format, size, title, and so on).

The second case examines two public sector educational resource providers (a social services one and a health one) that need to share data from their repositories. They're able to do this by using deposit services that use SWORD, so that every ingest with SWORD creates a set of metadata for discovery, search and display.

The third case elaborates upon direct deposit from the content creation tool, and the fourth demonstrates how one could receive various feeds from Feedforward, select a certain group, and deposit that into a repository.

The SWORD2 report essentially describes a successful interoperability update of the protocol and notes that advocacy efforts for SWORD should continue since it may be reaching "critical mass" and viability as a standard.

SWORD and the cases presented are interesting because I feel like it represents some of the infrastructure we've discussed. Particularly in the case OfficeSWORD, one could see how "invisible" it could become: just another button on the panel along with 'Print' and 'Save' that deposits the work for circulation in a repository. The cases where the user is able to deposit directly from the content creation tool are therefore really encouraging. It's obvious that a significant barrier to digital repositories is the click-count or other simple work that needs to be done for the deposit; if software or plugin developers only need to design a function for SWORD deposit, this could significantly change.

In terms of cyberinfrastructure, TCP/IP and HTTP are protocol and protocol sets that are so embedded in how "things just work" it's frequently difficult to imagine them actually being developed and adopted. No user ever has to deal with them. But it's clear we need similarly massively successful protocols fueling the cyberinfrastructure NSF calls for.

Tuesday, October 6, 2009

Science Big & Small

Salo, Dorothea. "Costs and Service Models for Data Curation." The Book of Trogool (Blog). Posted September 19, 2009. Available at http://scienceblogs.com/bookoftrogool/2009/09/cost_and_service_models_for_da.php (Accessed 10/06/2009).

This week I want to talk about a blog post (rather than an article) that caught my attention: Dorothea Salo's "Costs and Service Models for Data Curation." In her post, Salo discusses her concerns about the differences between how "Big Science" and "small science" will fare in matters of data curation and preservation. Perhaps because it is a blog post and not an article, Salo never quite defines "Big Science" and "small science," but from context and some reading around on the internet, it seems that "Big Science" generally refers to large multiple investigator projects, with big budgets, that tend to involve advanced (and expensive) technologies with large centralized labs (such as CERN), whereas "small science" involves a single investigator or a small team, usually working at a single lab on an independent research program whose funding is dependent on smaller limited-term grants. Salo sees the differences between big and small science as resulting in important disparities that affect data curation.

The problem as Salo, an academic librarian invested in e-research, sets it out is such: Big Science produces large, generally homogeneous data in huge quantities. However, because it's all part of one project, data standards evolve quickly and procedures can be institutionalized and set for the entire project. Thus once you get past the problem of storing huge amounts of data (and Salo is confident that they are coming up with solutions on this front), the human resources cost for curation is relatively small per terabyte of data. Salo characterizes small science, on the other hand, as tending to have ad hoc procedures without real data standards (since there are not enough people working on similar data who are willing to share that data and work together to generate standards). This is a problem because, without standards, each of small science's highly heterogeneous pieces of data will require "
individual attention if it is to be adequately described and future-proofed." Accordingly, small science has the potential to have a frighteningly high human resources cost per terabyte.

This is significant because Salo intuits that small science, as a whole, has the potential to generate more research data than Big Science. It is also significant because, as she argues, all research data is potentially equally important since you can't know in advance where the groundbreaking, paradigm-altering insights will come from (especially since those breakthroughs may come from subsequent mash-ups of data). So, small science, which is least equipped to pay for long-term data curation, is likely to have the highest costs.
This puts its "equally important" data at risk of getting lost in the data deluge (Salo explains in the subsequent comment posts that she has already seen instances of this type of loss). Salo is skeptical about the model of institutional cost-recovery cyberinfrastructure since she thinks Big Science will opt out (as she also explains in a subsequent comment) in favor of doing their own data curation with an embedded librarian whereas whatever money small science has to contribute may not be enough to adequately fund cost-recovery operations. The question that emerges for Salo is what type of business model will help us to address in a more equitable fashion the data curation disparities between Big and small science.

The comments that follow Salo's post also raise important questions. One of the respondents, Sayeed Choudhury, Director of digital curation at JHU (see Megan's email shout-out from earlier today), questions whether Big Science data really is as homogeneous as Salo presumes it to be. He also takes her to task in general (though, in the nicest way possible) for working from assumptions rather than concrete facts. He notes, for example that medical science is presumably "small science" but that the NIH has much more funding to distribute than the NSF. Choudhury suggests that libraries should become more involved in scientific data curation so that they can actually discover (rather than speculate about) what it entails in order to better build their infrastructures.

Choudhury's criticism that we need to actually compile and generate data on these issues is an important one. Salo's questions are provocative and, I think, important, which is why I've chosen to write about them here, but it is also unclear whether data curation in Big and small science really does work in the manner in which she supposes it does. Someone want to do a fact finding study?

In Atkins et al's NSF report "Revolutionizing Science and Engineering Through Cyberinfrastructure", the authors recognize that there is a "significant need [...] in many disciplines for long-term, distributed, and stable data and metadata repositories that institutionalize community data holdings" (42). In their imagining, these would provide tutorials on data formatting and quality control as well as offer tools for data preparation. This does seem like a potentially useful tactic for coping with data curation costs, but the question remains whether these tutorials and tools will really be equipped to cope with the heterogeneity of data. It also seems that these tools and standards would need to be engaged during the planning stages of researchers' projects if they wanted to avoid prohibitive human-resources costs further downstream. These are potentially important propositions, but we won't know how well such resources will serve small science until the tools are up and running and being widely used.

Cyberinfrastructurasaurus


This week I found an example of a scientist who is actively making use of cyberinfrastructure in his research. Allister Rees is a paleontologist/paleo-climatologist at the University of Arizona who coordinates the GEON PaleoIntegration Project (PIP). The PIP consists of a combination of five separate databases, available online, which can be applied individually or in any combination of the five at once. (The GEON PIP also features a free public portal, but you do have to request an account in order to have one authorized. What's really interesting is that the request form asks whether or not you want to be approved to post data to the databases, so there must be at least some data sharing going on. I did not request an account for fear of being an annoyance or time-waste, not being at all qualified to interpret paleontological data myself, but a preview tutorial is available here: PIP.) At any rate, the databases cover various eras of lithofacies (which I take to be referring to large patterns of sendimentary layers), the distribution of and data regarding certain types of climate sensitive rock and sediment from which the paleo-climate can be inferred, dinosaur distrabution, and plant fossil distrabution. The databases are then paired with paleographic maps and mapping tools.

In this brief article on the project, Rees asserts that his use of the PIP databases led him to conclude that dry, savanna-like climates tend to be the most likely to preserve dinosaur fossils, a conclusion which he claims Paleontologists had not yet realized since he was the first to directly look at paleo-vegetation diversity and the presence of dinosaur fossils in an overlay sort of function. Rees also feels like the PIP will be valuable in the study of mass extinctions and climate change.

I found the project interesting both because it illustrates the value for scientists in being able instantly to compare data from multiple specific fields courtesy of cyberinfrastructure and also because the presence of free, online databases does suggest that at least some scientists are willing to share their data.

NSF + COL = OOI (Ocean Observatories Initiative)

Once again the National Science Foundation is putting its hand into another effort to build scientific cyberinfrastructure, this time turning its attention to what lies under the sea. On Monday it was announced in a press release at the NSF website that the NSF and the Consortium for Ocean Leadership (COL) signed a cooperative agreement to build and operate the Ocean Observatories Initiative, which aims to “provide a network of undersea sensors” to observe the ocean and collect data like never before. From this article, it sounds like the NSF and the COL want to do for the seas something similar to what the National Virtual Observatory does for the sky.

This network of hundreds of sensors will be placed at several coastal, open-ocean, and seafloor locations, and will continuously collect data about “complex ocean processes” such as climate variability, ocean circulation, and ocean acidification. The Ocean Observatories Initiative looks to up the ante in data collection in terms of both rate and scale. And with this data, scientists studying the oceans hope to gain more of an understanding about how the oceans affect and interact with the earth and atmosphere.

This underwater initiative looks to take five or more years. The initial phase is the major construction phase, which includes “production engineering and prototyping of key coastal and open-ocean components (moorings, buoys, sensors), award of the primary seafloor cable contract, completion of a shore station for power and data, and software development for sensor interfaces to the network.” According to the article, the OOI should begin to see data flow by 2013, and the full system is projected to be up and running by 2015.

This is yet another effort that pours money into tools and technology to generate data, but the press release does not say much about what will be done with the data once it is collected. The article does say that the continuous data flow collected by the ocean sensors will be “integrated by a sophisticated computing network, and will be openly available to scientists, policy makers, students and the public.” Though the article fails to mention how exactly this will be made possible; I’m assuming that is because the project is still heavily in a development phase.

The OOI is a complex project, involving lots of cooperation from many different institutions, including Oregon State University, the Scripps Institution of Oceanography, the University of Washington, and University of California at San Diego (who is in charge of implementing the cyberinfrastructure component... though again the details are somewhat lacking).

So yet again the NSF is at the forefront of attempting to develop cyberinfrastructure, and once again a scientific field looks to be on the edge of a breakthrough when it comes to research and technology. It will be several years before we will start to really see the results of these efforts, and whether or not this project will be a success along the lines of the Virtual Observatory or another failed or incomplete attempt as we have encountered before in our readings. It sounds pretty exciting though, so I hope it works out.

Data-driven Scholarship in the Sciences and Humanities


In 2007, a piece entitled "The Virtual Observatory and the Roman de la Rose: Unexpected Relationships and the Collaborative Imperative" was posted on Academic Commons. It was authored by Sayeed Choudhury, the Director for Library Digital Programs and Hodson Director of the Digital Knowledge Center for the Sheridan Libraries at Johns Hopkins University, and Timothy Stinson, a post-doc fellow at Johns Hopkins with dual appointment in the Digital Research Curation Center and the Department of English. The article is concerned about the impact of cyberinfrastructure on the way that scholarship is conducted. The authors cite the Virtual Observatory and the Roman de la Rose Digital Library as two examples from different fields that exhibit the benefits afforded to scholarship thanks to cyberinfrastructure. The Virtual Observatory is a resource that brings together data from ground- and space-based telescopes, allowing scientists to find, retrieve, and analyze this astronomical data. The Roman de la Rose Digital Library, a collaborative effort between the Sheridan Libraries of Johns Hopkins University and the Bibliotheque national de France, is a digital library the goal of which is to create and house digital copies of all extant manuscripts of the Roman de la Rose.

Most would agree that the hard sciences are moving toward a data-driven model of scholarship, as evidenced by the Virtual Observatory project. However, the authors argue that the humanities, too, are beginning to be able to exploit (thanks to cyberinfrastructure) "data" for research in much that same way as the sciences. The parallel might at first seem not to fit, until one considers that data for the sciences and data for the humanities are two different things. For the sciences, data consist of sensor readings, scanner output, measurements of various kinds, and other things that are largely numerical in nature. However, humanities scholarship is based on a different, less straightforward type of data. As the authors of the Academic Commons article say, humanities materials can be considered data-rich in a variety of ways. For example, a single manuscript of a medieval text like the Roman de la Rose can contain many types of data in the form of illuminations, "artwork"/images, marginalia, annotations, as well as the semantic and linguistic data contained in the text itself. The creation of large digital libraries of this type is, for the humanities, analogous the creation of huge, complex datasets that make up digital curation projects like the Virtual Observatory, even though they seem comparatively data-poor.

The authors of this article further state that by exploiting resources such as the Roman de la Rose Digital Library, the humanities are poised to enter into the kind of collaborative scholarship that characterizes the hard sciences. "Digital tools," they assert, "are allowing us to capture, manipulate, and examine books and their data in ways that are revolutionizing the humanities." Digital versions of entire libraries can be created, and the contents of libraries that have been distributed over time can be virtually reassembled, creating new possibilities for scholarship. The authors recommend that humanities scholars work to define their own needs in the realm of cyberinfrastructure, and to create the specifications for and build the tools that they require for research.

The authors conclude by recognizing the imperative role that libraries and librarians will play in the future of these new data sets. While the role of the library in preserving the physical objects on which something like the Roman de la Rose Digital Library is based will remain vital, the library's role in preserving and making accessible the digital data that is created is one that is even more vital, all the more so because "the datasets are often not as highly regarded by libraries."

Monday, October 5, 2009

XML, JSON, and the importance of XSLT

This week Nature posted about their new search function meant for APIs. This new search interface will return results either using the tried and true (more later) XML or the newer and more streamlined JSON.

XML
I'm going to level with you guys. I secretly hate XML a little bit. When I heard about HTML5 and that XHTML was dead (sort of) I was psyched. Then I heard about JSON, and I had a bonafied I-told-you-so moment. (Quick summary so no panic ensues, HTML5 is combining many of the XHTML working groups deliverables into their own; you can still structure/nest tags neatly in HTML5, but it allows more flexibility; if there is one mistake on the page, the page won't fail.)

Then today I got the Flipsider email about the Federal Register in XML and I realized that XML isn't dead, even if aspects of it are going out of style. Not only are there many projects out there already using XML, but new ones will continue to pop up, no matter what w3c or JSON try.

In addition, XML is actually quite good at document mark-up (my fingers hurt just typing this.) I even like the strictness of syntax in XML (close inner tags then outer tags in the right order), but I think I (and, I venture to say, many information professionals) have confused XML as meant for something that it isn't meant for. XML is not a programming language. It is a document mark-up language that can be used to transport documents between programs. XML does not have a good way to deal with images or non-text data objects, although it claims to.

XML needs XSLT or other intermediary languages to be interpreted by browsers or other client-side applications.

JSON

JSON is a JavaScript data-interchange format. It is similar to XML in that it is meant as an envelope for items to be transported between programs, but it is simpler, and it was intended for use with programming languages. This makes it easier for use in browsers.

The best description of the difference between XML and JSON I found was at http://json.org/xml.html, "XML requires translating the structure of the data into a document structure. This mapping can be complicated. JSON structures are based on arrays and records. That is what data is made of. XML structures are based on elements (which can be nested), attributes (which cannot), raw content text, entities, DTDs, and other meta structures."

Both can be used, and XML can be transferred to JSON, the Nature approach is probably the best one.
Basically, I think I have to stop hating XML, but I can be grateful that there is something easier to program widgets with.

I apologize for the acronym party.