Wednesday, November 4, 2009

Ex(ml)tra, extra, read all about it!


This week I found a brief article detailing the metadata problems and solutions encountered by the Library of Congress when it participated in the National Digital Newspaper Program, a twenty-year initiative to create a national repository of historical newspapers in digital form. The end goal of the project was to have a complete and searchable online bibliography of newspapers published in the U.S. from 1690 onwards. In addition to this, select local newspapers of historical significance would be converted to digital form for instantaneous public access.

Several problems arose in creating a standard metadata for the project. First of all, print newspapers themselves do not fit well into any of the traditional cataloging standards. This is especially true when one tries to accurately and comprehensively describe newspapers from all fifty states and across drastically different historical eras. (17th-century spelling and abbreviation are particularly dicey for modern readers.) Second, most of the historical newspapers being covered by the project no longer exist in print form but were only saved as microfilm, introducing another form in the provenance. Finally, the NDNP wanted a system that was going to allow complete interoperability between the federal repository and individual state digital repositories as well as among the state repositories.

Ultimately, the NDNP elected to use the METS standard, entered in XML. Four separate METS elements were used: title document, issue document, page object, and reel document. The first three refer to the original physical form of the paper. XML is flexible enough to be able to incorporate different types of issue numbers while remaining interoperable, which was key for the program. Murray views the page object as the most important as he asserts that the page is the basic information-containing block of a newspaper. The reel document only applies to those papers digitized from microfilm but represents important administrative data.

I found this article interesting mainly for three reasons. First, it involves a system to provide more access to historical documents, including colonial era documents, which is a goal ever near to my heart. Second, I work in the serials unit at the Benson Collection and have much first hand experience (and frustration) at the incredible amount of variance in newspaper conventions. Third, I liked the fact that the NDNP came to a relatively straight-forward and simple solution, while maintaining interoperability. (Though, of course, it helps that there was central oversight in this project...) All in all, I thought it was a good example of the benefits conferred by standards as well as not making the solution harder than it has to be.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.