Stehno, Birgit, Alexander Egger, and Gregor Retti. "METAe-Automated Encoding of Digitized Texts." Literary and Linguistic Computing, 18, No.1 (2003): 77-88. Available at: http://llc.oxfordjournals.org/cgi/reprint/18/1/77. (Accessed 11/1/2009).
This article describes how the Austrian-based METAe project applied METS (Metadata Encoding and Transmission Standard) to encode automatically extracted metadata from page images, especially metadata describing document layout. Like OCR engines which extract text from image files, the METAe engine extracts layout elements (such text or graphic blocks) by relying on formal rules and syntactical principles. The project built and compiled a recognition model that would map elements in the physical structure of page layout (text blocks of differing sizes and graphics blocks) to logical ones (e.g. paragraphs, titles, footnotes). This would then create an ALTO ('Analysed layout and text object') file which is an XML file that consists of both the layout structures and the full text of a book page. The ALTO file and information formatted in a variety of metadata standards such as Dublin Core and DIG 35 are then incorporated into a METS schema, which (if I'm understanding correctly) serves as an outer wrapper that provides the structural map and holds all of the metadata for the object.
In devising the METAe engine, they decided to use METS rather than TEI (Text Encoding Initiative) for their encoding since TEI was "far too inexplicit for the purpose of automated recognition" (9). This struck me as an interesting point since if TEI is to be the encoding standard for text in the future, it will need to be more amenable to automatic application. METS was also preferred as it allowed for METAe to add metadata at any logical level and to include metadata from different formats and standards, including pointing to metadata external to the METS document. For example, the METAe project can use a "DMDID" (Descriptive MetaData Identifier) attribute to generate tags linking a journal issue to its appropriate MARC record on a web server while also using Dublin Core tags to provide descriptive metadata for each article/contribution.
What this article by Stehno, Egger, and Retti does not really address are the difficulties (and there must have been some) that the METAe project encountered in developing its engine. It seems extremely useful to have an engine that can map the physical structure of a page onto logical structures, but certainly there must be structures that it is better and worse at recognizing. I've also had some difficulty in discerning precisely what happened to METAe--it seems to have lived on as ALTO rather than as the METAe tool itself. As the METAe website (http://meta-e.aib.uni-linz.ac.at/index.html) details, the project ran from September 2000 - September 2003 with partial funding from the European Commission. The website also announces the metadata engine has been marketed as digitization software under the name docWorks/METAe Edition, however, though the docWorks page exists and seems to offer a product that performs the tasks that METAe does, there is very little mention of METAe on the site. The one mention I was able to locate was a notice that as of August 2009, the Library of Congress has taken over maintenance of the ALTO XML schema from CCS Content Conversion Specialists GmbH (the company which produces docWorks). This appears to be a sign at least of ALTO's success as a schema, as LC has created a new ALTO editorial board to "help shape and advocate usage of the standard."
I guess the questions that I'm left with after reading about ALTO and METAe are ones about the relationship between standards and schemas and the tools used implement them. To what extent do the constraints of the tools with which schemas were initially partnered affect the future use of those schemas, even after a given tool has been put aside or altered? Can schemas created within the context of automation equally useful when used for hand encoding or human quality control?
Monday, November 2, 2009
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.