In her article Innkeeper at the Roach Motel, Dorothea Salo writes about the problems of institutional repositories. While she discusses many different issues, I am going to focus on the causes she lists for insufficient metadata in institutional repositories.
First in her list of causes is the issue of trust. Many people are hesitant to trust librarians with their data because they believe their lack of subject area expertise will result in problems with metadata management. Salo does not discuss wether or not this belief has merit, only that it inhibits researchers in specific disciplines (such as the hard sciences) from trusting their work to librarians. Instead they trust their work with science experts who frequently are not working to solve long-term preservation issues.
Additionally, the different requirements of publishers make placing data into a repository difficult. One such requirement is that the meta-data have a set phrase which includes a link. The problem is that some of the popular institutional repository software, for example DSpace, does not have that particular functionality, further contributing to some author's hesitancy to use institutional repositories.
The information professional put in charge of an institutional repository is also at a disadvantage if they cannot program. They are left with the basic version of their repository without any patches or plugins that are not already a part of the system. Frequently the task of applying dublin core metadata to a particular work does not even fall to the information professional, but instead it falls to the author of a particular work. They information professional is left trusting the authors themselves to invest the time needed to ensure proper metadata entry, which rarely happens. The result is inconsistent use of metadata throughout the repository, greatly decreasing the usefulness of the works that are in it.
Repository systems also contribute to the problem of having good metadata in a repository because they are frequently difficult to use. They often overestimate the ability of the user (who is not always an information professional) and expects them to know what fields are required and even what some more obtuse fields mean. Salo does mention that EPrints does a better job of handling metadata than other repository systems because it displays metadata fields based on what type of content is being deposited, making the process more user-friendly.
The final cause of insufficient metadata in an institutional repository mentioned is that repositories are still not seen as a priority to many libraries. Without proper funding and sufficient staffing to maintain and oversee the repository and the metadata going into it, information professionals don't have a chance of ensuring the repository is useful. The library ensures the failure of a repository when they don't allot adequate time and money to its maintenance.
While I believe all of Salo's points are valid and the way metadata is handled should be reconsidered, I still believe the repositories currently face a larger problem, namely that authors do not want to put their work into the repository at all. Still, if we are going to put data into the repository, we should be able to find it again later and that necessitates that we take a good look at the way we are implementing metadata.
Wednesday, November 4, 2009
Ex(ml)tra, extra, read all about it!

This week I found a brief article detailing the metadata problems and solutions encountered by the Library of Congress when it participated in the National Digital Newspaper Program, a twenty-year initiative to create a national repository of historical newspapers in digital form. The end goal of the project was to have a complete and searchable online bibliography of newspapers published in the U.S. from 1690 onwards. In addition to this, select local newspapers of historical significance would be converted to digital form for instantaneous public access.
Several problems arose in creating a standard metadata for the project. First of all, print newspapers themselves do not fit well into any of the traditional cataloging standards. This is especially true when one tries to accurately and comprehensively describe newspapers from all fifty states and across drastically different historical eras. (17th-century spelling and abbreviation are particularly dicey for modern readers.) Second, most of the historical newspapers being covered by the project no longer exist in print form but were only saved as microfilm, introducing another form in the provenance. Finally, the NDNP wanted a system that was going to allow complete interoperability between the federal repository and individual state digital repositories as well as among the state repositories.
Ultimately, the NDNP elected to use the METS standard, entered in XML. Four separate METS elements were used: title document, issue document, page object, and reel document. The first three refer to the original physical form of the paper. XML is flexible enough to be able to incorporate different types of issue numbers while remaining interoperable, which was key for the program. Murray views the page object as the most important as he asserts that the page is the basic information-containing block of a newspaper. The reel document only applies to those papers digitized from microfilm but represents important administrative data.
I found this article interesting mainly for three reasons. First, it involves a system to provide more access to historical documents, including colonial era documents, which is a goal ever near to my heart. Second, I work in the serials unit at the Benson Collection and have much first hand experience (and frustration) at the incredible amount of variance in newspaper conventions. Third, I liked the fact that the NDNP came to a relatively straight-forward and simple solution, while maintaining interoperability. (Though, of course, it helps that there was central oversight in this project...) All in all, I thought it was a good example of the benefits conferred by standards as well as not making the solution harder than it has to be.
Tuesday, November 3, 2009
Metadata for Linguistics
In keeping with this week's topic I wanted to see what metadata has been implemented for linguistics data. As mentioned in some of our readings, the problem of local nomenclature and linguistic barriers are glaringly obvious when dealing with the study of language itself, so I wanted to see how this has been handled. As a side note, when dealing with the web, finding resources in non-roman scripts has been given a boost recently by the decision to approve the creation of top-level domains which use internationalized domain names (IDNs), although unofficial URLs in Thai script have been available for some time from ThaiURL.
As this paper describes, in December 2000 the OLAC (Open Language Archives Community) was founded to promote online sharing of language resources. Problems specific to this domain were becoming apparent online, as language names can have several different romanizations (one example: Fadicca, Fadicha, Fedija, Fadija, Fiadidja, Fiyadikkya, Feddica- all the same language!) and different names within the same country, as well as different names between countries (Deutsch vs. German). In addition, language names can change over time, and different social groups may have different preferred names. The end result was a searching nightmare with very low recall.
The solution proposed is simple and uses tools that should be familiar to us all. Linguists are using Dublin Core and the OAI to bring resources together. There are a limited number of extensions to DC such as allowing every tag to be tagged with the language it is written in (subject.language), as well as tagging the metadata record itself with multiple versions in other languages. In addition there are domain-specific tags that list the functionality of the resource (transcription, annotation, lexicon, etc.) and what kind of specification it includes (phonetic, prosodic, morphological, etc.)
Although RFC 3066 is the DC and web standard for specifying language names, SIL's Ethnologue is much more in-depth: for example, the Karen languages do not even appear in RFC 3066. In addition, Ethnologue provides the ability to identify different levels of language groupings, such as families.
We can see from this example that it is not always difficult to create useable metadata for a specific domain: using a few pre-existing domain-specific tools and some pre-existing standards with extensions we can accomplish a lot. To look at the current situation, development of the OLAC is on-going, and a tutorial on language archives was held in January of this year. I look forward to future developments in the intersection of these disciplines.
Pirates of Academia
This is the same old song of the digital age when you remove data from its physical container someone somewhere is going to put it online. This is true for music, books, and video games so we should not be surprised that scholarly publishing has been drawn in as well.
An article published on the Internet Scientific Publications www.ispub.com details a study conducted over a six month period examining the file sharing activities that occurred on a single site devoted to medical professionals. What they found was an Electronic Library section where users could go and request that papers be posted. Each user was allowed three paper requests a day and 83% of articles requested were fulfilled. Over the six month period of time 5,251 articles were posted with a total viewer ship of 23,461.
This is a very interesting finding and I really want to know how wide spread this sort of practice is and how it really affects the bottom line of scholarly journals. The journals in this case are in a different position in comparison to the record labels who have not fared so well in the fight against pirates and the emergence of new technology, but the audience for the journals is different than that of the labels.
I for one have never paid directly for an academic article, and I doubt seriously that a any of you reading this have either. I pay the university a good sum of money and then through the libraries the articles I am interested in are made available. The universities, in this model, are the prime customers and they could never pirate access to a journal. The article states that roughly 703,830.00 was not paid in access fees over the six month period that the study follows. This may be true but many may have used this forum to get access to article because it was easier than going to their local library and retrieving the same material there.
This is an interesting study but, in the end I don’t believe it tells us anything that we did not know. We live in an age where all digital data is shared and this has finally moved to the realm of scholarly publishing. Attempts to secure this data will be subverted and overcome, but my guess is in this case the big journals get the majority of their income from institutional sources and will not reconsider their current model based on this activity.
An article published on the Internet Scientific Publications www.ispub.com details a study conducted over a six month period examining the file sharing activities that occurred on a single site devoted to medical professionals. What they found was an Electronic Library section where users could go and request that papers be posted. Each user was allowed three paper requests a day and 83% of articles requested were fulfilled. Over the six month period of time 5,251 articles were posted with a total viewer ship of 23,461.
This is a very interesting finding and I really want to know how wide spread this sort of practice is and how it really affects the bottom line of scholarly journals. The journals in this case are in a different position in comparison to the record labels who have not fared so well in the fight against pirates and the emergence of new technology, but the audience for the journals is different than that of the labels.
I for one have never paid directly for an academic article, and I doubt seriously that a any of you reading this have either. I pay the university a good sum of money and then through the libraries the articles I am interested in are made available. The universities, in this model, are the prime customers and they could never pirate access to a journal. The article states that roughly 703,830.00 was not paid in access fees over the six month period that the study follows. This may be true but many may have used this forum to get access to article because it was easier than going to their local library and retrieving the same material there.
This is an interesting study but, in the end I don’t believe it tells us anything that we did not know. We live in an age where all digital data is shared and this has finally moved to the realm of scholarly publishing. Attempts to secure this data will be subverted and overcome, but my guess is in this case the big journals get the majority of their income from institutional sources and will not reconsider their current model based on this activity.
The Sloan Digital Sky Survey
This week I tried to stay on topic and I tried to find just one article and stick to it but that wasn't possible. I looked at the different places the Sloan Digital Sky Survey (SDSS) data is used. I've been very interested in the SDSS because it is the largest sky survey so it's possible to rely entirely on the data from SDSS to model large sections of the sky. However, the services I tried out this week sometimes pull data from other sky surveys to supplement SDSS.
On the SDSS and SDSS Sky Server site there are a few different ways to access images from the survey, there are many photos throughout the sites that can be enlarged including some image galleries and a listing of Famous Places. The way to look at the original data is to locate the original Flexible Image Transport System (FITS) files. There's currently seven data releases on the SDSS site and for every data release it's possible to either select coordinates for viewing or navigate through several levels in a directory to get to a list of FITS files. To view the FITS files download a FITS viewer. Not surprisingly, it's difficult to engage the public with the original data from this survey.
It's probably an understatement to say that astronomers have complicated ways to describe astronomical objects. After creating multiple metadata schemas to organize the massive amount of data from telescopes, astronomers are still having a hard time with descriptions. More accurately, there's images and data about the images, but it still isn't processed and understood by humans. Since there are so many astronomical objects that need to be classified, a team of astronomers, cosmologists and a few other experts created Galaxy Zoo. In order to understand how galaxies are formed and how to describe them, Galaxy Zoo has crowdsourced JPG images from SDSS. There is absolutely nothing technical on the Galaxy Zoo site, they don't mention FITS or metadata, all they are interested in are people's answers to simple questions about JPG images of galaxies.
At first I was skeptical of how successful Galaxy Zoo could be, but after reading what they've discovered I've embraced their practices.
Google Sky is another place where the SDSS images are used, the March 2008 NVO Newsletter provides some details on Google Sky. Google Sky started as an extension to Google Earth and in March 2008 it was also launched at google.com/sky. Google Sky is marked up with Keyhold Markup Language (KML), which means that anybody can participate in tagging. You can also convert a FITS file to KML and upload it or even convert a VOTable to KML with the NVO's Visual Integration and Mining Tools.
I looked around for information on how SDSS was ingested into Google Sky and if there's bad KMZ files in Google Sky, but I wasn't able to find anything.
A service that is similar to Google Sky is sky-map.org. I think this collection is curated even better than Google Sky, the maps of astronomical objects are much more compelling than the ones Google Sky selects and there's more insight to how the SDSS collects data. For instance, in this map the survey "stripes" are visible. There isn't a lot written about sky-map.org aside from what's found in their wiki and a few old articles from magazines.
A fourth place that uses SDSS is the Microsoft WorldWide Telescope (WWT). This software coplies with International Virtual Observatory Alliance standards which means that it uses the Astronomy Visualization Metadata (AVM) standard, so any user can tag in WWT. I haven't used WWT but I read on Wikipedia that it made an astronomer cry.
On the SDSS and SDSS Sky Server site there are a few different ways to access images from the survey, there are many photos throughout the sites that can be enlarged including some image galleries and a listing of Famous Places. The way to look at the original data is to locate the original Flexible Image Transport System (FITS) files. There's currently seven data releases on the SDSS site and for every data release it's possible to either select coordinates for viewing or navigate through several levels in a directory to get to a list of FITS files. To view the FITS files download a FITS viewer. Not surprisingly, it's difficult to engage the public with the original data from this survey.
It's probably an understatement to say that astronomers have complicated ways to describe astronomical objects. After creating multiple metadata schemas to organize the massive amount of data from telescopes, astronomers are still having a hard time with descriptions. More accurately, there's images and data about the images, but it still isn't processed and understood by humans. Since there are so many astronomical objects that need to be classified, a team of astronomers, cosmologists and a few other experts created Galaxy Zoo. In order to understand how galaxies are formed and how to describe them, Galaxy Zoo has crowdsourced JPG images from SDSS. There is absolutely nothing technical on the Galaxy Zoo site, they don't mention FITS or metadata, all they are interested in are people's answers to simple questions about JPG images of galaxies.
At first I was skeptical of how successful Galaxy Zoo could be, but after reading what they've discovered I've embraced their practices.
Google Sky is another place where the SDSS images are used, the March 2008 NVO Newsletter provides some details on Google Sky. Google Sky started as an extension to Google Earth and in March 2008 it was also launched at google.com/sky. Google Sky is marked up with Keyhold Markup Language (KML), which means that anybody can participate in tagging. You can also convert a FITS file to KML and upload it or even convert a VOTable to KML with the NVO's Visual Integration and Mining Tools.
I looked around for information on how SDSS was ingested into Google Sky and if there's bad KMZ files in Google Sky, but I wasn't able to find anything.
A service that is similar to Google Sky is sky-map.org. I think this collection is curated even better than Google Sky, the maps of astronomical objects are much more compelling than the ones Google Sky selects and there's more insight to how the SDSS collects data. For instance, in this map the survey "stripes" are visible. There isn't a lot written about sky-map.org aside from what's found in their wiki and a few old articles from magazines.
A fourth place that uses SDSS is the Microsoft WorldWide Telescope (WWT). This software coplies with International Virtual Observatory Alliance standards which means that it uses the Astronomy Visualization Metadata (AVM) standard, so any user can tag in WWT. I haven't used WWT but I read on Wikipedia that it made an astronomer cry.
Making DRM More Usable...or Profitable?
A. Arnab and A. Hutchison, “Fairer usage contracts for DRM,” in Proceedings of the 5th ACM workshop on Digital rights management, 2005, 7.
This article from two students at the University of Cape Town takes a look at the functionality of DRM.
DRM is widely discussed as a tool to enforce copyright as we move into the digital medium. The authors agree with many other critics of the various software technologies that constitute DRM in observing that DRM does not really enforce copyright. They argue that while copyright enforcement is theoretically possible, DRM lacks the sophistication to allow for fair use of digital objects. Fair use is an exception to a usually applicable copyright restriction and is frequently argued on a case-by-case basis; it's therefore extremely difficult to allow for this very important exception in a programmatic way. Consequently it's overlooked and rights holders instead are allowed to sidestep the whole issue by issuing some agreed contractual terms that usually favor themselves over the consumer or user.
The authors take some further time to observe how DRM is unable to restrict reproduction and distribution (the core protections of copyright) because of the nature of hardware and software. DRM at the application (iTunes) and operating system level (Windows, OS X) cannot prevent reproduction and distribution. Media-specific DRM (like CSS and AACS) can be compromised, moreover media is trending toward direct digital distribution. DRM at the chip level can also be skipped so long as computers continue to have removable components. Provided this continues (I certainly hope so) and systems can support multiple operating systems (again, let's hope so) DRM doesn't seem to have a future as a legitimate copyright enforcement tool.
The authors instead argue that DRM is used as a licensing tool. They present two improvements to the licensing model that might allow for more input from the user and more granularity in the rights holders' licensing terms.
Use licenses in DRM systems are typically explained to a computer through a Rights Expression Language (REL). They authors argue that RELs should be expanded to incorporate a more nuanced license-negotiation process that goes beyond a simple request-response that is used today (i.e., do you accept this license? Yes or no). In this model users would request a set of rights, the licensing server would evaluate the request and serve up a license with those rights and the terms, user accepts, denies or renegotiates.
This process could begin again after the user has purchased the digital object and the license if the user wants some new right. This could allow for better fair use control.
A second approach provides a credential construct in RELs that would allow users to identify themselves as various roles (a journalist, a university student, a researcher, etc.) and thereby request different rights.
While I agree that DRM is a poor copyright enforcement tool, I'm hesitant to embrace these more nuanced licensing models and processes right off the bat. As the authors note, their models allow for a lot of flexibility and "newer business models for the rights holders." Is that a good thing? I can easily see a business micromanaging their rights model to create a maximum profit. In fact they could read this article and wonder why it hadn't occur to them to create a pay-as-you-go licensing model that kept users coming back for more rights.
But, so long as digital objects like datasets, music, recordings, and so on have rights holders that need to exert control over their materials, a better model is needed, one that works. The authors here have faith that a more nuanced licensing model ultimately would create fairer deals because users/researchers have a say in what they'll accept and in what they want. To the extent that a model would facilitate such feedback, I'm for it. The DRM industry is looking for standards, I hope they settle on one that's reasonable and expandable.
This article from two students at the University of Cape Town takes a look at the functionality of DRM.
DRM is widely discussed as a tool to enforce copyright as we move into the digital medium. The authors agree with many other critics of the various software technologies that constitute DRM in observing that DRM does not really enforce copyright. They argue that while copyright enforcement is theoretically possible, DRM lacks the sophistication to allow for fair use of digital objects. Fair use is an exception to a usually applicable copyright restriction and is frequently argued on a case-by-case basis; it's therefore extremely difficult to allow for this very important exception in a programmatic way. Consequently it's overlooked and rights holders instead are allowed to sidestep the whole issue by issuing some agreed contractual terms that usually favor themselves over the consumer or user.
The authors take some further time to observe how DRM is unable to restrict reproduction and distribution (the core protections of copyright) because of the nature of hardware and software. DRM at the application (iTunes) and operating system level (Windows, OS X) cannot prevent reproduction and distribution. Media-specific DRM (like CSS and AACS) can be compromised, moreover media is trending toward direct digital distribution. DRM at the chip level can also be skipped so long as computers continue to have removable components. Provided this continues (I certainly hope so) and systems can support multiple operating systems (again, let's hope so) DRM doesn't seem to have a future as a legitimate copyright enforcement tool.
The authors instead argue that DRM is used as a licensing tool. They present two improvements to the licensing model that might allow for more input from the user and more granularity in the rights holders' licensing terms.
Use licenses in DRM systems are typically explained to a computer through a Rights Expression Language (REL). They authors argue that RELs should be expanded to incorporate a more nuanced license-negotiation process that goes beyond a simple request-response that is used today (i.e., do you accept this license? Yes or no). In this model users would request a set of rights, the licensing server would evaluate the request and serve up a license with those rights and the terms, user accepts, denies or renegotiates.
This process could begin again after the user has purchased the digital object and the license if the user wants some new right. This could allow for better fair use control.
A second approach provides a credential construct in RELs that would allow users to identify themselves as various roles (a journalist, a university student, a researcher, etc.) and thereby request different rights.
While I agree that DRM is a poor copyright enforcement tool, I'm hesitant to embrace these more nuanced licensing models and processes right off the bat. As the authors note, their models allow for a lot of flexibility and "newer business models for the rights holders." Is that a good thing? I can easily see a business micromanaging their rights model to create a maximum profit. In fact they could read this article and wonder why it hadn't occur to them to create a pay-as-you-go licensing model that kept users coming back for more rights.
But, so long as digital objects like datasets, music, recordings, and so on have rights holders that need to exert control over their materials, a better model is needed, one that works. The authors here have faith that a more nuanced licensing model ultimately would create fairer deals because users/researchers have a say in what they'll accept and in what they want. To the extent that a model would facilitate such feedback, I'm for it. The DRM industry is looking for standards, I hope they settle on one that's reasonable and expandable.
Is metadata data, or is data metadata, or are they different?
Dr. Winget posted this article on delicious, "Metadata vs Data: a wholly artificial distinction."
In this article, Terry Jones writes that it is wrong to distinguish between metadata and data. He opines that David Weinberger was right to state that all data is metadata. In saying this, what he means to get at is that people use metadata to search. In doing so, they don't necessarily only want to search by something that an authority has deemed metadata. Often, users wish to search full text in the same way they can search a title or an author.
He uses the example of the unix file system, which uses one command ("find") to search what has been pre-determined to be metadata, and another ("grep") to search inside what is deemed to be data. Although it is possible to combine these searches (a commenter on the blog provides the method), Jones' point is that separating the metadata and the data into two systems is problematic.
FluidDB, a cloud-based database, supposedly does not make this distinction, but holds all data together (all of it as metadata) in one searchable unit, made up of many pieces of data (or metadata). The author (and CEO) claims that this method is more robust because no data is more important than other data, meaning nothing will be lost or deleted, and anything can be added.
Functionally, I think it is a great idea to store metadata and data together, making either and both searchable. One of the advantages of digital storage is ease of search. It isn't really necessary to search the card catalog, get a reference to an item, and then go to it on the shelves. A federate search can directly pull on article from terms either in the title, in tags, or from the document.
However, I think conceptually, it could be very confusing to lump data together into an ever broadening category, or to lose the distinction between some metadata and an original item. For example, many collections add administrative (meta)data to items that should be kept distinct from what might be considered metadata but is also a part of the original item (title, date). For example, processing archivist is a piece of metadata that could be searched via a keyword search, but I would not want it to be confused as part of an original item.
Metadata is certainly data, and perhaps all data should be treated as metadata in search, but to me, strictly speaking, true metadata can only exist if the data to which it relates exists. A title with no article is just a phrase.
In this article, Terry Jones writes that it is wrong to distinguish between metadata and data. He opines that David Weinberger was right to state that all data is metadata. In saying this, what he means to get at is that people use metadata to search. In doing so, they don't necessarily only want to search by something that an authority has deemed metadata. Often, users wish to search full text in the same way they can search a title or an author.
He uses the example of the unix file system, which uses one command ("find") to search what has been pre-determined to be metadata, and another ("grep") to search inside what is deemed to be data. Although it is possible to combine these searches (a commenter on the blog provides the method), Jones' point is that separating the metadata and the data into two systems is problematic.
FluidDB, a cloud-based database, supposedly does not make this distinction, but holds all data together (all of it as metadata) in one searchable unit, made up of many pieces of data (or metadata). The author (and CEO) claims that this method is more robust because no data is more important than other data, meaning nothing will be lost or deleted, and anything can be added.
Functionally, I think it is a great idea to store metadata and data together, making either and both searchable. One of the advantages of digital storage is ease of search. It isn't really necessary to search the card catalog, get a reference to an item, and then go to it on the shelves. A federate search can directly pull on article from terms either in the title, in tags, or from the document.
However, I think conceptually, it could be very confusing to lump data together into an ever broadening category, or to lose the distinction between some metadata and an original item. For example, many collections add administrative (meta)data to items that should be kept distinct from what might be considered metadata but is also a part of the original item (title, date). For example, processing archivist is a piece of metadata that could be searched via a keyword search, but I would not want it to be confused as part of an original item.
Metadata is certainly data, and perhaps all data should be treated as metadata in search, but to me, strictly speaking, true metadata can only exist if the data to which it relates exists. A title with no article is just a phrase.
"Geekfest, you scoff?"

Trying to find an interesting article about metadata and knowledge sharing, I stumbled across this article: "Holy Knowledge Sharing, Batman!" on the site KMEdge.com (the KM is for Knowledge Sharing). While this article leaves something to be desired in terms of academic rigor and depth - it raises interesting points: Comic-Con brings together over 125,000 people, who are part of a major knowledge sharing community. This community exists in person for four days a year, and online all the time. The article's author notes that this community "influences marketing strategies and budgets for movies and TV shows. It shapes the direction of new video games and helps predict strategies for future toy sales..." One wonders how this major, largely informal, data-creating, data sharing community is being used by companies for marketing, production of new products and a host of other purposes. I don't know exactly how companies currently harvest the information of communities like this - if there are spiders crawling through comic-con message boards and blogs, searching out key terms, or if they have individuals pouring over such sites. It seems to me, however, that determining an effective way to gather and synthesize this information would be a profitable tool - making the most of data sharing in the for profit realm.
OAIster and OCLC
This week I found a blog entry that was actually on topic! In HangingTogether, the OCLC/RLG blog, the entry OAIster Update: More Access & No Conditions explains more about the impending changes to the Open Archives Initiative (OAI) as OCLC takes over stewardship of the mass of metadata harvested from OAI-compliant repositories. As a background, "OAIster is a union catalog of millions of records representing open archive digital resources that was built by harvesting from open archive collections worldwide using the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH)." Previously supported by the University of Michigan, the size of the aggregation of records became too large for UM to single-handedly care for. Enter OCLC.
In reviewing some of the older blog posts about this transition (in September), it seems there was a slight kerfluffle about the whole thing. OCLC originally laid out the plan as such:
Now, I'm not entirely sure what all this means, but people were not happy. Reading the comments from the initial entry, people were concerned OCLC was going to sell all kinds of data, not just provide access to the metadata, and that records previously not on WorldCat would suddenly be there against the author's wishes. And more.
A follow-up entry was necessary, and OCLC actually had to change their terms and conditions with OAIster to specifically state that only metadata will be harvested. This clarification brought more comments of general dis-satisfaction, correctly wondering why a license was necessary and why things would be changing so much, in terms of copyright and such. As never having really dealt with OAIster issues, I don't fully understand their grievances, but they sounded legitimate.
So, this leads us to the post this past week, summarizing the first month the change-over happened. First thing, the terms and conditions that concerned people is now gone. The post says: "In keeping with the open style of the Open Archives Initiative community, if you make your metadata available for harvesting, you must intend for it to be harvested. We will also feel free to index it, provide access to it, and allow Google to crawl it." The opt-out strategy is still in place. OCLC is also allowing OAIster records to be searched separately from other WorldCat items, because OCLC "decided that the OAIster aggregation is an important enough destination for finding open access content". Or because there was an uproar, whatever. The comments to this post are much more positive, so apparently OCLC either made the right changes or wore everyone down.
On the surface, it seems fantastic that a larger organization is taking over the harvesting and accessing of metadata for the sustainability of the information. Although on WorldCat.org, there isn't an easily identifiable place to just search OAIster records. But, as was mentioned in a blog post, harvesting is hard! Maybe we should be thankful that someone is doing it (assuming it's being done correctly with good intentions, no GoogleBook excuses). It will be interesting to see what happens in the next few months and the reaction from users/authors/institutions. I'm sure HangingTogether will have more on this subject.
In reviewing some of the older blog posts about this transition (in September), it seems there was a slight kerfluffle about the whole thing. OCLC originally laid out the plan as such:
- Continued collaboration with University of Michigan.
- Records will be freely discoverable along with all other content in WorldCat.org. However, it will not be possible to limit a search to OAIster records alone.
- Data providers must request their records not be harvested, otherwise they will (opt-out policy).
- OAIster contributers must request free access to records in FirstSearch.
Now, I'm not entirely sure what all this means, but people were not happy. Reading the comments from the initial entry, people were concerned OCLC was going to sell all kinds of data, not just provide access to the metadata, and that records previously not on WorldCat would suddenly be there against the author's wishes. And more.
A follow-up entry was necessary, and OCLC actually had to change their terms and conditions with OAIster to specifically state that only metadata will be harvested. This clarification brought more comments of general dis-satisfaction, correctly wondering why a license was necessary and why things would be changing so much, in terms of copyright and such. As never having really dealt with OAIster issues, I don't fully understand their grievances, but they sounded legitimate.
So, this leads us to the post this past week, summarizing the first month the change-over happened. First thing, the terms and conditions that concerned people is now gone. The post says: "In keeping with the open style of the Open Archives Initiative community, if you make your metadata available for harvesting, you must intend for it to be harvested. We will also feel free to index it, provide access to it, and allow Google to crawl it." The opt-out strategy is still in place. OCLC is also allowing OAIster records to be searched separately from other WorldCat items, because OCLC "decided that the OAIster aggregation is an important enough destination for finding open access content". Or because there was an uproar, whatever. The comments to this post are much more positive, so apparently OCLC either made the right changes or wore everyone down.
On the surface, it seems fantastic that a larger organization is taking over the harvesting and accessing of metadata for the sustainability of the information. Although on WorldCat.org, there isn't an easily identifiable place to just search OAIster records. But, as was mentioned in a blog post, harvesting is hard! Maybe we should be thankful that someone is doing it (assuming it's being done correctly with good intentions, no GoogleBook excuses). It will be interesting to see what happens in the next few months and the reaction from users/authors/institutions. I'm sure HangingTogether will have more on this subject.
Monday, November 2, 2009
Folksonomy
This week I've tried to blog at least close to on topic, and so have chosen to talk about a short article from the November 2006 issue of D-Lib Magazine by Elaine Peterson entitled "Beneath the Metadata: Some Philosophical Problems with Folksonomy."
Peterson begins her article with a review of the foundations of traditional classification. She states that traditional catalogers are part of the Aristotelian tradition and adhere to the same basic principles of contraries, particulars, and categories. Of particular importance is the idea of contraries. In traditional cataloging, an item cannot be both "A" and "not A." A photograph of a single horse, for example, cannot be assigned the subject term "white horse" and the subject term "black horse." It must be one or the other. Additionally, the priority of the author's intent in traditional cataloging is another one of its distinguishing characteristics.
In a folksonomy, however, philosophical relativism is the underlying principle. Because of this, instead of it being impossible for an object to be both "A" and "not A," the assignment of tags is entirely up to the judgement of the individual. Therefore, it is possible for an object to be both one thing and its opposite. The reader has the priority in a folksonomy, rather than the author, and contrary interpretations can exist. As Peterson points out, "if all interpretations be of equal worth, [and] users can continuously add tags to articles, at some point it is likely that the whole system will become unusable." She likens it to the story of the Chinese emperor who wanted an accurate map of China, and ended up getting a map that was very accurate but was also the size of China. Such a thing is accurate, but not useful.
Peterson also points out that meta noise is another problem with folksonomies. Meta noise results from inaccurate, irrelevant, and even inadvertently erroneous (such as spelling "white horse" as "whit horse") tagging. She quotes David Weinberger, who views folksonomic classifications of the Web as "'messy and inelegant and inefficient, but it will be Good Enough.'" Peterson holds, however, that while it may be good to allow users to supply their own tags, the resulting classification system will not produce an efficient index, and so is not really all that good.
Overall, Peterson believes that inconsistencies within the folksonomic classification system will always exist. In traditional cataloging, there is a right way and a wrong way to classify something. This is not true with a folksonomy, and this, it seems Peterson believes, is its main problem that can cause the system to break down. "Folksonomists," she says, "are confusing cataloging structure with personal opinions." It is good to have personal opinions, and to be able to express them. However, opinions do not lead to effective classification, and therefore do not provide for effective searching.
METAe and ALTO: Mapping Physical and Logical Structures
Stehno, Birgit, Alexander Egger, and Gregor Retti. "METAe-Automated Encoding of Digitized Texts." Literary and Linguistic Computing, 18, No.1 (2003): 77-88. Available at: http://llc.oxfordjournals.org/cgi/reprint/18/1/77. (Accessed 11/1/2009).
This article describes how the Austrian-based METAe project applied METS (Metadata Encoding and Transmission Standard) to encode automatically extracted metadata from page images, especially metadata describing document layout. Like OCR engines which extract text from image files, the METAe engine extracts layout elements (such text or graphic blocks) by relying on formal rules and syntactical principles. The project built and compiled a recognition model that would map elements in the physical structure of page layout (text blocks of differing sizes and graphics blocks) to logical ones (e.g. paragraphs, titles, footnotes). This would then create an ALTO ('Analysed layout and text object') file which is an XML file that consists of both the layout structures and the full text of a book page. The ALTO file and information formatted in a variety of metadata standards such as Dublin Core and DIG 35 are then incorporated into a METS schema, which (if I'm understanding correctly) serves as an outer wrapper that provides the structural map and holds all of the metadata for the object.
In devising the METAe engine, they decided to use METS rather than TEI (Text Encoding Initiative) for their encoding since TEI was "far too inexplicit for the purpose of automated recognition" (9). This struck me as an interesting point since if TEI is to be the encoding standard for text in the future, it will need to be more amenable to automatic application. METS was also preferred as it allowed for METAe to add metadata at any logical level and to include metadata from different formats and standards, including pointing to metadata external to the METS document. For example, the METAe project can use a "DMDID" (Descriptive MetaData Identifier) attribute to generate tags linking a journal issue to its appropriate MARC record on a web server while also using Dublin Core tags to provide descriptive metadata for each article/contribution.
What this article by Stehno, Egger, and Retti does not really address are the difficulties (and there must have been some) that the METAe project encountered in developing its engine. It seems extremely useful to have an engine that can map the physical structure of a page onto logical structures, but certainly there must be structures that it is better and worse at recognizing. I've also had some difficulty in discerning precisely what happened to METAe--it seems to have lived on as ALTO rather than as the METAe tool itself. As the METAe website (http://meta-e.aib.uni-linz.ac.at/index.html) details, the project ran from September 2000 - September 2003 with partial funding from the European Commission. The website also announces the metadata engine has been marketed as digitization software under the name docWorks/METAe Edition, however, though the docWorks page exists and seems to offer a product that performs the tasks that METAe does, there is very little mention of METAe on the site. The one mention I was able to locate was a notice that as of August 2009, the Library of Congress has taken over maintenance of the ALTO XML schema from CCS Content Conversion Specialists GmbH (the company which produces docWorks). This appears to be a sign at least of ALTO's success as a schema, as LC has created a new ALTO editorial board to "help shape and advocate usage of the standard."
I guess the questions that I'm left with after reading about ALTO and METAe are ones about the relationship between standards and schemas and the tools used implement them. To what extent do the constraints of the tools with which schemas were initially partnered affect the future use of those schemas, even after a given tool has been put aside or altered? Can schemas created within the context of automation equally useful when used for hand encoding or human quality control?
This article describes how the Austrian-based METAe project applied METS (Metadata Encoding and Transmission Standard) to encode automatically extracted metadata from page images, especially metadata describing document layout. Like OCR engines which extract text from image files, the METAe engine extracts layout elements (such text or graphic blocks) by relying on formal rules and syntactical principles. The project built and compiled a recognition model that would map elements in the physical structure of page layout (text blocks of differing sizes and graphics blocks) to logical ones (e.g. paragraphs, titles, footnotes). This would then create an ALTO ('Analysed layout and text object') file which is an XML file that consists of both the layout structures and the full text of a book page. The ALTO file and information formatted in a variety of metadata standards such as Dublin Core and DIG 35 are then incorporated into a METS schema, which (if I'm understanding correctly) serves as an outer wrapper that provides the structural map and holds all of the metadata for the object.
In devising the METAe engine, they decided to use METS rather than TEI (Text Encoding Initiative) for their encoding since TEI was "far too inexplicit for the purpose of automated recognition" (9). This struck me as an interesting point since if TEI is to be the encoding standard for text in the future, it will need to be more amenable to automatic application. METS was also preferred as it allowed for METAe to add metadata at any logical level and to include metadata from different formats and standards, including pointing to metadata external to the METS document. For example, the METAe project can use a "DMDID" (Descriptive MetaData Identifier) attribute to generate tags linking a journal issue to its appropriate MARC record on a web server while also using Dublin Core tags to provide descriptive metadata for each article/contribution.
What this article by Stehno, Egger, and Retti does not really address are the difficulties (and there must have been some) that the METAe project encountered in developing its engine. It seems extremely useful to have an engine that can map the physical structure of a page onto logical structures, but certainly there must be structures that it is better and worse at recognizing. I've also had some difficulty in discerning precisely what happened to METAe--it seems to have lived on as ALTO rather than as the METAe tool itself. As the METAe website (http://meta-e.aib.uni-linz.ac.at/index.html) details, the project ran from September 2000 - September 2003 with partial funding from the European Commission. The website also announces the metadata engine has been marketed as digitization software under the name docWorks/METAe Edition, however, though the docWorks page exists and seems to offer a product that performs the tasks that METAe does, there is very little mention of METAe on the site. The one mention I was able to locate was a notice that as of August 2009, the Library of Congress has taken over maintenance of the ALTO XML schema from CCS Content Conversion Specialists GmbH (the company which produces docWorks). This appears to be a sign at least of ALTO's success as a schema, as LC has created a new ALTO editorial board to "help shape and advocate usage of the standard."
I guess the questions that I'm left with after reading about ALTO and METAe are ones about the relationship between standards and schemas and the tools used implement them. To what extent do the constraints of the tools with which schemas were initially partnered affect the future use of those schemas, even after a given tool has been put aside or altered? Can schemas created within the context of automation equally useful when used for hand encoding or human quality control?
Wednesday, October 28, 2009
Improved Twitter feed curation
Was pleasantly surprised this evening as I logged into Twitter. They now have Lists in Beta - something I've been wanting for some time.
1) I can create lists for subjects that interest me, and move those that I follow into that list (functionally this is akin to creating labels in Gmail). Doing this will create a link on the right of the home-page Twitter feed that I can click and read only those in that list.
2) others can subscribe to this list - which creates a kind of sub-follower rating. Let's say for example that I compile a list of those who tweet about Contemporary Asian art: this would be its own list that interested people could subscribe to (essentially bypassing me and my tweets but reading those that I have aggregated - or perhaps just going to my list to harvest them for one's own list). This is an idea similar to reading filtered blog subscriptions or friends groups in Facebook as well as customized blog-rolls.
3) There are great implications for use for larger or more widely trusted entities whose lists could be very popular to subscribe to. The introduction of this to Twitter's networking capabilities is a tremendous improvement for ease of use and both content and contact searching.
Now, if only YouTube would similarly jump on board with this concept.
1) I can create lists for subjects that interest me, and move those that I follow into that list (functionally this is akin to creating labels in Gmail). Doing this will create a link on the right of the home-page Twitter feed that I can click and read only those in that list.
2) others can subscribe to this list - which creates a kind of sub-follower rating. Let's say for example that I compile a list of those who tweet about Contemporary Asian art: this would be its own list that interested people could subscribe to (essentially bypassing me and my tweets but reading those that I have aggregated - or perhaps just going to my list to harvest them for one's own list). This is an idea similar to reading filtered blog subscriptions or friends groups in Facebook as well as customized blog-rolls.
3) There are great implications for use for larger or more widely trusted entities whose lists could be very popular to subscribe to. The introduction of this to Twitter's networking capabilities is a tremendous improvement for ease of use and both content and contact searching.
Now, if only YouTube would similarly jump on board with this concept.
Get the gamers involved
I had intended to write about this topic earlier in the semester and had forgotten all about it until I saw a reference in this week's reading by Amy Friedlander, The Triple Helix: Cyberinfrastructure, Scholarly communication, and Trust. Friedlander discusses how problems that can be cleanly parsed into discrete tasks are well suited for a distributed capacity model which allows open, but still structured, participation from the public. She provides the example of protein folding, which is the topic of an April 2009 Wired Magazine article: Gamers Unravel the Secret Life of Protein.
The protein chemistry world has a biennial World Series competition to see who can predict the shape of a protein only knowing the sequence of its constitute parts (Community-Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction, or CASP). CASP surveys labs around the world to find proteins that are about to be solved, and compile a list of puzzles online.
David Baker, whose team had dominated the competition since 1998, had been using Rosetta@home, similar to SETI@home, which farmed out computations to volunteer PCs distributed globally - providing Baker with the equivalent of a supercomputer. However, the computers were unable to complete certain puzzles, which humans should be able to solve, having better spatial reasoning. Baker's friend David Salesin, a computer scientist, brought him together with Zoran Popovic, another computer scientist and graphics expert, and the three developed what turned into a massively multiplayer competition. Gamers are given a multicolored knot of spirals and clumps, which they fold and wiggle into its optimum shape.
Baker then entered potentially accurate CASP protein structures into the biennial competition. Of 15 submissions, 7 finished "in the money" and one took first place. The gamer team, led by a 13-year old, beat the best biochemists. Baker was also hoping to find prodigies... when "Cheese" (the 13-year old) was asked how he did it, he said, "it just looks right."
Baker has given the players a new challenge to design a new protein drug with the right size and binding properties. Baker will synthesize and test the most promising structures and if any have value in the real world, the gamers will share in the credit.
The article doesn't specifically address authenticity or trust, but the gamer submissions are not automatically deemed correct, even though the game is based on laws of physics. Submissions are reviewed by CASP, and/or tested in the lab. But, given the fact that there are more ways to fold protein than atoms in the universe, and they arrange in a fraction of a second, collective efforts are crucial.
Old Problem, New Scale
This article provided by George Mason University's Center for History and New Media (which is quite an interesting program in and of itself), notes that the problem of authenticating sources used for research is not a new one in the humanities. Scholars in fields such as art and history have always had to use their specialized knowledge to try to determine the authenticity of the works or documents which they examine. However, Bearman and Trant, the authors of the article, feel that the current digital era has led to an enormous increase in the scale of the problem. Whereas in the past it took much specialized skill to create a credible copy or forgery, the mere act of hitting a button or two can create a digital object nearly indistinguishable from the original (other than the lack of physicality, of course.) Also, more and more scholars are turning to digital research as a way of maximizing research time and for the thoroughness it affords. Yet, they no longer can use physical clues (for example, handwriting) to infer authenticity
Bearman and Trant see a number of possible solutions to the problem without really recommending one single method. They divide their methods into public methods, secret methods, and functionally dependent methods. Public methods may either be social (e.g. creating a collecting institution of record along the lines of a certified website which raises issues of who performs the certification) or technological, such as using public key encryption to create digital signatures for documents. Secret methods consist mainly of technological clues buried in empty space within files such as stegonography. Functionally dependent methods probably do the best job in authenticating because the very working of the file is tied to the end user's ability to authenticate it. However, this is much trickier to set up and probably a bit overkill. (After all, there is a certain amount of responsibility traditionally borne by the scholar...)
I found this article interesting largely because of my history background. When evaluating physical reprints or translations of primary sources, a historical scholar will check the publisher (kind of an early form of public authentication, I suppose.) But in digitizing, nearly anyone can be a publisher, so count me as convinced, scholarly responsibility notwithstanding, that some sort of authenticity determination scheme will be necessary as more and more primary documents are digitized
Bearman and Trant see a number of possible solutions to the problem without really recommending one single method. They divide their methods into public methods, secret methods, and functionally dependent methods. Public methods may either be social (e.g. creating a collecting institution of record along the lines of a certified website which raises issues of who performs the certification) or technological, such as using public key encryption to create digital signatures for documents. Secret methods consist mainly of technological clues buried in empty space within files such as stegonography. Functionally dependent methods probably do the best job in authenticating because the very working of the file is tied to the end user's ability to authenticate it. However, this is much trickier to set up and probably a bit overkill. (After all, there is a certain amount of responsibility traditionally borne by the scholar...)
I found this article interesting largely because of my history background. When evaluating physical reprints or translations of primary sources, a historical scholar will check the publisher (kind of an early form of public authentication, I suppose.) But in digitizing, nearly anyone can be a publisher, so count me as convinced, scholarly responsibility notwithstanding, that some sort of authenticity determination scheme will be necessary as more and more primary documents are digitized
Searching Twitter
So there have been some recent developments in the past week that have received a good deal of news. First off, Microsoft thought that it was going to make a decent cut into Google's mammoth share of the search market through a deal with Twitter which would allow tweets to show up in Bing searches. Unfortunately (for Microsoft) Google managed to make a deal of their own with Twitter so that tweets will also show up in Google searches. While Bing might still have the upper-hand since its Twitter search is already live, I want to take a moment to grieve for the older forms of searching Twitter.
The New York Times had a good article this week about Twitter's approach to innovation. At this point, the concept of the "democratization of innovation" is probably familiar to most people who know a thing or two about the interwebs, but Twitter was pretty exceptional in its adherence to a credo of bottom-up innovation. At the outset, the service afforded only the ability to write "micro-blogs" of 140 characters. Methods of searching these micro-blogs developed organically by twitterers themselves. Nearly every search convention:
1. The '@' before the screen name of another user
2. The letters 'RT' before a reproduction of something that already been tweeted (or a "re-tweet")
3. The '#' used before a certain topic so that topics with more than one word were easily searchable (e.g., #iranelection)
4. Even the term 'tweet' to refer to a single micro-blog
...all of these were developed by twitter users and later adopted as standard form by Twitter developers. Now that Microsoft and Google are in on the game....sigh. So much for democracy.
Admittedly, it's too early to tell what kind of affect this new way to search Twitter is going to have on things like the use of the hash tag, but it's pretty easy to see why Microsoft and Google want to get in on the Twitter action. Back in March, on the Tech Crunch blog, Michael Arrington wrote a post about just how useful Twitter is to businesses. Because tweets are so short, many people (and I mean MANY) use Twitter to do nothing more than gripe about things, especially their experiences with products. One can only imagine how useful this information is to businesses. In fact this information is SO useful, that Twitter, itself, has been valued at nearly $1 Billion, even though it has around 35 million users (compared to Facebook's 175 million). With that kind of valuation, it's surprising that Twitter hasn't decided to cash in and (like Facebook) allow more ad space. Since Twitter is so amenable to being searched, however, they may never have to get into the ad game in order to make serious money. Here's to hoping that Twitter holds-on to its original democratic, bottom-up approach to innovation despite all of the corporate interest.
The New York Times had a good article this week about Twitter's approach to innovation. At this point, the concept of the "democratization of innovation" is probably familiar to most people who know a thing or two about the interwebs, but Twitter was pretty exceptional in its adherence to a credo of bottom-up innovation. At the outset, the service afforded only the ability to write "micro-blogs" of 140 characters. Methods of searching these micro-blogs developed organically by twitterers themselves. Nearly every search convention:
1. The '@' before the screen name of another user
2. The letters 'RT' before a reproduction of something that already been tweeted (or a "re-tweet")
3. The '#' used before a certain topic so that topics with more than one word were easily searchable (e.g., #iranelection)
4. Even the term 'tweet' to refer to a single micro-blog
...all of these were developed by twitter users and later adopted as standard form by Twitter developers. Now that Microsoft and Google are in on the game....sigh. So much for democracy.
Admittedly, it's too early to tell what kind of affect this new way to search Twitter is going to have on things like the use of the hash tag, but it's pretty easy to see why Microsoft and Google want to get in on the Twitter action. Back in March, on the Tech Crunch blog, Michael Arrington wrote a post about just how useful Twitter is to businesses. Because tweets are so short, many people (and I mean MANY) use Twitter to do nothing more than gripe about things, especially their experiences with products. One can only imagine how useful this information is to businesses. In fact this information is SO useful, that Twitter, itself, has been valued at nearly $1 Billion, even though it has around 35 million users (compared to Facebook's 175 million). With that kind of valuation, it's surprising that Twitter hasn't decided to cash in and (like Facebook) allow more ad space. Since Twitter is so amenable to being searched, however, they may never have to get into the ad game in order to make serious money. Here's to hoping that Twitter holds-on to its original democratic, bottom-up approach to innovation despite all of the corporate interest.
Averting a Digital Katrina: Sustaining Trust in the Research Infrastructure (EDUCAUSE Review) | EDUCAUSE
Averting a Digital Katrina: Sustaining Trust in the Research Infrastructure (EDUCAUSE Review) | EDUCAUSE
Amy Friedlander, Director of Programs at the Council on Library and Information Resources, wrote a very helpful and insightful article on the issuing of sustaining trust in the research infrastructure. The article was published in the Educause Review website.
Friedlander starts off the article by explaining her Katrina anaology. When the Hurricane Katrina stormed over New Orleans and the levees broke, the victims placed a huge trust on the infrastructures to step in and help them out. However, the infrastructure themselves had broken down: social, engineering, and political, thus creating this huge Catch-22 with huge repercussions that are still felt today.
She states that trust in an infrastructure stems from familiarity of a system; when things are running along as expected we build trust on its reliability, which explains in part why traditional scholarly publication remains robust even with the development of other methods of communication are evolving. Friedlander quotes Christine Borgman in identifying the three major functions that traditional scholarly communication achieves:
1) legitimizes scholarly work
2) disseminates that work to an audience (or several)
3) provides access, preservation and citation
The writer and reader share the same expectations and the repeated "success" of such a system causes scholars to cling to the older methods of scholarly communication. Plus, the long tradition of building prestige from having your work published in a famous or prestigious journal is still in effect.
However, the electronic publishing format is going happen regardless and Friedlander points out some of inconsistencies and examples that display a lack of trust in such systems and then proceeds to explain why this is rightly so. Scientists rarely participate in social-networking or create blogs or wikis of their research, namely data, partly because there is a huge confusion of how to correctly cite e-works in a paper. Print is seen as more authentic and is easier to cite, to boot.
Friedlander spends some time discussing the issue of citation: print articles and journals are a form of efficiency in citation and also uphold the core values of research:
1) attribution and credit
2) reliability
3) persistance
4) validation of sources
5) integrity of audience
4) replication of results
She states that very interactivity and the dynamism of the digital medium undermine the core values of trust in scholarly research. For example, the ability to download data in order to verify the experiment means inconsistency in how the data is displayed, which introduces unexpected results in the infrastructure which is hardly conducive for one to relax in the cyberinfrastructure.
Friedlander ends the article by saying that "retreating to analog is hardly the answer," nor is giving up the standard methods of interpretation, however in building up trust in the cyberinfrastructure requires long-term management and preservation of the digital data along with the policies and methods that encourage discovery, community, and discussion is the key.
Friedlander's article echoed many of the sentiments that were discussed in the class readings in struggling to define where trust comes from and how the traditional methods of scholarly publication are still widely preferred because of familiarity and a well-established expectation of be able to gain scholarly prestige. She did not really come up with any solutions to the issue at hand, similar to the other articles. This issue is a philosophical one and will remain under discussion for quite some time because I feel that scholarly communication is under a huge transistion and will happen in spite of itself with issues more or less resolving itself in response to problems that may or may not occur.
Amy Friedlander, Director of Programs at the Council on Library and Information Resources, wrote a very helpful and insightful article on the issuing of sustaining trust in the research infrastructure. The article was published in the Educause Review website.
Friedlander starts off the article by explaining her Katrina anaology. When the Hurricane Katrina stormed over New Orleans and the levees broke, the victims placed a huge trust on the infrastructures to step in and help them out. However, the infrastructure themselves had broken down: social, engineering, and political, thus creating this huge Catch-22 with huge repercussions that are still felt today.
She states that trust in an infrastructure stems from familiarity of a system; when things are running along as expected we build trust on its reliability, which explains in part why traditional scholarly publication remains robust even with the development of other methods of communication are evolving. Friedlander quotes Christine Borgman in identifying the three major functions that traditional scholarly communication achieves:
1) legitimizes scholarly work
2) disseminates that work to an audience (or several)
3) provides access, preservation and citation
The writer and reader share the same expectations and the repeated "success" of such a system causes scholars to cling to the older methods of scholarly communication. Plus, the long tradition of building prestige from having your work published in a famous or prestigious journal is still in effect.
However, the electronic publishing format is going happen regardless and Friedlander points out some of inconsistencies and examples that display a lack of trust in such systems and then proceeds to explain why this is rightly so. Scientists rarely participate in social-networking or create blogs or wikis of their research, namely data, partly because there is a huge confusion of how to correctly cite e-works in a paper. Print is seen as more authentic and is easier to cite, to boot.
Friedlander spends some time discussing the issue of citation: print articles and journals are a form of efficiency in citation and also uphold the core values of research:
1) attribution and credit
2) reliability
3) persistance
4) validation of sources
5) integrity of audience
4) replication of results
She states that very interactivity and the dynamism of the digital medium undermine the core values of trust in scholarly research. For example, the ability to download data in order to verify the experiment means inconsistency in how the data is displayed, which introduces unexpected results in the infrastructure which is hardly conducive for one to relax in the cyberinfrastructure.
Friedlander ends the article by saying that "retreating to analog is hardly the answer," nor is giving up the standard methods of interpretation, however in building up trust in the cyberinfrastructure requires long-term management and preservation of the digital data along with the policies and methods that encourage discovery, community, and discussion is the key.
Friedlander's article echoed many of the sentiments that were discussed in the class readings in struggling to define where trust comes from and how the traditional methods of scholarly publication are still widely preferred because of familiarity and a well-established expectation of be able to gain scholarly prestige. She did not really come up with any solutions to the issue at hand, similar to the other articles. This issue is a philosophical one and will remain under discussion for quite some time because I feel that scholarly communication is under a huge transistion and will happen in spite of itself with issues more or less resolving itself in response to problems that may or may not occur.
The Information Bottleneck

In his article Institutional Repositories and Research Data Curation in a Distributed Environment, Michael Witt discusses the lack of a framework for organizing information digitally and the effect that has had on datasets.
Witt begins by talking about scientific research of the past and how carefully lab notebooks were kept by scientists and later preserved in archives as part of the scientific record. While I have trouble believing this was always the case, I would agree that changes in technology have changed the way the records are being kept by scientists.
The bulk of his article is spent talking about the existing repository infrastructure of the Purdue Libraries, but I found his ideas about an information bottleneck more interesting. Witt believes that the original data is narrowed for use within the scope of a particular article. This new data is all that most people ever see of the original data. Assuming that the long tail holds true, then that data is valuable for its many different future uses. Unfortunately, the nature of the information bottleneck has left the data stripped of much of its original information and perhaps also of its future value.
Witt insists that datasets need to be presented in context to remain meaningful and useful. And this seems to me a fairly obvious idea. So why aren't we getting the raw data into repositories along with the narrow-focus versions of the data? Based on some of the other readings that I have done, I can only suggest it is because we are still having trouble getting even the narrowly focused data into repositories.
Witt does not address how this is to be done, he just suggests that "...at some point in the future, the process and units of scholarly communication may be reconsidered to fully recognize and include research datasets. In some cases, such as the Human Genome Project, the value of a genome dataset itself is generally recognized to be greater than any single, published finding resulting from its analysis."
I agree with his sentiment, and can imagine many uses for such a collection of datasets. But while we are still struggling to have truly useful repositories, it seems like an idea that will have to remain in the future.
Current Open Access Income Models from the Scholarly Publishing and Academic Resources Coalition (SPARC)
Income Models for Open Access: An Overview of Current Practice - http://www.arl.org/sparc/bm~doc/incomemodels_v1.pdf
This is a lengthy document which describes in detail every tactic currently in use by open-access journals for financial support. SPARC compiled this overview to encourage existing journals who are thinking of becoming open-access as well as anybody who is thinking of starting an open-access journal from scratch. SPARC's definition of open-access is: "free and immediate online access to peer-reviewed journal literature".
Aside from wanting to participate in open scholarship, there are a few reasons why a journal might consider an open-access model for business purposes, these are those reasons according to SPARC:
• to increase access to its published research by lowering or eliminating market barriers to the content;
• to maximize market reach and support a new journal launch when the market will not support a traditional subscription model; or
• to implement a supply-side model (discussed below) in response to funder-mandated content deposit policies.
Legitimate advertising and/or sponsorship is one model for income for journals, these streams of revenue vary from Google Adsense (Open Government) to long-term corporate partnership (CERN Courier). I was particularly interested in this revenue source after reading about Elsevier's sheisty "sponsorship" so I looked at the examples that SPARC provides to see where the money is coming from. It turns out that for all of Elsevier's talk about open-access journals not being practical, they do sponsor at least one open-access journal. That journal is the The Journal of Electronic Publishing. JEP is interesting, these are the rest of their sponsors:
LexisNexis| O'Reilly Media| Newsbank Readex| Aptara| Wiley-Blackwell
JEP has a bunch of articles about open-access, one of them is Two Scenarios for How Scholarly Publishers Could Change Their Business Model to Open Access. How do they keep their corporate sponsors?
Corporate funding is probably the most problematic approach to revenue and SPARC provides only one example (Evidence-based Complementary and Alternative Medicine). In order for this model to work, the journal has to have an underwriting policy in order to prevent something like the Elsevier/MERCK situation from happening.
10% of open-access journals use the supply-side model of article processing fees for revenue, this is where authors pay the fee to publish in a journal. Most of these fees are subsidized by the author's institution or through a grant, only 5% of the fees are paid for by authors out-of-pocket. Some journals give authors the option to pay the fee to make their article open (BMJ Journals Unlocked).
PLoS is one journal that uses the article processing fee for income, but they do it slightly differently. Their model is based on an institutional membership fee which covers all of the papers submitted through an institution. PLoS even encourages consortial memberships to offset costs.
So far I've focused on some of the supply-side models for income. In terms of demand-side revenue, these are the models currently in place:
Versioning - Either selling an aggregate print copy of all issues at the end of the year or a print copy of each issue with additional non-research content throughout the year, it may also be considered the "special print" model.
Use-trigger Fees - In this model, the income is from institutions who use the journal heavily. There is a threshold of use, so most people will have access but when use is heavy from a single institution the fees are triggered. SPARC provides only one example of a journal who uses this model - Anthropological Index Online.
Convenience-Format License - This is what's currently happening with law journals. While the journals may be open access (Open Access Law Journals), they have an additional license with LexisNexis or WestLaw.
Value-Added Fee-Based Services - This is something I assume we're all familiar with, the web is full of sites like this.
Contextual E-Commerce - When a journal also has an e-commerce site, PLoS does this to a certain extent.
This is a lengthy document which describes in detail every tactic currently in use by open-access journals for financial support. SPARC compiled this overview to encourage existing journals who are thinking of becoming open-access as well as anybody who is thinking of starting an open-access journal from scratch. SPARC's definition of open-access is: "free and immediate online access to peer-reviewed journal literature".
Aside from wanting to participate in open scholarship, there are a few reasons why a journal might consider an open-access model for business purposes, these are those reasons according to SPARC:
• to increase access to its published research by lowering or eliminating market barriers to the content;
• to maximize market reach and support a new journal launch when the market will not support a traditional subscription model; or
• to implement a supply-side model (discussed below) in response to funder-mandated content deposit policies.
Legitimate advertising and/or sponsorship is one model for income for journals, these streams of revenue vary from Google Adsense (Open Government) to long-term corporate partnership (CERN Courier). I was particularly interested in this revenue source after reading about Elsevier's sheisty "sponsorship" so I looked at the examples that SPARC provides to see where the money is coming from. It turns out that for all of Elsevier's talk about open-access journals not being practical, they do sponsor at least one open-access journal. That journal is the The Journal of Electronic Publishing. JEP is interesting, these are the rest of their sponsors:
LexisNexis| O'Reilly Media| Newsbank Readex| Aptara| Wiley-Blackwell
JEP has a bunch of articles about open-access, one of them is Two Scenarios for How Scholarly Publishers Could Change Their Business Model to Open Access. How do they keep their corporate sponsors?
Corporate funding is probably the most problematic approach to revenue and SPARC provides only one example (Evidence-based Complementary and Alternative Medicine). In order for this model to work, the journal has to have an underwriting policy in order to prevent something like the Elsevier/MERCK situation from happening.
10% of open-access journals use the supply-side model of article processing fees for revenue, this is where authors pay the fee to publish in a journal. Most of these fees are subsidized by the author's institution or through a grant, only 5% of the fees are paid for by authors out-of-pocket. Some journals give authors the option to pay the fee to make their article open (BMJ Journals Unlocked).
PLoS is one journal that uses the article processing fee for income, but they do it slightly differently. Their model is based on an institutional membership fee which covers all of the papers submitted through an institution. PLoS even encourages consortial memberships to offset costs.
So far I've focused on some of the supply-side models for income. In terms of demand-side revenue, these are the models currently in place:
Versioning - Either selling an aggregate print copy of all issues at the end of the year or a print copy of each issue with additional non-research content throughout the year, it may also be considered the "special print" model.
Use-trigger Fees - In this model, the income is from institutions who use the journal heavily. There is a threshold of use, so most people will have access but when use is heavy from a single institution the fees are triggered. SPARC provides only one example of a journal who uses this model - Anthropological Index Online.
Convenience-Format License - This is what's currently happening with law journals. While the journals may be open access (Open Access Law Journals), they have an additional license with LexisNexis or WestLaw.
Value-Added Fee-Based Services - This is something I assume we're all familiar with, the web is full of sites like this.
Contextual E-Commerce - When a journal also has an e-commerce site, PLoS does this to a certain extent.
Tuesday, October 27, 2009
Digitizing Books: A Multi-Step Process
This week I looked at an article suggested by Prof. Winget, "A Mass Digitization Primer" by Juliet Sutherland, from the Summer 2008 issue of Library Trends. We've been discussing digitization projects throughout the class, and this article is a good place to look for a simple technical discussion of how these projects proceed.
Sutherland discusses the four levels of digitization, starting with page images. Page images have the advantage of replicating the actual experience of reading the book quite closely, as the layout is the same and images appear along with the text. They save future wear on the book by providing an easily browseable copy. However, they can be hard to view and require lots of scrolling on the screen. Sutherland mentions Google as the largest producer of scanned book images and states "Google's purpose is not to archive printed material, but rather to make it accessible via their expertise in search." This is an interesting perspective and makes it sound as if Google was merely looking for a large body of content to set their indexing and searching tools loose on, once they had conquered the web. Considered in this light their lack of interest in accuracy is not surprising.
Scanned images can be improved upon with OCR (optical character recognition) which produces the actual text, but here we run into problems based upon poor image quality. Intelligent OCR mechanisms that disambiguate words such as "modern" vs. "modem" based on the vocabulary contemporary to the text can reduce some of these problems.
Corrected OCR requires a lot of work, and there are a few different solutions. Double-keying, which is the practice of having two people type in the identical text to reduce errors, is costly and time-consuming. Some more intelligent solutions are Recaptcha and Distributed Proofreaders. The former relies on using CAPTCHAs (those scrambled text boxes required to log into sites to prevent automated spam) and human intelligence to read distorted words. The key trick here is Recaptcha always asks for two words: one it knows, and one it doesn't know, and can therefore judge whether a user made the correct judgement or not. This is a clever solution that utlizes lots of distributed bits of time that a user barely notices to add up, like SETI@Home. The second solution, Distributed Proofreaders, requires a little more active participation from users willing to proof scanned texts.
The final step in digitization which will take us beyond print books is semantically tagging books, such as identifying named entities ("Washington", "Europe"). This will likely require a combination of automation and intelligent intervention. Projects such as the Biodiversity Heritage Library are combining all of these steps to produce scanned books online which are useful and reliable. We can see that the distributed approach common to many online projects helps produce a higher-quality product. Also, the issues of trust and reliability come up with many of these projects: are we more likely to trust the Biodiversity Heritage Library because of the prestigious institutions involved? Or are we more likely to trust something that has been professionally proofread? At the very least we can know that most of the errors we might encounter are technical, and not the fault of a human.
Digitization Primer
I decided to check out Juliet Sutherland's "A Mass Digitization Primer" article from about a year ago. Sutherland is the head of the Distributed Proofreaders project, a site that allows volunteers to proofread pages from public domain books.
Sutherland provides a quick breakdown of mass digitization for books:
I'm surprised that Sutherland doesn't mention digitization of other parts of the book like the spine, headband, jacket, binding, endpapers and so on that would fall into this section. There must be some issues regarding the best way to proceed digitizing these, but I suppose the interested audience is not substantial for this article.
Shortcoming to page pictures are fairly obvious; there's nothing new discussed here, but Sutherland does point the wide range in quality page pictures encompass, Google marks the lower end of that spectrum.
Raw OCR is typically included as a next step after actual photographing of a page. Sutherland states that "commercial OCR of modem [sic] printed materials is now virtually perfect." She uses that typo to point out some of the shortcomings of OCR with older texts with their odd fonts, partial punctuation spacing, deteriorated paper and so on.
Corrected OCR fixes these problems by ensuring fidelity to the printed text. Humans can do this, but the labor is expensive. It's even expensive when it's cheap: double-key entry from workers in foreign countries is still costly for the amount text and books corrected.
Alternatives exist, and as Meg blogged a few weeks ago Recaptcha is one such solution. This service uses the ambiguous words an OCR device encounters as human-verification criteria for various sites. As users type the words they see in the image Recaptcha provides, they also provide a +1 for their reading of that word. Ultimately, this resolves the OCR's ambiguity. Recaptcha has important shortcomings though. It can't present words with old orthography or spacing, it will not present words the OCR device did not tag as ambiguous or words which it simply skipped on account of illegibility, and word context is stripped.
Sutherland's own Distributed Proofreaders is another solution. As expected accuracy is high, output is relatively low.
Semantic coding is another topic we've covered in class, particularly in regard to XML's and JSON's appropriateness for such marking. The semantic encoding mentioned here is fairly pedestrian as such things go: chapter headings or is Washington a city, state or person?, etc., but Sutherland leaves space for more interesting semantic markup in the future.
Sutherland concludes that digitization should be seen as a multistep process. I agree 100%, and I think it's clear Google similarly conceives of digitization as iterative, with it's "Getting there!" disposition toward accurate data and metadata (compared to OCA's metadata which is on the whole much much better). The slippery slope to this approach is that every step is a little less accomplished than it could be; raw OCR quality is contingent on page picture quality, corrected OCR needs somewhat accurate raw OCR for context. With that in mind I wonder if it would not be better to more closely consider first steps rather than applying so much functionality and completeness to later steps.
Sutherland provides a quick breakdown of mass digitization for books:
- Page Pictures
- Raw OCR
- Corrected OCR
- Semantic Coding
I'm surprised that Sutherland doesn't mention digitization of other parts of the book like the spine, headband, jacket, binding, endpapers and so on that would fall into this section. There must be some issues regarding the best way to proceed digitizing these, but I suppose the interested audience is not substantial for this article.
Shortcoming to page pictures are fairly obvious; there's nothing new discussed here, but Sutherland does point the wide range in quality page pictures encompass, Google marks the lower end of that spectrum.
Raw OCR is typically included as a next step after actual photographing of a page. Sutherland states that "commercial OCR of modem [sic] printed materials is now virtually perfect." She uses that typo to point out some of the shortcomings of OCR with older texts with their odd fonts, partial punctuation spacing, deteriorated paper and so on.
Corrected OCR fixes these problems by ensuring fidelity to the printed text. Humans can do this, but the labor is expensive. It's even expensive when it's cheap: double-key entry from workers in foreign countries is still costly for the amount text and books corrected.
Alternatives exist, and as Meg blogged a few weeks ago Recaptcha is one such solution. This service uses the ambiguous words an OCR device encounters as human-verification criteria for various sites. As users type the words they see in the image Recaptcha provides, they also provide a +1 for their reading of that word. Ultimately, this resolves the OCR's ambiguity. Recaptcha has important shortcomings though. It can't present words with old orthography or spacing, it will not present words the OCR device did not tag as ambiguous or words which it simply skipped on account of illegibility, and word context is stripped.
Sutherland's own Distributed Proofreaders is another solution. As expected accuracy is high, output is relatively low.
Semantic coding is another topic we've covered in class, particularly in regard to XML's and JSON's appropriateness for such marking. The semantic encoding mentioned here is fairly pedestrian as such things go: chapter headings or is Washington a city, state or person?, etc., but Sutherland leaves space for more interesting semantic markup in the future.
Sutherland concludes that digitization should be seen as a multistep process. I agree 100%, and I think it's clear Google similarly conceives of digitization as iterative, with it's "Getting there!" disposition toward accurate data and metadata (compared to OCA's metadata which is on the whole much much better). The slippery slope to this approach is that every step is a little less accomplished than it could be; raw OCR quality is contingent on page picture quality, corrected OCR needs somewhat accurate raw OCR for context. With that in mind I wonder if it would not be better to more closely consider first steps rather than applying so much functionality and completeness to later steps.
Subscribe to:
Posts (Atom)