Wednesday, November 4, 2009

Ontologies and data sharing

So what do y'all think of the idea that we could create one big ontology that describes everything? It's a pretty stereotypically librarian thing to do, but it might be the only way to realize Tim Berners-Lee's idea of the semantic web. Ephram Miles once told my Organizing Information class that he thought it might be overly ambitious and probably not possible to create an ontology that is broad yet descriptive enough to fully realize the semantic web in the way we dream about it today. In any case, the W3C just came out with the second version of its Web Ontology Language (OWL 2), which is a standard way of representing ontologies. The press release about this from the W3C doesn't mention giant ontologies, but it does talk a lot about discipline-specific ontologies. The tag line for the press release is "OWL 2 Connects the Web of Knowledge with the Web of Data." So I became curious about how ontologies might affect our discussions of data sharing.

The W3C lists a lot of semantic web case studies and use cases. I skimmed a few of these, and it appears most people are turning to ontologies to improve search. I'm interested, though, if ontologies might be able to improve data sharing. Tim Berners-Lee (TBL) talks about data sharing in his 2006 IEEE Intelligent Systems article, "The semantic web revisited." In fact, TBL suggests that in order for the larger semantic web to work, scientific communities have to first pioneer the way, much like they did in the early days of the web. There are a lot of ontologies out there. Just do a Google search for "biology ontology." I found many sites with many ontologies listed on them. So aren't we back where we started? There are lots of isolated "data sets" (read: ontologies) that don't really relate to each other. So we have the same problem with ontologies that we do with data: how do we assign authority? how do we create appropriate metadata? how do we make these islands talk to each other? TBL suggests we need more data mining, ways to infer knowledge from patterns, and distributed information systems to "crowd source" the ontologies. But the argument still seems circular to me. It seems like he's suggesting we use computer models to then build ontologies to then build computer models to connect the data.

Austin Forum on Science, Technology & Society

I had every intention of blogging on topic and even identified a couple of articles about metadata, but then I attended the Austin Forum on Science, Technology & Society. This month’s speaker, Gary Chapman, a Senior Lecturer in the LBJ School of Public Affairs, spoke on “The Internet and the Obama Administration – So Far.” Chapman worked with the Obama campaign and discussed the ways the Internet was used as a tool during the election, as well as the administration’s current web presence.

Chapman covered three primary topics during his talk: Obama’s online presence post-election, the issues of network neutrality, and online transparency. The third is the most related to topics of digital curation, but I thought I’d say a bit about the first two as well. The Obama administration used social networking sites like Facebook and Twitter to generate interest during the campaign and continue to use these sites post-election. According to Chapman, Obama was the first person to have over 1 million Facebook friends and Twitter followers. One primary aim was to raise funds using the Internet, which turned out to be enormously successful. Post-election, the administration has created a YouTube channel for the White House on which they broadcast a weekly address, and maintained Twitter and Facebook pages. Criticisms of these practices include questions about using of corporate companies for public purposes.

A primary topic of concern is over issues of net neutrality. Obama put Julius Genachowski into the position of Chairman of the FCC and charged him with enforcing network neutrality. The main question that arises however is whether the FCC has, or should have, the power to regulate the Internet. Chapman asserts that he does not believe this issue is a top priority for Obama at present, given the state of health care and international relations.

The third topic Chapman addressed was transparency on the web. Obama affirmed his commitment to freedom of information and web transparency on Inauguration Day and quickly launched sites that provided governmental information to the public, including data.gov (which Meg blogged about weeks ago) and Recovery.gov, which was launched in February to provide access to data related to the Recovery Act. The Obama administration is using these digital tools to increase transparency and access to government information. However, there has been some backlash, as the administration has not been completely open about everything, including negotiations with telecommunication agencies.

When reflecting back on issues of trust and authority, the government attempts at digital curation serve as interesting examples. Who owns the information generated by the government but made public through sites like Facebook and Twitter? Will Obama’s use of Twitter increase the need for or interest in finding a way to preserve tweets? Should the FCC be able to regulate the Internet? How will that change the free use of the technology? Will the transparency of the Obama administration continue beyond his presidency? What degree of openness and availability of information should we expect to have from the administration available online?

Let's Talk About OCLC

Katie's blog post got me thinking about OCLC and the power of metadata.

There are a few blog postings that might be useful to read:
I think that OCLC is interesting!

Innkeeper at the Roach Motel

Salo, Dorothea. (2008). Innkeeper at the Roach Motel. Library Trends. 57(2): 98-123. Retrieved from: http://muse.jhu.edu.ezproxy.lib.utexas.edu/content/crossref/journals/library_trends/v057/57.2.salo.html


Most of the articles we have read have largely focused on the wonderful possibilities of cyberinfrastructure and the attempts done so far to ease the transition from a largely analog world to a digital ruled environment. I have always been curious to know whether there were any dissenting opinions. Thanks to our class's Zotero account, I found Dorothea Salo's article "Innkeeper at the Roach Motel." It is certainly a dissenting voice amongst all the rosy and golden articles we've read this semester.* In fact, Salo compares institutional repositories to roach motels because according to her "whatever goes in there never comes out."

*Not that the golden and rosy articles are off the mark. Many of them are written very carefully and throuroughly, but Salo's article is so wary about the success of IR and open access that it made my eyes bug out.

Salo starts off bluntly: institutional repositories (IR) are not thriving, not catching on with libraries and university faculty in the way is needed for any kind of information transformation. Salo points out that the way people respond to the idea of an IR in their library has been of trepidition and negativity, citing a lack of trust and sdoubts of its credibility as major factors. She also states that initially, the idea of open access is appealing, but in the long run, it is not a major selling point for potential IR users.

Salo echoes many of the oft-discussed challenges of open access: issues of credibility, how to benefit from it monetarily and the general attitude that scholars have in regard to tenure and achieving academic fame. She states that, as wonderful as open access may sound, old habits take a long time to die and it may be a while before open access can catch on in the academic world.

Salo lists several factors that have led to IR's inability to revolutionize scholarly communication as promised and explains in detail of the inconsistencies and inefficiecies. IR and cyberinfrastructure requires participation from several groups: scholars/university faculty that submit work for IRs to collect, librarians or other staff members to manage and campaign the IR efforts, and the software developers to create and upkeep the technology required to run the IR.

The scholars and the university faculty members present a large barrier in the IR effort because not only do they not trust the credibility or status quo of submitting work in open access journals , but in general there is a distrust of any source digital as it is essentially contradictory to the tenure process. Also, in the area e-publishing with all its "pre-prints" and so on, not many people are familiar with the process and are not willing to educate themselves in this area. Librarians and other staff members that would be involved with the IR also display an open lack of trust of such a system and are reluctant to engage with it.

To put it bluntly, Salo calls IR "parasitic on existing research" because what they offer to people is not useful for them and what would be useful is not offered, because the program developers of IR technology is not collaborating with university staff and libraries. Essentially, everyone is in their little corner grousing about how this system with great potential is not working and ignoring the fact that if they collaborated, a newer and more useful system could emerge.

Salo examines different models that libraries have adopted in an attempt to create, manage and run an IR, lists their responsibilties and judges the advantages and disadvantages of each model. A IR manager is new area of expertise and there really no community of practice developed for this area of librarianship yet.

"Maverick Manager": librarian specifically in charge of the IR, has many different responsibilities and skills, but has no real place in the library's organizational structure. The maverick manager has freedom to experiment, but very little to no resources. These types of IR manager models have a large turnover due to lack of community, status and funds to do their job.

"No Accountability" Model: the IR is built by IT or an outside source is contracted to build an IR. Responsibilities are dispersed amongst the librarians for promoting and collecting the IR sources. This system is advantageous because it reduces the amount of groundwork that libraries have to do in order to create the IR, but this model also provides countless chances for miscommunication and requires librarians to travel a steep learning curve in getting familiar with the software.

"Consortial" Model: libraries share IR. This system is efficient, yet increases trouble with outreach, tech support, and content development.

Cooperative Model: this modle has been the most successful by far. Usually the IR is launched by a university administrator rather than a librarian. The librarians mediate content deposits, perform active searches through the Net and other sources for content to put in the IR and sometimes an automated workflow will automatically deposit work in the IR. The biggest problem with this system is that it is very expensive.

Salo ends this article with solutions that would increase the appeal of IRs and ensure a more positive rapport with IRs. First of all, support needs to be shown at home, libraries, library school programs, library organizations. IRs must be integrated into other library programs and priorities. IRs must take an active role in searching for content and be sensitive to faculty needs, digitize analog content, and develop relationships with software developers.

Dorothea Salo's article presents a different view on the whole idea of IR and outlines some of the challenges that prevent IRs from being truly successful. I think this answers a lot of my questions as to why I haven't heard more about IR until I started the Information Science program.

Social Meme Tracking: Here Come the Swedes

The first time you might have heard of Twingly was back in early 2007 with the launch of the Twingly Screensaver. The screensaver was an interesting innovation. For those of you who have not heard of it, the screensaver features a constantly updated list of the titles of every single blog post that is written anywhere in the world. The list scrolls in real time and there are links to each blog post which provide the URL, the name of the blog, the author, and the first two lines of the post. While an interesting development, from a digital curation perspective, the general criticism lobbed against the screensaver was that it just wasn't "useful." The critique goes something like - what is a user going to do with so much raw, unfiltered data?

Since the screensaver, Twingly has actually been looking at making the blogosphere more accessible and searchable. The Twingly Blog Search and Twingly Micro-Blog Search have become wildly popular in Europe. Anecdotally, I've been trying to find a way to get an RSS feed for one of my favorite Middle East opinion columnists, Robert Fisk, who writes for The Independent. After failing to find a RSS feed for Fisk alone (without the other Independent columnists) using Google Blog Search, Technorati, and even the RSS page on the website of the Independent, I was able to do so using the Twingly Blog Search. How they managed to provide a RSS feed for something that the source website does not provide is beyond my comprehension, but good for them!

Probably the most interesting development that the folks at Twingly are working on is what they are calling a "social meme-tracker" by the name of Twingly Channels. The web-based application is still in a limited-release beta version (you'll need to find an invite code, try: ALTSEARCHENGINES), and while there are more than a few kinks that need to be ironed-out, the idea seems pretty revolutionary. Essentially, Twingly Channels provides a social network based around memes rather than individuals (as is done on Facebook and Twitter). Anyone can create a meme and name it whatever they want (from as general as "Entertainment" to as specific as a particular person, place, or thing). Other users can "subscribe" to this meme the way that one "follows" a person on Twitter. These users can add things that are relevant to that meme - blog posts, tweets, articles, etc. and all of the subscribes can comment on these posts and specify whether they "like" this particular post. The more subscribers like the post, the higher the post is on the stream.

The only major issue I see at the moment is that Twingly has yet to determine an adequate way of organizing these memes (or "channels") apart from a simple list which orders the channels according to their popularity. The implications for a tool like this are huge. One can easily envision scientists creating "channels" which are relevant to a particular topic and using this tool to quickly share data, and comment and rank it. Unfortunately, at the moment, most of the channels relate to things in Sweden (in fact most of the comments are in Swedish) so it's too early to see whether something like this is catching-on globally, but there is a lot that I like about meme-centered social networking as opposed to person-centered social networking.

Possible Causes of Inadequate Metadata in the Institutional Repository

In her article Innkeeper at the Roach Motel, Dorothea Salo writes about the problems of institutional repositories. While she discusses many different issues, I am going to focus on the causes she lists for insufficient metadata in institutional repositories.
First in her list of causes is the issue of trust. Many people are hesitant to trust librarians with their data because they believe their lack of subject area expertise will result in problems with metadata management. Salo does not discuss wether or not this belief has merit, only that it inhibits researchers in specific disciplines (such as the hard sciences) from trusting their work to librarians. Instead they trust their work with science experts who frequently are not working to solve long-term preservation issues.
Additionally, the different requirements of publishers make placing data into a repository difficult. One such requirement is that the meta-data have a set phrase which includes a link. The problem is that some of the popular institutional repository software, for example DSpace, does not have that particular functionality, further contributing to some author's hesitancy to use institutional repositories.
The information professional put in charge of an institutional repository is also at a disadvantage if they cannot program. They are left with the basic version of their repository without any patches or plugins that are not already a part of the system. Frequently the task of applying dublin core metadata to a particular work does not even fall to the information professional, but instead it falls to the author of a particular work. They information professional is left trusting the authors themselves to invest the time needed to ensure proper metadata entry, which rarely happens. The result is inconsistent use of metadata throughout the repository, greatly decreasing the usefulness of the works that are in it.
Repository systems also contribute to the problem of having good metadata in a repository because they are frequently difficult to use. They often overestimate the ability of the user (who is not always an information professional) and expects them to know what fields are required and even what some more obtuse fields mean. Salo does mention that EPrints does a better job of handling metadata than other repository systems because it displays metadata fields based on what type of content is being deposited, making the process more user-friendly.
The final cause of insufficient metadata in an institutional repository mentioned is that repositories are still not seen as a priority to many libraries. Without proper funding and sufficient staffing to maintain and oversee the repository and the metadata going into it, information professionals don't have a chance of ensuring the repository is useful. The library ensures the failure of a repository when they don't allot adequate time and money to its maintenance.
While I believe all of Salo's points are valid and the way metadata is handled should be reconsidered, I still believe the repositories currently face a larger problem, namely that authors do not want to put their work into the repository at all. Still, if we are going to put data into the repository, we should be able to find it again later and that necessitates that we take a good look at the way we are implementing metadata.

Ex(ml)tra, extra, read all about it!


This week I found a brief article detailing the metadata problems and solutions encountered by the Library of Congress when it participated in the National Digital Newspaper Program, a twenty-year initiative to create a national repository of historical newspapers in digital form. The end goal of the project was to have a complete and searchable online bibliography of newspapers published in the U.S. from 1690 onwards. In addition to this, select local newspapers of historical significance would be converted to digital form for instantaneous public access.

Several problems arose in creating a standard metadata for the project. First of all, print newspapers themselves do not fit well into any of the traditional cataloging standards. This is especially true when one tries to accurately and comprehensively describe newspapers from all fifty states and across drastically different historical eras. (17th-century spelling and abbreviation are particularly dicey for modern readers.) Second, most of the historical newspapers being covered by the project no longer exist in print form but were only saved as microfilm, introducing another form in the provenance. Finally, the NDNP wanted a system that was going to allow complete interoperability between the federal repository and individual state digital repositories as well as among the state repositories.

Ultimately, the NDNP elected to use the METS standard, entered in XML. Four separate METS elements were used: title document, issue document, page object, and reel document. The first three refer to the original physical form of the paper. XML is flexible enough to be able to incorporate different types of issue numbers while remaining interoperable, which was key for the program. Murray views the page object as the most important as he asserts that the page is the basic information-containing block of a newspaper. The reel document only applies to those papers digitized from microfilm but represents important administrative data.

I found this article interesting mainly for three reasons. First, it involves a system to provide more access to historical documents, including colonial era documents, which is a goal ever near to my heart. Second, I work in the serials unit at the Benson Collection and have much first hand experience (and frustration) at the incredible amount of variance in newspaper conventions. Third, I liked the fact that the NDNP came to a relatively straight-forward and simple solution, while maintaining interoperability. (Though, of course, it helps that there was central oversight in this project...) All in all, I thought it was a good example of the benefits conferred by standards as well as not making the solution harder than it has to be.

Tuesday, November 3, 2009

Metadata for Linguistics

In keeping with this week's topic I wanted to see what metadata has been implemented for linguistics data. As mentioned in some of our readings, the problem of local nomenclature and linguistic barriers are glaringly obvious when dealing with the study of language itself, so I wanted to see how this has been handled. As a side note, when dealing with the web, finding resources in non-roman scripts has been given a boost recently by the decision to approve the creation of top-level domains which use internationalized domain names (IDNs), although unofficial URLs in Thai script have been available for some time from ThaiURL.

As this paper describes, in December 2000 the OLAC (Open Language Archives Community) was founded to promote online sharing of language resources. Problems specific to this domain were becoming apparent online, as language names can have several different romanizations (one example: Fadicca, Fadicha, Fedija, Fadija, Fiadidja, Fiyadikkya, Feddica- all the same language!) and different names within the same country, as well as different names between countries (Deutsch vs. German). In addition, language names can change over time, and different social groups may have different preferred names. The end result was a searching nightmare with very low recall.

The solution proposed is simple and uses tools that should be familiar to us all. Linguists are using Dublin Core and the OAI to bring resources together. There are a limited number of extensions to DC such as allowing every tag to be tagged with the language it is written in (subject.language), as well as tagging the metadata record itself with multiple versions in other languages. In addition there are domain-specific tags that list the functionality of the resource (transcription, annotation, lexicon, etc.) and what kind of specification it includes (phonetic, prosodic, morphological, etc.)

Although RFC 3066 is the DC and web standard for specifying language names, SIL's Ethnologue is much more in-depth: for example, the Karen languages do not even appear in RFC 3066. In addition, Ethnologue provides the ability to identify different levels of language groupings, such as families.

We can see from this example that it is not always difficult to create useable metadata for a specific domain: using a few pre-existing domain-specific tools and some pre-existing standards with extensions we can accomplish a lot. To look at the current situation, development of the OLAC is on-going, and a tutorial on language archives was held in January of this year. I look forward to future developments in the intersection of these disciplines.

Pirates of Academia

This is the same old song of the digital age when you remove data from its physical container someone somewhere is going to put it online. This is true for music, books, and video games so we should not be surprised that scholarly publishing has been drawn in as well.

An article published on the Internet Scientific Publications www.ispub.com details a study conducted over a six month period examining the file sharing activities that occurred on a single site devoted to medical professionals. What they found was an Electronic Library section where users could go and request that papers be posted. Each user was allowed three paper requests a day and 83% of articles requested were fulfilled. Over the six month period of time 5,251 articles were posted with a total viewer ship of 23,461.

This is a very interesting finding and I really want to know how wide spread this sort of practice is and how it really affects the bottom line of scholarly journals. The journals in this case are in a different position in comparison to the record labels who have not fared so well in the fight against pirates and the emergence of new technology, but the audience for the journals is different than that of the labels.

I for one have never paid directly for an academic article, and I doubt seriously that a any of you reading this have either. I pay the university a good sum of money and then through the libraries the articles I am interested in are made available. The universities, in this model, are the prime customers and they could never pirate access to a journal. The article states that roughly 703,830.00 was not paid in access fees over the six month period that the study follows. This may be true but many may have used this forum to get access to article because it was easier than going to their local library and retrieving the same material there.

This is an interesting study but, in the end I don’t believe it tells us anything that we did not know. We live in an age where all digital data is shared and this has finally moved to the realm of scholarly publishing. Attempts to secure this data will be subverted and overcome, but my guess is in this case the big journals get the majority of their income from institutional sources and will not reconsider their current model based on this activity.

The Sloan Digital Sky Survey

This week I tried to stay on topic and I tried to find just one article and stick to it but that wasn't possible. I looked at the different places the Sloan Digital Sky Survey (SDSS) data is used. I've been very interested in the SDSS because it is the largest sky survey so it's possible to rely entirely on the data from SDSS to model large sections of the sky. However, the services I tried out this week sometimes pull data from other sky surveys to supplement SDSS.

On the SDSS and SDSS Sky Server site there are a few different ways to access images from the survey, there are many photos throughout the sites that can be enlarged including some image galleries and a listing of Famous Places. The way to look at the original data is to locate the original Flexible Image Transport System (FITS) files. There's currently seven data releases on the SDSS site and for every data release it's possible to either select coordinates for viewing or navigate through several levels in a directory to get to a list of FITS files. To view the FITS files download a FITS viewer. Not surprisingly, it's difficult to engage the public with the original data from this survey.

It's probably an understatement to say that astronomers have complicated ways to describe astronomical objects. After creating multiple metadata schemas to organize the massive amount of data from telescopes, astronomers are still having a hard time with descriptions. More accurately, there's images and data about the images, but it still isn't processed and understood by humans. Since there are so many astronomical objects that need to be classified, a team of astronomers, cosmologists and a few other experts created Galaxy Zoo. In order to understand how galaxies are formed and how to describe them, Galaxy Zoo has crowdsourced JPG images from SDSS. There is absolutely nothing technical on the Galaxy Zoo site, they don't mention FITS or metadata, all they are interested in are people's answers to simple questions about JPG images of galaxies.

At first I was skeptical of how successful Galaxy Zoo could be, but after reading what they've discovered I've embraced their practices.

Google Sky is another place where the SDSS images are used, the March 2008 NVO Newsletter provides some details on Google Sky. Google Sky started as an extension to Google Earth and in March 2008 it was also launched at google.com/sky. Google Sky is marked up with Keyhold Markup Language (KML), which means that anybody can participate in tagging. You can also convert a FITS file to KML and upload it or even convert a VOTable to KML with the NVO's Visual Integration and Mining Tools.

I looked around for information on how SDSS was ingested into Google Sky and if there's bad KMZ files in Google Sky, but I wasn't able to find anything.

A service that is similar to Google Sky is sky-map.org. I think this collection is curated even better than Google Sky, the maps of astronomical objects are much more compelling than the ones Google Sky selects and there's more insight to how the SDSS collects data. For instance, in this map the survey "stripes" are visible. There isn't a lot written about sky-map.org aside from what's found in their wiki and a few old articles from magazines.

A fourth place that uses SDSS is the Microsoft WorldWide Telescope (WWT). This software coplies with International Virtual Observatory Alliance standards which means that it uses the Astronomy Visualization Metadata (AVM) standard, so any user can tag in WWT. I haven't used WWT but I read on Wikipedia that it made an astronomer cry.

Making DRM More Usable...or Profitable?

A. Arnab and A. Hutchison, “Fairer usage contracts for DRM,” in Proceedings of the 5th ACM workshop on Digital rights management, 2005, 7.

This article from two students at the University of Cape Town takes a look at the functionality of DRM.

DRM is widely discussed as a tool to enforce copyright as we move into the digital medium. The authors agree with many other critics of the various software technologies that constitute DRM in observing that DRM does not really enforce copyright. They argue that while copyright enforcement is theoretically possible, DRM lacks the sophistication to allow for
fair use of digital objects. Fair use is an exception to a usually applicable copyright restriction and is frequently argued on a case-by-case basis; it's therefore extremely difficult to allow for this very important exception in a programmatic way. Consequently it's overlooked and rights holders instead are allowed to sidestep the whole issue by issuing some agreed contractual terms that usually favor themselves over the consumer or user.

The authors take some further time to observe how DRM is unable to restrict reproduction and distribution (the core protections of copyright) because of the nature of hardware and software. DRM at the application (iTunes) and operating system level (Windows, OS X) cannot prevent reproduction and distribution. Media-specific DRM (like CSS and AACS) can be compromised, moreover media is trending toward direct digital distribution. DRM at the chip level can also be skipped so long as computers continue to have removable components. Provided this continues (I certainly hope so) and systems can support multiple operating systems (again, let's hope so) DRM doesn't seem to have a future as a legitimate copyright enforcement tool.

The authors instead argue that DRM is used as a licensing tool. They present two improvements to the licensing model that might allow for more input from the user and more granularity in the rights holders' licensing terms.

Use licenses in DRM systems are typically explained to a computer through a Rights Expression Language (REL). They authors argue that RELs should be expanded to incorporate a more nuanced license-negotiation process that goes beyond a simple request-response that is used today (i.e., do you accept this license? Yes or no). In this model users would request a set of rights, the licensing server would evaluate the request and serve up a license with those rights and the terms, user accepts, denies or renegotiates.

This process could begin again after the user has purchased the digital object and the license if the user wants some new right. This could allow for better fair use control.

A second approach provides a credential construct in RELs that would allow users to identify themselves as various roles (a journalist, a university student, a researcher, etc.) and thereby request different rights.

While I agree that DRM is a poor copyright enforcement tool, I'm hesitant to embrace these more nuanced licensing models and processes right off the bat. As the authors note, their models allow for a lot of flexibility and "newer business models for the rights holders." Is that a good thing? I can easily see a business micromanaging their rights model to create a maximum profit. In fact they could read this article and wonder why it hadn't occur to them to create a pay-as-you-go licensing model that kept users coming back for more rights.

But, so long as digital objects like datasets, music, recordings, and so on have rights holders that need to exert control over their materials, a better model is needed, one that works. The authors here have faith that a more nuanced licensing model ultimately would create fairer deals because users/researchers have a say in what they'll accept and in what they want. To the extent that a model would facilitate such feedback, I'm for it. The DRM industry is looking for standards, I hope they settle on one that's reasonable and expandable.

Is metadata data, or is data metadata, or are they different?

Dr. Winget posted this article on delicious, "Metadata vs Data: a wholly artificial distinction."

In this article, Terry Jones writes that it is wrong to distinguish between metadata and data. He opines that David Weinberger was right to state that all data is metadata. In saying this, what he means to get at is that people use metadata to search. In doing so, they don't necessarily only want to search by something that an authority has deemed metadata. Often, users wish to search full text in the same way they can search a title or an author.

He uses the example of the unix file system, which uses one command ("find") to search what has been pre-determined to be metadata, and another ("grep") to search inside what is deemed to be data. Although it is possible to combine these searches (a commenter on the blog provides the method), Jones' point is that separating the metadata and the data into two systems is problematic.

FluidDB, a cloud-based database, supposedly does not make this distinction, but holds all data together (all of it as metadata) in one searchable unit, made up of many pieces of data (or metadata). The author (and CEO) claims that this method is more robust because no data is more important than other data, meaning nothing will be lost or deleted, and anything can be added.

Functionally, I think it is a great idea to store metadata and data together, making either and both searchable. One of the advantages of digital storage is ease of search. It isn't really necessary to search the card catalog, get a reference to an item, and then go to it on the shelves. A federate search can directly pull on article from terms either in the title, in tags, or from the document.

However, I think conceptually, it could be very confusing to lump data together into an ever broadening category, or to lose the distinction between some metadata and an original item. For example, many collections add administrative (meta)data to items that should be kept distinct from what might be considered metadata but is also a part of the original item (title, date). For example, processing archivist is a piece of metadata that could be searched via a keyword search, but I would not want it to be confused as part of an original item.

Metadata is certainly data, and perhaps all data should be treated as metadata in search, but to me, strictly speaking, true metadata can only exist if the data to which it relates exists. A title with no article is just a phrase.


"Geekfest, you scoff?"


Trying to find an interesting article about metadata and knowledge sharing, I stumbled across this article: "Holy Knowledge Sharing, Batman!" on the site KMEdge.com (the KM is for Knowledge Sharing). While this article leaves something to be desired in terms of academic rigor and depth - it raises interesting points: Comic-Con brings together over 125,000 people, who are part of a major knowledge sharing community. This community exists in person for four days a year, and online all the time. The article's author notes that this community "influences marketing strategies and budgets for movies and TV shows. It shapes the direction of new video games and helps predict strategies for future toy sales..." One wonders how this major, largely informal, data-creating, data sharing community is being used by companies for marketing, production of new products and a host of other purposes. I don't know exactly how companies currently harvest the information of communities like this - if there are spiders crawling through comic-con message boards and blogs, searching out key terms, or if they have individuals pouring over such sites. It seems to me, however, that determining an effective way to gather and synthesize this information would be a profitable tool - making the most of data sharing in the for profit realm.

OAIster and OCLC

This week I found a blog entry that was actually on topic! In HangingTogether, the OCLC/RLG blog, the entry OAIster Update: More Access & No Conditions explains more about the impending changes to the Open Archives Initiative (OAI) as OCLC takes over stewardship of the mass of metadata harvested from OAI-compliant repositories. As a background, "OAIster is a union catalog of millions of records representing open archive digital resources that was built by harvesting from open archive collections worldwide using the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH)." Previously supported by the University of Michigan, the size of the aggregation of records became too large for UM to single-handedly care for. Enter OCLC.

In reviewing some of the older blog posts about this transition (in September), it seems there was a slight kerfluffle about the whole thing. OCLC originally laid out the plan as such:

  • Continued collaboration with University of Michigan.
  • Records will be freely discoverable along with all other content in WorldCat.org. However, it will not be possible to limit a search to OAIster records alone.
  • Data providers must request their records not be harvested, otherwise they will (opt-out policy).
  • OAIster contributers must request free access to records in FirstSearch.

Now, I'm not entirely sure what all this means, but people were not happy. Reading the comments from the initial entry, people were concerned OCLC was going to sell all kinds of data, not just provide access to the metadata, and that records previously not on WorldCat would suddenly be there against the author's wishes. And more.

A follow-up entry was necessary, and OCLC actually had to change their terms and conditions with OAIster to specifically state that only metadata will be harvested. This clarification brought more comments of general dis-satisfaction, correctly wondering why a license was necessary and why things would be changing so much, in terms of copyright and such. As never having really dealt with OAIster issues, I don't fully understand their grievances, but they sounded legitimate.

So, this leads us to the post this past week, summarizing the first month the change-over happened. First thing, the terms and conditions that concerned people is now gone. The post says: "In keeping with the open style of the Open Archives Initiative community, if you make your metadata available for harvesting, you must intend for it to be harvested. We will also feel free to index it, provide access to it, and allow Google to crawl it." The opt-out strategy is still in place. OCLC is also allowing OAIster records to be searched separately from other WorldCat items, because OCLC "decided that the OAIster aggregation is an important enough destination for finding open access content". Or because there was an uproar, whatever. The comments to this post are much more positive, so apparently OCLC either made the right changes or wore everyone down.

On the surface, it seems fantastic that a larger organization is taking over the harvesting and accessing of metadata for the sustainability of the information. Although on WorldCat.org, there isn't an easily identifiable place to just search OAIster records. But, as was mentioned in a blog post, harvesting is hard! Maybe we should be thankful that someone is doing it (assuming it's being done correctly with good intentions, no GoogleBook excuses). It will be interesting to see what happens in the next few months and the reaction from users/authors/institutions. I'm sure HangingTogether will have more on this subject.

Monday, November 2, 2009

Folksonomy

This week I've tried to blog at least close to on topic, and so have chosen to talk about a short article from the November 2006 issue of D-Lib Magazine by Elaine Peterson entitled "Beneath the Metadata: Some Philosophical Problems with Folksonomy."

Peterson begins her article with a review of the foundations of traditional classification. She states that traditional catalogers are part of the Aristotelian tradition and adhere to the same basic principles of contraries, particulars, and categories. Of particular importance is the idea of contraries. In traditional cataloging, an item cannot be both "A" and "not A." A photograph of a single horse, for example, cannot be assigned the subject term "white horse" and the subject term "black horse." It must be one or the other. Additionally, the priority of the author's intent in traditional cataloging is another one of its distinguishing characteristics.

In a folksonomy, however, philosophical relativism is the underlying principle. Because of this, instead of it being impossible for an object to be both "A" and "not A," the assignment of tags is entirely up to the judgement of the individual. Therefore, it is possible for an object to be both one thing and its opposite. The reader has the priority in a folksonomy, rather than the author, and contrary interpretations can exist. As Peterson points out, "if all interpretations be of equal worth, [and] users can continuously add tags to articles, at some point it is likely that the whole system will become unusable." She likens it to the story of the Chinese emperor who wanted an accurate map of China, and ended up getting a map that was very accurate but was also the size of China. Such a thing is accurate, but not useful.

Peterson also points out that meta noise is another problem with folksonomies. Meta noise results from inaccurate, irrelevant, and even inadvertently erroneous (such as spelling "white horse" as "whit horse") tagging. She quotes David Weinberger, who views folksonomic classifications of the Web as "'messy and inelegant and inefficient, but it will be Good Enough.'" Peterson holds, however, that while it may be good to allow users to supply their own tags, the resulting classification system will not produce an efficient index, and so is not really all that good.

Overall, Peterson believes that inconsistencies within the folksonomic classification system will always exist. In traditional cataloging, there is a right way and a wrong way to classify something. This is not true with a folksonomy, and this, it seems Peterson believes, is its main problem that can cause the system to break down. "Folksonomists," she says, "are confusing cataloging structure with personal opinions." It is good to have personal opinions, and to be able to express them. However, opinions do not lead to effective classification, and therefore do not provide for effective searching.

METAe and ALTO: Mapping Physical and Logical Structures

Stehno, Birgit, Alexander Egger, and Gregor Retti. "METAe-Automated Encoding of Digitized Texts." Literary and Linguistic Computing, 18, No.1 (2003): 77-88. Available at: http://llc.oxfordjournals.org/cgi/reprint/18/1/77. (Accessed 11/1/2009).

This article describes how the Austrian-based METAe project applied METS (Metadata Encoding and Transmission Standard) to encode automatically extracted metadata from page images, especially metadata describing document layout. Like OCR engines which extract text from image files, the METAe engine extracts layout elements (such text or graphic blocks) by relying on formal rules and syntactical principles. The project built and compiled a recognition model that would map elements in the physical structure of page layout (text blocks of differing sizes and graphics blocks) to logical ones (e.g. paragraphs, titles, footnotes). This would then create an ALTO ('Analysed layout and text object') file which is an XML file that consists of both the layout structures and the full text of a book page. The ALTO file and information formatted in a variety of metadata standards such as Dublin Core and DIG 35 are then incorporated into a METS schema, which (if I'm understanding correctly) serves as an outer wrapper that provides the structural map and holds all of the metadata for the object.

In devising the METAe engine, they decided to use METS rather than TEI (Text Encoding Initiative) for their encoding since TEI was "far too inexplicit for the purpose of automated recognition" (9). This struck me as an interesting point since if TEI is to be the encoding standard for text in the future, it will need to be more amenable to automatic application. METS was also preferred as it allowed for METAe to add metadata at any logical level and to include metadata from different formats and standards, including pointing to metadata external to the METS document. For example, the METAe project can use a "DMDID" (Descriptive MetaData Identifier) attribute to generate tags linking a journal issue to its appropriate MARC record on a web server while also using Dublin Core tags to provide descriptive metadata for each article/contribution.

What this article by Stehno, Egger, and Retti does not really address are the difficulties (and there must have been some) that the METAe project encountered in developing its engine. It seems extremely useful to have an engine that can map the physical structure of a page onto logical structures, but certainly there must be structures that it is better and worse at recognizing. I've also had some difficulty in discerning precisely what happened to METAe--it seems to have lived on as ALTO rather than as the METAe tool itself. As the METAe website (http://meta-e.aib.uni-linz.ac.at/index.html) details, the project ran from September 2000 - September 2003 with partial funding from the European Commission. The website also announces the metadata engine has been marketed as digitization software under the name docWorks/METAe Edition, however, though the docWorks page exists and seems to offer a product that performs the tasks that METAe does, there is very little mention of METAe on the site. The one mention I was able to locate was a notice that as of August 2009, the Library of Congress has taken over maintenance of the ALTO XML schema from CCS Content Conversion Specialists GmbH (the company which produces docWorks). This appears to be a sign at least of ALTO's success as a schema, as LC has created a new ALTO editorial board to "help shape and advocate usage of the standard."

I guess the questions that I'm left with after reading about ALTO and METAe are ones about the relationship between standards and schemas and the tools used implement them. To what extent do the constraints of the tools with which schemas were initially partnered affect the future use of those schemas, even after a given tool has been put aside or altered? Can schemas created within the context of automation equally useful when used for hand encoding or human quality control?

Wednesday, October 28, 2009

Improved Twitter feed curation

Was pleasantly surprised this evening as I logged into Twitter. They now have Lists in Beta - something I've been wanting for some time.

1) I can create lists for subjects that interest me, and move those that I follow into that list (functionally this is akin to creating labels in Gmail). Doing this will create a link on the right of the home-page Twitter feed that I can click and read only those in that list.

2) others can subscribe to this list - which creates a kind of sub-follower rating. Let's say for example that I compile a list of those who tweet about Contemporary Asian art: this would be its own list that interested people could subscribe to (essentially bypassing me and my tweets but reading those that I have aggregated - or perhaps just going to my list to harvest them for one's own list). This is an idea similar to reading filtered blog subscriptions or friends groups in Facebook as well as customized blog-rolls.

3) There are great implications for use for larger or more widely trusted entities whose lists could be very popular to subscribe to. The introduction of this to Twitter's networking capabilities is a tremendous improvement for ease of use and both content and contact searching.

Now, if only YouTube would similarly jump on board with this concept.

Get the gamers involved


I had intended to write about this topic earlier in the semester and had forgotten all about it until I saw a reference in this week's reading by Amy Friedlander, The Triple Helix: Cyberinfrastructure, Scholarly communication, and Trust. Friedlander discusses how problems that can be cleanly parsed into discrete tasks are well suited for a distributed capacity model which allows open, but still structured, participation from the public. She provides the example of protein folding, which is the topic of an April 2009 Wired Magazine article: Gamers Unravel the Secret Life of Protein.

The protein chemistry world has a biennial World Series competition to see who can predict the shape of a protein only knowing the sequence of its constitute parts (Community-Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction, or CASP). CASP surveys labs around the world to find proteins that are about to be solved, and compile a list of puzzles online.

David Baker, whose team had dominated the competition since 1998, had been using Rosetta@home, similar to SETI@home, which farmed out computations to volunteer PCs distributed globally - providing Baker with the equivalent of a supercomputer. However, the computers were unable to complete certain puzzles, which humans should be able to solve, having better spatial reasoning. Baker's friend David Salesin, a computer scientist, brought him together with Zoran Popovic, another computer scientist and graphics expert, and the three developed what turned into a massively multiplayer competition. Gamers are given a multicolored knot of spirals and clumps, which they fold and wiggle into its optimum shape.

Baker then entered potentially accurate CASP protein structures into the biennial competition. Of 15 submissions, 7 finished "in the money" and one took first place. The gamer team, led by a 13-year old, beat the best biochemists. Baker was also hoping to find prodigies... when "Cheese" (the 13-year old) was asked how he did it, he said, "it just looks right."

Baker has given the players a new challenge to design a new protein drug with the right size and binding properties. Baker will synthesize and test the most promising structures and if any have value in the real world, the gamers will share in the credit.

The article doesn't specifically address authenticity or trust, but the gamer submissions are not automatically deemed correct, even though the game is based on laws of physics. Submissions are reviewed by CASP, and/or tested in the lab. But, given the fact that there are more ways to fold protein than atoms in the universe, and they arrange in a fraction of a second, collective efforts are crucial.

Old Problem, New Scale

This article provided by George Mason University's Center for History and New Media (which is quite an interesting program in and of itself), notes that the problem of authenticating sources used for research is not a new one in the humanities. Scholars in fields such as art and history have always had to use their specialized knowledge to try to determine the authenticity of the works or documents which they examine. However, Bearman and Trant, the authors of the article, feel that the current digital era has led to an enormous increase in the scale of the problem. Whereas in the past it took much specialized skill to create a credible copy or forgery, the mere act of hitting a button or two can create a digital object nearly indistinguishable from the original (other than the lack of physicality, of course.) Also, more and more scholars are turning to digital research as a way of maximizing research time and for the thoroughness it affords. Yet, they no longer can use physical clues (for example, handwriting) to infer authenticity

Bearman and Trant see a number of possible solutions to the problem without really recommending one single method. They divide their methods into public methods, secret methods, and functionally dependent methods. Public methods may either be social (e.g. creating a collecting institution of record along the lines of a certified website which raises issues of who performs the certification) or technological, such as using public key encryption to create digital signatures for documents. Secret methods consist mainly of technological clues buried in empty space within files such as stegonography. Functionally dependent methods probably do the best job in authenticating because the very working of the file is tied to the end user's ability to authenticate it. However, this is much trickier to set up and probably a bit overkill. (After all, there is a certain amount of responsibility traditionally borne by the scholar...)

I found this article interesting largely because of my history background. When evaluating physical reprints or translations of primary sources, a historical scholar will check the publisher (kind of an early form of public authentication, I suppose.) But in digitizing, nearly anyone can be a publisher, so count me as convinced, scholarly responsibility notwithstanding, that some sort of authenticity determination scheme will be necessary as more and more primary documents are digitized

Searching Twitter

So there have been some recent developments in the past week that have received a good deal of news. First off, Microsoft thought that it was going to make a decent cut into Google's mammoth share of the search market through a deal with Twitter which would allow tweets to show up in Bing searches. Unfortunately (for Microsoft) Google managed to make a deal of their own with Twitter so that tweets will also show up in Google searches. While Bing might still have the upper-hand since its Twitter search is already live, I want to take a moment to grieve for the older forms of searching Twitter.

The New York Times had a good article this week about Twitter's approach to innovation. At this point, the concept of the "democratization of innovation" is probably familiar to most people who know a thing or two about the interwebs, but Twitter was pretty exceptional in its adherence to a credo of bottom-up innovation. At the outset, the service afforded only the ability to write "micro-blogs" of 140 characters. Methods of searching these micro-blogs developed organically by twitterers themselves. Nearly every search convention:

1. The '@' before the screen name of another user
2. The letters 'RT' before a reproduction of something that already been tweeted (or a "re-tweet")
3. The '#' used before a certain topic so that topics with more than one word were easily searchable (e.g., #iranelection)
4. Even the term 'tweet' to refer to a single micro-blog

...all of these were developed by twitter users and later adopted as standard form by Twitter developers. Now that Microsoft and Google are in on the game....sigh. So much for democracy.

Admittedly, it's too early to tell what kind of affect this new way to search Twitter is going to have on things like the use of the hash tag, but it's pretty easy to see why Microsoft and Google want to get in on the Twitter action. Back in March, on the Tech Crunch blog, Michael Arrington wrote a post about just how useful Twitter is to businesses. Because tweets are so short, many people (and I mean MANY) use Twitter to do nothing more than gripe about things, especially their experiences with products. One can only imagine how useful this information is to businesses. In fact this information is SO useful, that Twitter, itself, has been valued at nearly $1 Billion, even though it has around 35 million users (compared to Facebook's 175 million). With that kind of valuation, it's surprising that Twitter hasn't decided to cash in and (like Facebook) allow more ad space. Since Twitter is so amenable to being searched, however, they may never have to get into the ad game in order to make serious money. Here's to hoping that Twitter holds-on to its original democratic, bottom-up approach to innovation despite all of the corporate interest.