Tuesday, November 3, 2009

Metadata for Linguistics

In keeping with this week's topic I wanted to see what metadata has been implemented for linguistics data. As mentioned in some of our readings, the problem of local nomenclature and linguistic barriers are glaringly obvious when dealing with the study of language itself, so I wanted to see how this has been handled. As a side note, when dealing with the web, finding resources in non-roman scripts has been given a boost recently by the decision to approve the creation of top-level domains which use internationalized domain names (IDNs), although unofficial URLs in Thai script have been available for some time from ThaiURL.

As this paper describes, in December 2000 the OLAC (Open Language Archives Community) was founded to promote online sharing of language resources. Problems specific to this domain were becoming apparent online, as language names can have several different romanizations (one example: Fadicca, Fadicha, Fedija, Fadija, Fiadidja, Fiyadikkya, Feddica- all the same language!) and different names within the same country, as well as different names between countries (Deutsch vs. German). In addition, language names can change over time, and different social groups may have different preferred names. The end result was a searching nightmare with very low recall.

The solution proposed is simple and uses tools that should be familiar to us all. Linguists are using Dublin Core and the OAI to bring resources together. There are a limited number of extensions to DC such as allowing every tag to be tagged with the language it is written in (subject.language), as well as tagging the metadata record itself with multiple versions in other languages. In addition there are domain-specific tags that list the functionality of the resource (transcription, annotation, lexicon, etc.) and what kind of specification it includes (phonetic, prosodic, morphological, etc.)

Although RFC 3066 is the DC and web standard for specifying language names, SIL's Ethnologue is much more in-depth: for example, the Karen languages do not even appear in RFC 3066. In addition, Ethnologue provides the ability to identify different levels of language groupings, such as families.

We can see from this example that it is not always difficult to create useable metadata for a specific domain: using a few pre-existing domain-specific tools and some pre-existing standards with extensions we can accomplish a lot. To look at the current situation, development of the OLAC is on-going, and a tutorial on language archives was held in January of this year. I look forward to future developments in the intersection of these disciplines.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.