Thursday, September 10, 2009

Notes from Class - September 10

These are Megan's notes from today's class. They're necessarily brief - let's try putting addendums in the comments this week.

Readings covered this week:
We started out trying to define terms - This discussion, we decided, will be ongoing, as some of these terms are too broadly defined in the literature, or the definitions are conflicting.

But within the context of today's class:
  • Digital Curation: Refers or implies: collection development, selection and appraisal (whether we easily recognize the selection process - our experience is with humanities collections; selection might look very different in the sciences), organization, description, providing access models, management, and preservation. Primarily conveys a responsibility of stewardship.
  • Cyberinfrastructure: Refers to all of the background stuff that makes "the curated entity" (see below) possible. Hardware an software are the obvious elements, but also policy decisions, training, the culture of scholarship, and funding models for building and supporting the system. Infrastructure is needed to support...collaboration, scholarship, "the curated entity," and...what else?
  • Digital Libraries / Digital Repositories / Digital Archives: This was really sticky. My initial question was: "what do we call these large digital libraries that we've been talking about?" The term 'Digital curation,' while formally defined as being about management and preservation of digital collections of any kind (of any size), in practice usually refers to the management and preservation of huge scientific and humanities data sets. What do we call these huge data sets? What constitutes a cyberinfrastructure project?
We made distinctions between digital repositories / archives / libraries in that repositories' primary function is authentic and reliable preservation. They have interfaces, but those interfaces, so far, are secondary to the preservation function. Digital Libraries might use repositories on the back end, but include tools and services (bells and whistles) that allow people to interact with the collections in innovative or intuitive ways. Digital Archives - we only talked about in terms of whether they conform to paper-based archives (artifacts which no longer serve their original purpose)

This doesn't get us any closer to what we should call these huge data sets. We tried "Curated Data Sets," but the people who are working with humanities data didn't think that was appropriate (although humanities data is still data); "Curated Digital Repositories" although the word "repository" implied a "finished" artifact (you put finished materials in a repository; are repositories and archives synonymous?) and much of what we're talking about in this class is data that is very much in use; either by the creators, or by the community.

I think we need to come up with a new word, and for that matter, a new model of dealing with this stuff. All of our models are based on a paper system. This stuff is not paper; and the paper model is limiting our ability to meaningfully collect, describe, and provide access to these materials.

That was a lot of the discussion, in one form or another.

We also spoke at length about Personal Information Management, and how that relates to digital curation. Again, we discussed the challenges of coming up with a new not-paper-based-model, which is difficult, individuals' management techniques (folders v. chaos). I discussed my experience (Minds of Carolina and Managing the Digital University Desktop) and how this kind of research is showing that digital curation and cyberinfrastructure is very much dependent on personal information management techniques. Adding to the definition of cyberinfrastructure: developing a culture of organization (when we had paper, a culture of filing arose - we're more dependent on tool development now to find things - will a culture of organization of personal information make digital curation easier in the long run?)

Finally, we discussed the act of preserving these huge data sets - and whether we're talking about preserving the data, or are we also talking about preserving the actions that are performed on the data? My theory is that in the future, the data will be available, but the real value-added will be the algorithms written to make sense of, or organize, or hypothesize against (i.e., contextually retrieve) the data. So in the future, the valuable stuff will be the algorithms (or is this already the case? and are we preserving the algorithms? If so, how are they described, how is credit bestowed to the programmer, and what does the relationship between the programmer and the scientist look like?)

Finally, Finally, We talked about how many of the blog posts have the question, "why does x [usually Google] do y [usually, not allow users to add metadata]?" (for example, my question: "Why doesn't Google allow users to create layers (like in google maps or google earth) that include all of the correct metadata?" - so, for example, there'd be a UT Austin layer on top of Google Books (and Harvard, and Michigan...). We can't answer that question - " why won't x do y?" - so we tried to work out what might be a better set of questions, such as:
  • How does this project fit in with larger institutional goals?
  • Why has this institution decided to commit time, money and people to this project?
  • How do the characteristics of the project support their institutional goals?
  • What are the assumptions underlying the institution's implementation of "the entity"?
This sort of comes from our discussion last week (I did not provide notes) - where I argued that I think that expecting Google to provide metadata is not realistic and does not fit in with what they consider to be "their job." I'm not sure if I really believe that - they've been arguing that Google Books will be a useful resource, and it will demonstrably not be useful if the metadata is awful - but, on the other hand, Google Books represents Google stepping outside of their comfort zone, and the growing pains might be the lack of metadata. But that was last week's discussion.

Random Questions:
-Has anyone digitized documents written in shorthand? Court drawings?

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.