Readings covered this week:
- Beagrie, N. (2006). Digital Curation for Science, Digital Libraries, and Individuals. International Journal of Digital Curation, 1(1). Retrieved from http://www.ijdc.net/ijdc/article/view/6/5.
- Angevaare, I. (2009). Taking Care of Digital Collections and Data:‘Curation’and Organisational Choices for Research Libraries. Liber Quarterly, 19(1), 1 - 12. http://liber.library.uu.nl/publish/articles/000278/article.pdf
- Gold, A. (2007). Cyberinfrastructure, Data, and Libraries, Part 1: A Cyberinfrastructure Primer for Librarians. D-Lib Magazine. Retrieved April 1, 2008, from http://www.dlib.org/dlib/september07/gold/09gold-pt1.html.
- Gold, A. (2007). Cyberinfrastructure, Data, and Libraries, Part 2: Libraries and the Data Challenge: Roles and Actions for Libraries. D-Lib Magazine. Retrieved April 1, 2008, from http://www.dlib.org/dlib/september07/gold/09gold-pt2.html.
But within the context of today's class:
- Digital Curation: Refers or implies: collection development, selection and appraisal (whether we easily recognize the selection process - our experience is with humanities collections; selection might look very different in the sciences), organization, description, providing access models, management, and preservation. Primarily conveys a responsibility of stewardship.
- Cyberinfrastructure: Refers to all of the background stuff that makes "the curated entity" (see below) possible. Hardware an software are the obvious elements, but also policy decisions, training, the culture of scholarship, and funding models for building and supporting the system. Infrastructure is needed to support...collaboration, scholarship, "the curated entity," and...what else?
- Digital Libraries / Digital Repositories / Digital Archives: This was really sticky. My initial question was: "what do we call these large digital libraries that we've been talking about?" The term 'Digital curation,' while formally defined as being about management and preservation of digital collections of any kind (of any size), in practice usually refers to the management and preservation of huge scientific and humanities data sets. What do we call these huge data sets? What constitutes a cyberinfrastructure project?
This doesn't get us any closer to what we should call these huge data sets. We tried "Curated Data Sets," but the people who are working with humanities data didn't think that was appropriate (although humanities data is still data); "Curated Digital Repositories" although the word "repository" implied a "finished" artifact (you put finished materials in a repository; are repositories and archives synonymous?) and much of what we're talking about in this class is data that is very much in use; either by the creators, or by the community.
I think we need to come up with a new word, and for that matter, a new model of dealing with this stuff. All of our models are based on a paper system. This stuff is not paper; and the paper model is limiting our ability to meaningfully collect, describe, and provide access to these materials.
That was a lot of the discussion, in one form or another.
We also spoke at length about Personal Information Management, and how that relates to digital curation. Again, we discussed the challenges of coming up with a new not-paper-based-model, which is difficult, individuals' management techniques (folders v. chaos). I discussed my experience (Minds of Carolina and Managing the Digital University Desktop) and how this kind of research is showing that digital curation and cyberinfrastructure is very much dependent on personal information management techniques. Adding to the definition of cyberinfrastructure: developing a culture of organization (when we had paper, a culture of filing arose - we're more dependent on tool development now to find things - will a culture of organization of personal information make digital curation easier in the long run?)
Finally, we discussed the act of preserving these huge data sets - and whether we're talking about preserving the data, or are we also talking about preserving the actions that are performed on the data? My theory is that in the future, the data will be available, but the real value-added will be the algorithms written to make sense of, or organize, or hypothesize against (i.e., contextually retrieve) the data. So in the future, the valuable stuff will be the algorithms (or is this already the case? and are we preserving the algorithms? If so, how are they described, how is credit bestowed to the programmer, and what does the relationship between the programmer and the scientist look like?)
Finally, Finally, We talked about how many of the blog posts have the question, "why does x [usually Google] do y [usually, not allow users to add metadata]?" (for example, my question: "Why doesn't Google allow users to create layers (like in google maps or google earth) that include all of the correct metadata?" - so, for example, there'd be a UT Austin layer on top of Google Books (and Harvard, and Michigan...). We can't answer that question - " why won't x do y?" - so we tried to work out what might be a better set of questions, such as:
- How does this project fit in with larger institutional goals?
- Why has this institution decided to commit time, money and people to this project?
- How do the characteristics of the project support their institutional goals?
- What are the assumptions underlying the institution's implementation of "the entity"?
Random Questions:
-Has anyone digitized documents written in shorthand? Court drawings?
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.