Monday, August 31, 2009

A case study of a curated life sciences project

The case study I read was conducted by the Digital Curation Centre (DCC), and can be found at: http://www.dcc.ac.uk/docs/publications/case-studies/SCARP_EMAP.pdf

The case study was conducted on the digital curation of the Edinburgh Mouse Atlas Project (with the obligatory acronym EMAP.) EMAP's primary goal is to create and manage a mouse embryo gene database (which got its own acronym -- EMAGE.) The case study conducted by the Digital Curation Centre seemed to me to be a good example of the case studies we will be conducting during this class.

To find out how the EMAP was organized and curated, the author visited the website and conducted interviews, as well as attending meetings at the project. She compares their process and final product to the DCC Curation Lifecycle (DCCCL?).

The DCC Curation Lifecycle is a Digital Curation Centre diagram depicting the process of curating digital objects. From conceptualization of the item to transforming for preservation, curating, describing, and involving the community, the lifecycle is complex to say the least, and I'm unclear on how quantitative a comparison to such a chart could be. The image below depicts the cycle; you can view the original here: http://www.dcc.ac.uk/docs/publications/DCCLifecycle.pdf




Although the case study using the curation lifecycle of EMAP states that the EMAGE database is the first scientific visualization of gene-expression (what -- no other spatial maps of Mouse genes?) the study also reported several areas that need further problem-solving.
1. Copyright inhibits use of needed images.
2. New ways of displaying data need some standardizing.
3. Data entry methods are slow and at times unreliable.
4. Quality Assurance is not, in fact, assured
These four puzzles struck me as being probably very common problems in creating infrastructure for projects, and carrying out almost any digital curation project successfully.
Quality Assurance in particular has historically been a -- perhaps purposefully -- overlooked area in hardware and software development.

The author and editors of this paper lean toward believing that data can and will be shared in the life sciences, and they write, "Recent life sciences studies, such as the Joint Data Standards Study5 (2005), have demonstrated the value of sharing and re‐using data." They believe that publicly funded projects will therefore continue to build tools with this data sharing, and that the main problem to overcome is the planning for creating the "cyberinfrastructure" for all that data.

While they may not be entirely correct that scientist will willingly share their data, I agree that standardization and planning must continue before useful tools and projects such as EMAP can spread to a community beyond determined and/or specialist users.

Post1: RepositoryMan

RepositoryMan is the blog of Les Carr, a researcher and lecturer in the School of Electronics and Computer Science at the University if Southampton, Southampton, Hampshire, UK. His area of research, for the purposes of this blog is e-research, and specifically how researchers use the repository he manages. As of 2007 this repository contained “about 10,000 records and gets about 600 new deposits per year. There are 2400 papers published since 2004, of which 1400 have open access full texts,” its serves as the “bibliographic record of all school output” (post “Some Background,” 10 July 2007). This repository meets the criteria of digital curation as it in breadth: it attempts to collect all the published material of the school’s faculty and utilizes tools. The repository creates a bibliography of the institution’s work, and is an important tool for academic networking (keep track of citations, sharing your research via social networking). Furthermore, Carr explains the value of the repository to the faculty in terms of how they might utilize other tools with the repository to describe their work (using cites like dipity.com to create a timeline of an individual’s publications) so that a depositor does not feel like placing their work into the repository is equivalent to a dead end.

Likewise Carr uses visualization tools to inform the research and faculty community about what is contained in the repository and how it works. When a New Head of School was hired, Carr created a 'thumbnail wall' of all the files stored in" the repository that he could print in poster format to emphasize the importance of the repository.

Furthermore, one of the most interesting entries on Carr’s blog was a recent posting titled "Institutional Visualization" 2 July 2009, wherein Carr uses tools to show “the spread of research on a particular topic across the institution.”



Monday, August 17, 2009

Welcome to INF 385T - Digital Curation

Hello: This blog has been set up for students of INF385T - Digital Curation at the School of Information at the University of Texas at Austin. Students in this class will make 10 blog postings throughout the semester on topics pertaining to digital curation.