Tuesday, September 22, 2009
Cyberinfrastructure for Liberal Arts
Liberal arts scholarship is largely viewed as a solitary process, with the individual responsible for creating and managing their own content and tools to disseminate their research to a larger scholarly community. In this article, the authors raise the question of what will drive a shift toward collaboration. They summarize many of the issues we have discussed in recent weeks, from the issue of the publication process and the link to recognition, tenure, and promotion, the need to finance such initiatives, underuse by scholars, and the ever-present intellectual property concerns.
Early issues the authors explore include digitization happening piecemeal, with little in the way of standards and metadata. Resources were scanned as needed by an individual scholar, but without consideration of a broader collection (or infrastructure). The implication here is that large community or discipline-based steps are needed to create an infrastructure that will be useful and utilized by the community that contributes to it.
The authors identify two broad approaches for conceiving of cyber scholarship. The first is data-driven and focused on building and growing digital repositories, using the power of computers to store, sort, search, and analyze large amounts of data in disparate repositories. The second is concerned with harnessing the technological capabilities and developments, such as social networking, to grow a new genre of web activity for scholarship that is defined by the participants. In these two approaches, there is a tension between idea of computer as an instrument for manipulating and mining data or as a tool for making social and intellectual connections. The two approaches are not in opposition to one another however, and both can be developed to change the current state of scholarship, allowing for more shared data and collaboration, as well as the incorporation of more data formats into humanities research deliverables.
Predominantly, the authors present the idea that it comes down to getting people involved and finding a way to finance the growth and spread of resources to as large a community as possible. The development of tools and products, whether created for profit or using an open-source model, is integral to the success of a humanities cyberinfrastructure. One goal the authors offer up is to stop thinking about issues within an institution and to shift to considering connections between institutions, which opens the idea of a community working to create a collaborative research and teaching environment. To this end, they offer up eight recommendations for institutions, which were published in Our Cultural Commonwealth (one of our readings for this week). The overarching message to liberal arts schools is to get involved with the development of cyber scholarship in the humanities: invest, encourage, develop polices, cooperate, cultivate leaders, establish centers, create standards, and produce digital collections.
This article does a nice job of laying out the key issues for any cyberinfrastructure efforts. I think it functions as a reminder for the liberal arts community to engage themselves, in ways that perhaps the scientific community is already advancing. I am interested in the distinctions between liberal arts and other scholarly groups, particularly in the types of data, formats, and collaborative possibilities for researchers in this community.
Google Books. This time as a data set.
This is what started it:
"We’d expect libraries to have information like the place of publication and author birthdates. But it’s harder than you think – some libraries don’t know when books originated and simply used '1899' as a placeholder for any book they didn’t know. You see mentions of 'Internet' in publications dated 1620 – that’s because the mentions were in 1997 from a journal that’s been published since 1620."
Um, libraries are at fault? Not Google, with their secretive scanning practices? I can't tell if he's paraphrasing Google Books' Matthew Gray here or musing himself, and I don't know if that makes a difference. The two were at IBM’s Transparent Text symposium, which is/was a 2-day affair this week with the tagline "Text is Data". Zuckerman and Gray spoke on a panel titled "Analyzing the Written Record" but only Zuckerman has an attached abstract.
Never fear, however - the solution is at hand. A person "can analyze language models from texts published in different years and then make intelligent guesses about when a book with bad or no metadata was published." Excellent! So a stereotypical college freshman, up late working on a paper due the next day, can perform some simple algorithmic tests to make sure the source he/she is citing is correct. Of course.
I think Zuckerman is trying to illustrate how Google's vast data sets can lead to some interesting investigations by comparing information. This could be very advantageous to someone studying language evolution or printing history - until the dates aren't correct. Or will there be enough examples that the amount of incorrect metadata will be at the end of the curve? But do we know how many books vary from their physical counterparts? Unfortunately, the example in the blog is an incomplete link, so I'm not exactly sure how it's supposed to work.
To give this author a break, he seems quite well-rounded, as his blog covers Africa, international topics, and technology AND he's a researcher at Harvard's Berkman Center for Internet and Society. With that pedigree I'm not sure I understand how a basic understanding of library cataloging procedures (WorldCat anyone?) was overlooked. And I'm very interested in seeing new research come from the datasets of text and for now Google is the only one with that set, flawed or not.
Making Data Mining Sexy (Even to Literary Scholars)
It is no accident that this article by Catherine Plaisant et al. bears the provocative title "Exploring Erotics in Emily Dickinson's Correspondence with Text Mining and Visual Interfaces." Its authors, a multidisciplinary group of researchers associated with the Nora Project (www.noraproject.org), are intentionally trying to make the text mining of digital collections both sexier to and more accessible for literary scholars. Their project is rooted in the premise that to impact research in the humanities, computational methods must facilitate and participate in the process of interpretation. Accordingly, the Nora Project seeks to develop a digital architecture that will enable literary scholars with no special IT expertise to perform text-mining on an electronic corpus of 18th and 19th century literature.
Plaisant et al.'s specific case study reports on their use of computational tools to classify individual Emily Dickinson letters as "hot" or "not hot" (that is, erotically-charged or non-erotic). The user first manually rates a representative set of sample documents on a five point color-coded scale of "hot-ness." These ratings then serve as "training" that permits a multinomial naive Bayes (NB) algorithm (executed by a D2K data mining tool) to rate the remainder of the letters in the collection on the hot-ness scale. The user then can examine some of the automatically-rated letters and accept, reject, or modify ratings before re-running the process in order to increase the accuracy of the rating predictions. The better the training set, the better the classification results.

As you can see in the above figure, which depicts the interface, color is used to help visualize the classification. The tool also suggests a list of words (see right side of figure) that it posits may be indicators of hotness or not-hotness. To help the literary scholars "read" the results, the interface offers several pop-up FAQs (E.g. "What do the purple squares mean"). The tool also enables the user to create scatter plots to visualize the rated documents with regards to variables such as date (no correlation between letter date and eroticism was found). A video demonstration of the interface is available through the Nora Project website at http://noraproject.org/nora_ol_video/. Though the study does not appear to have been large, the authors report that an Emily Dickinson expert who served as a test-user found that the results of the data-mining genuinely shed new interpretive light on familiar texts. Additional user feedback led to ideas for improving the interface, such as allowing users to assign extra weight to certain words.
Though Plaisant et al concede that the merit of the project will depend on how successful literary scholars are at publishing papers that draw on their tools, they assert that their initial findings are positive: their architecture and interfaces are usable by those without experience in data-mining, and their tools appear to yield provocative and inspiring insights for literary scholars.
Both this article and Blackwell and Crane's "Conclusion: Cyberinfrastructure, the Scaife Digital Library and Classics in a Digital Age" constructively provide clear examples of how computational and analytic tools applied to curated digital collections might advance scholarship in the humanities. However, features such as citation identification, syntactic and metrical analysis, or interpretive classification all rely on unfettered access to digital collections. Does this mean that literary scholarship will advance unevenly across periodizing lines? Since much of modernist and all of post-modernist literature still falls under copyright, there's no available digital corpus on which to apply these new computational tools. Will Sappho and Emily Dickinson scholarship benefit from added digital functionality while Denise Levertov scholarship remains mired in analog methodologies? One solution might be to make data-mining and other textual analysis services more portable. This could allow a motivated scholar to scan his or her own personal library of relevant books and, using OCR, create his or her own digital research collection. Data-mining could then be performed on this mirco-collection. It's not an ideal solution, but neither is excepting most post-1923 literature from the methodological innovations of digital scholarship!
Digital Curator, re-defined, again.
While many of the comments left on the post commend Rubel for his ideas, I have doubts. Does the internet need subject curating? One interesting comment likened digg to a corner poster shop with Hannah Montana posters, and Mahalo as the Louvre. Basically – popular doesn’t equal quality. Another commenter called for Digital Archeologists to come and dig up all the buried material before the curator can make an assessment. Finally, a poster mentioned that they had in fact been a curator for years, though her title was “Renaissance Person,” fastcompany.com .
Perhaps the most interesting thing about this post is that I cam upon the article by looking at Megan's delicious page, tag: digital curation...
Monday, September 21, 2009
Digital Repository Sustainability - frustrations, but no answers
In this blog entry, Salo vents her frustrations about repositories that lack preparing for the future, mostly in terms of long term funding plans. One of the concerns she mentions is that arXiv is looking for a new funding model because Cornell University is tired of paying for the entire tab to host and operate the site. If Cornell no longer wants to pay to keep arXiv running, and no one else comes along to pitch in, what happens then?
The second frustration Salo talks about (which she expanded upon last week in another blog entry called “Object lesson: when researchers run repositories”) is trying to find a long-term disciplinary repository outside of her institution and running into more sustainability problems. In Salo’s search for an appropriate disciplinary repository she found two possibilities – one repository restricted by geography and the other... which made her angry. The second repository, which she discusses more in last week’s blog entry, was apparently cobbled together and then abandoned by a pair of researchers. She lists other examples of institutional and disciplinary repositories that have failed to plan for the future and wind up needing funding rescues to keep going. Sometimes the repositories get rescued, but sometimes they don’t, and when the latter happens, these repositories end up abandoned or shut down, and data ends up lost. Salo calls this lack of succession planning “flagrantly irresponsible.” The italics are hers, used for added emphasis to convey her mounting frustration.
The point she makes in both of these blog entries that our field’s “continuing error... is that we have no infrastructure or plan for accomplishing these rescues [of repositories].” This is one of many major concerns with digital repositories. Not only is there the problem that we keep mentioning in class with getting people to share their documents and data in repositories, there is also the concern of whether the repositories will continue to exist even in the short term, let alone for the future. How can we expect people to deposit their information in repositories if there is the risk that the repository will be shut down or abandoned due to poor planning? As more and more information becomes and is born digital, repositories will keep popping up online, and Salo predicts that more and more digital repository rescues will be needed in the future because of a failure to plan long-term. This prediction seems likely and highly frustrating.
One thing Salo fails to offer in her blog entries is any sort of real solution (which her critics point out somewhat harshly in the blog comments), but she does call for an open discussion of this issue in a “cross-institutional fora”. From what we keep reading in class, sustainability is a huge issue for all digital information, not just for digital repositories. But, though people keep writing about sustainability and talking about it, there doesn't yet seem to be many solutions for it. I suppose this is part of what is challenging, frustrating, and a little exciting about all of today's (and tomorrow's) digital development.
Poetry means more XML so that new tech can break the lines

Related to the text mark up that we read about this week, I found this article: Transcendental Data: Toward a Cultural History and Aesthetics of the New Encoded Discourse from 2004 useful in conceptualizing just how fundamentally important metadata and mark up is for every piece of raw electronic data.
Not even a simple word is just what it appears on the surface, the layers supporting the output of that word are complex, including the mark up and software that is used to create and then view it. Every "bit" of information must be represented, for example, many of us have seen what happens to Microsoft Word "smart quotes" when transferred to html.
Both Microsoft Word and html are almost ubiquitous, and yet the two do not really communicate well. Proper standards of metadata and encoding should ensure that this communication disconnect is kept to a minimum. Liu writes in this article that even creation, or authorship, is fundamentally changed into a new experience by today's technology. But in addition, machines, the interpreters of this new medium, must know how to present data.
In the instance of a poem, XML, for example, can be used to code each aspect of that poem with a commonly (interoperable) understood mark-up, rather than exact procedural instructions. Procedural instructions are too specific, and fallible. They often do not extend to unforeseen circumstances. For example, giving the exact size in pixels of a word will not necessarily translate to a small screen like those on mobile devices.
This principle of scalable metadata (in this case XML mark-up) should be applied to any data sharing online project (or whatever we are calling them) that needs to be future-proof and robust.
That means less linear instructions like, "put this letter next to the left, make this 24 pixels"
and more breaking up of text or data into cross-platform recognizable objects like, "this is a line, this is a stanza, this is a capital letter, this is twice as big as whatever the normal size is"
Scientists aren't sharing data
by Caroline J. Savage, Andrew J. Vickers
The PLoS data sharing policy states that “publication is conditional upon the agreement of authors to make freely available any materials and information associated with their publication that are reasonably requested by others for the purpose of academic, non-commercial research.” - That's a pretty clear guideline. The authors of this article requested data from ten researchers who published in PLoS but they were only able to get data from one of them.
Click here to see the chart that illustrates the responses they got from researchers.
The people who gathered the requested data don't appear to be doing research for private industry, I'm assuming that they are doing research at universities or for non-profits. So in theory they are working for the greater good of scholarship. Perhaps I'm wrong about this, but I do think that data sharing in a field like private pharmaceutical research is very different from sharing data that's gathered at a university.
This article by Savage and Vickers reinforces the one insurmountable problem with sharing data - it takes a long time to annotate data so that other people can understand it. Somehow the one person who sent them data had it all annotated ahead of time. Otherwise, they were able to do it within a few hours, but that doesn't seem possible.
The other problem with sharing data is that it is uncharacteristic of scientists, so to expect this of them is a shift in how they think about their work. One of the scientists that was contacted responded that they just didn't know this was a guideline of submitting to PLoS and that if they had known ahead of time they never would have submitted. It turns out that their data was not supposed to be released to a third party.
As if this article isn't already enough bad news, Savage and Vickers mention a previous similar study done in which requests for data were sent to authors who published in American Psychological Association (APA) journals. The APA study had a 25% return on data requests. The reason why this is discouraging is that the APA only requires researchers to share data to "verify the substantive claims through reanalysis", which means that when researchers are asked for their data, they're essentially being put on trial. When Savage and Vickers requested data, they claimed to be looking for a new hypothesis based on the data, so it was clearly for their own research. Also, the APA data sets were merely survey responses while Savage and Vickers were requesting data from medical studies. Not surprisingly, it's easier to acquire raw data from surveys than data from health sciences, but this is likely a reflection of how long it takes to annotate complicated data. It's just unfortunate that data which is detrimental to health can't be shared as easily.
In order for digital curation to function, researchers need to deposit their data into repositories. If they aren't willing to share data with their fellow researchers, there isn't much hope for them to hand the data over to strangers.
Sunday, September 20, 2009
The Open Dinosaur Project: Crowd-Sourcing Dinosaur Science!
While poking around in Michael Nielsen's bookmarks on delicious, I found this one for The Open Dinosaur Project, and I thought it was a good example both of a digital curation project and of the use of crowd-sourcing. The goals of The Open Dinosaur Project are " to involve scientists and the public alike in developing a comprehensive database of dinosaur limb bone measurements, to investigate questions of dinosaur function and evolution." Phase I of the project will involve using data submitted by the public to discover patterns of limb bone evolution in ornithischian dinosaurs, and how this relates to the evolution of locomotion in this group. They accept submissions of measurements harvested from scholarly literature or measurements obtained first hand (though scholarly literature is their primary source). And, they accept submissions from anyone, regardless of age, education, or qualifications. Preliminary results will be blogged on the website, and the final paper will be submitted to a scholarly journal for peer review. All of those who contribute to data collection will be listed as junior authors. The projected dates for completion are as follows: completion of data collection by 1 February 2010, completion of data analysis by 1 March 2010, and submission of final paper by 1 April 2010. Friday, September 18, 2009
In the news: Google & reCAPTCHA; researcher scooped after publishing data
In other news, Science just reported that a researcher who posted her data online was scooped before she was able to publish a paper from that data. Laura Bierut deposited her data related to genetic studies of addiction in dbGaP, NIH's database of genotypes and phenotypes. The NIH policy says that she should have had an embargo period of 9 to 12 months during which no one could use the data for publication. Her embargo period ends on September 23rd. But back in March, Heping Zhang, a Yale researcher, submitted a paper based on the data to the Proceedings of the National Academy of Sciences (PNAS). Zhang's paper was published on August 31st, and Bierut quickly responded by sending emails to Yale, PNAS, NIH, and colleagues. She was even polite, considering the circumstances: "'[T]his was likely an unintentional act, [but] this incident remains very concerning,' she wrote, adding that it 'sends a very chilling message to investigators.'" Yale took down a press release about the study and NIH froze Zhang's access to dbGaP. On September 9th, they retracted the paper. Currently, they are investigating the situation, and Zhang won't comment until that investigation is completed.
Bierut is quoted in the Science article as saying, “I think NIH and PNAS moved very quickly to resolve the issues" and "The good news is I think the system worked." But, the paper is still on the online version of PNAS because of the "very awkward consequences - librarians got confused about papers being cited that no long exist," (which is a pretty ungracious way for the PNAS editor-in-cheif to put it - "confused"? c'mon now). Bierut has been invited to submit her paper to PNAS, but one has to wonder if it will receive the same response as if her paper had been first. As one researcher quoted in the article put it, the situation leaves "the gate open for predators with no investment in the data to do quick-and-dirty analyses that pick the eyes out of the data without looking at any of the subtleties." It's worrisome to me that horror stories like this will discourage researchers from publishing their data.
The other day in class I was speculating that perhaps some of the problems with sharing data might be alleviated if we rewarded scholars for publishing data, not just papers. I think it might be less scary for researchers to put their data out there if they knew that, if all else fails, they'd at least get some credit for the data. But I'm starting to understand that publishing data probably just wouldn't have the same impact as publishing the analysis of that data, especially if it's a breakthrough. Data is just data until someone gives it meaning. And it would probably take a while for the community to realize the impact of a particular dataset (like if the number of publications using the data were tracked). The amount of credit someone should receive for a particular dataset would take a while to assess. In the end, I think the ultimate goal of open science is such a worthwhile one that we have to keep trying, but it feels like it might be a long road.
Wednesday, September 16, 2009
Flickr launches Flickr Galleries
Please explore the Galleries that people have created here: http://www.flickr.com/galleries. One can comment on the gallery, but not rate it or add it (which I think would be great features). Two of my favorites are Moleskinerie and tea. I would love the ability to enter metadata on gallery collections as a whole - as part of cataloging a curated exhibit goes beyond each individual image to address the relationship the images have to each other and how they interact.
The architect that I code metadata in his image database for - has been explaining to me his vision for a metadata hierarchy or schema. His vision seems to be one that is personal, intuitive and based upon his own use of images as visual inspiration. It is fascinating to work through and help develop this kind of metadata functionality in a workable way. For example: He wants a category called "Activities" - in which would include courtyards, outdoor cafes, public squares, people. He would like a category "Skylines" which would include architectural facades against the sky.
Going back to the Galleries on Flickr - I am inspired by seeing the arrangements and limited selections selected by individuals, named brief and poetic titles. I find these to be creative and mental exercises that not only allow others a glimpse into how others see the world, but allow others the chance to craft and curate one's own expression through the pairing, juxtoposition, arranging and poetic guidance of images that speak to us.
Flickr user "nonac" has created numerous Galleries on Flickr and what I especially love about them is that they contain written text - the voice of the curator. The eyes have it.
I expect to follow great things on Flickr because of this new capability - but would like to see gallery-specific tagging in addition to the ability to favorite these collections. The potential to comment on and discuss these arrangements as well as the possibility for artists to collaborate via this application is very exciting.
Can this creative use of digital/social-tagging and curating be extended to other creative uses of digital content online? I find YouTube ripe with creative usability problems. Amazon.com allows one to create book lists. I am very interested in how one might move beyond the "accumulation" level of online content collection to the "curating" level - and how that might play out in other creative online platforms.
Digital collection in the wild
The team that Developed the MVP examined the various digital storage strategies for scholarly digital resources that were already in practice around the campus to see how the users interacted with these collections in order to ascertain their needs. Unsurprisingly they found that a large number of users worked with systems that had little metadata consistency, were run by faculty that lacked the skills to manage digital collections or digital infrastructures and had little way of making their digital assets available to others. The Media Vault Program is a campus wide solution to these problems.
The MVP program gives users the tools to up load datasets adjust the permissions and the licensing as necessary for each digital object. The team has created a tiered ingest work flow system that allows the user to prepare media assets in one stage then set them to a submit area. Following submission the assets are available to the user in the Archive section. In this stage the use cannot effect changes to the asset. The user can then publish the asset in websites or galleries or in other forms. This tiered system allows the user to work with their data and not worry about damaging any files.
The collection is backed-up of site. For added and added level of protection. A couple personal favorite functions of the system are; the system assigns permanent URLs (PURL) to the assets and the administrators of the collection assist in the creation of metadata during the ingest process if need be. These two steps should go a long way in assuring that the assets will be locatable.
The MVP group has also developed strategies for managing the data through out its life cycle . This includes data appraisal and preservative storage to encourage reuse.
This Project provides a excellent model to how a comprehensive data storage solution should look. I find it heartening that individual groups (or in this case campuses) are meeting with the problems of digital collections head on. These individual collections will hopefully provide the experience and practice that is need so that we can use the digital deluge rather than be washed away by it.
Data Sharing: Empty archives (if you build it... will they come?)
I'm very interested in the traditions, notions, and expectations of data sharing that we've been discussing the past couple weeks. Even with the best intentions and adequate incentive, the complexities appear to outweigh the ultimate benefit (at least at this point in time).
This article opens with a brief history of the University of Rochester, NY digital archive. Launched in 2003, the university had spent six months of research and marketing to determine that a university repository would be supported and used - both internally and by the public. Today, the repository is mostly empty. Researchers were very supportive of the idea, but once available, couldn't find their data, didn't know how to use the archive, or complained of lack of time and resources to properly transfer their data.
Nelson observes that like institutional repositories, "if you build it, they will come" also does not yet apply to other data sharing efforts. The concept is widely embraced, but the advantages often do not outweigh the concerns of the researchers. Repositories such as arXiv.org, the Protein Data Bank and the International Virtual Observatory Alliance should be considered exceptions, rather than the norm; their respective fields have established traditions of open access and data requirements.
Mark Parsons, an open-data advocate, manages a global program aiming to preserve and organize data from the International Polar Year project. Data needed, and was mandated, to be made available quickly. "Part of what is driving that is the rapidness of change in the poles," says Parsons. "If we're going to wait five years for data to be released, the Arctic is going to be a completely different place." Unfortunately, he found that the infrastructure needed to be created, and this was not available through his program's resources. His team was able to delegate the collection work to national coordinators, "data-wranglers", who contact the investigators and prepare the data for the databank. Sweden has become one of the most successful at data-wrangling - it formed a subcommittee to correct the lag in data collecting and is now housing data from smaller projects that may not reach the international databanks. However, unlike many countries, there is no practice that requires funded projects to submit project data to data centers.
The complexities to open data vary wildly. Parsons points out that even if the data is available, it is not always clear where the data belongs. Furthermore, funded data centers may be forced to reject data from projects without a related funding source. Other fields grapple with the concerns over quantity and quality. Do data centers keep only the information most likely to be of value (as understood by current science), or keep the vast quantities of data accumulated? Should they be concerned with naive users making premature discoveries, or allow for a fresh perspective? Szabolcs Márka, the lead scientist for LIGO (Laser Interferometer Gravitational-Wave Observatory)
asserts that data centers are not in the business of data analysis, but are charged with the task of provide access to accurate information. Other issues involve proper citation of data and credit given to researchers.
Infrastructure, standards, and culture all need to be created or adapted in order for open data to be as effective as the ideals imply. It is clear that it will take global initiatives, financial support, and likely many test cases to create adequate middleware and workflow techniques to push the data to the centers.
Crowd Sourcing Art Exhibitions
Kasprzak contends that "collections and curated exhibitions are about creating links, developing narratives, and composing responses to perennial questions and ideas. These collections and groupings are then presented in ways so that they will effectively reach audiences." She goes on to state that "it is the work of curators and all cultural workers to perform extensive research on who is or could be the audience for a particular exhibit or collection." While crowd-sourcing may be a useful and effective way to control a repository of widely-agreed upon facts (e.g., Wikipedia) is crowd-sourcing a useful way to think about curating things that require the creation of something as intricate as a narrative? I honestly don't know. I haven't seen either of the exhibits. I'm tempted to say that I can't imagine the the collections would amount to anything more than bunch of random stuff. Which is not to say that the collection couldn't be interesting in some way, but it probably wouldn't be anything like what a curator would assemble.
This article also made me think about this week's reading and about what happens when "old" notions about authority in the sciences and humanities are broken down because of some sort of technological revolution. In the latter part of the article, Kasprzak quotes Clay Shirky's book Here Comes Everybody "As with the printing press, the loss of professional control will be bad for many of society's core institutions...the printing press broke more things than it fixed, plunging Europe into a period of intellectual and political chaos that ended only in the 1600s." I would argue that the chaos in Europe created by the printing press lasted much longer. I'm certainly not the first person who would make a link between the French Revolution and the revolutions of 1848 with the intellectual freedom that the printing press afforded. These revolutions (or at least the French one) produced great upheavals in our ideas about authority. I would say that the Internet has, while less dramatically, done something similar. Do curators have the right to "subject" museum goers to their ideas about what is "noteworthy, good, and compelling?" The question here, is one of expertise. Museum curators are trained as curators and, one would assume, have a better idea about how to create an exhibit that is "noteworthy, good, and compelling." However, while it's easy to talk about what is "good" in the natural sciences (it either "works" or it doesn't), "good" in the humanities is clearly much more subjective. This is a huge topic, and I'm not going to pretend to make any final statements about it. Suffice it to say, that I think that the digital revolution may ultimately force humanists to rethink their ideas about authority and about what is "good."
Scientific Workflow Systems: a way to relieve the programmers?
Last week we talked about how important database administrators and programmers are to making data sets usable by scientists. It's a big enough job to preserve and ensure access to data. Who is responsible for performing the complex manipulations that makes the data useful? Does it require subject specialty? Do scientists themselves have to become programmers? I started wondering if there is a way to automate some of the most routine things that scientists might do to data in database. I discovered a new term, scientific workflow, which means breaking down a scientific problem into a series of steps. I think it works like this: take raw data from source, perform X process, take that data, perform Y process, take that data, perform Z process... etc, until you reach the desired output. Basically, a workflow is a way to describe a set of computer processes so that they are understandable by humans. It turns out there are such things as scientific workflow systems (an example is Kepler) that allow scientists to design, implement, reuse, evolve, archive, and share these workflows without knowing much about computer programming. Scientists can select computational tasks from drop down menus and the workflow is represented in graphical form. The application interprets the scientists' workflow into a programming language (in the case of Kepler, java).

A phylogenetics workflow implemented in the Kepler system.
There has even been some work on finding ways to make workflows and workflow systems standardized and efficient. Researchers at UC Davis have devised a system “to make it easier for scientists to design workflows, to clearly show how workflow products were derived, to automatically optimize the performance of workflow execution, and otherwise make scientific workflow automation both accessible and practical for scientists” (McPhillips, Bowers, Zinn, and Ludascher, “Scientific workflow design for mere mortals.” Future Generation Computer Systems, 25:1, May 2009, 541-551. doi:10.1016/j.future.2008.06.013). They argue that scientific workflow systems should provide: well-formedness, clarity, predictability, recordability, reportability, reusability, scientific data modeling, and automatic optimization. They propose the “collection oriented modeling and design” (COMAD) framework. Just as a workflow intends to make what was once computer code understandable by humans, COMAD attempts to make a framework even more transparent and reusable. COMAD is pretty complex to understand without ever using a scientific workflow system, and COMAD isn’t the only framework being proposed. But, it’s heartening that people are trying to go beyond solving problems locally and are attempting to create standardized ways of modeling data and data processes. In addition, if workflows can be shared, or if parts of them can be shared, then the task of data manipulation might become easier.
Creating a culture of data sharing
The complexities of a data sharing system is something that I kept thinking about while reading Borgman as well, and something that has come up briefly in class a few times (is someone writing a paper on this?) To find out more about why researchers do and don’t share data, I looked at the Nature issue addressing this. http://www.nature.com/news/specials/datasharing/index.html
The article, Data Sharing: Empty Archives http://www.nature.com/news/2009/090909/full/461160a.html discusses why researchers have disincentives to share, and how a community of sharing could be fostered.
In Data Sharing: Empty Archives the author, Nelson, writes that while data sharing is always touted as a good idea, and researchers seem to support it in theory, in practice projects to create data repositories remain empty. The example he gives is a University of Rochester project that cost $200,000 in 2003 and still does not have much in it.
The reasons he gives for supporting data in theory aren’t much different from Borgman’s or from other familiar pro-data-sharing sources, “It opens up observations to independent scrutiny, fosters new collaborations and encourages further discoveries in old data sets.” But they also suggest several questions that complicate the matter, “What will keep work from being scooped, poached or misused? What rights will the scientists have to relinquish? Where will they get the hours and money to find and format everything?”
So why do some open-data projects succeed? The article uses the example of Cornell’s arXiv among others, which is for physicists, mathematicians and computer-scientists. Open-source programs are very successful online today as well, so there are some sharing projects that work, but why can’t most projects fill their databases while others can?
Researchers may be worried about the safety of their data, or may feel that they have intellectual rights to their data, or may simply not have the time to format anything for a database or project. Even when projects “wrangle data” as Nelson puts it, often there is no clear way to categorize it or place it in an infrastructure.
Perhaps our inability to plan how to create standards for data is the obstacle to open-data that we should focus on first. Successful projects such as the human genome project are projects where creating standards is perhaps more straightforward. It does seem that open-source programs may be shared to avoid what would clearly be duplication of effort. Many times, a javascript code to create an alert pop-up would be written very similarly by two completely different programmers. With some research data, it may be more difficult to see how sharing would circumvent that duplication of effort (though I still believe it is often the case.) In fact, many of the more creative computer programming solutions (though not all) are proprietary.
Secondly, as Borgman suggests, Nelson also writes that funds and rewards should be based on data sharing for researchers to see the benefit of it. Nelson writes that sharing is more common when the expectations are clear, and I believe this is true for both standards and rewards.
We do have some new tools that probably will encourage data sharing such as Creative Commons, but we will need a cultural shift to continue to go in that direction.
Data and the Institutional Repository
Dorothea Salo wrote her blog post in response to Bryn Nelson's article "Data Sharing: Empty archives" which appeared in the special data issue of Nature.
In the article, Nelson's definition of data seems to refer to anything digital, including "dissertations, preprints, working papers, photographs, music scores", and Salo believes Nelson has confused two very different definitions of data in his article. Salo argues that data are "the stuff coming out of the research process that isn't prose aimed at a human audience." Which means dissertations, preprints, working papers, musical scores and most photographs are not data. Salo argues that most institutional repositories had't even thought about the complexities of data (by her definition) when they began.
According to Salo, institutional repositories decided to deal with their empty repositories by expanding their collection policies to include data. When they expanded in this way, institutional repositories failed to consider how poorly suited they are for data because they were designed to deal with documents.
But Salo remains optimistic that data repositories will not suffer from the same emptiness that document repositories have been suffering from. This is because there was already an accepted workflow for documents in most disciplines, one that most people do not believe needs to be repaired. Changing that workflow to include institutional repositories is an uphill battle. In contrast, data has never had an effective workflow to allow for scholarly communication. The data repositories that allowed scholarly communication would help researchers much more than document repositories can and as a result would be filed more readily.
As difficult as it is to clearly define something like data, failure to do so inhibits our ability to have an effective discussion about the field of digital curation. Whether we choose to accept Nelson's definition of data as anything digital or Salo's more exclusive definition of data changes the nature of the problems being discussed and in turn, the solutions to those problems. So what is data?
The discovery of 2003 UB313 Eris, the 10th planet largest known dwarf planet
This article was passed on to me from pipkin after I talked with him about researchers not wanting to share data in early stages of analysis.
Eris is a dwarf planet in the Kuiper belt that is somewhat larger than Pluto. It is the furthest object yet to be found orbiting the sun and it takes twice as long as Pluto to orbit the sun. Eris was found in the same week of July 2005 as two other bright objects in the Kuiper belt.
These objects were found in Southern California at the Palomar Observatory using the robotic Samuel Oschin Telescope and the Palomar QUEST camera. Every night photos of the sky are sent from the observatory to Pasadena and 10 computers at Caltech analyze the photos for moving objects. All photos that may contain moving objects are flagged for human analysis. In the case of Eris, the initial computer analysis did not detect Eris because the dwarf planet was moving so slowly and was so far away. It wasn't until a year and a half later when the photos were being reanalyzed for very distant objects that Eris was discovered.
On page 199 of Scholarship for the Digital Age Borgman mentions the controversy over finding Eris. What happened is an example of why scientists guard their data until their papers have been released and credit has been explicitly assigned to the original researcher or team of researchers.
What Happened: In July 2005 the astronomers who found Eris in Pasadena released an abstract of an upcoming talk and in the abstract they mention the object K40506A, which is what their automated software named the object. What the astronomers did not know was that the data from some of their telescopes was not secure and that it was possible to search the web for K40506A and discover the object. By the end of July, computers at the Instituto de Astrofisica in Spain accessed the telescope records containing information on K40506A. Two days after they accessed the data an e-mail was sent from the same IP address at the Instituto de Astrofisica to the International Astronomical Union Minor Planet Center with news that P. Santos-Sanz and J.-L. Ortiz had discovered a distant object in the Kuiper belt.
Somehow the scientists in Pasadena were lucky enough to discover that their records were public, this prompted them to officially announce their findings without an official scientific paper.
After all of the excitement of Eris and after scientists broke their protocol of writing scientific papers to release findings, there was a backlash against the Pasadena astronomers for not officially sharing Eris sooner. The accusations that the astronomers were keeping secrets are surprising but they have prompted an interesting discussion of the scientific process. The rest of the article is in defense of the scientific process and explains why scientists have such clandestine practices.
This article is a good case study of when and why scientists share their data and it illustrates the conflict between open access to data and science as we know it. My remaining questions: Is this event an anomaly? Are there other contemporary occurrences of scientists being scooped out of their own discoveries? Also, after an event like this, how can we be surprised that scientists are slow to deposit their data into a collective repository?
Data Anonymization
Ohm, P. (2009). Broken promises of privacy: Responding to the surprising failure of anonymization. SSRN.
In this article the author investigates the merits of anonymization. Anonymization is a now-standard practice where datasets are stripped of personally identifiable information (PII) such that they are agreed to be suitable for sharing, whether with a private party or the public. Ohm argues that anonymization is inadequate in the face of reidentification, a strategy where datasets are combined with external data and information to yield unique data fingerprints of individuals.
The author highlights key advantages to anonymization in a social, political, and legal framework. It allows a discrete action to take place before data can be shared. It allows this action to be enforced and legislated, and so consequently allows an easy division between responsible data-sharers and irresponsible parties. Anonymization balances interests between researchers, legislators, and the public.
Key to anonymization is the idea of PII. When PII are identified in a record (perhaps a name, ZIP, birth date, etc.), those values are either suppressed (removed) or generalized (for instance a full birth date is translated to just the year of birth) so that they are no longer deemed personally identifiable values.
Ohm presents several instances where datasets anonymized in this fashion were used to very accurately identify a single person or demonstrate the ease of this task: the 2006 AOL release of 20 million search queries that identified user 4417749 as Thelma Arnold, the release of Massachusetts state employee health records by the Group Insurance Commission that revealed the governor’s own health records (diagnoses and prescriptions included), and Netflix’s release of 100 million movie ratings records in an open contest to develop a better recommendation algorithm. In the latter case UT researchers were able to demonstrate how little outside data an investigator needed to identify individual’s profiles amongst these records.
The author goes on to suggest that “anonymization” and PII are not realistic terms. The former implies absolute success without varying degrees, and the latter fails to account for the accretion of non-PII fields to constitute a unique identity, which is the primary strategy of reidentification. The author also demonstrates how attempts to extend PII fields invariably decrease the value of data: the more anonymous data becomes, the more value it loses with researchers.
On this struggle the author contrasts the US’s approach to privacy-data legislation and the EU’s. The US largely functions sector-by-sector, addressing the health profession, scientific communities, state departments, etc., in turn. EU maintains a Data Protection Directive which seeks to regulate all data everywhere for PII elements. Both have serious drawbacks, since US legislation is liable to skip or under-regulate whole industries with the idea that its data cannot be used harmfully. The EU meanwhile faces a rising tide of reidentification instances that place more data in the PII camp, thus severely limiting sharing.
Last is a consideration of alternatives to legislating privacy in datasets. The author argues that comprehensive and sector-by-sector approaches need to be combined, and that a risk-assessment of data can guide the sector-by-sector guidelines. Some of the risk factors the author lists are: data handling techniques, whether data is moving to a private or public party, the sheer quantity of the data (since reidentification benefits considerably from large datasets), and motive (academic researchers have little motive to reidentify because of rules and professionalism).
We’ve mentioned privacy briefly in class, by I think this review of anonymization is relevant to digital curation since the field hinges on the idea of mass sharing of data that’s not necessarily tracked carefully. Fortunately the field consists of academic researchers and graduate students so it seems unlikely that in private-to-private data transactions much malevolent intent will occur. Anonymization will probably be used quite a bit in curation since it allows one to “share and forget.” So long as that’s the case it’s good to acknowledge that anonymized data has only been scrubbed to a certain degree, and in conjunction with other external datasets could achieve a lot more specificity than originally intended.
Tuesday, September 15, 2009
Scholarly Communications: Using big words to do things right
In his brief literature review, Burton mentions some of the documents we've read for class, such as the NSF report and Christine Borgman's book, Scholarship in the Digital Age, as important resources and foundations to the ideas of cyberinfrastructure. He thinks all scholars should think about these issues because scholarship is becoming more interdisciplinary and collaborative. The amount of data and how it's being used are also reasons everyone should care about cyberinfrastructure.
Building on a past blog entry, Scholarly Communications must be Syndicated, Burton states that traditional print journals are becoming increasingly outdated due to their "static and isolated nature" when compared to linking, rich media, etc. of digital communications. He suggests syndicating scholarly communications, making it as easy as getting updates from a favorite blog. In the Big Word shout-out, Burton states, in the earlier post, that "the print paradigm cannot win against the ease of use and the timeliness of RSS-supplied content" and "what is starting to matter most with scholarly work does not align neatly with those quantifiables of the print paradigm, publications in scholarly journals" (in the most recent entry). Information sharing will come through different modes, from social media to research tools.
A main point Burton makes is that new attitudes within digital scholarly communications will value the building of an cyberinfrastructure, not just what the infrastructure holds and its interpretive possibilities. Burton discusses how Perseus and related projects help the evolution of cyberinfrastructure to organize accessible information, just as scholars use previous works upon which to build their research. This echoes Borgman's discussion of "infrastructure for information" and not "infrastructure of information" (p42).
I especially like that Burton mentions these issues are beyond technical concerns, that curating, organizing, and hosting are larger issues, which I include in 'cyberinfrastructure'. I also hadn't considered cyberinfrastructure as scholarly communication previously, but do buy Burton's argument that projects that do it right are creating important research and 'publications' (in the very new paradigm-shifting sense). While the technical aspects of these topics overwhelm me at times, I can understand the need for good cyberinfrastructure (since we've all seen bad examples), and I like that Burton pushes the scholarly community - the authors themselves - to understand this as well.
The Rome Agenda - Mouse research and data sharing
With regard to access to data associated with formal publications, the Rome meeting had four suggestions. The first was that on publication, the relevant mice and embryonic stem cells should be deposited in a public repository within a specified time period, and that funders should be willing to cover to cost of such deposits. Secondly, the Rome meeting recommended that the explanation of where and how to access data resources related to a publication be made a mandatory component of that publication. Thirdly, journals and funding agencies should take a firmer stand with regard to existing data sharing policies, clearly explaining said policies and the consequences of noncompliance, and take consistent action to enforce those policies. Finally, plans for data and materials-sharing should be made requisite parts of funding proposals and a consideration in the award of funding.
With regard to data sharing infrastructure and standards, the Rome meeting had less concrete suggestions. The Rome agenda "encourage[s] further investment and recommend[s] that public database coverage and stability be looked at in a coordinated way by funding organizations and the community with increased urgency," but had no more specific suggestions as to how resource-sharing infrastructures could be improved. They also highlighted the importance of data standardization and metadata in making data openly accessible, and called for the creation of better tools for data retrieval, mining, and computation.
The issues raised by the Rome meeting touch on many of the same issues that have been present in our readings: issues of metadata, standardized data formatting, the underuse of data depositories, and the need for a return to the principles of open science. Indeed, the Rome meeting asserted that the solutions to many of the problems underlying data sharing could be found in the creation of a research commons, defined as "a set of resources available to all scientists, either as part of the public domain or on standard terms and conditions that facilitate scientific collaborations, efficient reuse of materials and data, and the dissemination of knowledge." In essence this is what all proponents of data sharing in the sciences have called for, showing that the same basic principles need to underlie the infrastructures of the different branches of the scientific community, and that the major barriers to data sharing are common to all.