Monday, November 30, 2009

DRAMBORA

In April 2008, DRAMBORA Interactive was released. Digital Repository Audit Method Based on Risk Assessment (DRAMBORA) is a methodology and online tool that includes a form-based interface, peer comparisons, reporting mechansisms, and maturity tracking. The tool requires repository personnel to describe its characteristics. The goal is to create an ever-evolving ontology of repository attributes through which repositories can be compared to similar repositories. The toolkit was developed by the Digital Curation Centre (DCC) and DigitalPreservationEurope (DPE). It is not competitive with audit checklists, such as nestor or TRAC, but encourages repositories to use these tools in conjunction with DRAMBORA.

DRAMBORA is a two phase system self-assessment tool. The first phase results in a comprehensive organizational overview. The second phase involves risk identification, assessment, and management. Through the bottom-up approach repositories evaluate themselves based on the repositories own contextual environment and can then compare their characteristics to other repositories. According to this article, DRAMBORA’s creators are also developing “key lines of enquiry” which are sets of questions that will guide auditors within an organization to focus on significant issues or risk factors.

The overall idea is that a successful digital repository is one that plans for uncertainties, converts them into risks, then manages these risks. It involves a cyclical process by which a repository reduces the level of risk after each iteration. Participation is not rewarded by a certification or endorsement. Repositories benefit by strengthening their self-awareness and ability to identify and manage risks. This enables them to present information about their repository that makes them more approachable and trustworthy. The creators argue that DRAMBORA “offers benefits to repositories both individually and collectively” in that it opens up lines of communication between repositories. DRAMBORA facilitates classification of digital repositories so that services and characteristics are more easily communicated to an audience.

The website states that he purpose of the DRAMBORA toolkit is to facilitate the auditor in:
- Defining the mandate and scope of functions of the repository
- Identifying the activities and assets of the repository
- Identifying the risks and vulnerabilities associated with the mandate, activities and assets
- Assessing and calculating the risks
- Defining risk management measures
- Reporting on the self-audit

There is an offline version of DRAMBORA that you can download from the website, but both options require registration. While it is a European initiative, DRAMBORA has been implemented at numerous repositories in the US and Europe.

Tuesday, November 24, 2009

Two Views on Digital Curation

The two articles I looked at for this post both concern digital curation. One was an article in Advertising Age called “Digital Curation Is a Key Service in Attention-Strapped Economy” and the second article is a response to the first article at a blog called Resource Shelf, which calls itself a “daily newsletter with resources of interest to information professionals, educators and journalists.”

The first article is written by Steve Rubel, who is a marketing strategist and a blogger. In his article he talks about the internet offers an endless supply of choices that far exceeds our ability to pay attention to it all (basically he is saying that supply is outstripping demand). He mentions how dominant both Facebook and Google have become. Rubel then goes on to say that no matter how dominant these two sites are they can’t hold our complete attention, in part because they “are often a mile wide and half-an-inch deep.” This leads him into bringing up digital curation, and discussing how brands (like IBM, UPS, and Microsoft, which are some of the example sites and "curation" project he gives) are starting to curate digital information to help people "find the good stuff." He doesn't spend much time discussing his thoughts on brands being curators, though, which is disappointing because the idea is intriguing.

Rubel closes by saying that both human-powered and automated digital curation “will be the next big thing to shake the web.”

I enjoyed seeing a non-information professional’s take on the internet and curation. And it was especially interesting to me how Rubel was using the term digital curation. He never really gives a definition, and seems to make it sound like a fairly simple thing, when in fact it is pretty confusing and complex. I also find it perplexing that in Rubel’s take on digital curation it is journalists who will be playing a key part in it all. I have a journalism background (and a degree in the subject that I will never use), and I can’t imagine journalists doing half the things we have discussed in class about curating information. Rubel says that journalists won’t be the only ones taking part, but he completely fails to mention information professionals at all in his article, which is what the Resource Shelf blog entry responded to.

The Resource Shelf writer says they were sad to see that librarians and information professionals were left out of Rubel’s article. Sadder still, Resource Shelf says that librarians and information professionals are often forgotten by those outside of our field when it comes to discussions like these – which seems ridiculous because who would be better at curating information than information professionals? This response article also goes on to talk about how librarians have been “curating” digital information for years (which we’ve learned) and also about how collection development will become a form of digital curation in the future, which I given much thought to before.

Apparently the blogger at Resource Shelf actually emailed Rubel about his article to invite him “a virtual tour of some of the resources librarians have been curating for years.” Hopefully Rubel responds to the blog because that could result in an enlightening discussion.

Anyway, it was just interesting to read these two different takes on digital curation. Both writers agree that digital curation is a worthy goal to work towards.

Shared Names Project: Linking Biomedical Databases

I stumbled across an interview of science blogger Walter Jessen that is full of interesting gems related to this class. I'm going to pick the Shared Names Project to talk about in depth here. The Shared Names Project is attempting to assign URIs for publicly available biomedical database records and publish RDF documentation about those records. In doing this, they hope to make it easier to link data sets and tools across projects. Without shared URIs, it is necessary to do a lot of mapping to get two data sets to link up. The shared URIs will make it so that pieces of a data set can easily be extracted into another.

Each URI will point to an RDF document that is hosted on servers maintained by the project. Within the RDF document, there will be documentation about what the URI denotes (what database it is), links to various versions of the records (XML, ASN, HTML), links to corresponding resources that use the shared naming scheme, links to external resources ("For example, the RDF for PubMed record 15456405 could link out to the iHOP page for the article described by the PubMed record."), and other information that might be useful to humans or computers.

The project has a wiki that discusses the cyberinfrastructure (technical and administrative) issues they're working out. One technical issue I think I teased out from the Use Cases page is that some "records" may be contained in other records (just as an image may be contained inside an article but is still be a resource in itself), so how do you indicate the level of granularity you want? They also discuss what metadata they should provide for each URI. I don't fully understand how one would use these shared names in a practical sense, but I thought it might be worth pointing out since we've talked a bit about (albeit vaguely) about data sets talking to each other. Similar projects mentioned on the Shared Names wiki are MIRIAM and Integr8, though I don't know how this project and those are similar or different.

Monday, November 23, 2009

Digital Curation & User Testing

Marchionni, Paola. “Why Are Users So Useful?: User Engagement and the Experience of the JISC Digitisation Programme.” Ariadne, no. 61 (October 2009). http://www.ariadne.ac.uk/issue61/marchionni/.


This recent article by JISC's Paola Marchionni refocuses attention on the purpose of curated digital collections: that is, their use by different user groups. Marchionni begins by noting that many digitization projects are still not paying enough attention to their users that their users' needs. They become so caught up in trying to make their content accessible online, that they don't adequately research their key users. As a result many publicly-funded projects are going un- (or under-) used. Marchionni illustrates the insight users can provide by presenting two case studies of projects that incorporated users into their development process: the British Library's Archival Sound Recordings 2 (ASCR2) project (a collection consisting of over 25,000 recordings) and Oxford University's First World War Poetry Digital Archive (WW1PDA) (a collection that contains over 7000 items pertaining to WWI poets, including digitized images of materials held at UT's own Harry Ransom Center).

Though Marchionni's article helpfully reminds digitization (and digital curation) projects to keep their electronic eyes on the prize and really take their users into account, I'm not sure that much of what Marchionni presents in her list of suggestions for user engagement is particularly surprising. She recommends first recognizing the importance of interacting with users and even having an "Engagement Officer" position as the ASR2 project did. She also advises establishing an early and on-going relationship with users. The WW1PDA project, for example, developed a typology of users, with a steering committee of scholars in the field of WWI literature advising which materials should be digitized and participating in quality control, and a separate group of secondary school and higher education instructors helping to develop and offer feedback on the education section of the project. Marchionni also emphasizes the importance of knowing what to do with user feedback. When users expressed anxieties about the integration of Web 2.0 tools out of fear that they might undermine the authority of the WW1PDA archive, the project decided to integrate such functionality in a way that made the lines between the archivists' and the users contributions more clear.

Some of the more interesting lessons regarding users came from the WW1PDA's approach to educational resources. The project held workshops for teachers in order to discover what functionality this group would like to see on the site. In a rather ballsy move, the project then asked the workshop members to help author a number of learning resources for the website. Though this did result in the creation of some resources, ultimately the project realized that perhaps it had overreached in what it was asking their busy users to produce. (Frankly, I would be a little annoyed if I agreed to participate in a workshop on a new resource and then came away having been assigned the time-consuming "homework" of creating a bunch of resources for that project).

Though asking users to create lessons plans and other teaching materials was not as successful as the WW1PDA project might have hoped, users were willing (and excited!) to contribute materials from their own familial archives to the project. In fact, the project received such a high level of response to their requests that they held extra workshops to help the public digitize their items.

Some of Marchionni's suggestions seem to blend user engagement and marketing. For example, the WW1PDA's teachers' workshops seemed to have functioned in part as a source of user feedback, but also as a forum for promoting and publicizing the resource. Teachers were seen as the key to two user groups: teachers and students. Similarly, Marchionni also suggests targeting any information dissemination activities at specific user groups. The ASR project, for example, publicized its Holocaust collection by contacting networks for Historians and those in the field of Jewish and Theological Studies. Though his may seem more like advertising than user engagement, it's nevertheless important to remember that we sometimes need to market our resources if we want them to be used. Finally, after highlighting the importance of user engagement, Marchionni ends with a reminder not to lose sight of the project's mission: though it's important to listen to user feedback, we shouldn't be bullied by it. Focus on the the needs of one's primary users and keep in mind that you can't satisfy everyone.

One thing that I wish Marchionni had addressed in greater detail is the expense involved in maintaining a high level of user engagement. Obviously it's more expensive in the long run to pour money into a resource that doesn't get used than to devote some money to engaging users, but nonetheless creating sustained relationships with users can be a drain on already strained budgets and staff schedules. I'd love to hear more about how small projects or ones meager financial resources might effectively develop ongoing relationships with users.

Friday, November 20, 2009

The Relevance of Twitter

This has been a Twitter heavy semester for me, and admittedly for a while I was thinking something along the lines of "you know, does anyone actually use this tool for anything beyond mindless banter?" A few people have questioned the relevance of Twitter to my face. The usual refrain goes something like "you know Twitter is talked about a lot, but I don't think anyone actually uses it for anything useful." Then there might be a mention of the much lower retention rate of Twitter vis-a-vis Facebook (the two are always compared though I'm not entirely sure why considering that Facebook is a social networking tool and Twitter is a micro-blogging tool. Different things.)

Leslie Carr, on his blog RepositoryMan recently confessed to having similar doubts about the utility of Twitter asking himself if it wasn't just "some gratuitous teenager technology?" So he conducted a study. At a recent CETIS conference, Carr used the Twitter API to aggregate all of the tweets from the conferees. From these tweets he wanted to determine how many of the tweets were either, on the one hand, "technical/academic/professional," and on the other, "personal/informal/gossipy." Although he created other categories for the tweets, Carr was clearly interested in quantitatively studying the relative "significance" of the informational value of the tweets from the CETIS conference.

From his analysis, Carr determined that 70% of the tweets provided the sort of "informational" value that he was looking for and that about 41% of the conference attendees contributed tweets that were either "entirely" or "mainly" informational. Carr doesn't go too in-depth into what the criteria according to which he determined the relative informational value of the tweets, though he did admit that "useful information" was information that was useful "to him."

It's an interesting study, though after reading this post, I have to say that it seems that even the tweets which Carr did not think had "informational value" (e.g., tweets about the poor quality of wireless connectivity at the conference) could be very useful for other audiences. The conference organizers might have found the gripey tweets about the wireless issues very useful, as do businesses which are now mining Twitter for information regarding reactions to their products. The take-away from reading this study, for me, was actually that tweets can provide useful information to any Twitter user given the right circumstances and conditions. This is not to say that I think every tweet is useful. Tweets about Megan Fox can almost always be ignored. Although, I don't know. Maybe in twenty years, a Media Studies or Gender Studies researcher might even find the Megan Fox tweets to be of some value.

Thursday, November 19, 2009

Google Swirl

I am consistently excited by new developments in visual browsing and searching on the web. Google's new development to come out of Google Labs, Google Image Swirl is very provocative and is very close to something that I have been imagining a need for. It would be fascinating it the application could be employed on pre-curated collections of images (ie. to "Google Swirl" a collection of visual art, such as ARTstor). But first, what is Google Swirl?

As a bridging of Picasa Face Recognition and Similar Images, search results depend upon both image metadata and computer vision research. There are comparisons to Google's Wonder Wheel (which displays search results graphically) and Visual Thesaurus. One enters a search term and 12 groupings of images appear visualized as photo stacks. One chooses a particular image and the Flash experimental interface "swirls" to display your image and branches to numerous other images with varying degrees of relationship to that image.

Wednesday, November 18, 2009

Automated Data Processing: Too Big for Our Puny Brains

Over my past few blog entries I've been thinking more and more about automated processing of scientific data, as well as distributed efforts to handle massive backlogs of observations. Science is now generating such huge data sets that our ability to intelligently query them in any manual fashion is impossible. For example, CERN's Large Hadron Collider will be producing 40 TB of data every single day. For these reasons, if we hope to make any progress with scientific data, we'll need to employ AI (artificial intelligence) as our new scientific tools. Maybe this topic will seem to be at the far reaches of what we've been discussing in class, but I think it is a useful look at just how these datasets will actually be used. Instead of human searchers crawling the databases, we'll see automated systems trying to distill the knowledge from data bits.

An article from this month's Communications of the ACM describes the efforts of two computer scientists from Cornell University to build a new machine learning system. Whereas older machine learning systems sought to build predictions based on data, this new system was looking for basic invariant relationships such as the conservation of energy, that are scientific constants. This kind of system could formulate scientific laws as basic as the law of gravitation.

The breakthrough in this case comes from giving computers relatively little starting information and instead allowing them to derive rules and theories as they proceed, testing different ones and weighing their success in describing the situation. Other scientists have expressed interest in this system, because it is also scalable across domains. The ACM article on this mentions that this system automates the final part of the traditional scientific paradigm: from observational data to model formulation, to predictions, to laws, to explanatory theories.

Interestingly, the end of the article hints that the results of this system may be beyond human understanding. Interpretation of the results and their meaning may not be possible in human terms. This raises the question, with all this data, what good is it? How does it further our knowledge? How do we even know if the results are correct, if we can't understand them?

Digital Scholarly Communication Projects List

I mentioned a couple of weeks ago that I did an independent research project last spring about new forms of scholarly communication. In the spirit of academic sharing, here's my googledocs spreadsheet with about 75 or so different projects I found and related notes. Halfway through the semester I decided that there was too much to talk about with any depth, so I narrowed my project to be about emerging forms of peer review. I discussed things like rating systems, online comments, different kinds of blinded review, and legitimization and trust. I think the conclusion of the paper was a little weak because it turns out (gasp!) that the promotion and tenure system isn't really capable of handling new forms of peer review. I came to realize that the projects in the spreadsheet containing some element of peer review, whether formal (formal as in, I guess, journal sanctioned) or informal, are really just isolated experiments without any measurable influence in the larger scheme of things. It would have been more useful if I had been able to talk to some of the people that contribute to these projects to see if they found them valuable or not. Here are a few selected projects you guys might be interested in.
  • gpeerreview: Google's unfinished answer to peer review, involving getting endorsements from "endorsement organizations," graphing those endorsements, and then providing some kind of credibility ranking
  • Faculty of 1000: online research tool that highlights the most interesting papers in biology, based on the recommendations of over 1000 "leading scientists"
  • MONK: an open source "digital environment" that humanities scholars can use to analyze patterns in texts
  • www.myexperiment.org: scientists contribute their scientific workflows (I assume this is like sharing their lab notebooks) so that others can use them, also see UsefulChem
  • SciLink: Facebook for scientists, but instead of using email contacts to mine your connections like Facebook does, it uses bibliographies from articles

PLANETS: integrated services for digital preservation

PLANETS (Preservation and Long-term Access through NETworked Services) is a European initiative to provide long-term access to large digital collections. I originally came across an article and followed up by looking at the website for more information. PLANETS seems to be predominantly about tool development with the goal of creating a sustainable framework for “increasing Europe's ability to ensure access in perpetuity to its digital information.” The project began in June 2006 and is funded by the European Union under the Sixth Framework Programme.

Planets is not a repository project but expects each participating institution to maintain storage for their digital data. The goal is to work toward preserving entire collections, not just creating stand-alone applications that can handle one aspect of data preservation, such as migration or emulation. The project is a collaborative effort based on the idea that no single institution is going to be able to handle the level of development needed. The initiative is drawing from the expertise and experience of numerous partners in different countries.

The website explains the deliverables:

Preservation Planning
services that empower organisations to define, evaluate, and execute preservation
Methodologies, tools and services for the Characterisation of digital objects
Innovative solutions for Preservation Actions tools which will transform and emulate obsolete digital assets
An Interoperability Framework to seamlessly integrate tools and services in a distributed service network
A Testbed to provide a consistent and coherent evidence-base for the objective evaluation of different protocols, tools, services and complete preservation plans
A comprehensive Dissemination and Takeup program to ensure vendor adoption and effective user training.

After hearing a presentation on digital preservation initiatives in the US by classmates in another course, I am interested in the differences between the systems in place in the EU that make this kind of large scale collaborative project possible. Reflecting on the Larsen article, On the Threshold of Cyberscholarship, I realize that PLANETS falls clearly into the research stage of activity, where tool development is key to the success of a collective infrastructure for access to digital materials. I get the impression that European institutions, and maybe scholars, are more likely to achieve success in the area of cyberinfrastructure development. Is this because of funding, social behavior, scholarly expectations?

Free Access to the Web

For my final blog post, I found an old (by internet standards) article from 2000 that views the entire web as one low-cost library and discusses the marvel that much of the access to it is free. Arms, writing at Cornell, envisions the web as replacing libraries in many people's lives. Whereas libraries, and especially research libraries, are incredibly expensive and limit access to their members, the web is self-funding (in that individual "publishers" pay for that privilege) with free access for anyone with an internet connection. (Of course, one could argue that this doesn't really constitute free...) Furthermore, Arms concludes that although the expected model of information provision on the internet was fee-based subscriptions, it turned out that there is enough free information of quality out there to obtain genuine substitutes. The example he gives is that Cornell provides legal sources via the web, including case reports, hitherto only available with an exorbitantly expensive Westlaw subscription, and one need not be affiliated with Cornell to access it.

Another point that Arms makes is that digital libraries can be nearly completely automated. He asserts that a brute force search, such as that provided by Google, with enough information and in the hands of a good researcher, can actually be much more powerful than an intelligent search by trained librarians. (While I don't like the implications for library services, it does seem like a lot of the focus on the need for reference librarians is in terms of not-good or amateur researchers...) This automation actually increases access as you no longer need to work through a small group of homogeneously trained elites.

Arms identifies two issues or potential problems for further research. One is insuring quality of information. This role was traditionally performed by the publishing process, but with the self-publishing afforded by the web, we can no longer count on good publishing practices. The second is permanence. Flip a switch on a server, and its information vanishes.

I found this article interesting mostly because of its now somewhat historic outlook. Some things have not turned out as Arms saw them in 2000, namely the level of free access. As we see with the Google Books project, proprietary interests are finding their way into the new cyber-reality, and the Great Copyright War has yet to be fought. Interestingly, though, Arms did identify two key issues that continue to be relevant: trusting found information, and ensuring its permanence. I suspect that the best answer to the former is via education of the public. At some point, the onus has to be on the searcher. The second problem, in my mind, is much more problematic, and it is one that plagues the physical as well as virtual information worlds. All in all, I found this early article quite interesting.

Tuesday, November 17, 2009

Aquatic and Riparian Effectiveness Monitoring Program

After talking about the problems facing ecological data, I wanted to read a little about the data collecting work my friend did over the summer for Aquatic and Riparian Effectiveness Monitoring Program (AREMP) in Oregon and see how that data fits in with the Long Term Ecological Research program. AREMP surveys 250 watersheds in the northwest and their collection practices (including what photos to take, how to record coordinates and how to use site markers) are described on their site. To collect the data they have to enlist a number of people like my friend to collect watershed samples using GIS. Although 250 is a lot of watersheds, it only represents about 10% of the watersheds in the area that is being sampled.

The goal of AREMP is to use a decision support model to evaluate watersheds for overall watershed condition. A number of attributes are assigned to each watershed and once all the attributes for a watershed are sampled the data is aggregated to determine a watershed score. To aggregate the data and find a score, AREMP uses software called Ecosystem Management Decision Support (EMDS) which creates the model and then assesses the condition of the watersheds based on the data. AREMP says that they would be happy to share their data with anybody who would like to see it.

EMDS is pretty interesting, the EMDS document says, "EMDS does contain tools for conducting “what if” scenarios. For example, one can estimate how watershed condition will improve if 500 pieces of large wood were added to the stream".

I wasn't able to find anything connecting AREMP with LTER but I did discover more problems with ecological data. For instance, the EPA and AREMP both use probability sampling designs but, "indicator and sampling methods differ from those used by the EPA, and these differences hinder collaboration and data comparison" (Hughes, 2008, p. 853).

I would still like to know more about AREMP's data and how their data collection methods compare to other ecological studies.

Hughes, R., Peck, D. (2008). Acquiring data for large aquatic resource surveys: The art of compromise among science, logistics, and reality. Journal of the North American Benthological Society (27)4, 837-859.

OCRIS: Online Catalogue and Repository Interoperability Study

JISC: Online Catalog and Repository Interoperability Study (OCRIS): Final Report

This study reviewed Library Management Systems (and the associated OPAC) with the Institutional Repositories (IR) of Higher Education Institutions in the UK.

The goals: determine whether the repository content within the scope of the institutional OPAC (and extent it is recorded in the OPAC); examine interoperability of OPAC and repository software; list services offered by OPAC's and repositories; identify potential for improvement in links to other institutional services; make recommendations for development of further links between OPAC's and repositories.

The primary findings are distressing. Only 2 percent of the study respondents stated that their systems were definitely interoperable, and 14 percent stated that interoperability was pending. There was an 81 percent overlap in scope for all items in IRs and OPACs; generally, IR's contained bibliographic data and OPACs contained full text.

Clearly, differences in scope/policy are not clear and there is either uncoordinated effort, hindering interoperability, or duplication of effort and/or redundant information. In order to provide a more feasible and appropriate long-term vision for IRs and OPACs, institutions should take a structured look at the goals of each service (IR vs OAPC) and coordinate efforts to best provide interoperability and reduce duplication of effort.

This was interesting... OPACs and IRs arguably have different intents - generically speaking, circulating collections versus long-term preservation and access, but they are really not so different. While both may not store information, each provides a service in locating information. In light of the volumes of money spent on IR development/OPACs/interoperability, and the number of available/established entities available to study, it is a very good time to step back and consider how these services might be better managed and coordinated to provide the best available service to the end user and the institution. I would like to see a corresponding study of American institutions.

Competing Requirements for Self-archiving

Stevan Harnad wrote a response blog to a letter to the editor that appeared in D-Lib Magazine. The original letter to the editor was a hypothetical dialog between an author whose work was recently accepted and an open-access consultant. In the dialog, the author is unsure about how to deal with a funder who requires the article to be deposited in an online repository and a publisher who may or may not allow the article to be deposited in a location they have no control over. In his post, Harnad takes excerpt from the dialog and writes more information about what the open-access consultant is saying, sometimes correcting him.
First, Harnad clarifies the opt-out clauses in self-archiving mandates. The opt-out clause is about whether or not you need to persuade the journal to accept your addendum thereby formalizing your right to deposit your article. While this is worthwhile, it is not essential so authors can choose to opt-out if they cannot persuade the journal or they simply don't wish to try. But regardless of the presence of an opt-out clause, authors still deposit their articles immediately. It is not necessary to find another publisher if the publisher denies your request.
Second, if the publisher has an embargo period, and the author wishes to honor it, they simply deposit the article as closed access for the appropriate period of time. There is no need to talk to the publisher about the embargo period.
Harnad suggests that depositing an article is not as confusing as it seems. An author simply needs to deposit all drafts as soon as an article is accepted for publication. According to his post, 63% of the top journals allow such a deposit. If your journal is one of the other 37%, simply set the article to closed access.
The issues surrounding depositing a published article are complex and probably result in many authors choosing to do nothing rather than try to navigate through all the competing requirements. Though closed access is not ideal, having a copy in a repository is better than not having anything stored. It seems that the more an author knows about the process of depositing an article and what they are allowed to do, the more likely it is that they will go to the trouble of depositing their article. The way things currently stand, most authors choose not to mess with online repositories because they don't even know where to find out what they are allowed to do.

Monday, November 16, 2009

JoVE (Journal of Visualized Experiments) is a new approach to peer reviewed journals : specifically devoted to the biological sciences (life sciences), JoVE is indexed by PubMed. The goal of JoVE aid the transmission of information - particularly that which is not sufficiently represented by static text and images. JoVE claims that their approach - using video publishing - "promotes efficiency and performance" by eliminating time wasted on learning and perfecting new techniques based on traditionally published works.

The Editorial Board for JoVE boasts members of the scientific community from the best institutions in the world... Harvard,
Mount Sinai School of Medicine, Princeton,
University of Zurich -the list illustrates an impressive number of highly qualified board members. The project was begun at Harvard in 2006 by a post-doc, Moishe Pritsker, now CEO and editor-in-chief of JoVE.

Access:
Initially, JoVE was conceived of as an open-access project, however, that model proved itself impossible, given the high costs of producing the videos. According to Pritsker,

"The reason is simple: we have to survive. To cover costs of our operations, to break even, we have to charge $6,000 per video article. This is to cover costs of the video-production and technological infrastructure for video-publication, which are higher than in traditional text-only publishing. Academic labs cannot pay $6,000 per article, and therefore we have to find other sources to cover the costs." (http://scholarlykitchen.sspnet.org/2009/04/06/jove/)

Thus, a pricing structure was created to cover these costs: "$1,000 for small colleges to $2,400 for PhD-granting institutions, prices which are in league with other commercial scientific journals. In addition, authors are charged $1,500 per article for video production services ($500 without), and there are open access options: $3,000/article with production services ($2,000 without)." (http://scholarlykitchen.sspnet.org/2009/04/06/jove/)

Taking a look JoVE's Press section, I was surprised to find they had posted articles criticizing their decision to go closed-access. While these criticisms are valid, so are JoVE's explanations of why open-access wasn't a possibility. Given the newness of this "product" and the fact that there is no existing model, it makes sense that charging for access is the only way to offset the price of video production without all of the costs falling on the research institute.

How it looks:
Videos are very high quality and accompanied by a complete scholarly article that acts as a transcript of sorts to the video.

Overall, I'm very impressed by JoVE, and while for now it is closed access it seems that as more stakeholders buy in and the model becomes more widely accepted it may be possible in the future.

Monkeying Around with Twitter Data

This week I was looking for articles about access to data and found a new interesting development. An Austin-based company called InfoChimps.org, that offers data sets for download, last week announced they were selling Twitter datasets. For quite a bit of money.

InfoChimps' mission "is to increase the world's access to structured data". The company appears to offer much data for free and prefers to be a platform where people may "post data under an open license". The data is available for browsing, and if the site doesn't actually have the data, it will point the user to where they can get the data for free. This is excellent for data sharing and access to large data sets. The site's homepage lists "Interesting Datasets" for perusal. I clicked on the first one, which was College Enrollment of Recent High School Completers 1960-2005. There is an intro paragraph to the data, where it's from, and an example of it. This set was prefaced with the caveat that the "files have data mixed with notes and references, multiple tables per sheet, and, worst of all, the table headers are not easily matched to their rows and columns." Kind of funny, yet good information to know!

So, Infochimps offers free datasets - fabulous. But their recent announcement to sell Twitter data met with some skeptical questions. The Read Write Web blog discusses this development in depth, describing the data, which isn't the full tweets, but hashtags, RTs, @ messages, and other associated info. This is apparently really useful and great information to have, and the developers at InfoChimps are hoping that people create interesting apps with this data. InfoChimps mined this data themselves, by hitting Twitter Developer API 20,000 times per hour (I almost know what this means). That's a LOT of data. Marshall Kirkpatrick, the author at RWW, questions the complete legality of selling this data, and worries Twitter is going to come a-knockin'. Many commenters on the blog entry also thought so. InfoChimps (last week) swore they were on solid legal ground, but a new post from yesterday on their blog revealed that Twitter had asked them to remove the datasets. While InfoChimps swears their data had nothing personally revealing, privacy concerns came from commenters and apparently Twitter, who claims they just want to prevent any 'malicious use' of the data.

I gleaned from the blog entries related to this issue that Twitter isn't very forthcoming with their data, which is making them a new Bad Guy of social media and data sharing. It's interesting that people are upset that Twitter won't share, but probably don't care as much if a university won't share some research datasets? Shouldn't data be data? Who cares if one is more sexy than another? Access is still important. InfoChimps sounds a little naive to think that selling Twitter datasets for $9000 wouldn't cause a stir, but maybe that's what they actually intended! At least it brought some mild attention to the issues at play here.

Sunday, November 15, 2009

Gordon

On November 4th, the San Diego Supercomputing Center (SDSC) announced that they have been awarded $20 million by the National Science Foundation to develop a new supercomputer aimed at "solving critical science and societal problems now overwhelmed by the avalanche of data generated by the digital devices of our era." This new computer is known as Gordon.

The SDSC has been part of the National Science Foundation's plan for cyberinfrastructure in the sciences for a number of years, as we read earlier in the semester. The development of Gordon is part of the NSF's continuing effort to keep up with the computing needs of the sciences. According to Jose L. Munoz, the deputy director and senior science advisor for the NSF's Office of Cyberinfrastructure, "'Gordon will do for data-driven science what tera-/peta-scale systems have done for the simulation and modeling communities, and provides a new tool to conduct transformative research.'"

Gordon, which is scheduled to be installed by Appro International, Inc. in 2011, will employ flash memory to speed solutions to data-intensive problems, such the analysis of individual genomes, much faster than spinning disk technology. Gordon is the follow up to the Dash system, another SDSC project, which was the first computer to use flash devices. Gordon will feature 245 teraflops of total compute power, 64 terabytes of DRAM, 256 terabytes of flash memory, and four petabytes of disk storage.

Another feature of Gordon will be 32 so-called "supernodes," each consisting of 32 compute nodes capable of 240 gigaflops/node and 64 gigabytes of DRAM. Linked together by virtual shared memory, these "supernodes" each have to potential of 7.7 teraflops of compute power and 10 terabytes of memory. The "supernodes" will be linked together with an InfiniBand network, which is capable of 16 gigabits per second of bidirectional bandwidth, which, apparently, is eight times faster than some of the most recently developed supercomputers. Gordon will be made available to researchers via an open-access national grid.

The point of Gordon is to make possible the complex applications of data-driven science. According to the press release, Gordon will have potential benefits in both the academic and industrial settings. Gordon should also be useful for predictive science, which seeks to make models of real-life phenomena. Because of the large-scale memory on a single node and the consequent increase in computation speeds, Gordon should allow for the creation of models that more accurately mimic these phenomena. With such capabilities, Gordon will be the next step in the development of data-driven science.



Wednesday, November 11, 2009

Articles for Rent

In my blog last week I addressed the issue of journal piracy, people actively hunting down journal article and re-posting them on forums so that they could be down loaded by others. This is a relatively predictable outcome to the shift to digital content, and I tried to make the point that in the end it probably did not effect the journal publishing system to much.

This week I am continuing with the journal access them by looking at the DeepDyve journal access system brought to my attention by this post which has a shout out to a U.T. librarian.

This company offers limited time access to journals articles through a “rental” system. Users can “rent” articles for .99 per day. Users do not have the ability to save off these article or to print them and they only have access to each article for 24 hours, at least at this level, there are other levels that allow users to access article for a week or a month with the price appropriately tiered. Users can browse the holdings of the company and view abstracts for free. On the page with the abstract the user has the option to purchase the article from the supplying institution. I decided to follow this link and the journal who published the article I was looking at also had a daily use option, but it was 8.00, DeepDyve was beginning to look like a better deal.

The system also has a set of widgets that collect articles and alert you via RSS feeds, this allows you to set up a query and have it alert you when new articles which match your query are available.

This program is aimed at those who have need of journal articles but do not have access to an institutional license, I really have no idea how big of a market this is, but I don't imagine that it is huge still I have been wrong before.

So how does it work as an article finder? Turns out pretty well, actually. This article reminded me of another article I read while putting together my blog for this week, it concerned searching abstracts for conclusion material and and extracting that data via semantic web / Linked data Format in order to print out summaries of relevant research, but I could not remember where I found it or who wrote it. So I typed what I could remember into the search bar and the article was on the first page. By the way the article can be found here , enjoy.

This is a different model for journal access than any we have seen yet. There are some things that I really like about it, and other things that I am for some reason hesitant to embrace.

Collection Space

Collection Space is another digital repository project supported by the Andrew Mellon Foundation, like JSTOR and ARTstor, but with very important differences. CollectionSpace emerged from scholarly institutional collaboration "with the common goal of developing and deploying an open-source, web-based software application for the description, management, and dissemination of museum collections information." Which is to say - It's an open-source software that will allow for museum collection management and dissemination but with many Web 2.0 twists. This is not your regular museum registrar database. CollectionSpace's project team currently includes web-developers, designers, and architects from Cambridge, Museum for the Moving Image, UC Berkeley School of Information and the University of Toronto.

This venture was inspired by the large information gap that existed for 1/3 to 1/2 of collecting institutions in the U.S. (historical societies, archeological repositories, museums and others) having neither their collections online or in many cases even catalog records. CollectionSpace addresses these needs through working to develop open-source software solutions providing stable, authoritative but flexible collection architecture "from which interpretive materials and experiences - from printed catalogs to mobile gallery guides - may be efficiently developed, and that can
serve as a cost-effective alternative to proprietary collections management systems for museums in need, regardless of size or scope." Think perhaps ARTstor for museums owning, managing (and publicizing) their own rotating collections.

The project began as a series of collaborative workshops with museum, archival and library professionals followed by a two-year period of software development complete by beta-testing by the same communities that helped advise its design.

CollectionSpace Release 0.2, debuted October 6, 2009. This version allows storing "multiple record types with flexible schemas." They have an extensive Project Wiki filled with a variety of documentation, handouts and powerpoint presentation.

Funding is supported by the Andrew W. Mellon Foundation, Program in Research in Information Technology which "supports the creation of "enterprise" administrative and infrastructural software by means of distributed, collaborative open-source development projects."

It aims to support registrars, curators, educators, collections managers, and administrators. It seeks to change earlier system models which seem to be digital representations of paper models. CollectionSpace promises to bring processes like cataloging, loans, media handling, location tracking "into the Web 2.0+ era." It seeks to move from "records-based navigation" to "action-based navigation."

Distributed Computing: It's OK To Share Data with Lots of People

Discussion on the separation between those who generate models and those who generate experimental data in Birnholtz and Bietz article this week led me to thinking about larger collaborative linkages in research. We could also consider the linkages between hardware designers, software designers, and the different demands and constraints placed on these groups. One public example that shows the collaboration between researchers generating experimental data and those who process it can be seen in the many distributed computing projects such as SETI@Home, Folding@Home, and others.

As a brief introduction, Folding@Home and projects like it rely on thousands, if not millions (SETI@Home has 3 million users currently) of users around the globe to download experimental data to their computer and process the data. In the case of SETI@Home, the data consists of radio telescope data that is analyzed for signs of extraterrestrial life, but other projects investigate biochemistry (protein folding) and many other scientific queries. Processing is usually automatic, but since users are downloading the data, the experimenters must trust that users will not tamper with the data in any way.

A posting from August 20 of this year on the Folding@Home blog mentioned the danger of tampering with data downloaded from the Folding@Home project. Users may tamper with data with completely benign intentions, such as maximizing their processor time, but we can see that data transformations might taint the overall data, and for this reason it is not allowed by this project. This is a simple example of something we have seen again and again: researchers need to maintain control over their data, especially in this case, where they are collecting results from processing done on their data with the future intention of making discoveries. This is a model for the way science will be done in the years to come: data is collected, stored, and then farmed out to processing teams, and fits in with the "stream" model in the aforementioned paper, instead of the discrete event data model. The fact that the data can be broken up for parallel processing enables this sort of massive crowdsourcing. Despite the huge gains in processor power over the past few decades, some problems are just too complex to solve overnight, no matter how many processor cores you have.
There is one additional benefit to projects like SETI@Home and Folding@Home: they provide an incentive to participate for users in the form of groups that can compete for prestige by contributing the largest number of processed hours, and there is a feeling that one is "helping out" to solve complex scientific problems. We have seen that one of the problems in getting scientists to share their data is figuring out how to tie the social and scientific rewards together, and I feel that these projects do this well.

Welcome to the caBIG� Community Website —

Welcome to the caBIG� Community Website —

Pertaining to this week's theme of knowledge sharing, I checked out a huge example of knowledge sharing: the National Cancer Institute's caBIG. CaBig defines itself as a "a virtual network of interconnected data, individuals, and organizations whose goal is to redefine how research is conducted, how care is provided, and to interact with others in biomedical research.

The goals of caBIG is to adapt or build tools for interacting wth information with cancer research, connect with other in cancer research, deploy and extend standards, rules, and common languages to facilitate interoperability. The website underlines that all this is needed because there is a "critical problem facing basic and clinical researchers today: the explosion of data that requires new and different approaches in sharing data."

Each of the different areas of cancer research, such as clinical trials, tissue banks and in vivo imaging, have their own workgroup that allows like minded people to post and share their information.

caBIG also has an entire area within their network devoted to data sharing to add on additional information features that researchers may need in order to share and display their research, such as patient privacy protections and intellectual property interests. The data sharing and security framework is designed to help researchers overcome the common barriers that obstruct the ease of sharing information. As it is, this particular feature is not fully developed as people are still working on defining policies, best practices, model documents and the trust fabric. Users can select levels of proprietary values, levels of privacy and restrictions and be able to adjust amount of data sharing based on the metadata that the scientist adds to it.

The caGrid is caBIG's underlying network architecture and platform that enables the connectivity within the network. It promotes interoperability within the site and all the biomedical research tools and facilitates communication between the users/researchers to render data sharing more efficient.

All in all, the caBIG network looks promising since the ease of sharing data about cancer research is the very core of its motivation. The researchers in the field also realize the importance of sharing data and are looking for ways to improve scholarly communication.

Legal Knowledge Sharing

This week I found an article written by the librarian of a private law firm on the value of internal knowledge sharing to firms and how this can be achieved via an intranet. I liked this article for a couple of reasons. First, it emphasizes the value of knowledge sharing even in non-scientific environments. Second, it rags on the fad of naming things such-and-such-2.0, which I also happen to think is pretty stupid. But, I digress.

Eiseman's argument is based on two premises. First, that one cannot necessarily determine what knowledge will qualify as useful to share. The example he gives is that of employers not wanting comany wikis or other public fora because they fear an inundation of private knowledge such as restaurant recommendations. Eiseman points out, however, that many legitimate business functions occur in restaurants, so someone looking for a new place to take a client would benefit from that knowledge. Second, Eiseman argues that the entire organization can benefit by collectively sharing knowledge formerly possessed solely by specialists. His example for this is posting reference questions and their answers to a wiki instead of to a two person email. This is efficient, in that it (in theory) prevents repetition of the same questions. However, it also spreads general knowledge of specific legal situations throughout the entire firm, which will enable the various lawyers to better serve their clients rather than referring them to a second attorney.

Especially in a setting such as a law firm, user-generated content on the wikis (or other knowledge sharing apps... Eiseman discusses a couple) serves an essential role in knowledge sharing, as each firm partner, associate, and librarian will have an indvidualized body of expertise.

The one thing that I did wonder about from this article, is why limit knowledge sharing on the law to one's own firm. I suspect it is a mix of liability fears (only trust knowledge with a known provenance!) along with the competetive nature of the legal field. Although, an idealist might note that the more determinant a given legal situation is, the better off everyone operating within that field would be. (Obviously, I'd like to see a bit more legal knowledge sharing than via firm intranets...) All in all, though, I thought the article provided good perspective on the importance of knowledge sharing in non-scientific areas.

EveryBlock: data about your neighborhood

Adrian Holovaty leads the team that has created a new website called Everyblock. The site, which is relatively new and only has information available for 15 cities, pulls information from a variety of sources and creates a kind of activity stream for specific neighborhoods within these cities. Everyblock is designed to keep track of everything that is going on in your neighborhood, from local news stories to civic information like building permits, restaurant inspections and even changes in liqueur licenses, as well as activity from social networking site like photos of the neighborhood that have shown up on Flickr, Craigslist postings, and user reviews of local businesses.
Site construction on EveryBlock began in 2007, with the launch of its first 3 cities in early 2008. Since then, it has continued to expand the number of cities it covers and EveryBlock was acquired by MSNBC this past August. Designed to tell users what is happening in their own neighborhood, each city block has its own news feed that lists information that applies to its location. It does not show information about the city in general. For example, a feed on the Hollywood neighborhood in LA will contain information specific to that area but will not contain information about LA in general, even if it has implications for the neighborhood in Hollywood. They also do not put in directory type information about an area so you won't find information about who lives down the street from you. The site also offers email alerts and RSS feeds so users don't have to visit the site to know what is going on in their area.
The combination of different sources (news, civic and social) was at first a bit surprising. While it seems a bit less far-fetched after having looked at the site, I still can't help but wonder if anyone would use EveryBlock for anything other than entertainment purposes. Do people really want to know every time someone takes a picture of the local flower shop or every time the restaurant down the street is inspected? I don't actually know. But this particular combination of data from different sources is something I haven't seen before and I've certainly never thought about all the different types of data generated about a particular area as being something that could (or should) be combined. If nothing else, EveryBlock helps users to visualize the incredible amount of data generated about a particular location and offers a way to make that data easier to access.

Policy: NIH Data Sharing for Genome Research

This week I came across the Notice on Development of Data Sharing Policy for Sequence and Related Genomic Data from the NIH. This notice says that the NIH is considering new additions to their policy on genome research. In class we've discussed what the motivations of the genome investigators might have been for being open with their data and so far there doesn't seem to be a single answer. However, this ongoing revision of policy from the NIH indicates that this area of science has unique requirements for data sharing because of the resources that are needed to acquire genome data. The NIH mentions that these resource needs "necessarily limits the number of projects that can be supported for any disease". In addition to these factors, the NIH says that these data sets have "added scientific value when combined with other large data sets".

The policy that the NIH is working on right now is basically a continuation of the Policy for Sharing of Data Obtained in NIH Supported or Conducted Genome-Wide Association Studies (GWAS) which went into effect January 2008. Basically, when researchers have genome data they have to put it in the central genome-wide associated studies (GWAS) database which makes all genome data accessible for combining with other data sets as early as possible. Although other people can see it, the investigator who deposits data has exclusive rights to publication based on that data for some period of time, the NIH encourages investigators to keep this period of time short and they limit the exclusivity period to 12 months.

Is the NIH the only organization that has this kind of exclusivity period? I remember a discussion in class about an investigator who deposited their data into a repository and then somebody else published a paper during the exclusivity period. I also recall that the organization with the policy and database (again, I'm pretty sure it was the NIH) didn't seem very alarmed by the situation.

One reason why the NIH is updating their policy might be because of the confidential nature of the data and the need for participant privacy. In the original policy the NIH predicted that in a few years technology would "make the identification of specific individuals from raw genotype-phenotype data feasible and increasingly straightforward", which is problematic for data in a federal repository because that data is accessible through FOIA. At that time they decided that they would have to deny FOIA requests for unredacted data sets. About eight months after that policy went into effect the new developments on how to identify people based on genome data was made public (NIH reigns in genome access) and the NIH has to further restrict access to the data - NIH Modifications to Genome-Wide Association Studies (GWAS) Data Access (sorry, that's a PDF).

This article in PLoS Genetics explores the issue of confidentiality with human genome research in more detail: Public access to genome-wide data: Five views on balancing research with privacy and protection. One person basically says we can't trust anybody to keep data confidential, even though any researcher accessing GWAS would have to agree to the policy agreeing that they would "not attempt to identify individual participants from whom data within a dataset were obtained".

Tuesday, November 10, 2009

Facebook for Natural Historians? This is getting silly...

Article: "Darwin Meets Facebook: Social Networking Tool Lets Natural Historians Share Data" from ScienceDaily (Nov. 9, 2009).

As with other types of data we have mentioned this semester, curating biological data is a challenge. Apparently, taxonomic data (which is data about the classification of living things) has grown in volume over the years, while the number of taxonomic experts has decreased, and researchers have found it increasingly more difficult to find, access, and publish taxonomic data. To help with these problems, the idea for a social networking tool for natural historians to share their data was born. This tool, called Scratchpads, is being developed by researchers at the Natural History Museum of London.

Scratchpads got its name because it “reflects the fact that taxonomy is in a state of 'perpetual beta', constantly changing and reforming to reflect our current knowledge of a particular group.” The sites and “virtual workbenches” users create act as online notepads; the goal of Scratchpads is that users will work together to 'scratch out' their ideas and share taxonomic information. Scratchpad sites provide an open access to data in ways that traditional publications can’t.

Scientists and researchers can use Scratchpads to manage information about classifications, phylogenies, bibliographies documents, image galleries, custom data, specimen records, and even maps. Registration is free for any interested scientist; all that is needed to use scratchpads is to complete a registration form.

As far as technology is concerned, Scratchpads uses a Content Management System called Drupal, which provides the underlying architecture that Scratchpads relies on. Drupal allows Scratchpads to offer features such blogs, collaborative authoring, forums, peer-to-peer networking, newsletters, picture galleries, file uploads and downloads, and more.

The goal of Scratchpads is to be user-friendly and encourage users to generate, organize, and share their data. According to the Scratchpads website, its “infrastructure combines databases, network protocols and computational services to bring people, information and computational tools together to perform and publish natural history.” Scratchpads can help users collaborate via the web on projects that might have taken individual users a lifetime to complete alone by providing a space on the web for “communities to bring taxonomic information together without the limitations of traditional paper based publications.”

I blogged recently about the Facebook for Scientists which is being funded by the NIH. Now with Scratchpads, we apparently have a Facebook for Natural Historians. I’m not sure how I feel about this. The name is cute, and the idea is trendy, but how truly useful is this? On one hand, it is great that researchers are trying to collaborate and spread open access to information, but at the same time it feels like people keep trying to reinvent the wheel, developing similar projects that will probably fail to live up to expectations.

Why Isaac Newton is influencing science sharing today // how to make scientists share

As I was looking for an article related to data sharing for this weeks blog I came back to the September Nature special on Open Access, and Bryan Nelson's article "Data sharing: Empty archives." I know that this particular issue of Nature was widely read by our class, so in an attempt to limit redundancy I followed the Comments of Nelson's article to a similarly themed article.

The article in question "Doing science in the open" by Michael Nielson appeared on physicsworld.com in May. In the early years of science - the really early years - scientists made every effort to keep their discoveries secret. Going so far as to publishing ideas as anagrams in order to have time to work out the kinks of a theory while keeping the idea itself a mystery. This practice, it seems was rather common in the 17th and 18th centuries. Scientists were, presumably, motivated in large part by personal gain. This culture began to change when governments became the patrons of scientists - scientific discovery benefited the public good and the reputation of both the scientists and their patron governments. Since this time some 300 years ago which brought about the practices of scholarly publishing we know today little has changed.

As Nelson noted in "Data sharing: Empty archives," the scientific field has been slow, in some views, to adopt digital data sharing. Nielson, however, notes several successes in scientific data sharing: the "physics preprint server arXiv, which lets physicists share preprints of their papers without the months-long delay typical of a conventional journal, and GenBank, an online database where biologists can deposit and search for DNA sequences." Also, Nielson describes (what I consider to be the coolest data sharing in science I've heard of so far) the "Journal of Visualized Experiments, [JoVE] which lets scientists upload videos that show how their experiments work."

(At the point I reached this part of the article, I wanted to jump ship and blog about JoVE - a site well worth a look!)

But back to the data sharing matters at hand: Nielson goes on to examine some of the failures science online: comments sites, which lack a lot of buy ins; and wikipedia, which hasn't appreciated the contributions of the scientific community as expected.

Collaboration, Nielson argues, is quite natural in the scientific community. No scientists can answer all of the questions, even Albert Einstein occasionally asked a colleague and friend for help. The problem is translating this natural collaboration into the realm of the internet. Nielson closes his article with a sentiment often mentioned in class: "We still require a cultural change that embraces an open scientific culture. This will include new metrics that acknowledge online collaboration as a genuine scientific contribution — something that will act as an incentive for scientists to share their problems online."

Data and the Greater Good

"Sacrifice for the greater good?" Editorial. Nature. 421:6926 (February 27, 2003).

For this week I read an editorial on the topic of data sharing in Nature. The article commented on a recent (2003) closed meeting held in Fort Lauderdale regarding data sharing in the field of genomics. The editorial reported that the genomics researchers at this meeting expressed the desire for immediate deposition of data sets and unconditional access to those data sets, even if this means that the originator of the data loses priority over that data. In other words, once data was submitted anyone could access and download it, and anyone could publish from it, even if they scooped the originators. (This relates to data produced by publicly funded projects). Any attempt to impose licensing agreements to prevent the originators to lose publishing priority would not be allowed, and centers that tried to impose such licenses would not be eligible for funding. Such immediate and open access, they contend, is in the interest of the greater good and the advancement of science.

Not only do these genome researchers want these principles to apply to their own corner of the scientific community, they desire that these ideals of data sharing apply throughout the world of biology. The Fort Lauderdale proposal suggests "that any project where a 'community resource' is the objective should subject itself to the same principles of immediate release without conditions." The author of the editorial reports that these ideals are definitely not accepted by all in the scientific community. These principles, the author writes, are great if you're a top name in your field and don't really have to worry about losing out on funding, publication priority, and reputation if your data is scooped, but not so great if you are a scientist who is less favored, or just starting out in your field, and need the benefits that the traditional system of peer-reviewed publication gives. The Fort Lauderdale proposal does offer a way to give credit to data originators, however. They suggest that gene sequencers publish statements of intent describing the analysis they planned to do with the data. This would provide "a citable means of giving them credit" even if they are scooped when it comes to publishing a full-length article. The author of the editorial questions whether such a system would compensate for the loss of publication priority protection, and whether there would still be any incentive to engage in data generation.

The author also takes issue with the idea that funding agencies would insist on the adoption of the principle of immediate data deposition with loss of priority. The author fears that such "coercion" may cause gene sequencers to shy away from public funding. With regard to journals, the author feels that it is the job of journal editors not to take part in coercing scientists into depositing their data, as, for example, by refusing to review articles whose authors have not deposited their data at the time of submission. The editorial also asserts that it is the job of editors and peer reviewers to "do what they can" to ensure that sufficient credit is given to data originators, with the caveat that "if a good piece of whole-genome analysis arrives on our desks, we'll publish it whoever it comes from."

What drives Continued Knowledge Sharing?

This week we read about all the reasons why scientists do not share data, and why they should do so. I agree that data sharing is important. Unlike the depressing article we read by Sterling, T., & Weinkam, J., I even think that it may be possible to encourage more data sharing.

In response to the articles we read this week, I read this article, in which the authors make the point that in sustained knowledge sharing, "a distinction between knowledge-contribution and knowledge-seeking behaviors and an adequate emphasis on their variance in terms of user belief is needed." The authors argue that a Knowledge Management System should support both rolls, so that people will trust it as a place to put their knowledge. This will encourage knowledge sharing because people will seek prestige. I think that is true, infrastructure makes things easier, but I still have lingering questions about why scientists would share their data freely, and if they should have to.

There are reasons for not sharing data that I sympathize with, and reasons that I fundamentally think are wrong. For example, a scientist who will not share their data because they do not want anyone to disprove their finding is wrong. I think that disproving a theory is a reason to force data to be open, not the opposite. Authors are not protected from criticism, and I think it is equally important that scientists should not be protected from assessment.

The privacy of data can be protected. Anonymize the data, and make available what you can. Yes, it takes more work, but we have a responsibility to say, this is part of being a good scientist. You must share your data, so if some information needs to be protected, that should be planned for from the start.

The reasons that I can understand are more difficult to address. For example, if a scientist works long and hard to acquire data, they should have the right to not only use that data first and for whatever they can think of to do with it, but also to get credit when someone else uses the gathered data. However, data is a slippery slope of definitions. And getting credit for data seems to be a worse tangle.

According to the Stanford Fair Use Overview, "There are some things that copyright law will not protect. Copyright will not protect the titles of a book or movie, nor will it protect short phrases such as "Make my day." Copyright protection also doesn't cover facts, ideas or theories. These things are free for all to use without authorization."

There are others laws, like patent laws, that may cover some of these, but patent law can be a bit misunderstood as well, take the famous patented peanut butter and jelly sandwich. In addition, a phrase could be trademarked, but the likelihood is that scientists are not using anything in their data set as an advertising tool (I hope...)

But facts, ideas, and theories are not protected. So if data is a fact, then that data is not meant to be protected by copyright law. The way that fact is presented may be, but wouldn't that be the paper written about the data, not the data itself?

According to one article we read this week, Data at Work, there are both experimentalists, and theoretical modelers. If we need both (and I hope we do, because I consider myself an experimentalist) then we need to be certain that the experimentalist is encouraged to contribute. Regathering data seems awfully inefficient.
If I collect data, and share it, is there a way that I can be repaid for the time and effort I put into that? Or do I just need to hoard my data and keep publishing using it?
Is collecting data a public good, like utilities, and therefore should it be funded, alleviating the need for sustained reward as a motivation for researchers?

Sunday, November 8, 2009

Who's Who in Digital Humanities Collaboration

Spiro, Lisa. "Examples of Collaborative Digital Humanities Projects." Digital Scholarship in the Humanities (blog). Posted June 1, 2009. Available at: http://digitalscholarship.wordpress.com/2009/06/01/examples-of-collaborative-digital-humanities-projects/. (Accessed November 8, 2009).

In a long and wonderfully detailed blog post (yes, it even has footnotes!), Lisa Spiro provides a descriptive overview of collaboration in digital humanities projects. She begins by noting that historically collaboration has not been a part of the publication model in the humanities. As evidence, she cites her own finding that between 2004 and 2008 only 2% or the articles published in American Literary History were co-authored. Similarly, she notes that Cronin et al found in their longer-ranging survey that only 2% of articles published between 1900 and 2000 in the philosophy journal Mind had more than one author. Spiro, however, observes that the humanities do have an extensive tradition of circulating and providing feedback on one another's work, and that new digital technologies such as CommentPress and Zotero are helping to facilitate the exchange of ideas. For Spiro (and John Unsworth, whom she cites), collaboration in the digital humanities holds much potential: "Through online collaboration, scholars can divide labor (whether in making a translation, developing software, or building a digital collection), exchange and refine ideas (via blogs, wikis, listservs, virtual worlds, etc.), engage multiple perspectives, and work together to solve complex problems." Indeed she suggests that the incidence of humanities collaboration in the digital environment is higher than in the paper and ink world.

Spiro's discussion of collaboration only sometimes overlaps with the notion of data sharing in the sciences. I wonder if the differences between what counts as raw "data" in the humanities (i.e. primary sources--books, artworks, historical documents) vs. in sciences (i.e. observed and experimental data) means that collaboration is a more useful concept for the humanities than is data sharing.

Spiro spends the lion's share of her blog post providing detailed examples of different types of collaboration in the humanities. She explains that she was having difficulty articulating how collaboration functions in humanities research until she began exploring concrete examples. She divides the types of collaboration into three main categories: "facilitating communication and knowledge building," "sharing and aggregating content," and "collaborative annotation, transcription, and knowledge production." Classified under each heading are more discrete project types. Here is a skeleton of how she schematizes collaboration in the digital humanities (though, as she notes, there is inevitably some overlap among the types of collaboration:

Facilitating communication and knowledge building:
  • Online communities/virtual organizations (e.g. listservs, online forums, online communities, advanced video conferencing)
  • Collaboratories (which are virtual research environments that use advanced networking, remote instrumentation, databases, and digital libraries to foster "communication, collaboration, resource sharing, and research regardless of physical distance.")
Sharing and aggregating content:
  • Digital memory banks/user-contributed content (i.e. various projects to which users can contribute their own content, sometimes with the help of flickr and youtube, such as The Hurricane Digital Memory Bank and the Oxford-sponsored Great War Archive.)
  • Content aggregation and integration (i.e. federated digital collections which draw from a variety of archives to overcome the "silo" effect that can plague individual institutional collections. Two examples Spiro gives are the Walt Whitman Archive’s Finding Aids for Poetry Manuscripts and the the Quilt Index.)
  • Data sharing (e.g.Open Context , an archaeology project that permits researchers to upload, tag, analyze and share data sets).
Collaborative annotation, transcription, and knowledge production:
  • Crowdsourcing transcription (these include efforts that attempt to crowdsource transcription, just as Project Gutenberg and Project Madurai are crowdsourcing the proofreading of OCR texts).
  • Collaborative translation (e.g. Suda Online (SOL), which "brings together classicists to collaborate in translating into English the Suda, a tenth century encyclopedia of ancient learning written by a committee of Byzantine scholars.")
  • Collaborative editing (wherein collaborative online editions of texts are made)
  • Social bibliographies, collaborative filtering, and annotation (including platforms like zotero and eComma, which enable sharing bibliographies and collaborative annotation respectively)
  • Collaborative writing (e.g. subject wikis such as the Pynchon Wiki.)
  • Gaming: "collaborative play" and games as research (wherein "games provide motivation and a structure for collaboration" and "teamwork enables puzzles to be solved more rapidly.")
  • Publishing (Spiro's examples include posting materials online for peer-to-peer reviews prior to print publication)
  • Social learning (wherein participating in digital projects is a form of apprenticeship for undergraduates and graduate students, as they digitize materials, provide metadata, do programming, or contribute to a wiki).
I'd actually like to pause on this last item for a minute since it raises some of the same questions regarding the changing nature of authorship in the digital environment that we've been discussing throughout the course of the semester. Spiro's entry on "social learning" makes me a little uneasy since she frames the students' contributions as an interactive mode of "learning" rather than "authoring"--sure, they're learning, but they are also generating content as well. One thing I'd really like to see is a discussion of how digital collaboration is transforming notions of authorship in the humanities. Lisa Spiro recently blogged on the topic, but it wasn't quite as down and dirty as I wanted it to be. It is, however, an area in which she's conducting ongoing research.

In general, Spiro's blog entry on humanities collaboration is more descriptive than analytic. She doesn't really delve into the logic of her classifications or the problems associated with "social scholarship" (though she does link to an earlier blog of hers addressing this second topic). Accordingly, her piece is useful for familiarizing oneself with the types of collaboration going in the humanities, rather than thinking through the meaty issues associated with collaboration and sharing. Perhaps it might be worthwhile to consider such matters from a humanities perspective in class.

Saturday, November 7, 2009

Secrecy in Industry and Academic Science

W. Hong and J. P Walsh, “For Money or Glory?: Commercialization, Competition and Secrecy in the Entrepreneurial University” (2008).


We've talked about the various incentives or lack thereof for scientists to share pre- and post-publication data and knowledge, both in terms of the workload and the secrecy/competition factor. This article from Sociological Quarterly is interesting to that discussion as it examines secrecy in academic science with regards to two primary variables: its relations to industry science and the general level of competitiveness.

The study is based off two surveys: one conducted in 1966 by Warren Hagstrom, who surveyed a national random sample of 1,947 academic scientists in six fields (mathematics, experimental physics, theoretical physics, experimental biology, other biology, and chemistry), the other conducted in 1998 with a national random sample of 399 scientists from four fields (experimental biology, mathematics, physics, and sociology).

The authors use the 30+ years between the surveys to try and demonstrate a long-term change in the level of secrecy practiced in academic science. The surveys are appended to the article, but the authors sum up their content best:

The two surveys measure secrecy by asking how safe scientists feel in discussing their current research with others doing similar work. The surveys also include a measure of scientific competition, asking respondents how concerned they are about being anticipated in their current research. The later survey also includes measures of patenting, industry funding, industry collaboration, gender, institution type, seniority, and publication productivity.

Finally, the authors restrict their analysis to three fields: mathematics, physics and experimental biology.

Their results are unsurprising in some aspects (though not all) but very interesting regardless. Broadly they find that secrecy has increased among academic scientists, but that there has been an overemphasis on the effects of commercialization (collaboration, funding and so on from industry science) over a general rise in the competitiveness of science. The authors note competition is endemic to both realms of science and that as competition for priority (the first to discover, patent, publish, etc.) increases, antisocial behavior rises. This can be as benign as forgetting to answer data requests to deliberate concealment of knowledge to data fabrication. As one frequently cited author
(Robert K. Merton) summarizes, "The culture of science is, in this measure, pathogenic."

The authors found a slightly positive relationship between academic endeavors with industry funding and the level of secrecy. Interesting here too is a negative relationship between industry collaboration and secrecy. In this case ties with the commercial or industrial sector benefitted openness and data sharing. The authors also note that the field of experimental biology has seen the sharpest rise in secrecy, likely because of the commercial and public pressure for marketable products and a corresponding heavy involvement by the industry scientists and companies.

In resolving this problem the authors suggest a shift to a more secure funding stream along with a much reduced emphasis on immediate or short-term results. Of course such a change is not easy; the authors note that the open, communal knowledge-sharing principles of science is presently couched in a capitalist environment that values nondisclosure of knowledge to gain lead time to priority discoveries, products and patents. It's important to note that the authors don't believe there has been a change in the scientific model. Individuals still value sharing and openness but the constraints on sharing (priority, rewards, funding, advancement) have intensified to a degree that has become problematic.

So how can we reduce the competitiveness of academic science? So much work it seems is contingent on funding and the individual scientist. Ever-increasing competitiveness is not conducive to data curation, in fact I think it directly opposes it (as well as good science). The long scope of this study (1966-1998) has left me pretty worried.