Wednesday, November 18, 2009

Free Access to the Web

For my final blog post, I found an old (by internet standards) article from 2000 that views the entire web as one low-cost library and discusses the marvel that much of the access to it is free. Arms, writing at Cornell, envisions the web as replacing libraries in many people's lives. Whereas libraries, and especially research libraries, are incredibly expensive and limit access to their members, the web is self-funding (in that individual "publishers" pay for that privilege) with free access for anyone with an internet connection. (Of course, one could argue that this doesn't really constitute free...) Furthermore, Arms concludes that although the expected model of information provision on the internet was fee-based subscriptions, it turned out that there is enough free information of quality out there to obtain genuine substitutes. The example he gives is that Cornell provides legal sources via the web, including case reports, hitherto only available with an exorbitantly expensive Westlaw subscription, and one need not be affiliated with Cornell to access it.

Another point that Arms makes is that digital libraries can be nearly completely automated. He asserts that a brute force search, such as that provided by Google, with enough information and in the hands of a good researcher, can actually be much more powerful than an intelligent search by trained librarians. (While I don't like the implications for library services, it does seem like a lot of the focus on the need for reference librarians is in terms of not-good or amateur researchers...) This automation actually increases access as you no longer need to work through a small group of homogeneously trained elites.

Arms identifies two issues or potential problems for further research. One is insuring quality of information. This role was traditionally performed by the publishing process, but with the self-publishing afforded by the web, we can no longer count on good publishing practices. The second is permanence. Flip a switch on a server, and its information vanishes.

I found this article interesting mostly because of its now somewhat historic outlook. Some things have not turned out as Arms saw them in 2000, namely the level of free access. As we see with the Google Books project, proprietary interests are finding their way into the new cyber-reality, and the Great Copyright War has yet to be fought. Interestingly, though, Arms did identify two key issues that continue to be relevant: trusting found information, and ensuring its permanence. I suspect that the best answer to the former is via education of the public. At some point, the onus has to be on the searcher. The second problem, in my mind, is much more problematic, and it is one that plagues the physical as well as virtual information worlds. All in all, I found this early article quite interesting.

Tuesday, November 17, 2009

Aquatic and Riparian Effectiveness Monitoring Program

After talking about the problems facing ecological data, I wanted to read a little about the data collecting work my friend did over the summer for Aquatic and Riparian Effectiveness Monitoring Program (AREMP) in Oregon and see how that data fits in with the Long Term Ecological Research program. AREMP surveys 250 watersheds in the northwest and their collection practices (including what photos to take, how to record coordinates and how to use site markers) are described on their site. To collect the data they have to enlist a number of people like my friend to collect watershed samples using GIS. Although 250 is a lot of watersheds, it only represents about 10% of the watersheds in the area that is being sampled.

The goal of AREMP is to use a decision support model to evaluate watersheds for overall watershed condition. A number of attributes are assigned to each watershed and once all the attributes for a watershed are sampled the data is aggregated to determine a watershed score. To aggregate the data and find a score, AREMP uses software called Ecosystem Management Decision Support (EMDS) which creates the model and then assesses the condition of the watersheds based on the data. AREMP says that they would be happy to share their data with anybody who would like to see it.

EMDS is pretty interesting, the EMDS document says, "EMDS does contain tools for conducting “what if” scenarios. For example, one can estimate how watershed condition will improve if 500 pieces of large wood were added to the stream".

I wasn't able to find anything connecting AREMP with LTER but I did discover more problems with ecological data. For instance, the EPA and AREMP both use probability sampling designs but, "indicator and sampling methods differ from those used by the EPA, and these differences hinder collaboration and data comparison" (Hughes, 2008, p. 853).

I would still like to know more about AREMP's data and how their data collection methods compare to other ecological studies.

Hughes, R., Peck, D. (2008). Acquiring data for large aquatic resource surveys: The art of compromise among science, logistics, and reality. Journal of the North American Benthological Society (27)4, 837-859.

OCRIS: Online Catalogue and Repository Interoperability Study

JISC: Online Catalog and Repository Interoperability Study (OCRIS): Final Report

This study reviewed Library Management Systems (and the associated OPAC) with the Institutional Repositories (IR) of Higher Education Institutions in the UK.

The goals: determine whether the repository content within the scope of the institutional OPAC (and extent it is recorded in the OPAC); examine interoperability of OPAC and repository software; list services offered by OPAC's and repositories; identify potential for improvement in links to other institutional services; make recommendations for development of further links between OPAC's and repositories.

The primary findings are distressing. Only 2 percent of the study respondents stated that their systems were definitely interoperable, and 14 percent stated that interoperability was pending. There was an 81 percent overlap in scope for all items in IRs and OPACs; generally, IR's contained bibliographic data and OPACs contained full text.

Clearly, differences in scope/policy are not clear and there is either uncoordinated effort, hindering interoperability, or duplication of effort and/or redundant information. In order to provide a more feasible and appropriate long-term vision for IRs and OPACs, institutions should take a structured look at the goals of each service (IR vs OAPC) and coordinate efforts to best provide interoperability and reduce duplication of effort.

This was interesting... OPACs and IRs arguably have different intents - generically speaking, circulating collections versus long-term preservation and access, but they are really not so different. While both may not store information, each provides a service in locating information. In light of the volumes of money spent on IR development/OPACs/interoperability, and the number of available/established entities available to study, it is a very good time to step back and consider how these services might be better managed and coordinated to provide the best available service to the end user and the institution. I would like to see a corresponding study of American institutions.

Competing Requirements for Self-archiving

Stevan Harnad wrote a response blog to a letter to the editor that appeared in D-Lib Magazine. The original letter to the editor was a hypothetical dialog between an author whose work was recently accepted and an open-access consultant. In the dialog, the author is unsure about how to deal with a funder who requires the article to be deposited in an online repository and a publisher who may or may not allow the article to be deposited in a location they have no control over. In his post, Harnad takes excerpt from the dialog and writes more information about what the open-access consultant is saying, sometimes correcting him.
First, Harnad clarifies the opt-out clauses in self-archiving mandates. The opt-out clause is about whether or not you need to persuade the journal to accept your addendum thereby formalizing your right to deposit your article. While this is worthwhile, it is not essential so authors can choose to opt-out if they cannot persuade the journal or they simply don't wish to try. But regardless of the presence of an opt-out clause, authors still deposit their articles immediately. It is not necessary to find another publisher if the publisher denies your request.
Second, if the publisher has an embargo period, and the author wishes to honor it, they simply deposit the article as closed access for the appropriate period of time. There is no need to talk to the publisher about the embargo period.
Harnad suggests that depositing an article is not as confusing as it seems. An author simply needs to deposit all drafts as soon as an article is accepted for publication. According to his post, 63% of the top journals allow such a deposit. If your journal is one of the other 37%, simply set the article to closed access.
The issues surrounding depositing a published article are complex and probably result in many authors choosing to do nothing rather than try to navigate through all the competing requirements. Though closed access is not ideal, having a copy in a repository is better than not having anything stored. It seems that the more an author knows about the process of depositing an article and what they are allowed to do, the more likely it is that they will go to the trouble of depositing their article. The way things currently stand, most authors choose not to mess with online repositories because they don't even know where to find out what they are allowed to do.

Monday, November 16, 2009

JoVE (Journal of Visualized Experiments) is a new approach to peer reviewed journals : specifically devoted to the biological sciences (life sciences), JoVE is indexed by PubMed. The goal of JoVE aid the transmission of information - particularly that which is not sufficiently represented by static text and images. JoVE claims that their approach - using video publishing - "promotes efficiency and performance" by eliminating time wasted on learning and perfecting new techniques based on traditionally published works.

The Editorial Board for JoVE boasts members of the scientific community from the best institutions in the world... Harvard,
Mount Sinai School of Medicine, Princeton,
University of Zurich -the list illustrates an impressive number of highly qualified board members. The project was begun at Harvard in 2006 by a post-doc, Moishe Pritsker, now CEO and editor-in-chief of JoVE.

Access:
Initially, JoVE was conceived of as an open-access project, however, that model proved itself impossible, given the high costs of producing the videos. According to Pritsker,

"The reason is simple: we have to survive. To cover costs of our operations, to break even, we have to charge $6,000 per video article. This is to cover costs of the video-production and technological infrastructure for video-publication, which are higher than in traditional text-only publishing. Academic labs cannot pay $6,000 per article, and therefore we have to find other sources to cover the costs." (http://scholarlykitchen.sspnet.org/2009/04/06/jove/)

Thus, a pricing structure was created to cover these costs: "$1,000 for small colleges to $2,400 for PhD-granting institutions, prices which are in league with other commercial scientific journals. In addition, authors are charged $1,500 per article for video production services ($500 without), and there are open access options: $3,000/article with production services ($2,000 without)." (http://scholarlykitchen.sspnet.org/2009/04/06/jove/)

Taking a look JoVE's Press section, I was surprised to find they had posted articles criticizing their decision to go closed-access. While these criticisms are valid, so are JoVE's explanations of why open-access wasn't a possibility. Given the newness of this "product" and the fact that there is no existing model, it makes sense that charging for access is the only way to offset the price of video production without all of the costs falling on the research institute.

How it looks:
Videos are very high quality and accompanied by a complete scholarly article that acts as a transcript of sorts to the video.

Overall, I'm very impressed by JoVE, and while for now it is closed access it seems that as more stakeholders buy in and the model becomes more widely accepted it may be possible in the future.

Monkeying Around with Twitter Data

This week I was looking for articles about access to data and found a new interesting development. An Austin-based company called InfoChimps.org, that offers data sets for download, last week announced they were selling Twitter datasets. For quite a bit of money.

InfoChimps' mission "is to increase the world's access to structured data". The company appears to offer much data for free and prefers to be a platform where people may "post data under an open license". The data is available for browsing, and if the site doesn't actually have the data, it will point the user to where they can get the data for free. This is excellent for data sharing and access to large data sets. The site's homepage lists "Interesting Datasets" for perusal. I clicked on the first one, which was College Enrollment of Recent High School Completers 1960-2005. There is an intro paragraph to the data, where it's from, and an example of it. This set was prefaced with the caveat that the "files have data mixed with notes and references, multiple tables per sheet, and, worst of all, the table headers are not easily matched to their rows and columns." Kind of funny, yet good information to know!

So, Infochimps offers free datasets - fabulous. But their recent announcement to sell Twitter data met with some skeptical questions. The Read Write Web blog discusses this development in depth, describing the data, which isn't the full tweets, but hashtags, RTs, @ messages, and other associated info. This is apparently really useful and great information to have, and the developers at InfoChimps are hoping that people create interesting apps with this data. InfoChimps mined this data themselves, by hitting Twitter Developer API 20,000 times per hour (I almost know what this means). That's a LOT of data. Marshall Kirkpatrick, the author at RWW, questions the complete legality of selling this data, and worries Twitter is going to come a-knockin'. Many commenters on the blog entry also thought so. InfoChimps (last week) swore they were on solid legal ground, but a new post from yesterday on their blog revealed that Twitter had asked them to remove the datasets. While InfoChimps swears their data had nothing personally revealing, privacy concerns came from commenters and apparently Twitter, who claims they just want to prevent any 'malicious use' of the data.

I gleaned from the blog entries related to this issue that Twitter isn't very forthcoming with their data, which is making them a new Bad Guy of social media and data sharing. It's interesting that people are upset that Twitter won't share, but probably don't care as much if a university won't share some research datasets? Shouldn't data be data? Who cares if one is more sexy than another? Access is still important. InfoChimps sounds a little naive to think that selling Twitter datasets for $9000 wouldn't cause a stir, but maybe that's what they actually intended! At least it brought some mild attention to the issues at play here.

Sunday, November 15, 2009

Gordon

On November 4th, the San Diego Supercomputing Center (SDSC) announced that they have been awarded $20 million by the National Science Foundation to develop a new supercomputer aimed at "solving critical science and societal problems now overwhelmed by the avalanche of data generated by the digital devices of our era." This new computer is known as Gordon.

The SDSC has been part of the National Science Foundation's plan for cyberinfrastructure in the sciences for a number of years, as we read earlier in the semester. The development of Gordon is part of the NSF's continuing effort to keep up with the computing needs of the sciences. According to Jose L. Munoz, the deputy director and senior science advisor for the NSF's Office of Cyberinfrastructure, "'Gordon will do for data-driven science what tera-/peta-scale systems have done for the simulation and modeling communities, and provides a new tool to conduct transformative research.'"

Gordon, which is scheduled to be installed by Appro International, Inc. in 2011, will employ flash memory to speed solutions to data-intensive problems, such the analysis of individual genomes, much faster than spinning disk technology. Gordon is the follow up to the Dash system, another SDSC project, which was the first computer to use flash devices. Gordon will feature 245 teraflops of total compute power, 64 terabytes of DRAM, 256 terabytes of flash memory, and four petabytes of disk storage.

Another feature of Gordon will be 32 so-called "supernodes," each consisting of 32 compute nodes capable of 240 gigaflops/node and 64 gigabytes of DRAM. Linked together by virtual shared memory, these "supernodes" each have to potential of 7.7 teraflops of compute power and 10 terabytes of memory. The "supernodes" will be linked together with an InfiniBand network, which is capable of 16 gigabits per second of bidirectional bandwidth, which, apparently, is eight times faster than some of the most recently developed supercomputers. Gordon will be made available to researchers via an open-access national grid.

The point of Gordon is to make possible the complex applications of data-driven science. According to the press release, Gordon will have potential benefits in both the academic and industrial settings. Gordon should also be useful for predictive science, which seeks to make models of real-life phenomena. Because of the large-scale memory on a single node and the consequent increase in computation speeds, Gordon should allow for the creation of models that more accurately mimic these phenomena. With such capabilities, Gordon will be the next step in the development of data-driven science.



Wednesday, November 11, 2009

Articles for Rent

In my blog last week I addressed the issue of journal piracy, people actively hunting down journal article and re-posting them on forums so that they could be down loaded by others. This is a relatively predictable outcome to the shift to digital content, and I tried to make the point that in the end it probably did not effect the journal publishing system to much.

This week I am continuing with the journal access them by looking at the DeepDyve journal access system brought to my attention by this post which has a shout out to a U.T. librarian.

This company offers limited time access to journals articles through a “rental” system. Users can “rent” articles for .99 per day. Users do not have the ability to save off these article or to print them and they only have access to each article for 24 hours, at least at this level, there are other levels that allow users to access article for a week or a month with the price appropriately tiered. Users can browse the holdings of the company and view abstracts for free. On the page with the abstract the user has the option to purchase the article from the supplying institution. I decided to follow this link and the journal who published the article I was looking at also had a daily use option, but it was 8.00, DeepDyve was beginning to look like a better deal.

The system also has a set of widgets that collect articles and alert you via RSS feeds, this allows you to set up a query and have it alert you when new articles which match your query are available.

This program is aimed at those who have need of journal articles but do not have access to an institutional license, I really have no idea how big of a market this is, but I don't imagine that it is huge still I have been wrong before.

So how does it work as an article finder? Turns out pretty well, actually. This article reminded me of another article I read while putting together my blog for this week, it concerned searching abstracts for conclusion material and and extracting that data via semantic web / Linked data Format in order to print out summaries of relevant research, but I could not remember where I found it or who wrote it. So I typed what I could remember into the search bar and the article was on the first page. By the way the article can be found here , enjoy.

This is a different model for journal access than any we have seen yet. There are some things that I really like about it, and other things that I am for some reason hesitant to embrace.

Collection Space

Collection Space is another digital repository project supported by the Andrew Mellon Foundation, like JSTOR and ARTstor, but with very important differences. CollectionSpace emerged from scholarly institutional collaboration "with the common goal of developing and deploying an open-source, web-based software application for the description, management, and dissemination of museum collections information." Which is to say - It's an open-source software that will allow for museum collection management and dissemination but with many Web 2.0 twists. This is not your regular museum registrar database. CollectionSpace's project team currently includes web-developers, designers, and architects from Cambridge, Museum for the Moving Image, UC Berkeley School of Information and the University of Toronto.

This venture was inspired by the large information gap that existed for 1/3 to 1/2 of collecting institutions in the U.S. (historical societies, archeological repositories, museums and others) having neither their collections online or in many cases even catalog records. CollectionSpace addresses these needs through working to develop open-source software solutions providing stable, authoritative but flexible collection architecture "from which interpretive materials and experiences - from printed catalogs to mobile gallery guides - may be efficiently developed, and that can
serve as a cost-effective alternative to proprietary collections management systems for museums in need, regardless of size or scope." Think perhaps ARTstor for museums owning, managing (and publicizing) their own rotating collections.

The project began as a series of collaborative workshops with museum, archival and library professionals followed by a two-year period of software development complete by beta-testing by the same communities that helped advise its design.

CollectionSpace Release 0.2, debuted October 6, 2009. This version allows storing "multiple record types with flexible schemas." They have an extensive Project Wiki filled with a variety of documentation, handouts and powerpoint presentation.

Funding is supported by the Andrew W. Mellon Foundation, Program in Research in Information Technology which "supports the creation of "enterprise" administrative and infrastructural software by means of distributed, collaborative open-source development projects."

It aims to support registrars, curators, educators, collections managers, and administrators. It seeks to change earlier system models which seem to be digital representations of paper models. CollectionSpace promises to bring processes like cataloging, loans, media handling, location tracking "into the Web 2.0+ era." It seeks to move from "records-based navigation" to "action-based navigation."

Distributed Computing: It's OK To Share Data with Lots of People

Discussion on the separation between those who generate models and those who generate experimental data in Birnholtz and Bietz article this week led me to thinking about larger collaborative linkages in research. We could also consider the linkages between hardware designers, software designers, and the different demands and constraints placed on these groups. One public example that shows the collaboration between researchers generating experimental data and those who process it can be seen in the many distributed computing projects such as SETI@Home, Folding@Home, and others.

As a brief introduction, Folding@Home and projects like it rely on thousands, if not millions (SETI@Home has 3 million users currently) of users around the globe to download experimental data to their computer and process the data. In the case of SETI@Home, the data consists of radio telescope data that is analyzed for signs of extraterrestrial life, but other projects investigate biochemistry (protein folding) and many other scientific queries. Processing is usually automatic, but since users are downloading the data, the experimenters must trust that users will not tamper with the data in any way.

A posting from August 20 of this year on the Folding@Home blog mentioned the danger of tampering with data downloaded from the Folding@Home project. Users may tamper with data with completely benign intentions, such as maximizing their processor time, but we can see that data transformations might taint the overall data, and for this reason it is not allowed by this project. This is a simple example of something we have seen again and again: researchers need to maintain control over their data, especially in this case, where they are collecting results from processing done on their data with the future intention of making discoveries. This is a model for the way science will be done in the years to come: data is collected, stored, and then farmed out to processing teams, and fits in with the "stream" model in the aforementioned paper, instead of the discrete event data model. The fact that the data can be broken up for parallel processing enables this sort of massive crowdsourcing. Despite the huge gains in processor power over the past few decades, some problems are just too complex to solve overnight, no matter how many processor cores you have.
There is one additional benefit to projects like SETI@Home and Folding@Home: they provide an incentive to participate for users in the form of groups that can compete for prestige by contributing the largest number of processed hours, and there is a feeling that one is "helping out" to solve complex scientific problems. We have seen that one of the problems in getting scientists to share their data is figuring out how to tie the social and scientific rewards together, and I feel that these projects do this well.

Welcome to the caBIG� Community Website —

Welcome to the caBIG� Community Website —

Pertaining to this week's theme of knowledge sharing, I checked out a huge example of knowledge sharing: the National Cancer Institute's caBIG. CaBig defines itself as a "a virtual network of interconnected data, individuals, and organizations whose goal is to redefine how research is conducted, how care is provided, and to interact with others in biomedical research.

The goals of caBIG is to adapt or build tools for interacting wth information with cancer research, connect with other in cancer research, deploy and extend standards, rules, and common languages to facilitate interoperability. The website underlines that all this is needed because there is a "critical problem facing basic and clinical researchers today: the explosion of data that requires new and different approaches in sharing data."

Each of the different areas of cancer research, such as clinical trials, tissue banks and in vivo imaging, have their own workgroup that allows like minded people to post and share their information.

caBIG also has an entire area within their network devoted to data sharing to add on additional information features that researchers may need in order to share and display their research, such as patient privacy protections and intellectual property interests. The data sharing and security framework is designed to help researchers overcome the common barriers that obstruct the ease of sharing information. As it is, this particular feature is not fully developed as people are still working on defining policies, best practices, model documents and the trust fabric. Users can select levels of proprietary values, levels of privacy and restrictions and be able to adjust amount of data sharing based on the metadata that the scientist adds to it.

The caGrid is caBIG's underlying network architecture and platform that enables the connectivity within the network. It promotes interoperability within the site and all the biomedical research tools and facilitates communication between the users/researchers to render data sharing more efficient.

All in all, the caBIG network looks promising since the ease of sharing data about cancer research is the very core of its motivation. The researchers in the field also realize the importance of sharing data and are looking for ways to improve scholarly communication.

Legal Knowledge Sharing

This week I found an article written by the librarian of a private law firm on the value of internal knowledge sharing to firms and how this can be achieved via an intranet. I liked this article for a couple of reasons. First, it emphasizes the value of knowledge sharing even in non-scientific environments. Second, it rags on the fad of naming things such-and-such-2.0, which I also happen to think is pretty stupid. But, I digress.

Eiseman's argument is based on two premises. First, that one cannot necessarily determine what knowledge will qualify as useful to share. The example he gives is that of employers not wanting comany wikis or other public fora because they fear an inundation of private knowledge such as restaurant recommendations. Eiseman points out, however, that many legitimate business functions occur in restaurants, so someone looking for a new place to take a client would benefit from that knowledge. Second, Eiseman argues that the entire organization can benefit by collectively sharing knowledge formerly possessed solely by specialists. His example for this is posting reference questions and their answers to a wiki instead of to a two person email. This is efficient, in that it (in theory) prevents repetition of the same questions. However, it also spreads general knowledge of specific legal situations throughout the entire firm, which will enable the various lawyers to better serve their clients rather than referring them to a second attorney.

Especially in a setting such as a law firm, user-generated content on the wikis (or other knowledge sharing apps... Eiseman discusses a couple) serves an essential role in knowledge sharing, as each firm partner, associate, and librarian will have an indvidualized body of expertise.

The one thing that I did wonder about from this article, is why limit knowledge sharing on the law to one's own firm. I suspect it is a mix of liability fears (only trust knowledge with a known provenance!) along with the competetive nature of the legal field. Although, an idealist might note that the more determinant a given legal situation is, the better off everyone operating within that field would be. (Obviously, I'd like to see a bit more legal knowledge sharing than via firm intranets...) All in all, though, I thought the article provided good perspective on the importance of knowledge sharing in non-scientific areas.

EveryBlock: data about your neighborhood

Adrian Holovaty leads the team that has created a new website called Everyblock. The site, which is relatively new and only has information available for 15 cities, pulls information from a variety of sources and creates a kind of activity stream for specific neighborhoods within these cities. Everyblock is designed to keep track of everything that is going on in your neighborhood, from local news stories to civic information like building permits, restaurant inspections and even changes in liqueur licenses, as well as activity from social networking site like photos of the neighborhood that have shown up on Flickr, Craigslist postings, and user reviews of local businesses.
Site construction on EveryBlock began in 2007, with the launch of its first 3 cities in early 2008. Since then, it has continued to expand the number of cities it covers and EveryBlock was acquired by MSNBC this past August. Designed to tell users what is happening in their own neighborhood, each city block has its own news feed that lists information that applies to its location. It does not show information about the city in general. For example, a feed on the Hollywood neighborhood in LA will contain information specific to that area but will not contain information about LA in general, even if it has implications for the neighborhood in Hollywood. They also do not put in directory type information about an area so you won't find information about who lives down the street from you. The site also offers email alerts and RSS feeds so users don't have to visit the site to know what is going on in their area.
The combination of different sources (news, civic and social) was at first a bit surprising. While it seems a bit less far-fetched after having looked at the site, I still can't help but wonder if anyone would use EveryBlock for anything other than entertainment purposes. Do people really want to know every time someone takes a picture of the local flower shop or every time the restaurant down the street is inspected? I don't actually know. But this particular combination of data from different sources is something I haven't seen before and I've certainly never thought about all the different types of data generated about a particular area as being something that could (or should) be combined. If nothing else, EveryBlock helps users to visualize the incredible amount of data generated about a particular location and offers a way to make that data easier to access.

Policy: NIH Data Sharing for Genome Research

This week I came across the Notice on Development of Data Sharing Policy for Sequence and Related Genomic Data from the NIH. This notice says that the NIH is considering new additions to their policy on genome research. In class we've discussed what the motivations of the genome investigators might have been for being open with their data and so far there doesn't seem to be a single answer. However, this ongoing revision of policy from the NIH indicates that this area of science has unique requirements for data sharing because of the resources that are needed to acquire genome data. The NIH mentions that these resource needs "necessarily limits the number of projects that can be supported for any disease". In addition to these factors, the NIH says that these data sets have "added scientific value when combined with other large data sets".

The policy that the NIH is working on right now is basically a continuation of the Policy for Sharing of Data Obtained in NIH Supported or Conducted Genome-Wide Association Studies (GWAS) which went into effect January 2008. Basically, when researchers have genome data they have to put it in the central genome-wide associated studies (GWAS) database which makes all genome data accessible for combining with other data sets as early as possible. Although other people can see it, the investigator who deposits data has exclusive rights to publication based on that data for some period of time, the NIH encourages investigators to keep this period of time short and they limit the exclusivity period to 12 months.

Is the NIH the only organization that has this kind of exclusivity period? I remember a discussion in class about an investigator who deposited their data into a repository and then somebody else published a paper during the exclusivity period. I also recall that the organization with the policy and database (again, I'm pretty sure it was the NIH) didn't seem very alarmed by the situation.

One reason why the NIH is updating their policy might be because of the confidential nature of the data and the need for participant privacy. In the original policy the NIH predicted that in a few years technology would "make the identification of specific individuals from raw genotype-phenotype data feasible and increasingly straightforward", which is problematic for data in a federal repository because that data is accessible through FOIA. At that time they decided that they would have to deny FOIA requests for unredacted data sets. About eight months after that policy went into effect the new developments on how to identify people based on genome data was made public (NIH reigns in genome access) and the NIH has to further restrict access to the data - NIH Modifications to Genome-Wide Association Studies (GWAS) Data Access (sorry, that's a PDF).

This article in PLoS Genetics explores the issue of confidentiality with human genome research in more detail: Public access to genome-wide data: Five views on balancing research with privacy and protection. One person basically says we can't trust anybody to keep data confidential, even though any researcher accessing GWAS would have to agree to the policy agreeing that they would "not attempt to identify individual participants from whom data within a dataset were obtained".

Tuesday, November 10, 2009

Facebook for Natural Historians? This is getting silly...

Article: "Darwin Meets Facebook: Social Networking Tool Lets Natural Historians Share Data" from ScienceDaily (Nov. 9, 2009).

As with other types of data we have mentioned this semester, curating biological data is a challenge. Apparently, taxonomic data (which is data about the classification of living things) has grown in volume over the years, while the number of taxonomic experts has decreased, and researchers have found it increasingly more difficult to find, access, and publish taxonomic data. To help with these problems, the idea for a social networking tool for natural historians to share their data was born. This tool, called Scratchpads, is being developed by researchers at the Natural History Museum of London.

Scratchpads got its name because it “reflects the fact that taxonomy is in a state of 'perpetual beta', constantly changing and reforming to reflect our current knowledge of a particular group.” The sites and “virtual workbenches” users create act as online notepads; the goal of Scratchpads is that users will work together to 'scratch out' their ideas and share taxonomic information. Scratchpad sites provide an open access to data in ways that traditional publications can’t.

Scientists and researchers can use Scratchpads to manage information about classifications, phylogenies, bibliographies documents, image galleries, custom data, specimen records, and even maps. Registration is free for any interested scientist; all that is needed to use scratchpads is to complete a registration form.

As far as technology is concerned, Scratchpads uses a Content Management System called Drupal, which provides the underlying architecture that Scratchpads relies on. Drupal allows Scratchpads to offer features such blogs, collaborative authoring, forums, peer-to-peer networking, newsletters, picture galleries, file uploads and downloads, and more.

The goal of Scratchpads is to be user-friendly and encourage users to generate, organize, and share their data. According to the Scratchpads website, its “infrastructure combines databases, network protocols and computational services to bring people, information and computational tools together to perform and publish natural history.” Scratchpads can help users collaborate via the web on projects that might have taken individual users a lifetime to complete alone by providing a space on the web for “communities to bring taxonomic information together without the limitations of traditional paper based publications.”

I blogged recently about the Facebook for Scientists which is being funded by the NIH. Now with Scratchpads, we apparently have a Facebook for Natural Historians. I’m not sure how I feel about this. The name is cute, and the idea is trendy, but how truly useful is this? On one hand, it is great that researchers are trying to collaborate and spread open access to information, but at the same time it feels like people keep trying to reinvent the wheel, developing similar projects that will probably fail to live up to expectations.

Why Isaac Newton is influencing science sharing today // how to make scientists share

As I was looking for an article related to data sharing for this weeks blog I came back to the September Nature special on Open Access, and Bryan Nelson's article "Data sharing: Empty archives." I know that this particular issue of Nature was widely read by our class, so in an attempt to limit redundancy I followed the Comments of Nelson's article to a similarly themed article.

The article in question "Doing science in the open" by Michael Nielson appeared on physicsworld.com in May. In the early years of science - the really early years - scientists made every effort to keep their discoveries secret. Going so far as to publishing ideas as anagrams in order to have time to work out the kinks of a theory while keeping the idea itself a mystery. This practice, it seems was rather common in the 17th and 18th centuries. Scientists were, presumably, motivated in large part by personal gain. This culture began to change when governments became the patrons of scientists - scientific discovery benefited the public good and the reputation of both the scientists and their patron governments. Since this time some 300 years ago which brought about the practices of scholarly publishing we know today little has changed.

As Nelson noted in "Data sharing: Empty archives," the scientific field has been slow, in some views, to adopt digital data sharing. Nielson, however, notes several successes in scientific data sharing: the "physics preprint server arXiv, which lets physicists share preprints of their papers without the months-long delay typical of a conventional journal, and GenBank, an online database where biologists can deposit and search for DNA sequences." Also, Nielson describes (what I consider to be the coolest data sharing in science I've heard of so far) the "Journal of Visualized Experiments, [JoVE] which lets scientists upload videos that show how their experiments work."

(At the point I reached this part of the article, I wanted to jump ship and blog about JoVE - a site well worth a look!)

But back to the data sharing matters at hand: Nielson goes on to examine some of the failures science online: comments sites, which lack a lot of buy ins; and wikipedia, which hasn't appreciated the contributions of the scientific community as expected.

Collaboration, Nielson argues, is quite natural in the scientific community. No scientists can answer all of the questions, even Albert Einstein occasionally asked a colleague and friend for help. The problem is translating this natural collaboration into the realm of the internet. Nielson closes his article with a sentiment often mentioned in class: "We still require a cultural change that embraces an open scientific culture. This will include new metrics that acknowledge online collaboration as a genuine scientific contribution — something that will act as an incentive for scientists to share their problems online."

Data and the Greater Good

"Sacrifice for the greater good?" Editorial. Nature. 421:6926 (February 27, 2003).

For this week I read an editorial on the topic of data sharing in Nature. The article commented on a recent (2003) closed meeting held in Fort Lauderdale regarding data sharing in the field of genomics. The editorial reported that the genomics researchers at this meeting expressed the desire for immediate deposition of data sets and unconditional access to those data sets, even if this means that the originator of the data loses priority over that data. In other words, once data was submitted anyone could access and download it, and anyone could publish from it, even if they scooped the originators. (This relates to data produced by publicly funded projects). Any attempt to impose licensing agreements to prevent the originators to lose publishing priority would not be allowed, and centers that tried to impose such licenses would not be eligible for funding. Such immediate and open access, they contend, is in the interest of the greater good and the advancement of science.

Not only do these genome researchers want these principles to apply to their own corner of the scientific community, they desire that these ideals of data sharing apply throughout the world of biology. The Fort Lauderdale proposal suggests "that any project where a 'community resource' is the objective should subject itself to the same principles of immediate release without conditions." The author of the editorial reports that these ideals are definitely not accepted by all in the scientific community. These principles, the author writes, are great if you're a top name in your field and don't really have to worry about losing out on funding, publication priority, and reputation if your data is scooped, but not so great if you are a scientist who is less favored, or just starting out in your field, and need the benefits that the traditional system of peer-reviewed publication gives. The Fort Lauderdale proposal does offer a way to give credit to data originators, however. They suggest that gene sequencers publish statements of intent describing the analysis they planned to do with the data. This would provide "a citable means of giving them credit" even if they are scooped when it comes to publishing a full-length article. The author of the editorial questions whether such a system would compensate for the loss of publication priority protection, and whether there would still be any incentive to engage in data generation.

The author also takes issue with the idea that funding agencies would insist on the adoption of the principle of immediate data deposition with loss of priority. The author fears that such "coercion" may cause gene sequencers to shy away from public funding. With regard to journals, the author feels that it is the job of journal editors not to take part in coercing scientists into depositing their data, as, for example, by refusing to review articles whose authors have not deposited their data at the time of submission. The editorial also asserts that it is the job of editors and peer reviewers to "do what they can" to ensure that sufficient credit is given to data originators, with the caveat that "if a good piece of whole-genome analysis arrives on our desks, we'll publish it whoever it comes from."

What drives Continued Knowledge Sharing?

This week we read about all the reasons why scientists do not share data, and why they should do so. I agree that data sharing is important. Unlike the depressing article we read by Sterling, T., & Weinkam, J., I even think that it may be possible to encourage more data sharing.

In response to the articles we read this week, I read this article, in which the authors make the point that in sustained knowledge sharing, "a distinction between knowledge-contribution and knowledge-seeking behaviors and an adequate emphasis on their variance in terms of user belief is needed." The authors argue that a Knowledge Management System should support both rolls, so that people will trust it as a place to put their knowledge. This will encourage knowledge sharing because people will seek prestige. I think that is true, infrastructure makes things easier, but I still have lingering questions about why scientists would share their data freely, and if they should have to.

There are reasons for not sharing data that I sympathize with, and reasons that I fundamentally think are wrong. For example, a scientist who will not share their data because they do not want anyone to disprove their finding is wrong. I think that disproving a theory is a reason to force data to be open, not the opposite. Authors are not protected from criticism, and I think it is equally important that scientists should not be protected from assessment.

The privacy of data can be protected. Anonymize the data, and make available what you can. Yes, it takes more work, but we have a responsibility to say, this is part of being a good scientist. You must share your data, so if some information needs to be protected, that should be planned for from the start.

The reasons that I can understand are more difficult to address. For example, if a scientist works long and hard to acquire data, they should have the right to not only use that data first and for whatever they can think of to do with it, but also to get credit when someone else uses the gathered data. However, data is a slippery slope of definitions. And getting credit for data seems to be a worse tangle.

According to the Stanford Fair Use Overview, "There are some things that copyright law will not protect. Copyright will not protect the titles of a book or movie, nor will it protect short phrases such as "Make my day." Copyright protection also doesn't cover facts, ideas or theories. These things are free for all to use without authorization."

There are others laws, like patent laws, that may cover some of these, but patent law can be a bit misunderstood as well, take the famous patented peanut butter and jelly sandwich. In addition, a phrase could be trademarked, but the likelihood is that scientists are not using anything in their data set as an advertising tool (I hope...)

But facts, ideas, and theories are not protected. So if data is a fact, then that data is not meant to be protected by copyright law. The way that fact is presented may be, but wouldn't that be the paper written about the data, not the data itself?

According to one article we read this week, Data at Work, there are both experimentalists, and theoretical modelers. If we need both (and I hope we do, because I consider myself an experimentalist) then we need to be certain that the experimentalist is encouraged to contribute. Regathering data seems awfully inefficient.
If I collect data, and share it, is there a way that I can be repaid for the time and effort I put into that? Or do I just need to hoard my data and keep publishing using it?
Is collecting data a public good, like utilities, and therefore should it be funded, alleviating the need for sustained reward as a motivation for researchers?

Sunday, November 8, 2009

Who's Who in Digital Humanities Collaboration

Spiro, Lisa. "Examples of Collaborative Digital Humanities Projects." Digital Scholarship in the Humanities (blog). Posted June 1, 2009. Available at: http://digitalscholarship.wordpress.com/2009/06/01/examples-of-collaborative-digital-humanities-projects/. (Accessed November 8, 2009).

In a long and wonderfully detailed blog post (yes, it even has footnotes!), Lisa Spiro provides a descriptive overview of collaboration in digital humanities projects. She begins by noting that historically collaboration has not been a part of the publication model in the humanities. As evidence, she cites her own finding that between 2004 and 2008 only 2% or the articles published in American Literary History were co-authored. Similarly, she notes that Cronin et al found in their longer-ranging survey that only 2% of articles published between 1900 and 2000 in the philosophy journal Mind had more than one author. Spiro, however, observes that the humanities do have an extensive tradition of circulating and providing feedback on one another's work, and that new digital technologies such as CommentPress and Zotero are helping to facilitate the exchange of ideas. For Spiro (and John Unsworth, whom she cites), collaboration in the digital humanities holds much potential: "Through online collaboration, scholars can divide labor (whether in making a translation, developing software, or building a digital collection), exchange and refine ideas (via blogs, wikis, listservs, virtual worlds, etc.), engage multiple perspectives, and work together to solve complex problems." Indeed she suggests that the incidence of humanities collaboration in the digital environment is higher than in the paper and ink world.

Spiro's discussion of collaboration only sometimes overlaps with the notion of data sharing in the sciences. I wonder if the differences between what counts as raw "data" in the humanities (i.e. primary sources--books, artworks, historical documents) vs. in sciences (i.e. observed and experimental data) means that collaboration is a more useful concept for the humanities than is data sharing.

Spiro spends the lion's share of her blog post providing detailed examples of different types of collaboration in the humanities. She explains that she was having difficulty articulating how collaboration functions in humanities research until she began exploring concrete examples. She divides the types of collaboration into three main categories: "facilitating communication and knowledge building," "sharing and aggregating content," and "collaborative annotation, transcription, and knowledge production." Classified under each heading are more discrete project types. Here is a skeleton of how she schematizes collaboration in the digital humanities (though, as she notes, there is inevitably some overlap among the types of collaboration:

Facilitating communication and knowledge building:
  • Online communities/virtual organizations (e.g. listservs, online forums, online communities, advanced video conferencing)
  • Collaboratories (which are virtual research environments that use advanced networking, remote instrumentation, databases, and digital libraries to foster "communication, collaboration, resource sharing, and research regardless of physical distance.")
Sharing and aggregating content:
  • Digital memory banks/user-contributed content (i.e. various projects to which users can contribute their own content, sometimes with the help of flickr and youtube, such as The Hurricane Digital Memory Bank and the Oxford-sponsored Great War Archive.)
  • Content aggregation and integration (i.e. federated digital collections which draw from a variety of archives to overcome the "silo" effect that can plague individual institutional collections. Two examples Spiro gives are the Walt Whitman Archive’s Finding Aids for Poetry Manuscripts and the the Quilt Index.)
  • Data sharing (e.g.Open Context , an archaeology project that permits researchers to upload, tag, analyze and share data sets).
Collaborative annotation, transcription, and knowledge production:
  • Crowdsourcing transcription (these include efforts that attempt to crowdsource transcription, just as Project Gutenberg and Project Madurai are crowdsourcing the proofreading of OCR texts).
  • Collaborative translation (e.g. Suda Online (SOL), which "brings together classicists to collaborate in translating into English the Suda, a tenth century encyclopedia of ancient learning written by a committee of Byzantine scholars.")
  • Collaborative editing (wherein collaborative online editions of texts are made)
  • Social bibliographies, collaborative filtering, and annotation (including platforms like zotero and eComma, which enable sharing bibliographies and collaborative annotation respectively)
  • Collaborative writing (e.g. subject wikis such as the Pynchon Wiki.)
  • Gaming: "collaborative play" and games as research (wherein "games provide motivation and a structure for collaboration" and "teamwork enables puzzles to be solved more rapidly.")
  • Publishing (Spiro's examples include posting materials online for peer-to-peer reviews prior to print publication)
  • Social learning (wherein participating in digital projects is a form of apprenticeship for undergraduates and graduate students, as they digitize materials, provide metadata, do programming, or contribute to a wiki).
I'd actually like to pause on this last item for a minute since it raises some of the same questions regarding the changing nature of authorship in the digital environment that we've been discussing throughout the course of the semester. Spiro's entry on "social learning" makes me a little uneasy since she frames the students' contributions as an interactive mode of "learning" rather than "authoring"--sure, they're learning, but they are also generating content as well. One thing I'd really like to see is a discussion of how digital collaboration is transforming notions of authorship in the humanities. Lisa Spiro recently blogged on the topic, but it wasn't quite as down and dirty as I wanted it to be. It is, however, an area in which she's conducting ongoing research.

In general, Spiro's blog entry on humanities collaboration is more descriptive than analytic. She doesn't really delve into the logic of her classifications or the problems associated with "social scholarship" (though she does link to an earlier blog of hers addressing this second topic). Accordingly, her piece is useful for familiarizing oneself with the types of collaboration going in the humanities, rather than thinking through the meaty issues associated with collaboration and sharing. Perhaps it might be worthwhile to consider such matters from a humanities perspective in class.

Saturday, November 7, 2009

Secrecy in Industry and Academic Science

W. Hong and J. P Walsh, “For Money or Glory?: Commercialization, Competition and Secrecy in the Entrepreneurial University” (2008).


We've talked about the various incentives or lack thereof for scientists to share pre- and post-publication data and knowledge, both in terms of the workload and the secrecy/competition factor. This article from Sociological Quarterly is interesting to that discussion as it examines secrecy in academic science with regards to two primary variables: its relations to industry science and the general level of competitiveness.

The study is based off two surveys: one conducted in 1966 by Warren Hagstrom, who surveyed a national random sample of 1,947 academic scientists in six fields (mathematics, experimental physics, theoretical physics, experimental biology, other biology, and chemistry), the other conducted in 1998 with a national random sample of 399 scientists from four fields (experimental biology, mathematics, physics, and sociology).

The authors use the 30+ years between the surveys to try and demonstrate a long-term change in the level of secrecy practiced in academic science. The surveys are appended to the article, but the authors sum up their content best:

The two surveys measure secrecy by asking how safe scientists feel in discussing their current research with others doing similar work. The surveys also include a measure of scientific competition, asking respondents how concerned they are about being anticipated in their current research. The later survey also includes measures of patenting, industry funding, industry collaboration, gender, institution type, seniority, and publication productivity.

Finally, the authors restrict their analysis to three fields: mathematics, physics and experimental biology.

Their results are unsurprising in some aspects (though not all) but very interesting regardless. Broadly they find that secrecy has increased among academic scientists, but that there has been an overemphasis on the effects of commercialization (collaboration, funding and so on from industry science) over a general rise in the competitiveness of science. The authors note competition is endemic to both realms of science and that as competition for priority (the first to discover, patent, publish, etc.) increases, antisocial behavior rises. This can be as benign as forgetting to answer data requests to deliberate concealment of knowledge to data fabrication. As one frequently cited author
(Robert K. Merton) summarizes, "The culture of science is, in this measure, pathogenic."

The authors found a slightly positive relationship between academic endeavors with industry funding and the level of secrecy. Interesting here too is a negative relationship between industry collaboration and secrecy. In this case ties with the commercial or industrial sector benefitted openness and data sharing. The authors also note that the field of experimental biology has seen the sharpest rise in secrecy, likely because of the commercial and public pressure for marketable products and a corresponding heavy involvement by the industry scientists and companies.

In resolving this problem the authors suggest a shift to a more secure funding stream along with a much reduced emphasis on immediate or short-term results. Of course such a change is not easy; the authors note that the open, communal knowledge-sharing principles of science is presently couched in a capitalist environment that values nondisclosure of knowledge to gain lead time to priority discoveries, products and patents. It's important to note that the authors don't believe there has been a change in the scientific model. Individuals still value sharing and openness but the constraints on sharing (priority, rewards, funding, advancement) have intensified to a degree that has become problematic.

So how can we reduce the competitiveness of academic science? So much work it seems is contingent on funding and the individual scientist. Ever-increasing competitiveness is not conducive to data curation, in fact I think it directly opposes it (as well as good science). The long scope of this study (1966-1998) has left me pretty worried.