Wednesday, November 11, 2009

Articles for Rent

In my blog last week I addressed the issue of journal piracy, people actively hunting down journal article and re-posting them on forums so that they could be down loaded by others. This is a relatively predictable outcome to the shift to digital content, and I tried to make the point that in the end it probably did not effect the journal publishing system to much.

This week I am continuing with the journal access them by looking at the DeepDyve journal access system brought to my attention by this post which has a shout out to a U.T. librarian.

This company offers limited time access to journals articles through a “rental” system. Users can “rent” articles for .99 per day. Users do not have the ability to save off these article or to print them and they only have access to each article for 24 hours, at least at this level, there are other levels that allow users to access article for a week or a month with the price appropriately tiered. Users can browse the holdings of the company and view abstracts for free. On the page with the abstract the user has the option to purchase the article from the supplying institution. I decided to follow this link and the journal who published the article I was looking at also had a daily use option, but it was 8.00, DeepDyve was beginning to look like a better deal.

The system also has a set of widgets that collect articles and alert you via RSS feeds, this allows you to set up a query and have it alert you when new articles which match your query are available.

This program is aimed at those who have need of journal articles but do not have access to an institutional license, I really have no idea how big of a market this is, but I don't imagine that it is huge still I have been wrong before.

So how does it work as an article finder? Turns out pretty well, actually. This article reminded me of another article I read while putting together my blog for this week, it concerned searching abstracts for conclusion material and and extracting that data via semantic web / Linked data Format in order to print out summaries of relevant research, but I could not remember where I found it or who wrote it. So I typed what I could remember into the search bar and the article was on the first page. By the way the article can be found here , enjoy.

This is a different model for journal access than any we have seen yet. There are some things that I really like about it, and other things that I am for some reason hesitant to embrace.

Collection Space

Collection Space is another digital repository project supported by the Andrew Mellon Foundation, like JSTOR and ARTstor, but with very important differences. CollectionSpace emerged from scholarly institutional collaboration "with the common goal of developing and deploying an open-source, web-based software application for the description, management, and dissemination of museum collections information." Which is to say - It's an open-source software that will allow for museum collection management and dissemination but with many Web 2.0 twists. This is not your regular museum registrar database. CollectionSpace's project team currently includes web-developers, designers, and architects from Cambridge, Museum for the Moving Image, UC Berkeley School of Information and the University of Toronto.

This venture was inspired by the large information gap that existed for 1/3 to 1/2 of collecting institutions in the U.S. (historical societies, archeological repositories, museums and others) having neither their collections online or in many cases even catalog records. CollectionSpace addresses these needs through working to develop open-source software solutions providing stable, authoritative but flexible collection architecture "from which interpretive materials and experiences - from printed catalogs to mobile gallery guides - may be efficiently developed, and that can
serve as a cost-effective alternative to proprietary collections management systems for museums in need, regardless of size or scope." Think perhaps ARTstor for museums owning, managing (and publicizing) their own rotating collections.

The project began as a series of collaborative workshops with museum, archival and library professionals followed by a two-year period of software development complete by beta-testing by the same communities that helped advise its design.

CollectionSpace Release 0.2, debuted October 6, 2009. This version allows storing "multiple record types with flexible schemas." They have an extensive Project Wiki filled with a variety of documentation, handouts and powerpoint presentation.

Funding is supported by the Andrew W. Mellon Foundation, Program in Research in Information Technology which "supports the creation of "enterprise" administrative and infrastructural software by means of distributed, collaborative open-source development projects."

It aims to support registrars, curators, educators, collections managers, and administrators. It seeks to change earlier system models which seem to be digital representations of paper models. CollectionSpace promises to bring processes like cataloging, loans, media handling, location tracking "into the Web 2.0+ era." It seeks to move from "records-based navigation" to "action-based navigation."

Distributed Computing: It's OK To Share Data with Lots of People

Discussion on the separation between those who generate models and those who generate experimental data in Birnholtz and Bietz article this week led me to thinking about larger collaborative linkages in research. We could also consider the linkages between hardware designers, software designers, and the different demands and constraints placed on these groups. One public example that shows the collaboration between researchers generating experimental data and those who process it can be seen in the many distributed computing projects such as SETI@Home, Folding@Home, and others.

As a brief introduction, Folding@Home and projects like it rely on thousands, if not millions (SETI@Home has 3 million users currently) of users around the globe to download experimental data to their computer and process the data. In the case of SETI@Home, the data consists of radio telescope data that is analyzed for signs of extraterrestrial life, but other projects investigate biochemistry (protein folding) and many other scientific queries. Processing is usually automatic, but since users are downloading the data, the experimenters must trust that users will not tamper with the data in any way.

A posting from August 20 of this year on the Folding@Home blog mentioned the danger of tampering with data downloaded from the Folding@Home project. Users may tamper with data with completely benign intentions, such as maximizing their processor time, but we can see that data transformations might taint the overall data, and for this reason it is not allowed by this project. This is a simple example of something we have seen again and again: researchers need to maintain control over their data, especially in this case, where they are collecting results from processing done on their data with the future intention of making discoveries. This is a model for the way science will be done in the years to come: data is collected, stored, and then farmed out to processing teams, and fits in with the "stream" model in the aforementioned paper, instead of the discrete event data model. The fact that the data can be broken up for parallel processing enables this sort of massive crowdsourcing. Despite the huge gains in processor power over the past few decades, some problems are just too complex to solve overnight, no matter how many processor cores you have.
There is one additional benefit to projects like SETI@Home and Folding@Home: they provide an incentive to participate for users in the form of groups that can compete for prestige by contributing the largest number of processed hours, and there is a feeling that one is "helping out" to solve complex scientific problems. We have seen that one of the problems in getting scientists to share their data is figuring out how to tie the social and scientific rewards together, and I feel that these projects do this well.

Welcome to the caBIG� Community Website —

Welcome to the caBIG� Community Website —

Pertaining to this week's theme of knowledge sharing, I checked out a huge example of knowledge sharing: the National Cancer Institute's caBIG. CaBig defines itself as a "a virtual network of interconnected data, individuals, and organizations whose goal is to redefine how research is conducted, how care is provided, and to interact with others in biomedical research.

The goals of caBIG is to adapt or build tools for interacting wth information with cancer research, connect with other in cancer research, deploy and extend standards, rules, and common languages to facilitate interoperability. The website underlines that all this is needed because there is a "critical problem facing basic and clinical researchers today: the explosion of data that requires new and different approaches in sharing data."

Each of the different areas of cancer research, such as clinical trials, tissue banks and in vivo imaging, have their own workgroup that allows like minded people to post and share their information.

caBIG also has an entire area within their network devoted to data sharing to add on additional information features that researchers may need in order to share and display their research, such as patient privacy protections and intellectual property interests. The data sharing and security framework is designed to help researchers overcome the common barriers that obstruct the ease of sharing information. As it is, this particular feature is not fully developed as people are still working on defining policies, best practices, model documents and the trust fabric. Users can select levels of proprietary values, levels of privacy and restrictions and be able to adjust amount of data sharing based on the metadata that the scientist adds to it.

The caGrid is caBIG's underlying network architecture and platform that enables the connectivity within the network. It promotes interoperability within the site and all the biomedical research tools and facilitates communication between the users/researchers to render data sharing more efficient.

All in all, the caBIG network looks promising since the ease of sharing data about cancer research is the very core of its motivation. The researchers in the field also realize the importance of sharing data and are looking for ways to improve scholarly communication.

Legal Knowledge Sharing

This week I found an article written by the librarian of a private law firm on the value of internal knowledge sharing to firms and how this can be achieved via an intranet. I liked this article for a couple of reasons. First, it emphasizes the value of knowledge sharing even in non-scientific environments. Second, it rags on the fad of naming things such-and-such-2.0, which I also happen to think is pretty stupid. But, I digress.

Eiseman's argument is based on two premises. First, that one cannot necessarily determine what knowledge will qualify as useful to share. The example he gives is that of employers not wanting comany wikis or other public fora because they fear an inundation of private knowledge such as restaurant recommendations. Eiseman points out, however, that many legitimate business functions occur in restaurants, so someone looking for a new place to take a client would benefit from that knowledge. Second, Eiseman argues that the entire organization can benefit by collectively sharing knowledge formerly possessed solely by specialists. His example for this is posting reference questions and their answers to a wiki instead of to a two person email. This is efficient, in that it (in theory) prevents repetition of the same questions. However, it also spreads general knowledge of specific legal situations throughout the entire firm, which will enable the various lawyers to better serve their clients rather than referring them to a second attorney.

Especially in a setting such as a law firm, user-generated content on the wikis (or other knowledge sharing apps... Eiseman discusses a couple) serves an essential role in knowledge sharing, as each firm partner, associate, and librarian will have an indvidualized body of expertise.

The one thing that I did wonder about from this article, is why limit knowledge sharing on the law to one's own firm. I suspect it is a mix of liability fears (only trust knowledge with a known provenance!) along with the competetive nature of the legal field. Although, an idealist might note that the more determinant a given legal situation is, the better off everyone operating within that field would be. (Obviously, I'd like to see a bit more legal knowledge sharing than via firm intranets...) All in all, though, I thought the article provided good perspective on the importance of knowledge sharing in non-scientific areas.

EveryBlock: data about your neighborhood

Adrian Holovaty leads the team that has created a new website called Everyblock. The site, which is relatively new and only has information available for 15 cities, pulls information from a variety of sources and creates a kind of activity stream for specific neighborhoods within these cities. Everyblock is designed to keep track of everything that is going on in your neighborhood, from local news stories to civic information like building permits, restaurant inspections and even changes in liqueur licenses, as well as activity from social networking site like photos of the neighborhood that have shown up on Flickr, Craigslist postings, and user reviews of local businesses.
Site construction on EveryBlock began in 2007, with the launch of its first 3 cities in early 2008. Since then, it has continued to expand the number of cities it covers and EveryBlock was acquired by MSNBC this past August. Designed to tell users what is happening in their own neighborhood, each city block has its own news feed that lists information that applies to its location. It does not show information about the city in general. For example, a feed on the Hollywood neighborhood in LA will contain information specific to that area but will not contain information about LA in general, even if it has implications for the neighborhood in Hollywood. They also do not put in directory type information about an area so you won't find information about who lives down the street from you. The site also offers email alerts and RSS feeds so users don't have to visit the site to know what is going on in their area.
The combination of different sources (news, civic and social) was at first a bit surprising. While it seems a bit less far-fetched after having looked at the site, I still can't help but wonder if anyone would use EveryBlock for anything other than entertainment purposes. Do people really want to know every time someone takes a picture of the local flower shop or every time the restaurant down the street is inspected? I don't actually know. But this particular combination of data from different sources is something I haven't seen before and I've certainly never thought about all the different types of data generated about a particular area as being something that could (or should) be combined. If nothing else, EveryBlock helps users to visualize the incredible amount of data generated about a particular location and offers a way to make that data easier to access.

Policy: NIH Data Sharing for Genome Research

This week I came across the Notice on Development of Data Sharing Policy for Sequence and Related Genomic Data from the NIH. This notice says that the NIH is considering new additions to their policy on genome research. In class we've discussed what the motivations of the genome investigators might have been for being open with their data and so far there doesn't seem to be a single answer. However, this ongoing revision of policy from the NIH indicates that this area of science has unique requirements for data sharing because of the resources that are needed to acquire genome data. The NIH mentions that these resource needs "necessarily limits the number of projects that can be supported for any disease". In addition to these factors, the NIH says that these data sets have "added scientific value when combined with other large data sets".

The policy that the NIH is working on right now is basically a continuation of the Policy for Sharing of Data Obtained in NIH Supported or Conducted Genome-Wide Association Studies (GWAS) which went into effect January 2008. Basically, when researchers have genome data they have to put it in the central genome-wide associated studies (GWAS) database which makes all genome data accessible for combining with other data sets as early as possible. Although other people can see it, the investigator who deposits data has exclusive rights to publication based on that data for some period of time, the NIH encourages investigators to keep this period of time short and they limit the exclusivity period to 12 months.

Is the NIH the only organization that has this kind of exclusivity period? I remember a discussion in class about an investigator who deposited their data into a repository and then somebody else published a paper during the exclusivity period. I also recall that the organization with the policy and database (again, I'm pretty sure it was the NIH) didn't seem very alarmed by the situation.

One reason why the NIH is updating their policy might be because of the confidential nature of the data and the need for participant privacy. In the original policy the NIH predicted that in a few years technology would "make the identification of specific individuals from raw genotype-phenotype data feasible and increasingly straightforward", which is problematic for data in a federal repository because that data is accessible through FOIA. At that time they decided that they would have to deny FOIA requests for unredacted data sets. About eight months after that policy went into effect the new developments on how to identify people based on genome data was made public (NIH reigns in genome access) and the NIH has to further restrict access to the data - NIH Modifications to Genome-Wide Association Studies (GWAS) Data Access (sorry, that's a PDF).

This article in PLoS Genetics explores the issue of confidentiality with human genome research in more detail: Public access to genome-wide data: Five views on balancing research with privacy and protection. One person basically says we can't trust anybody to keep data confidential, even though any researcher accessing GWAS would have to agree to the policy agreeing that they would "not attempt to identify individual participants from whom data within a dataset were obtained".

Tuesday, November 10, 2009

Facebook for Natural Historians? This is getting silly...

Article: "Darwin Meets Facebook: Social Networking Tool Lets Natural Historians Share Data" from ScienceDaily (Nov. 9, 2009).

As with other types of data we have mentioned this semester, curating biological data is a challenge. Apparently, taxonomic data (which is data about the classification of living things) has grown in volume over the years, while the number of taxonomic experts has decreased, and researchers have found it increasingly more difficult to find, access, and publish taxonomic data. To help with these problems, the idea for a social networking tool for natural historians to share their data was born. This tool, called Scratchpads, is being developed by researchers at the Natural History Museum of London.

Scratchpads got its name because it “reflects the fact that taxonomy is in a state of 'perpetual beta', constantly changing and reforming to reflect our current knowledge of a particular group.” The sites and “virtual workbenches” users create act as online notepads; the goal of Scratchpads is that users will work together to 'scratch out' their ideas and share taxonomic information. Scratchpad sites provide an open access to data in ways that traditional publications can’t.

Scientists and researchers can use Scratchpads to manage information about classifications, phylogenies, bibliographies documents, image galleries, custom data, specimen records, and even maps. Registration is free for any interested scientist; all that is needed to use scratchpads is to complete a registration form.

As far as technology is concerned, Scratchpads uses a Content Management System called Drupal, which provides the underlying architecture that Scratchpads relies on. Drupal allows Scratchpads to offer features such blogs, collaborative authoring, forums, peer-to-peer networking, newsletters, picture galleries, file uploads and downloads, and more.

The goal of Scratchpads is to be user-friendly and encourage users to generate, organize, and share their data. According to the Scratchpads website, its “infrastructure combines databases, network protocols and computational services to bring people, information and computational tools together to perform and publish natural history.” Scratchpads can help users collaborate via the web on projects that might have taken individual users a lifetime to complete alone by providing a space on the web for “communities to bring taxonomic information together without the limitations of traditional paper based publications.”

I blogged recently about the Facebook for Scientists which is being funded by the NIH. Now with Scratchpads, we apparently have a Facebook for Natural Historians. I’m not sure how I feel about this. The name is cute, and the idea is trendy, but how truly useful is this? On one hand, it is great that researchers are trying to collaborate and spread open access to information, but at the same time it feels like people keep trying to reinvent the wheel, developing similar projects that will probably fail to live up to expectations.

Why Isaac Newton is influencing science sharing today // how to make scientists share

As I was looking for an article related to data sharing for this weeks blog I came back to the September Nature special on Open Access, and Bryan Nelson's article "Data sharing: Empty archives." I know that this particular issue of Nature was widely read by our class, so in an attempt to limit redundancy I followed the Comments of Nelson's article to a similarly themed article.

The article in question "Doing science in the open" by Michael Nielson appeared on physicsworld.com in May. In the early years of science - the really early years - scientists made every effort to keep their discoveries secret. Going so far as to publishing ideas as anagrams in order to have time to work out the kinks of a theory while keeping the idea itself a mystery. This practice, it seems was rather common in the 17th and 18th centuries. Scientists were, presumably, motivated in large part by personal gain. This culture began to change when governments became the patrons of scientists - scientific discovery benefited the public good and the reputation of both the scientists and their patron governments. Since this time some 300 years ago which brought about the practices of scholarly publishing we know today little has changed.

As Nelson noted in "Data sharing: Empty archives," the scientific field has been slow, in some views, to adopt digital data sharing. Nielson, however, notes several successes in scientific data sharing: the "physics preprint server arXiv, which lets physicists share preprints of their papers without the months-long delay typical of a conventional journal, and GenBank, an online database where biologists can deposit and search for DNA sequences." Also, Nielson describes (what I consider to be the coolest data sharing in science I've heard of so far) the "Journal of Visualized Experiments, [JoVE] which lets scientists upload videos that show how their experiments work."

(At the point I reached this part of the article, I wanted to jump ship and blog about JoVE - a site well worth a look!)

But back to the data sharing matters at hand: Nielson goes on to examine some of the failures science online: comments sites, which lack a lot of buy ins; and wikipedia, which hasn't appreciated the contributions of the scientific community as expected.

Collaboration, Nielson argues, is quite natural in the scientific community. No scientists can answer all of the questions, even Albert Einstein occasionally asked a colleague and friend for help. The problem is translating this natural collaboration into the realm of the internet. Nielson closes his article with a sentiment often mentioned in class: "We still require a cultural change that embraces an open scientific culture. This will include new metrics that acknowledge online collaboration as a genuine scientific contribution — something that will act as an incentive for scientists to share their problems online."

Data and the Greater Good

"Sacrifice for the greater good?" Editorial. Nature. 421:6926 (February 27, 2003).

For this week I read an editorial on the topic of data sharing in Nature. The article commented on a recent (2003) closed meeting held in Fort Lauderdale regarding data sharing in the field of genomics. The editorial reported that the genomics researchers at this meeting expressed the desire for immediate deposition of data sets and unconditional access to those data sets, even if this means that the originator of the data loses priority over that data. In other words, once data was submitted anyone could access and download it, and anyone could publish from it, even if they scooped the originators. (This relates to data produced by publicly funded projects). Any attempt to impose licensing agreements to prevent the originators to lose publishing priority would not be allowed, and centers that tried to impose such licenses would not be eligible for funding. Such immediate and open access, they contend, is in the interest of the greater good and the advancement of science.

Not only do these genome researchers want these principles to apply to their own corner of the scientific community, they desire that these ideals of data sharing apply throughout the world of biology. The Fort Lauderdale proposal suggests "that any project where a 'community resource' is the objective should subject itself to the same principles of immediate release without conditions." The author of the editorial reports that these ideals are definitely not accepted by all in the scientific community. These principles, the author writes, are great if you're a top name in your field and don't really have to worry about losing out on funding, publication priority, and reputation if your data is scooped, but not so great if you are a scientist who is less favored, or just starting out in your field, and need the benefits that the traditional system of peer-reviewed publication gives. The Fort Lauderdale proposal does offer a way to give credit to data originators, however. They suggest that gene sequencers publish statements of intent describing the analysis they planned to do with the data. This would provide "a citable means of giving them credit" even if they are scooped when it comes to publishing a full-length article. The author of the editorial questions whether such a system would compensate for the loss of publication priority protection, and whether there would still be any incentive to engage in data generation.

The author also takes issue with the idea that funding agencies would insist on the adoption of the principle of immediate data deposition with loss of priority. The author fears that such "coercion" may cause gene sequencers to shy away from public funding. With regard to journals, the author feels that it is the job of journal editors not to take part in coercing scientists into depositing their data, as, for example, by refusing to review articles whose authors have not deposited their data at the time of submission. The editorial also asserts that it is the job of editors and peer reviewers to "do what they can" to ensure that sufficient credit is given to data originators, with the caveat that "if a good piece of whole-genome analysis arrives on our desks, we'll publish it whoever it comes from."

What drives Continued Knowledge Sharing?

This week we read about all the reasons why scientists do not share data, and why they should do so. I agree that data sharing is important. Unlike the depressing article we read by Sterling, T., & Weinkam, J., I even think that it may be possible to encourage more data sharing.

In response to the articles we read this week, I read this article, in which the authors make the point that in sustained knowledge sharing, "a distinction between knowledge-contribution and knowledge-seeking behaviors and an adequate emphasis on their variance in terms of user belief is needed." The authors argue that a Knowledge Management System should support both rolls, so that people will trust it as a place to put their knowledge. This will encourage knowledge sharing because people will seek prestige. I think that is true, infrastructure makes things easier, but I still have lingering questions about why scientists would share their data freely, and if they should have to.

There are reasons for not sharing data that I sympathize with, and reasons that I fundamentally think are wrong. For example, a scientist who will not share their data because they do not want anyone to disprove their finding is wrong. I think that disproving a theory is a reason to force data to be open, not the opposite. Authors are not protected from criticism, and I think it is equally important that scientists should not be protected from assessment.

The privacy of data can be protected. Anonymize the data, and make available what you can. Yes, it takes more work, but we have a responsibility to say, this is part of being a good scientist. You must share your data, so if some information needs to be protected, that should be planned for from the start.

The reasons that I can understand are more difficult to address. For example, if a scientist works long and hard to acquire data, they should have the right to not only use that data first and for whatever they can think of to do with it, but also to get credit when someone else uses the gathered data. However, data is a slippery slope of definitions. And getting credit for data seems to be a worse tangle.

According to the Stanford Fair Use Overview, "There are some things that copyright law will not protect. Copyright will not protect the titles of a book or movie, nor will it protect short phrases such as "Make my day." Copyright protection also doesn't cover facts, ideas or theories. These things are free for all to use without authorization."

There are others laws, like patent laws, that may cover some of these, but patent law can be a bit misunderstood as well, take the famous patented peanut butter and jelly sandwich. In addition, a phrase could be trademarked, but the likelihood is that scientists are not using anything in their data set as an advertising tool (I hope...)

But facts, ideas, and theories are not protected. So if data is a fact, then that data is not meant to be protected by copyright law. The way that fact is presented may be, but wouldn't that be the paper written about the data, not the data itself?

According to one article we read this week, Data at Work, there are both experimentalists, and theoretical modelers. If we need both (and I hope we do, because I consider myself an experimentalist) then we need to be certain that the experimentalist is encouraged to contribute. Regathering data seems awfully inefficient.
If I collect data, and share it, is there a way that I can be repaid for the time and effort I put into that? Or do I just need to hoard my data and keep publishing using it?
Is collecting data a public good, like utilities, and therefore should it be funded, alleviating the need for sustained reward as a motivation for researchers?

Sunday, November 8, 2009

Who's Who in Digital Humanities Collaboration

Spiro, Lisa. "Examples of Collaborative Digital Humanities Projects." Digital Scholarship in the Humanities (blog). Posted June 1, 2009. Available at: http://digitalscholarship.wordpress.com/2009/06/01/examples-of-collaborative-digital-humanities-projects/. (Accessed November 8, 2009).

In a long and wonderfully detailed blog post (yes, it even has footnotes!), Lisa Spiro provides a descriptive overview of collaboration in digital humanities projects. She begins by noting that historically collaboration has not been a part of the publication model in the humanities. As evidence, she cites her own finding that between 2004 and 2008 only 2% or the articles published in American Literary History were co-authored. Similarly, she notes that Cronin et al found in their longer-ranging survey that only 2% of articles published between 1900 and 2000 in the philosophy journal Mind had more than one author. Spiro, however, observes that the humanities do have an extensive tradition of circulating and providing feedback on one another's work, and that new digital technologies such as CommentPress and Zotero are helping to facilitate the exchange of ideas. For Spiro (and John Unsworth, whom she cites), collaboration in the digital humanities holds much potential: "Through online collaboration, scholars can divide labor (whether in making a translation, developing software, or building a digital collection), exchange and refine ideas (via blogs, wikis, listservs, virtual worlds, etc.), engage multiple perspectives, and work together to solve complex problems." Indeed she suggests that the incidence of humanities collaboration in the digital environment is higher than in the paper and ink world.

Spiro's discussion of collaboration only sometimes overlaps with the notion of data sharing in the sciences. I wonder if the differences between what counts as raw "data" in the humanities (i.e. primary sources--books, artworks, historical documents) vs. in sciences (i.e. observed and experimental data) means that collaboration is a more useful concept for the humanities than is data sharing.

Spiro spends the lion's share of her blog post providing detailed examples of different types of collaboration in the humanities. She explains that she was having difficulty articulating how collaboration functions in humanities research until she began exploring concrete examples. She divides the types of collaboration into three main categories: "facilitating communication and knowledge building," "sharing and aggregating content," and "collaborative annotation, transcription, and knowledge production." Classified under each heading are more discrete project types. Here is a skeleton of how she schematizes collaboration in the digital humanities (though, as she notes, there is inevitably some overlap among the types of collaboration:

Facilitating communication and knowledge building:
  • Online communities/virtual organizations (e.g. listservs, online forums, online communities, advanced video conferencing)
  • Collaboratories (which are virtual research environments that use advanced networking, remote instrumentation, databases, and digital libraries to foster "communication, collaboration, resource sharing, and research regardless of physical distance.")
Sharing and aggregating content:
  • Digital memory banks/user-contributed content (i.e. various projects to which users can contribute their own content, sometimes with the help of flickr and youtube, such as The Hurricane Digital Memory Bank and the Oxford-sponsored Great War Archive.)
  • Content aggregation and integration (i.e. federated digital collections which draw from a variety of archives to overcome the "silo" effect that can plague individual institutional collections. Two examples Spiro gives are the Walt Whitman Archive’s Finding Aids for Poetry Manuscripts and the the Quilt Index.)
  • Data sharing (e.g.Open Context , an archaeology project that permits researchers to upload, tag, analyze and share data sets).
Collaborative annotation, transcription, and knowledge production:
  • Crowdsourcing transcription (these include efforts that attempt to crowdsource transcription, just as Project Gutenberg and Project Madurai are crowdsourcing the proofreading of OCR texts).
  • Collaborative translation (e.g. Suda Online (SOL), which "brings together classicists to collaborate in translating into English the Suda, a tenth century encyclopedia of ancient learning written by a committee of Byzantine scholars.")
  • Collaborative editing (wherein collaborative online editions of texts are made)
  • Social bibliographies, collaborative filtering, and annotation (including platforms like zotero and eComma, which enable sharing bibliographies and collaborative annotation respectively)
  • Collaborative writing (e.g. subject wikis such as the Pynchon Wiki.)
  • Gaming: "collaborative play" and games as research (wherein "games provide motivation and a structure for collaboration" and "teamwork enables puzzles to be solved more rapidly.")
  • Publishing (Spiro's examples include posting materials online for peer-to-peer reviews prior to print publication)
  • Social learning (wherein participating in digital projects is a form of apprenticeship for undergraduates and graduate students, as they digitize materials, provide metadata, do programming, or contribute to a wiki).
I'd actually like to pause on this last item for a minute since it raises some of the same questions regarding the changing nature of authorship in the digital environment that we've been discussing throughout the course of the semester. Spiro's entry on "social learning" makes me a little uneasy since she frames the students' contributions as an interactive mode of "learning" rather than "authoring"--sure, they're learning, but they are also generating content as well. One thing I'd really like to see is a discussion of how digital collaboration is transforming notions of authorship in the humanities. Lisa Spiro recently blogged on the topic, but it wasn't quite as down and dirty as I wanted it to be. It is, however, an area in which she's conducting ongoing research.

In general, Spiro's blog entry on humanities collaboration is more descriptive than analytic. She doesn't really delve into the logic of her classifications or the problems associated with "social scholarship" (though she does link to an earlier blog of hers addressing this second topic). Accordingly, her piece is useful for familiarizing oneself with the types of collaboration going in the humanities, rather than thinking through the meaty issues associated with collaboration and sharing. Perhaps it might be worthwhile to consider such matters from a humanities perspective in class.

Saturday, November 7, 2009

Secrecy in Industry and Academic Science

W. Hong and J. P Walsh, “For Money or Glory?: Commercialization, Competition and Secrecy in the Entrepreneurial University” (2008).


We've talked about the various incentives or lack thereof for scientists to share pre- and post-publication data and knowledge, both in terms of the workload and the secrecy/competition factor. This article from Sociological Quarterly is interesting to that discussion as it examines secrecy in academic science with regards to two primary variables: its relations to industry science and the general level of competitiveness.

The study is based off two surveys: one conducted in 1966 by Warren Hagstrom, who surveyed a national random sample of 1,947 academic scientists in six fields (mathematics, experimental physics, theoretical physics, experimental biology, other biology, and chemistry), the other conducted in 1998 with a national random sample of 399 scientists from four fields (experimental biology, mathematics, physics, and sociology).

The authors use the 30+ years between the surveys to try and demonstrate a long-term change in the level of secrecy practiced in academic science. The surveys are appended to the article, but the authors sum up their content best:

The two surveys measure secrecy by asking how safe scientists feel in discussing their current research with others doing similar work. The surveys also include a measure of scientific competition, asking respondents how concerned they are about being anticipated in their current research. The later survey also includes measures of patenting, industry funding, industry collaboration, gender, institution type, seniority, and publication productivity.

Finally, the authors restrict their analysis to three fields: mathematics, physics and experimental biology.

Their results are unsurprising in some aspects (though not all) but very interesting regardless. Broadly they find that secrecy has increased among academic scientists, but that there has been an overemphasis on the effects of commercialization (collaboration, funding and so on from industry science) over a general rise in the competitiveness of science. The authors note competition is endemic to both realms of science and that as competition for priority (the first to discover, patent, publish, etc.) increases, antisocial behavior rises. This can be as benign as forgetting to answer data requests to deliberate concealment of knowledge to data fabrication. As one frequently cited author
(Robert K. Merton) summarizes, "The culture of science is, in this measure, pathogenic."

The authors found a slightly positive relationship between academic endeavors with industry funding and the level of secrecy. Interesting here too is a negative relationship between industry collaboration and secrecy. In this case ties with the commercial or industrial sector benefitted openness and data sharing. The authors also note that the field of experimental biology has seen the sharpest rise in secrecy, likely because of the commercial and public pressure for marketable products and a corresponding heavy involvement by the industry scientists and companies.

In resolving this problem the authors suggest a shift to a more secure funding stream along with a much reduced emphasis on immediate or short-term results. Of course such a change is not easy; the authors note that the open, communal knowledge-sharing principles of science is presently couched in a capitalist environment that values nondisclosure of knowledge to gain lead time to priority discoveries, products and patents. It's important to note that the authors don't believe there has been a change in the scientific model. Individuals still value sharing and openness but the constraints on sharing (priority, rewards, funding, advancement) have intensified to a degree that has become problematic.

So how can we reduce the competitiveness of academic science? So much work it seems is contingent on funding and the individual scientist. Ever-increasing competitiveness is not conducive to data curation, in fact I think it directly opposes it (as well as good science). The long scope of this study (1966-1998) has left me pretty worried.






Wednesday, November 4, 2009

Ontologies and data sharing

So what do y'all think of the idea that we could create one big ontology that describes everything? It's a pretty stereotypically librarian thing to do, but it might be the only way to realize Tim Berners-Lee's idea of the semantic web. Ephram Miles once told my Organizing Information class that he thought it might be overly ambitious and probably not possible to create an ontology that is broad yet descriptive enough to fully realize the semantic web in the way we dream about it today. In any case, the W3C just came out with the second version of its Web Ontology Language (OWL 2), which is a standard way of representing ontologies. The press release about this from the W3C doesn't mention giant ontologies, but it does talk a lot about discipline-specific ontologies. The tag line for the press release is "OWL 2 Connects the Web of Knowledge with the Web of Data." So I became curious about how ontologies might affect our discussions of data sharing.

The W3C lists a lot of semantic web case studies and use cases. I skimmed a few of these, and it appears most people are turning to ontologies to improve search. I'm interested, though, if ontologies might be able to improve data sharing. Tim Berners-Lee (TBL) talks about data sharing in his 2006 IEEE Intelligent Systems article, "The semantic web revisited." In fact, TBL suggests that in order for the larger semantic web to work, scientific communities have to first pioneer the way, much like they did in the early days of the web. There are a lot of ontologies out there. Just do a Google search for "biology ontology." I found many sites with many ontologies listed on them. So aren't we back where we started? There are lots of isolated "data sets" (read: ontologies) that don't really relate to each other. So we have the same problem with ontologies that we do with data: how do we assign authority? how do we create appropriate metadata? how do we make these islands talk to each other? TBL suggests we need more data mining, ways to infer knowledge from patterns, and distributed information systems to "crowd source" the ontologies. But the argument still seems circular to me. It seems like he's suggesting we use computer models to then build ontologies to then build computer models to connect the data.

Austin Forum on Science, Technology & Society

I had every intention of blogging on topic and even identified a couple of articles about metadata, but then I attended the Austin Forum on Science, Technology & Society. This month’s speaker, Gary Chapman, a Senior Lecturer in the LBJ School of Public Affairs, spoke on “The Internet and the Obama Administration – So Far.” Chapman worked with the Obama campaign and discussed the ways the Internet was used as a tool during the election, as well as the administration’s current web presence.

Chapman covered three primary topics during his talk: Obama’s online presence post-election, the issues of network neutrality, and online transparency. The third is the most related to topics of digital curation, but I thought I’d say a bit about the first two as well. The Obama administration used social networking sites like Facebook and Twitter to generate interest during the campaign and continue to use these sites post-election. According to Chapman, Obama was the first person to have over 1 million Facebook friends and Twitter followers. One primary aim was to raise funds using the Internet, which turned out to be enormously successful. Post-election, the administration has created a YouTube channel for the White House on which they broadcast a weekly address, and maintained Twitter and Facebook pages. Criticisms of these practices include questions about using of corporate companies for public purposes.

A primary topic of concern is over issues of net neutrality. Obama put Julius Genachowski into the position of Chairman of the FCC and charged him with enforcing network neutrality. The main question that arises however is whether the FCC has, or should have, the power to regulate the Internet. Chapman asserts that he does not believe this issue is a top priority for Obama at present, given the state of health care and international relations.

The third topic Chapman addressed was transparency on the web. Obama affirmed his commitment to freedom of information and web transparency on Inauguration Day and quickly launched sites that provided governmental information to the public, including data.gov (which Meg blogged about weeks ago) and Recovery.gov, which was launched in February to provide access to data related to the Recovery Act. The Obama administration is using these digital tools to increase transparency and access to government information. However, there has been some backlash, as the administration has not been completely open about everything, including negotiations with telecommunication agencies.

When reflecting back on issues of trust and authority, the government attempts at digital curation serve as interesting examples. Who owns the information generated by the government but made public through sites like Facebook and Twitter? Will Obama’s use of Twitter increase the need for or interest in finding a way to preserve tweets? Should the FCC be able to regulate the Internet? How will that change the free use of the technology? Will the transparency of the Obama administration continue beyond his presidency? What degree of openness and availability of information should we expect to have from the administration available online?

Let's Talk About OCLC

Katie's blog post got me thinking about OCLC and the power of metadata.

There are a few blog postings that might be useful to read:
I think that OCLC is interesting!

Innkeeper at the Roach Motel

Salo, Dorothea. (2008). Innkeeper at the Roach Motel. Library Trends. 57(2): 98-123. Retrieved from: http://muse.jhu.edu.ezproxy.lib.utexas.edu/content/crossref/journals/library_trends/v057/57.2.salo.html


Most of the articles we have read have largely focused on the wonderful possibilities of cyberinfrastructure and the attempts done so far to ease the transition from a largely analog world to a digital ruled environment. I have always been curious to know whether there were any dissenting opinions. Thanks to our class's Zotero account, I found Dorothea Salo's article "Innkeeper at the Roach Motel." It is certainly a dissenting voice amongst all the rosy and golden articles we've read this semester.* In fact, Salo compares institutional repositories to roach motels because according to her "whatever goes in there never comes out."

*Not that the golden and rosy articles are off the mark. Many of them are written very carefully and throuroughly, but Salo's article is so wary about the success of IR and open access that it made my eyes bug out.

Salo starts off bluntly: institutional repositories (IR) are not thriving, not catching on with libraries and university faculty in the way is needed for any kind of information transformation. Salo points out that the way people respond to the idea of an IR in their library has been of trepidition and negativity, citing a lack of trust and sdoubts of its credibility as major factors. She also states that initially, the idea of open access is appealing, but in the long run, it is not a major selling point for potential IR users.

Salo echoes many of the oft-discussed challenges of open access: issues of credibility, how to benefit from it monetarily and the general attitude that scholars have in regard to tenure and achieving academic fame. She states that, as wonderful as open access may sound, old habits take a long time to die and it may be a while before open access can catch on in the academic world.

Salo lists several factors that have led to IR's inability to revolutionize scholarly communication as promised and explains in detail of the inconsistencies and inefficiecies. IR and cyberinfrastructure requires participation from several groups: scholars/university faculty that submit work for IRs to collect, librarians or other staff members to manage and campaign the IR efforts, and the software developers to create and upkeep the technology required to run the IR.

The scholars and the university faculty members present a large barrier in the IR effort because not only do they not trust the credibility or status quo of submitting work in open access journals , but in general there is a distrust of any source digital as it is essentially contradictory to the tenure process. Also, in the area e-publishing with all its "pre-prints" and so on, not many people are familiar with the process and are not willing to educate themselves in this area. Librarians and other staff members that would be involved with the IR also display an open lack of trust of such a system and are reluctant to engage with it.

To put it bluntly, Salo calls IR "parasitic on existing research" because what they offer to people is not useful for them and what would be useful is not offered, because the program developers of IR technology is not collaborating with university staff and libraries. Essentially, everyone is in their little corner grousing about how this system with great potential is not working and ignoring the fact that if they collaborated, a newer and more useful system could emerge.

Salo examines different models that libraries have adopted in an attempt to create, manage and run an IR, lists their responsibilties and judges the advantages and disadvantages of each model. A IR manager is new area of expertise and there really no community of practice developed for this area of librarianship yet.

"Maverick Manager": librarian specifically in charge of the IR, has many different responsibilities and skills, but has no real place in the library's organizational structure. The maverick manager has freedom to experiment, but very little to no resources. These types of IR manager models have a large turnover due to lack of community, status and funds to do their job.

"No Accountability" Model: the IR is built by IT or an outside source is contracted to build an IR. Responsibilities are dispersed amongst the librarians for promoting and collecting the IR sources. This system is advantageous because it reduces the amount of groundwork that libraries have to do in order to create the IR, but this model also provides countless chances for miscommunication and requires librarians to travel a steep learning curve in getting familiar with the software.

"Consortial" Model: libraries share IR. This system is efficient, yet increases trouble with outreach, tech support, and content development.

Cooperative Model: this modle has been the most successful by far. Usually the IR is launched by a university administrator rather than a librarian. The librarians mediate content deposits, perform active searches through the Net and other sources for content to put in the IR and sometimes an automated workflow will automatically deposit work in the IR. The biggest problem with this system is that it is very expensive.

Salo ends this article with solutions that would increase the appeal of IRs and ensure a more positive rapport with IRs. First of all, support needs to be shown at home, libraries, library school programs, library organizations. IRs must be integrated into other library programs and priorities. IRs must take an active role in searching for content and be sensitive to faculty needs, digitize analog content, and develop relationships with software developers.

Dorothea Salo's article presents a different view on the whole idea of IR and outlines some of the challenges that prevent IRs from being truly successful. I think this answers a lot of my questions as to why I haven't heard more about IR until I started the Information Science program.

Social Meme Tracking: Here Come the Swedes

The first time you might have heard of Twingly was back in early 2007 with the launch of the Twingly Screensaver. The screensaver was an interesting innovation. For those of you who have not heard of it, the screensaver features a constantly updated list of the titles of every single blog post that is written anywhere in the world. The list scrolls in real time and there are links to each blog post which provide the URL, the name of the blog, the author, and the first two lines of the post. While an interesting development, from a digital curation perspective, the general criticism lobbed against the screensaver was that it just wasn't "useful." The critique goes something like - what is a user going to do with so much raw, unfiltered data?

Since the screensaver, Twingly has actually been looking at making the blogosphere more accessible and searchable. The Twingly Blog Search and Twingly Micro-Blog Search have become wildly popular in Europe. Anecdotally, I've been trying to find a way to get an RSS feed for one of my favorite Middle East opinion columnists, Robert Fisk, who writes for The Independent. After failing to find a RSS feed for Fisk alone (without the other Independent columnists) using Google Blog Search, Technorati, and even the RSS page on the website of the Independent, I was able to do so using the Twingly Blog Search. How they managed to provide a RSS feed for something that the source website does not provide is beyond my comprehension, but good for them!

Probably the most interesting development that the folks at Twingly are working on is what they are calling a "social meme-tracker" by the name of Twingly Channels. The web-based application is still in a limited-release beta version (you'll need to find an invite code, try: ALTSEARCHENGINES), and while there are more than a few kinks that need to be ironed-out, the idea seems pretty revolutionary. Essentially, Twingly Channels provides a social network based around memes rather than individuals (as is done on Facebook and Twitter). Anyone can create a meme and name it whatever they want (from as general as "Entertainment" to as specific as a particular person, place, or thing). Other users can "subscribe" to this meme the way that one "follows" a person on Twitter. These users can add things that are relevant to that meme - blog posts, tweets, articles, etc. and all of the subscribes can comment on these posts and specify whether they "like" this particular post. The more subscribers like the post, the higher the post is on the stream.

The only major issue I see at the moment is that Twingly has yet to determine an adequate way of organizing these memes (or "channels") apart from a simple list which orders the channels according to their popularity. The implications for a tool like this are huge. One can easily envision scientists creating "channels" which are relevant to a particular topic and using this tool to quickly share data, and comment and rank it. Unfortunately, at the moment, most of the channels relate to things in Sweden (in fact most of the comments are in Swedish) so it's too early to see whether something like this is catching-on globally, but there is a lot that I like about meme-centered social networking as opposed to person-centered social networking.

Possible Causes of Inadequate Metadata in the Institutional Repository

In her article Innkeeper at the Roach Motel, Dorothea Salo writes about the problems of institutional repositories. While she discusses many different issues, I am going to focus on the causes she lists for insufficient metadata in institutional repositories.
First in her list of causes is the issue of trust. Many people are hesitant to trust librarians with their data because they believe their lack of subject area expertise will result in problems with metadata management. Salo does not discuss wether or not this belief has merit, only that it inhibits researchers in specific disciplines (such as the hard sciences) from trusting their work to librarians. Instead they trust their work with science experts who frequently are not working to solve long-term preservation issues.
Additionally, the different requirements of publishers make placing data into a repository difficult. One such requirement is that the meta-data have a set phrase which includes a link. The problem is that some of the popular institutional repository software, for example DSpace, does not have that particular functionality, further contributing to some author's hesitancy to use institutional repositories.
The information professional put in charge of an institutional repository is also at a disadvantage if they cannot program. They are left with the basic version of their repository without any patches or plugins that are not already a part of the system. Frequently the task of applying dublin core metadata to a particular work does not even fall to the information professional, but instead it falls to the author of a particular work. They information professional is left trusting the authors themselves to invest the time needed to ensure proper metadata entry, which rarely happens. The result is inconsistent use of metadata throughout the repository, greatly decreasing the usefulness of the works that are in it.
Repository systems also contribute to the problem of having good metadata in a repository because they are frequently difficult to use. They often overestimate the ability of the user (who is not always an information professional) and expects them to know what fields are required and even what some more obtuse fields mean. Salo does mention that EPrints does a better job of handling metadata than other repository systems because it displays metadata fields based on what type of content is being deposited, making the process more user-friendly.
The final cause of insufficient metadata in an institutional repository mentioned is that repositories are still not seen as a priority to many libraries. Without proper funding and sufficient staffing to maintain and oversee the repository and the metadata going into it, information professionals don't have a chance of ensuring the repository is useful. The library ensures the failure of a repository when they don't allot adequate time and money to its maintenance.
While I believe all of Salo's points are valid and the way metadata is handled should be reconsidered, I still believe the repositories currently face a larger problem, namely that authors do not want to put their work into the repository at all. Still, if we are going to put data into the repository, we should be able to find it again later and that necessitates that we take a good look at the way we are implementing metadata.

Ex(ml)tra, extra, read all about it!


This week I found a brief article detailing the metadata problems and solutions encountered by the Library of Congress when it participated in the National Digital Newspaper Program, a twenty-year initiative to create a national repository of historical newspapers in digital form. The end goal of the project was to have a complete and searchable online bibliography of newspapers published in the U.S. from 1690 onwards. In addition to this, select local newspapers of historical significance would be converted to digital form for instantaneous public access.

Several problems arose in creating a standard metadata for the project. First of all, print newspapers themselves do not fit well into any of the traditional cataloging standards. This is especially true when one tries to accurately and comprehensively describe newspapers from all fifty states and across drastically different historical eras. (17th-century spelling and abbreviation are particularly dicey for modern readers.) Second, most of the historical newspapers being covered by the project no longer exist in print form but were only saved as microfilm, introducing another form in the provenance. Finally, the NDNP wanted a system that was going to allow complete interoperability between the federal repository and individual state digital repositories as well as among the state repositories.

Ultimately, the NDNP elected to use the METS standard, entered in XML. Four separate METS elements were used: title document, issue document, page object, and reel document. The first three refer to the original physical form of the paper. XML is flexible enough to be able to incorporate different types of issue numbers while remaining interoperable, which was key for the program. Murray views the page object as the most important as he asserts that the page is the basic information-containing block of a newspaper. The reel document only applies to those papers digitized from microfilm but represents important administrative data.

I found this article interesting mainly for three reasons. First, it involves a system to provide more access to historical documents, including colonial era documents, which is a goal ever near to my heart. Second, I work in the serials unit at the Benson Collection and have much first hand experience (and frustration) at the incredible amount of variance in newspaper conventions. Third, I liked the fact that the NDNP came to a relatively straight-forward and simple solution, while maintaining interoperability. (Though, of course, it helps that there was central oversight in this project...) All in all, I thought it was a good example of the benefits conferred by standards as well as not making the solution harder than it has to be.