Friday, September 18, 2009

In the news: Google & reCAPTCHA; researcher scooped after publishing data

I'm blogging early this week, so by the time class rolls around this will probably be old news, but here goes anyway. First, I forgot to mention in class that yesterday that Google just bought reCAPTCHA, which is the company that makes those little boxes with the funny letters that prove you're human when you're filling out an online form. We use the technology a lot on the UT Libraries web site. So what does Google want with reCAPTCHA? Turns out those funny characters come from scanned archival newspapers and old books. When someone uses the captcha form, they teach computers how to read that text. Distributed OCR; pretty neat. Google plans to use this for Google Books. It won't solve the awful metadata problem, but I thought it was a good example of an innovative solution to a common problem.

In other news, Science just reported that a researcher who posted her data online was scooped before she was able to publish a paper from that data. Laura Bierut deposited her data related to genetic studies of addiction in dbGaP, NIH's database of genotypes and phenotypes. The NIH policy says that she should have had an embargo period of 9 to 12 months during which no one could use the data for publication. Her embargo period ends on September 23rd. But back in March, Heping Zhang, a Yale researcher, submitted a paper based on the data to the Proceedings of the National Academy of Sciences (PNAS). Zhang's paper was published on August 31st, and Bierut quickly responded by sending emails to Yale, PNAS, NIH, and colleagues. She was even polite, considering the circumstances: "'[T]his was likely an unintentional act, [but] this incident remains very concerning,' she wrote, adding that it 'sends a very chilling message to investigators.'" Yale took down a press release about the study and NIH froze Zhang's access to dbGaP. On September 9th, they retracted the paper. Currently, they are investigating the situation, and Zhang won't comment until that investigation is completed.

Bierut is quoted in the Science article as saying, “I think NIH and PNAS moved very quickly to resolve the issues" and "The good news is I think the system worked." But, the paper is still on the online version of PNAS because of the "very awkward consequences - librarians got confused about papers being cited that no long exist," (which is a pretty ungracious way for the PNAS editor-in-cheif to put it - "confused"? c'mon now). Bierut has been invited to submit her paper to PNAS, but one has to wonder if it will receive the same response as if her paper had been first. As one researcher quoted in the article put it, the situation leaves "the gate open for predators with no investment in the data to do quick-and-dirty analyses that pick the eyes out of the data without looking at any of the subtleties." It's worrisome to me that horror stories like this will discourage researchers from publishing their data.

The other day in class I was speculating that perhaps some of the problems with sharing data might be alleviated if we rewarded scholars for publishing data, not just papers. I think it might be less scary for researchers to put their data out there if they knew that, if all else fails, they'd at least get some credit for the data. But I'm starting to understand that publishing data probably just wouldn't have the same impact as publishing the analysis of that data, especially if it's a breakthrough. Data is just data until someone gives it meaning. And it would probably take a while for the community to realize the impact of a particular dataset (like if the number of publications using the data were tracked). The amount of credit someone should receive for a particular dataset would take a while to assess. In the end, I think the ultimate goal of open science is such a worthwhile one that we have to keep trying, but it feels like it might be a long road.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.