This week I came across the Notice on Development of Data Sharing Policy for Sequence and Related Genomic Data from the NIH. This notice says that the NIH is considering new additions to their policy on genome research. In class we've discussed what the motivations of the genome investigators might have been for being open with their data and so far there doesn't seem to be a single answer. However, this ongoing revision of policy from the NIH indicates that this area of science has unique requirements for data sharing because of the resources that are needed to acquire genome data. The NIH mentions that these resource needs "necessarily limits the number of projects that can be supported for any disease". In addition to these factors, the NIH says that these data sets have "added scientific value when combined with other large data sets".
The policy that the NIH is working on right now is basically a continuation of the Policy for Sharing of Data Obtained in NIH Supported or Conducted Genome-Wide Association Studies (GWAS) which went into effect January 2008. Basically, when researchers have genome data they have to put it in the central genome-wide associated studies (GWAS) database which makes all genome data accessible for combining with other data sets as early as possible. Although other people can see it, the investigator who deposits data has exclusive rights to publication based on that data for some period of time, the NIH encourages investigators to keep this period of time short and they limit the exclusivity period to 12 months.
Is the NIH the only organization that has this kind of exclusivity period? I remember a discussion in class about an investigator who deposited their data into a repository and then somebody else published a paper during the exclusivity period. I also recall that the organization with the policy and database (again, I'm pretty sure it was the NIH) didn't seem very alarmed by the situation.
One reason why the NIH is updating their policy might be because of the confidential nature of the data and the need for participant privacy. In the original policy the NIH predicted that in a few years technology would "make the identification of specific individuals from raw genotype-phenotype data feasible and increasingly straightforward", which is problematic for data in a federal repository because that data is accessible through FOIA. At that time they decided that they would have to deny FOIA requests for unredacted data sets. About eight months after that policy went into effect the new developments on how to identify people based on genome data was made public (NIH reigns in genome access) and the NIH has to further restrict access to the data - NIH Modifications to Genome-Wide Association Studies (GWAS) Data Access (sorry, that's a PDF).
This article in PLoS Genetics explores the issue of confidentiality with human genome research in more detail: Public access to genome-wide data: Five views on balancing research with privacy and protection. One person basically says we can't trust anybody to keep data confidential, even though any researcher accessing GWAS would have to agree to the policy agreeing that they would "not attempt to identify individual participants from whom data within a dataset were obtained".
Wednesday, November 11, 2009
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.