Wednesday, September 30, 2009
A Real Solution to High Textbook Prices
PLoS and the new Metrics
In the tradition model of academic publishing the key metric to a “successful” article was the number of times it was cited in the literature. It should be noted here what a very poor metric this is to judge the influence of an article, the information that is contained in the publication could be well, suspect but if it provokes conversation or is used as a counter example then it will still be cited and still be seen as influential when counting citations. While perhaps there is no system that is perfect it is obvious that this traditional model could use a bit of up-dating.
The folks over at PloS have implemented a new set of article level metrics that go beyond the tradition of counting only citations and have begun counting other ways that articles are accessed and used.
The new metrics count the number of times the article is access, book marks, comments, ratings blog links as well as citation. The ability to collect all of this information is one of the exiting changes that online journals can offer. This collection of way that the journal is used can tell more about the quality of the work than the one dimensional view provided by examining citations alone.
Of course there are some know issues already when it comes to using these new metrics. For the most part PloS is up front about the issues inherent in these metrics and lists the steps that they are taking to address them. We should remember that each of these metrics has strengths and weakness but taken as a whole they reinforce each other and potentially have the power to change the way journals are used.
Now if we can just change the tenure paradigm in tandem.
Government Data
Cyberinfrastructure = Hardware + Software + Bandwidth + People
This week I was really not sure what I could write about, but then I came across Michael Roy's article in the Academic Commons that discussed the idea of cyberinfrastructure and building a computer research center at Northeastern University. I'm still working out my own ideas of what a cyberinfrastructure exactly is and this article featured a graspable array of issues and ideas that make up the necessity of establishing an infrastructure for research.
Roy begins his article by stating that computers were once thought to be for research, but then they became known to be teaching aids, but now the pendulum is swinging back to the idea that computers are used for research. He divides his article up into three parts that each address particular areas of establishing a IT research center: the practical issues, such as equipment, servers, staffing, the politics and the process and then the academic research process.
The first section elaborates on how Northeastern University decided to establish a Computer Research center. A committee was created and they interviewed the faculty to see what kind of research projects were currently being done and also asked about future projects to gain a picture of research needs. They decided at the end to build a cluster with features that focused on job management, size and interactvity. Then they had to set up a proposal to find the vendors that would be able to provide Northeastern with the technological equipment needed.
Needless to say, the IT Computer Research Center unleashed a series of critiques and skepticism amongst the faculty and within the academic world (mainly from the liberal arts world, predictably); however Leslie Hitch, Director of Academic Technology at Northeastern University, pointed out that Northeastern had a made an organizational goal to increase the quanity and quality of faculty research and the establishment of the Computing Research Center fell within the parameters of the university's mission. Then Roy began to discuss some of the challenges that they've had: staffing--how to staff the computing center, how much control to allow for, what defines productivity and so on.
Essentially, this article was useful for me to read because it provided some kind of framework of an individual institution's efforts to bank on digital research and clearly outlined some of the challenges we continue to have with defining the parameters of computing research and exactly what to do with it.
Twitter, Curation, and Horizontality
We are already seeing some web-sites and applications which are doing exactly that. Breaking Tweets lists links to people are are tweeting about global "news worthy" events. Alltop lets you know about the celebrities and bloggers whom you "should" be following on Twitter. Applications like TweetDeck allow you to create your own filters so that you can organize your tweets and Facebook updates into categories like "news" and "friends" and "ridiculous" (there's a lot of that). From this, one can extrapolate ways in which other Twitter users could use websites to access tweets about various subjects: botanists could more easily access tweets about plant research, and journalists could access tweets about global events. Right now, the only way to filter tweets is by keyword searching for hash tags, but that can be a time-consuming and frustrating process.
While I can definitely appreciate the increased accessibility that such websites afford to Twitter users who have a specific interest, I also think there is something to be said for the horizontality of the current state of Twitter. All Twitter users have only 140 characters in which to express themselves, and no one is given more or less space based on their social status or another other factor. There's something egalitarian about the current state of Twitter and I think we're losing something important if we start imposing a hierarchical system of selection.
While I stated in last week's blog post that I like the idea of individual Twitter users curating web-content and posting it on Twitter, the idea of placing filters on Twitter itself seems to miss the point of Twitter altogether. In an interview for the Los Angeles Times, the co-creator of Twitter, Jack Dorsey, talks about the original concept of Twitter (this is a lengthy quote, but I feel like it's worth reproducing in its entirety): "The concept is so simple and so open-ended that people can make of it whatever they wish. They seek value and they add value. I've always said that Twitter is whatever you make of it. Because the first complaint we hear from everyone is: Why would I want to join this stupid useless thing and know what my brother's eating for lunch? But that really misses the point because Twitter is fundamentally recipient-controlled -- you choose to listen and you choose to leave. But you also choose what to put down and what to share. So if you decide to hook your plants up to Twitter and have it report when it needs to be watered, then that's a valid usage, or if you just decide to report what you're eating for lunch, that's a valid usage too." I think this gets at what I think is the uniqueness of Twitter - it's a giant behemoth of little snippets of information which the user must wade through to determine what he or she does or does not find interesting or useful. I don't know that we've really seen anything like this before, and I'm worried that if we start using websites that filter tweets for us, Twitter is going to become just another news website.
Twitter recently announced that it is, in fact, saving everyone's tweets. I like the idea of having one long, historical feed of 140 character blurbs with no form of organization apart from a date stamp. I think it will afford a unique, unfiltered way of looking at history.
Digitizing Historical Documents at the Mint
The article, posted in 2001, briefly describes a digitization effort at the Mint that involved 225k pages, 5k photographs, 2.6k microfiche reels, and over 400 monographs being scanned and stored on servers as tif files and pdf files for display. Two separate databases were used to keep track of the resulting digital output, one for accessioning and item records, and a second for increased cataloging and description. The databases each possessed both internal and external interfaces. Finally, finding aids were put on the web via EAD.
What struck me about this article is that it lists specific difficulties encountered in digitizing historical records, many of which still plague us nearly a decade later. For example, Rothfeld bemoans the lack of funding for large projects such as at the Mint. While electronic storage space has become remarkably cheaper, scanning unique items by hand remains a labor intensive, costly process. Furthermore, software costs and limitations often determine what precisely a collection can do. Rothfeld also found it problematic to be attempting to instill "tech people" such as the database designers with the full idea of how historical archives get used. Conversely, her archivists did not program. Finally, Rothfeld noted that the big challenge that remained for her project was achieving interoperability with other, outside valuable sources and incorporating outside materials into the Mint's online collection.
The answer to those problems, of course, is to further the cyber-infrastructure, not just in terms of the goods (i.e. digital replications of historical or other social scientific data) but also in terms of channels of information. Of particular importance, I think, is the education aspect of building cyber-infrastructure. Many of the problems Rothfeld encountered would have been mitigated if the programming and historical/archival planning skills had been vested in the same individuals. With a higher level of 21st century information literacy, cross-trained individuals would also be able to better take advantage of some of the low-cost open source software that can be customized to be at least somewhat field specific.
In conclusion, while it is encouraging that much historical (and other social scientific) data is finding its way into a digital format, further effort does need to be spent on developing cyber-infrastructure within the field.
UK asks data developers for feedback
The UK may change that. The UK's Cabinet Office website government website has posted an open call of sorts to data developers. While the details of the project are vague, the site states that there are "over 1000 existing data sets, from 7 departments (brought together in re-useable form for the first time) and community resources".
They are asking developers to offer feedback on their current site design and ideas about how to transform their data into a usable, single point of access for government-held public data. They are interested in any features that should be added or changes that should be made to the site or even any data that should be there but isn't. To participate, developers need only to join the Google Group.
The post states: "From today we are inviting developers to show government how to get the future public data site right - how to find and use public sector information."
I am having trouble envisioning how this kind of crowdsourcing could be bad, but I have no trouble imagining the good that could come from it. The added input may help anticipate broader uses of the data in addition to providing design changes that will make the public data usable. I am excited to see this kind of openness in the site creation process and look forward to seeing whether or not it proves to be successful. All the digitally stored data in the world isn’t going to do us any good if it isn’t usable.
How Last.fm inspired a scientific breakthrough
Mendeley has launched a research paper organizing service with software that is similar to the existing PDF organizing tools (Zotero, CiteULike, Connotea). Not only does Mendeley work really well to help organize citations, it also keeps statistics on what people are reading and from what field of study those readers come from. This way, users can look at the statistics for what is most popular in their field.
The article from the Guardian is particularly enthusiastic about Mendeley being the revolutionizing thing that the internet/researchers/science have been waiting for. They say Mendeley "shows the second phase of the dotcom boom is throwing up great, practical ideas". This is from a technology blog. After reading a few statements like this I downloaded Mendeley and tried it out. It is pretty interesting to log in and look at the statistics they provide, but I'm not sure if it will go much further than being interesting to look at. The have a lot of people joining very quickly, so it might be possible that Mendeley will change the face of science (this is what Dr. Werner Vogels of Amazon predicts).
In Mendeley authors can upload PDFs of their work and include it in the Mendeley database. Mendeley has a very interesting take on the legality of this in their Mendeley FAQ. They say that most journals are usually okay with authors "self-archiving", which is not defined. Mendeley's simple answer to self-archiving is "Simply self-archive your preprint as well as your postprint, and wait to see whether the publisher ever requests removal." After downloading and using Mendeley a little bit I am still not clear on how much access is provided to the articles that authors upload.
There's one thing that the author of this column wrote that I do not understand:
"Mendeley says that instead of waiting for papers to be published after a lengthy procedure of acquiring citations, they could move to a regime of "real-time" citations, thereby greatly reducing the time taken for research to be applied in the real world and actually boost economic growth."
What is this lengthy procedure of acquiring citations? I've asked a couple of people about this and nobody knows what he's talking about. Is he somehow confusing this with the lengthy procedure of peer review? He doesn't actually mention peer review in the blog at all. Maybe the technology column of the Guardian isn't the most reliable source on how scholarly publishing works.
Blogs (hangingtogether.org, Open Access News) have mentioned that the metadata in Mendeley is horrible. This looks like another case of why won't fill in organization start fill in the practice we like. The metadata might not be that great, but providing stellar metadata is probably outside the scope of Mendeley's goals. One component of Mendely is to crowdsource articles, which points the argument back to folksonomies vs. cataloging. While it would be great to have a service that uses both tagging and quality metadata, I'm not sure what that kind of service would look like or who would actually do the work to create it.
A data model for interoperability between scholarly repositories
- a shared data model for the digital objects
- a surrogate format that serializes the digital object according to the digital model
- support for a universal way of handling repository materials as surrogates: obtain, harvest and put.
Tuesday, September 29, 2009
Mendeley - not Manderley
One entry, found on Savage Minds, an anthropology blog, very logically lays out the impediments of several different bibliographic softwares, coming to the conclusion that Mendeley works best for the author, probably a student or scholar, since it is "both a web application and stand alone desktop software". This means the tool combines the best aspects of Zotero, Sente, Evernote, and others. The blogger realizes Mendeley is still evolving but has high hopes.
Another post on Hanging Together, a blog written by OCLC/RLG staff, quotes an article from The Guardian and explains Mendeley's sharing techniques, comparing it to last.fm's profile building and recommendation system. The blog rightly assumes librarians will wonder how a "web application based on an entertainment model should have proved so much more attractive than the painstakingly built repositories we have been holding under the noses of our academic authors over the last several years?". This is answered by the fact that Mendeley runs statistics about who is looking at what paper, offers cross referencing and recommendation systems, and will integrate with programs like Zotero. "Mendeley has grabbed the attention of users because it understands what they like. They like simplicity. They like instantaneous results". Mendeley apparently very lightly touches upon copyright and metadata (using OpenSource ethics and 'change it if it's wrong' metadata concept), which means "the requirements libraries often put up front are almost dismissed as non-issues" by Mendeley.
Its real time editing abilities are what grabbed my attention. I'm not sure if users can actually comment on others' papers, but by posting pre-published papers, feedback on use will be available, at least. This is just another tool that is changing the academic scholarship traditions. I'm interested to see where this bibliographic tool goes in the near future, and if it delivers on the so-far limited yet excited hype. Maybe someone will upload Rebecca, and we can all go to Manderley. Easily.
PLoS Currents
Once again this week I've found myself drawn to the topic of data sharing. In this case, I've discovered PLoS Currents, a beta version of a service that is described as "a moderated collection for rapid and open sharing of useful new scientific data, analyses, and ideas." This particular PLoS Currents forum is dedicated to the H1N1 influenza virus (aka swine flu). The purpose of PLoS currents is to allow for expeditious communication between experts on pertinent issues to enable rapid advancements in these areas. A forum such as this is a novel way to approach time-sensitive issues such as the swine flu epidemic by presenting a setting in which scientists can share their data, research, and theories without the high level of peer review required by scholarly journals.Shifting roles, shifting space
September 24, 2009
In light of our discussion last week about similarities and differences between traditional physical libraries and digital libraries, I found “Libraries of the Future” from September 23, 2009 to be an interesting read. The article summarizes statements made by Daniel Greenstein, vice provost for academic planning and programs at the University of California System, to a group of academic librarians meeting to discuss sustainable scholarship. Steve Kolowich, the author of the article, states that Greenstein’s vision is one in which, “the university library of the future will be sparsely staffed, highly decentralized, and have a physical plant consisting of little more than special collections and study areas.” In effect, the digital repository would become the defining entity for library, while storage and maintenance of physical materials would be shared between institutions, and if possible, outsourced.
The primary motivation for this line of thinking seems to be economic, with budget shifts requiring university libraries to restructure and reallocate resources. Greenstein points to a decrease in individual libraries staff, collections, and services, but an increase in collaboration with other repositories. The goal would be for libraries to become drastically different over the next decade, which some librarians present at the talk, as well as those who posted comments, thought was an unreasonable amount of time for such a shift in the library world.
Most comments made seemed primarily concerned with the changing role of the librarian who, perhaps in the eyes of Greenstein, has just lost their job. I think the questions that arise are about what the roles of librarians might be in a system that outsources most of the services they performed. Might librarians move to other positions in which the take on the roles within the companies performing the tasks needed? Who remains a part of the small staff at the university library? Is it really cheaper to outsource cataloging, conservation, etc?
I am not sure if the answer to the recent economic problems is to dismantle the system and depend on outsourcing to provide services. I think the changes that will happen to the university library system will involve many shifts in the various roles and potential employment opportunities, of librarians. Still, I find the physical shift suggested unsettling. Where do all the books go? How does the space shift?
Disappearing Digital Archives
“Digital Archives That Disappear” (April 22, 2009) is a reaction article to the sudden disappearance of the online newspaper archive “Paper of Record.” Began in 1999, Paper of Record is a digital collection of newspapers (consisting of 21 million images), with a large portion of the collection consisting of Mexican newspapers.
Paper of Record was secretly bought from its former owner in 2006, however, this purchase was not publicly revealed until 2008, when the archive suddenly disappeared. Amidst an outcry from historians who relied on the materials available through paperofrecord.com, Google agreed to let the site resume services under the previous owner, however, its future is still uncertain. For example, there is concern over the possibility of Google charging prohibitive institutional subscription fees for access.
This case raises many interesting questions about the future of digital archives and scholarly research. First, in the case of paperofrecord.com, the Mexican government as well as Mexican universities paid for many of the Mexican newspapers to be digitized. This has upset many members of Mexico’s historian community. Finally, this raises questions about the continued accessibility of online archival collections.
But no really, be practical, HOW will we do all that?
Here's one example from this week: "Data at both the individual and firm level, as well as national and international levels, must be integrated. Complex qualitative and quantitative data from a variety of modes must be combined." (p. 20 of http://ucdata.berkeley.edu/pubs/CyberInfrastructure_FINAL.pdf)
While I am aware that technology changes quickly, I thought it might help to find some examples of how current digital library projects are structured. I began to hunt for articles about social science databases. Megan's comment about persistent identifiers was timely; 404 not found and I got to know each other quite well.

Finally, I found an article about the Exploratorium, an online science museum. http://saturn.exploratorium.edu/partner/nsdl/pubs.html
The article, From Playful Exhibits to LOM: Lessons from Building an Exploratorium Digital Library, can be found here. The authors describe the metadata standards adopted to create the interactive Exploratorium.
Specifically, they designed their metadata for use with "Learning Objects." They knew in advance what their mission was, to provide educational material, so they knew the metadata should be structured to support that. The standard they chose (IEEE LOM) also maps to Dublin Core, so I think it is extensible in the way that Berman and Brady were suggesting from my earlier quote.
They assessed what they had to put in the collection before starting, and then used an assets management solution called Canto, which supports scalability and APIs.
Should they have tried to combine more projects and possibly be more extensible? Do they have enough tools, enough scope, or enough added information to be classified as a curation project? I'm not sure. I know I found their website much easier to understand than some others. An image of one view of their structure and metadata is below.

The question of "how?" deserves more consideration, but I believe this site is a good example of knowing what you want in the collection, setting goals and standards, and adapting when necessary. The authors acknowledge that their use of URLs as stable locations may need to change, but currently they do not have a better option.
(I did not receive a single 404 error while exploring their site.)
Who's Afraid of Google Scholar?
The latest critique of Google's metadata comes from Peter Jasco writing on Google Scholar for Library Journal. In his analysis, Jasco illustrates how ill-trained parsers and web crawlers are responsible for creating a variety of metadata errors (particularly as pertains to authorship) which in turn produce incorrect publication and citation counts.
Jasco begins by noting that whereas Google Books blamed the embarrassing metadata errors exposed by Geoffrey Nunberg on wonky data from libraries and publishers, they do not have this excuse for errors in Google Scholar since Google chose not to use the metadata offered them by publishers and instead rely on its crawlers and parsers. The result, Jasco asserts, is that Google's algorithms replace real authors with ersatz ones like "P Login" (here Google has misconstrued a login prompt on a publisher's page, "Please Login," as an article's author). Alternately, Google Scholar also inflates publication and citation rates by counting citations as records in author searches and by generating multiple master records for the same article. Jasco notes that a search for his own name as author (e.g. author:jasco) yielded 578 records, but 403 (or 79%) of these records were for articles that cite his papers, not articles he authored. Such problems are compounded when Google Scholar's figures are naively adopted by tools such as the Google Scholar Citation Count gadget or Publish or Perish (PoP) software.
I'm torn as to how to respond to Jasco's article. On the one hand, it is clear that Google Scholar's metadata is error-riddled. One can't help but wince at some of Jasco's examples: the Google parser identified as author names things such as headers, section titles, and publication information, including Methods (42,700 author records) Contents (25,200), Limited (234,000) and Ltd (452,000). On the other hand, any university or department that's willing to make tenure decisions based on a cursory search of Google Scholar has such deep-rooted problems that inaccuracies on the part of Google Scholar can only be the least of these. It is to be expected that an unwitting high school or college student might be led astray by mis-attributed authorship or incorrect publication dates in Google Books, but it's less understandable for professionals in the field to outsource their committee work to Google Scholar or (even more perplexing) for an individual to assume that Google Scholar knows more about how many articles s/he has published than s/he does. It is part of being a professional in the field to keep an up-to-date CV and be cognizant of the publications one has authored or co-authored. Citation indexes are admittedly a trickier issue since such information is not as easily compiled by individuals, however shouldn't we expect (demand?) that academics and administrators would vet any results Google Scholar yields?
Jasco is not optimistic about Google Scholar's future. He sees their refusal of publisher/indexer metadata in favor of their own under-trained crawlers and parsers as a "lethal mix of ignorance and arrogance." Though Google Scholar has corrected some widely publicized errors, Jasco asserts "[t]he parsers [themselves] have not improved much in the past five years despite much criticism."
In addition to the actively "malevolent" threats to cyberinfrasctructure posed by hackers, phishers, and other cybertransgressors that are discussed in the Workshop on Cyberinfrastructure for the Social and Behavioral Sciences Final Report, cyberinfrastructure is at risk from the perhaps even more damaging effects of incompetence, laziness, and insouciance--factors which undermine our confidence in cyberinfrastructures and the collections they support. There is plenty of blame to go around. Google needs to be more concerned about inaccuracies in its products. Perhaps the responsible thing would be for Google Scholar either to preface certain functionalities, such as citation counts, with clear "user beware" warnings or to suspend them until it can ensure that its records are more accurate (it would still yield record results for the user to vet and tabulate). However, it is also incumbent upon users to educate themselves better regarding the tools they are adopting for important tasks. After all, Google Scholar announces itself as "Beta" in its logo. Hiring, promotion, and funding decisions should not be based on numbers drawn from tools that are still very much works-in-progress simply because it's easier to get your numbers from Google Scholar than to crunch them yourselves.
Monday, September 28, 2009
NSF awards $7 Million to Texas Advanced Computing Center
I wasn’t sure what to blog about this week, so on a whim I went to Google news and searched for the words “science” and “cyberinfrastructure”; one of the articles I found was actually pretty exciting. The news broke today – according to the University of Texas public affairs news site, the National Science Foundation awarded the Texas Advanced Computing Center here at UT a 7 MILLION dollar grant to fund a three-year project that aims to "provide a new computing resource and the largest, most comprehensive suite of visualization and data analysis (VDA) services to the open science community."
According to the article, this new computing resource, unsurprisingly named "Longhorn", will help science communities both nationally and internationally. Longhorn will enable scientists to "interactively visualize and analyze datasets of near petabyte scale." I didn't know this before now, but a petabyte is apparently equal to one quadrillion bytes or 1,000 terabytes... which is quite a large number.
The article goes on to give many more technical system capabilities and components that Longhorn will have, which all sound really impressive, but I have no idea what any of it really means because I am not a deeply technical person. This project is being touted as “unprecedented” though, so the technology must be pretty spectacular. Check out the article for the specifics. To sum up the important part that I got out of this article: basically, this three-year project, which is slated to “enter full production” on January 4, 2010, will “provide the framework for leading researchers to collaborate and analyze their terascale datasets while still offering low barriers to entry and usability.” The 7 million dollar grant aims to "stimulate and support" new research in the area of visualization and data analysis. The grant will expire July 2012.
John Mullen, the vice president and general manager of Education, State and Local Government at Dell, praises this “long-term strategic investment” and says that Longhorn’s new advancements in technologies (particularly in CPU, GPU and networking technologies) will allow scientists to take on new, complex challenges. What TACC is developing could also have an impact on other fields. Kelly Gaither, who is the principal investigator and director of data and information analysis at TACC, says that Longhorn’s developments in interactive visualization and data analysis can help solve problems not only in science but also in engineering, medicine, and in the areas of national security and safety. The humanities aren't mentioned in this article, but I imagine it and other similar fields could probably benefit as well.
Basically the entire article mentions points we keep discussing in class: science and technology are both evolving rapidly; the internet and high performance computing are creating more and more data than people can handle at once; new technologies must be created to handle such large datasets and help link together those who are working on complex projects; and, finally, the fact that strategic investing is important in building parts of a cyberinfrastructure.
This seems like an exciting opportunity for the university, and I’m interested in finding out more about the project. As I mentioned, this was the first article that I found on the subject, and it came out only this morning. UT's much quoted motto is "what starts here changes the world"... we will have to see during the next three years if Longhorn will live up to this lofty goal.
Wednesday, September 23, 2009
Visible Archive Series Browser
I am very excited about the prospects these kinds of visualization technologies can have on digital scholarship, particularly in the areas of the arts and history. The browser is interactive, with visualized archival series (represented by squares) arranged both chronologically and in terms of relations to other series. When one scrolls over a series a bubble appears with general scope and content information. The size of the series is illustrated by how big the square is. An inner square represents the volume and the outer square represents the physical shelf space that the series occupies. By selecting and zooming in on a series - lines demonstrating relationships with other series appear (different kind of link relationships are categorized by different colors).
One can also browse and research by agencies outside of series to gain broader context. All this is handled through interaction with the visualized data.
Three hours ago Mitchell launched a video demo describing the A1 Explorer, a similarly interactive visualization of the A1 series in the National Archives of Australia. In this we are looking at what appears to be a kind of tag cloud - as a word-frequency visualization of the contents based upon the item titles. When you hover over one of the words, lines linking it to related words appear. The thickness of the connecting lines is also an additional layer of information. Selecting a text item displays a list of items belonging to that category.
Another aspect of this example is the histogram which displays visually the number of items appearing in the collection per year. This demonstrate the distribution over time but also works as another port of entry into exploring the archive. The text cloud and histogram are also linked in that hovering over a term will display how many of such items appear in the collection per year.
An additionally cool functionality of Mitchell's visualization is that if there is a term that is dominating the text cloud - if you select and click on it you can remove it from your view and the cloud will adjust accordingly, allowing you to see and explore the remaining items in the collection. If you select an item from the list it will load a digital image of that item from the archive.
I couldn't help but think about how cool this application would be for navigating my links in delicious.com, but really - the applicability of this for all manner of archives is immense. I hope to see this technology catch on!
Not To Share
It should probably come as no surprise that out of the ten researchers contacted only one was willing or able to actually produce the dataset, and then only after talking with the requester about what would be done with the data.
This raises so many of the key issues concerning scientific data sharing that it could almost act as a summary of why instituting a data sharing methodology has been so problematic. Some researchers were “too Busy” to fulfill the data requests, even though “we pre-specified that any request must not create undue work for the investigators”. Others claimed that they were “forbidden” to release the dataset. Some simply did not respond to the request at all.
The reasons given for not sharing data are easy to sympathize with, gathering a data set, annotating it so that others will find it useful takes an investment in time. It is difficult to persuade anyone to take a few hours out of their busy day to give you something they feel attached to, when they have no idea of your intentions or how important the data is or even who you are. It is illuminating that the one researcher who complied with the request did so after an interview process; almost like a vetting procedure to see if the requesters intentions were true.
It must be pointed out that however reasonable the denials are they come from a group of people whom have published in a journal which explicitly states the requirement that datasets need to be accessible. This active participation in the journal should nullify any claims of being “too busy”. It does not, and this demonstrates one of the key faults with data sharing as it is: journals lack the teeth and methodology to ensure that datasets be made available.
The main reason for this is that journals are not and should not be concerned with enforcement, journals should not be the police ensuring that everyone play nice with data. The best a journal could do is to deny further publication. The real and effective enforcement should come from the community in general. There needs to be a ground up call for data sharing. This sort of change in the way science works will only take place slowly and if there is a proper system in place to reward data sharing.
A Library Where No One Came
Issue 461 of Nature magazine is devoted to issues surrounding open access to data, with a feature on what happens when researchers don’t take advantage of data repositories and keep their data to themselves. This article is a good reminder that despite all the hard work that information professionals often put into building tools, it is all for naught if the tools aren’t used. It is important here to distinguish sharing of data from sharing of publications in online open access. While the latter is becoming much more common to the point of being the norm in some fields, the former lags behind.
In this article Bryn Nelson explains that different groups of researchers have differing cultures that affect their views on sharing data. Some disciplines may have rigidly structured data, such as genome sequences, while others may have data that doesn’t fit any predefined schema, such as climate data. In addition, researchers are often concerned about misuse of their data or lack of attribution and losing the original rights they have to the data. Some data can’t be understood without looking at a whole other group of related data: an example given in the article is a group who is trying to track down weak gravitational fields and must gather a huge amount of background data such as ocean currents and atmospheric conditions to use as filters on their data. Releasing such primary data to the public runs the risk of people taking it out of context and thinking they have found spurious results.
Nelson suggests that three authorities in these fields exercise their control to help corral the data: publishers grant funders, and scientific societies. Grant agencies can make it a requirement that the projects they fund will have their data collected in a repository, and publishers can ask that authors share their data, although they must be careful not to scare away authors who would turn elsewhere. In addition, Nelson says that large government efforts are required to build the kinds of data standards that will work across scientific communities: it’s not enough to think that ad hoc solutions will work in the long run.
On the one hand, the problems of the researchers are all ones we have seen before: they worry about losing the rights to their work, misattribution, and people drawing unwarranted conclusions from the data. While this is nothing new and can occur with any work put out into the public, there is an expectation of objectivity to data that makes releasing it especially dangerous. The traditional gatekeepers of journals should be able to hold at bay the worst examples of pseudoscience that draw on public data, but perhaps it would be beneficial to build in dependencies to data streams that link data that can’t be separated. Data could be provided as a package, not independent streams.
Listening In
What intrigued me about this article is that it describes both practices of popular curation and shared curation of digital objects. In terms of popular curation, I-tunes technology allows users to share their playlists with co-workers. However, you can choose which specific playlists you want to share, and you retain control over the naming of the list as well as the metadata tags of the individual songs (i-tunes retrieves song info via the net but the user has final authority to change the info on his or her individual machine.) What is really interesting about this, is that the study found that people used i-tunes' sharing features as a way of projecting identity, and thus the choice of which playlists to share (exhibit?) became tantamount. For instance one participant worried about some of the songs he purchased for his wife being spotted by co-workers:
I mean if people are looking at my playlist to get a picture of
the kind of music I like and don’t like, you know. Or to get a
little insight into what I’m about, it’d be kind of inaccurate
‘cuz there’s, you know, there’s Justin Timberlake and
there’s another couple of artists on here that…Michael
McDonald, you know. Some of this stuff I would not, you
know, want to be like kind of associated with it.…I guess
part of it is it wouldn’t be bad if, you know, people thought I
was kind of hip and current with my music instead of like an
old fuddy duddy with music. I mean I sort of like to
experiment a little bit with stuff. I mean I’m not like totally
wild but I like to experiment with, you know, some newer
stuff. So I guess it would be okay if people thought that I
had good taste. It wouldn’t be so good if they said “God! He
likes Justin Timberlake? That sucks!” (P1).
The whole concept reminded me of our discussion of the popular curation project at the Brooklyn museum but on a much more informal and individualistic scale. The fact that informal, individualistic curation regularly occurs also highlights the importance of creating infrastructure in the terms of knowledge, skill sets, and education on digital technology. Or perhaps, such knowledge is already ingrained but not formally articulated? I'm not sure.
However, the other thing that struck me here is that the curation responsibility is shared by i-tunes. As i mentioned earlier, the metadata tags associated with songs are initially included with their purchase or are downloaded via song recognition software if being imported from cd. Also, the default for sharing is to share nothing, so i-tunes helps enforce privacy. Finally, the availability of certain fields only as tags in someway informs the curation through software constraints.
All in all, though, I think popular curation as an expression of identity is a trend that will grow to encompass ever more fields in the future.
Data Sharing
The paper details a study in which 10 authors of PLoS articles were asked to share their data for use in a Master educational project. PLoS authors were chosen because of the journal’s policy that “publication is conditional upon the agreement of authors to make freely available any materials and information associated with their publication that are reasonably requested by others for the purpose of academic, non-commercial research.”
It stands to reason then, that the authors of articles published by PLoS would be willing to share their data. That was not the case. Only 1 of the 10 authors provided their data. Although the study was small, the results suggest a greater problem. Why are authors so reluctant to share their data even when the publication requires it?
Neylon suggests that the findings of this study are not new, and the reasons the authors gave for not sharing their data are not new either. Some authors did not reply, others were unable to be contacted, some said it was “too hard” or their institution “forbids it.” She goes on to suggest that none of these excuses are acceptable.
Neylon insists that it is time for the publications to step up and take responsibility for ensuring their data sharing policies are followed. She believes that the most effective way to go about it is to pull the papers of authors who refuse to share their data. By making an example of a few authors, the data sharing problems for a particular journal will quickly become a thing of the past as authors choose to either share their information or publish elsewhere. Moreover, the publication will benefit from being seen as a source of quality data, further differentiating it from other journals in the field.
And for those who say sharing their data is “too hard” Neylon states that if an author cannot maintain their data so that it can be used again, they shouldn’t be publishing anyway. Neylon does allow that there are occasions when data must be kept confidential, but she argues that these should be the exceptions- exceptions made only after careful consideration of the individual circumstances.
In theory, requiring publications to follow up on the data sharing practices of their authors sounds like a good idea. But I can’t help but wonder if we are putting too much weight on the publishing industry. Is it really their job to make sure their authors share their data? It seems that the authors should step up and make data sharing happen, but it is also clear that they are not currently doing it. What kinds of fundamental changes need to occur to ensure data sharing?
The Virtual Observatory and the Roman de la Rose: Unexpected Relationships and the Collaborative Imperative | Academic Commons
I found this article in the December 2007 Academic Commons issue. It intrigued me because it took two digital projects, seemingly as opposite on the academia spectrum as they possibly could be and drew some strong parallels to one another, and by doing so, was able to pinpoint cyberinfrastructure within a humanist point of view. This was very helpful me, as an old fashioned English major used to stuffy texts and houndstooth jackets, to be able to apply the idea of cyberinfrastructure into the study of literature and other humanities areas.
The article, by Sayeed Choudhury and Timothy Stinson, both from the John Hopkins University, starts off by picking two digital libraries: The Virtual Observatory (VO) and the Roman De La Rose Digital Library. The VO is a typical area, structure, application that one would expect from a digital project: large complex datasets shared, visualized and analyzed by a group of astronomers. The Roman de la Rose Digital Library is a collection of digitized medieval manuscripts chronicling the many different editions, write-ups, illuminations, citations and so on and so forth of the Roman de la Rose.
The traditional view is that these two areas of academic discipline have nothing in common: one is numbers and the other is a qualitative analysis of a medieval work of literature, but Choudhury and Stinston draw some strong parallels in order to establish the shape that cyberinfrastructure can take place in the humanities by suggesting that we expand our defintion of the term "data-rich." According to Choudhury and Stinson, "data-rich" does not have to refer to technology, but rather the ease and difficulty that the practitioners are able to generate, acquire, and process data. For example, in 1627, the Roman de la Rose, had been written, illuminated and studied many times over and had a wealth of data that could be transmitted and used to create new data, and this was way before computers.
Then article goes on to talk about the data deluge: it is not just happening in the more quantitative disciplines, but also in the humanities because manuscripts, being "data-rich," still retain their original values and meanings, but also acquire new interpretations and values with each new generation of academia. This also includes metadata generated. Therefore, the humanities field has just as much presence and need for a cyberinfrastructre to keep it all from losing the progress of information through the ages.
In a sense, the scientific datasets are a modern equivalent of medieval manuscripts. Choudhury and Stinston point out that the many different versions of the Roman de la Rose that build off of one another mirror the intertextuality of World Wide Web and digital media provides an opportunitu reflect more accurately on forms of medieval textuality and transmission of information that may have been lost during the print era.
I felt that Stinston and Choudhury's article heavily drew on ideas and theories of new forms of scholarship and how digital curation can make that possible. It also exposes the need for humanities computing and gives it an adequate place within the information spectrum of the humanities discipline by doing something inherently "liberal arts," comparing two distinct opposites and drawing parallels to show that its not all that different!
Online archive opens access to research
The University of Western Ontario has launched a new initiative, Scholarship@Western, meant to provide open access to articles and other scholarly materials produced by faculty and students. Like many institutions world-wide, Canadian funding agencies are beginning to require open access to published findings. For example, the Canadian Institute of Health Research will withdraw funding if open access to peer-reviewed articles is not provided within six months of publication. This is not only an effort to distribute knowledge, but also to promote transparency and public accountability. Launched ten months ago, over 460 papers have been added to Scholarship@Western.
Individuals create a homepage using a Western template which may include a profile, c.v., and contact information. The page may provide abstracts and citations, links to published and non-published articles, contributions, books, conference proceedings, presentations, etc., and PDFs may be uploaded. Site visitors are invited to join a mailing list to receive notifications of additions to the page. All information is maintained by the individual, and they are "encouraged" to review journals' copyright agreements before posting reprints to the repository. If the individual leaves Western, the page is not removed, but the template will be reverted to a generic version - it will continue to be accessible.
Scholarship@Western is intended to showcase Western achievements, increase the researcher and university profile, connect researchers and create collaborations, prevent duplicate research efforts, and advance new research developments. Additionally, it serves to collect, disseminate, archive, and preserve materials created or sponsored by Western. It is an institutional repository that provides a simple user interface allowing browsing and searching, and provides access to the institutions scholarly output via a single portal.
Western is admittedly behind the curve with providing open access to materials and does not currently require faculty to submit materials to an open access repository. Scholarship@Western was created as a response to the growing need for open access. The Compact for Open Access Publishing Equity is yet another response to this need, and has provided a framework for institutions to support open access journals as an alternative to traditional subscription journals. The agreement commits institutions to “the timely establishment of durable mechanisms for underwriting reasonable publication charges for articles written by its faculty and published in fee-based open-access journals and for which other institutions would not be expected to provide funds.”
Open access is increasingly necessary and increasingly prohibited by the traditional subscription journal system. Institutional repositories are one method of preservation of the achievements of an institution and it's researchers, but do not adequately provide open access. Wide-spread support of the Compact, in addition to open access mandates by funding and other sponsoring agencies, should serve to create and encourage a new environment for open access scholarly communication and publishing.
Tuesday, September 22, 2009
Flickr Commons and the Smithsonian
Journalism as Digital Curation
The Chicago Tribune, however, is upending the traditional reportage hierarchy by engaging in what is being referred to in the journalism world as a "curated approach" to their Twitter feed. The Tribune actually has two Twitter feeds: @ChicagoTribune, which is a standard RSS feed of headlines from the online edition of the newspaper, and @ColonelTribune which is something entirely different. The idea for ColonelTribune came about because headlines would often be cut-off due to Twitter's 140 character limit. The Tribune decided to create another feed which didn't simply parrot the headline but created something more succinct and Twitter-friendly. Instead of the headline, they might lift an interesting quote or statistic from the article. Then the Tribune started doing something really revolutionary with ColonelTribune - they started tweeting about blog posts that were relevant to news stories in the Chicago metropolitan area. They would still post articles from their own website, but they would also post articles from blogs, websites of local Chicago television stations, and even the website of the Chicago Sun-Times. Readers were also allowed to respond to tweets and reader tweets are included in the ColonelTribune stream. I've even seen tweets where ColonelTribune responds to reader tweets. The interest in the "curated" vs. "automated" models is indicated by the relative number of followers of the two feeds. While 19,520 people follow @ChicagoTribune, 584,779 follow @ColonelTribune.
Obviously, there are downsides to the curated model. While the automated model requires very little effort, the curated model has less of a return on investment because the Tribune has to pay someone to run the feed (that being the case - I have to say that tweeting about interesting things going on in Chicago sounds like the best job in the world!) But the implications of the curated model are much more interesting. The Tribune is essentially curating a daily history of relevant digital articles relating to the Chicago metropolitan area from a number of sources outside of their own newspaper. That seems like a project that is worth the cost. If only they could find a way to make it profitable. Any ideas?
