Tuesday, October 27, 2009

Collaboration and Trust in Digital Preservation

In Toward Distributed Infrastructures for Digital Preservation: The Roles of Collaboration and Trust, Michael Day explores the connections between collaboration and trust within the international scientific community. He outlines the growing need for collaboration, especially in scientific disciplines that deal with expensive and sophisticated instruments. Collaboration between several research teams over many years is often necessary to complete thorough and meaningful research. Drawing from another study, he outlines 4 types of scientific collaborations:

▪ Bureaucratic collaborations
▪ Leaderless collaborations
▪ Non-specialized collaborations
▪ Participatory collaborations

He suggests that data curation facilities are more likely to grow out of participatory collaboration efforts, which tend to be egalitarian in nature. His work also underscores one of the main themes we have been discussing all semester: that data curation projects are often created within disciplines or sub-disciplines and are confined to small communities that share data standards. Day acknowledged the importance of dealing with the social and legal implications of developing infrastructures for sharing data and collaborating. He takes up the subject of institutional repositories, but encourages collaboration across institutions and with other parties to increase preservation efforts. Day presents international initiatives to study and implement shared preservation environments.

Collaboration for digital preservation of scientific data requires a network of institutions that share trust. Day defines the relationship by stating that, “trust is at least partly about participants accepting a level of vulnerability in exchange for certain perceived benefits.” Trust often develops over time as institutions work together over a period of years and reach a level of expectation about the other’s reliability. Day relates trust to notions of control, describing the relationship as interdependent. Formal control mechanisms include policies, rules, and procedures, while informal controls can be simple shared institutional cultures. It is important to develop relationships of trust between institutions participating in digital curation initiatives. The trustworthiness of digital repositories is based on the strength of the institutions contributing data and participating in preservation efforts.

Finally, Day reviews various certification measures that may help establish a level of trust. He mentions the responsibilities set forth as the OAIS Reference model, discusses the RLG & OCLC published criteria for “trusted digital repositories,” and outlines the purpose of the RLG-NARA Digital Repository Certification Task Force.

The underlying theme of these initiatives is self-evaluation and definition. Repositories are encouraged to document policies, understand their own strengths and responsibilities, and consider more formal certification. The Digital Curation Centre has created a self-assessment toolkit, DCC Digital Repository Audit Method Based on Risk Assessment (DRAMBORA). The tool is designed to help institutions examine their strengths and weaknesses with an eye toward digital collaboration and risk management. The overarching theme of the article is that while collaboration is important and trust is necessary to adequately collaborate, an institution must thoroughly understand their contributions and roles in collaborating and must be aware of risks.

Renting Scholarship

This week I read an interesting post on the Read Write Web blog titled "Netflix for Researchers: DeepDyve Launches Rental Service for Researchers". DeepDyve is a business that already indexes articles from thousands of journals (and apparently also thinks mis-spelling their name is cute and different). The company will now offer articles for rent. 99 cents for 1 day. DeepDyve rightly recognizes the increasing costs of online journal subscription as prohibitive to scholarship and learning, I think most libraries across the country are feeling that. But I'm not positive this will happen (from their press release headline): "Service Gives Unprecedented Access to Wealth of Research Data, Offers Publishers New Channel to Attract Customers and Monetize Content". Pretty much my new favorite byline (even though I didn't have a favorite before).

Basically DeepDyve's plan is to offer any article in their database for rent, at various prices for various time limits. $1 can get you 1 day, but no printing allowed! $9.99 to $19.99 per month will get you more articles and longer time limits, and a special yearly package allows printing. Read Write Web says that printing isn't an issue "
because most users are just looking at these articles for a few facts or a bibliography and don't need them for extended periods of time". DeepDyve calls these users "knowledge workers" who are people that "use the web for research for their education, health and careers". Hm. Firstly, I would think that if someone is doing research, she would actually want some article she downloaded for a little while, possibly printed out in order for later reference. Secondly, the term "knowledge worker" is silly, although I don't doubt that someday we could see the iSchool turn into the kwSchool. DeepDyve sees "knowledge workers" as professionals needing up-to-date research but not actually doing the research, I think. An untapped consumer base.

I understand that for the Average Joe or Jane, accessing articles must be a real pain, especially if one isn't associated with an institution with a library of any type. But does the Average Joe or Jane actually look for scientific articles about how "High Nucleotide Divergence in Developmental Regulatory Genes Contrasts With the Structural Elements of Olfactory Pathways in Caenorhabditis" (an article listed on DeepDyve's homepage)? I might guess not, unless a serious medical situation is at hand. Perhaps a researcher's institution is private without interlibrary loan and does not have a particular journal, then I could see a person needing this service. I'm trying to envision a nurse or doctor trying to figure out a medical problem or wants the latest information on a topic not in JAMA (or maybe I'm just envisioning House). Then DeepDyve could be quite useful.

The part of this new service that I find particularly intriguing, however, is the strange middle-ground this takes between traditional journal access and open access. I don't think this model flies in the face of the current paradigm of scholarly publishing, since the company is only offering the product of the publishing, not changing the process. But I do like how extremely technical scholarship can now leave the hallowed halls of academia and its libraries and can slum it with knowledge workers for just a dollar. Even if for only a day.

VIVOweb - Facebook for Scientists

I read this week (in a post on DigitalKoans from last week) that the National Institutes for Health (NIH) has awarded a grant of 12.2 million for two years to develop what is being referred to as "Facebook for scientists." Three universities, the University of Florida in partnership with Cornell University and Indiana University, are to develop VIVOweb which will be a social networking software for scientists and researchers. A goal of VIVOweb is to "foster alliances" in hopes of speeding development in biomedical research.

According to the AP press release, presently scientists can use search engines and educated guessing to connect with others needed for their research. The new Facebook-style professional network would use "emerging" Semantic Web technology (hasn't the Semantic Web been supposedly emerging for years now?) to make information from institutions, academic journals, and researchers themselves available to scientists.

VIVOweb will be based on an open-source software developed by Cornell in 2003 called VIVO (which sounds like it should be an acronym, but actually isn't). At Cornell, VIVO is a tool that aids research and discovery by creating a single access point to find scholarly information. VIVO was created to "bring together in one site publicly available information on the people, departments, graduate fields, facilities, and other resources that collectively make up the research and scholarship environment in all disciplines." This technology uses the Vitro integrated ontology editor and semantic web application software, and was created in part to ease the mounting frustrations from Cornell faculty members who had a difficult time finding collaborators in different disciplines across campus because there was no simple, single place to go to look. The VIVO website boasts that users can search the VIVO website to find information about “anything related to academic and research pursuits at Cornell.”

VIVO has since branched out to other institutions.

The VIVOweb project that the NIH is funding for the sciences will use Cornell’s VIVO as a basis to "build a multi-institutional platform for the biomedical community." By creating this multi-institutional Facebook for scientists, VIVOweb will make it easier for scientists to extend their research networks, and to find existing projects or fellow scientists with similar research interests. Cornell will be in charge of developing the multi-dimensional aspect of the project, while the University of Florida will be focusing on creating and developing technology to keep the VIVOweb data current, and the University of Indiana will be responsible for developing the social networking tools that the scientists will need to make VIVOweb useful.

Trying to emulate something as successful and popular as Facebook, but for a more specific purpose, sounds like a pretty good idea. Based on what little I’ve read, VIVO seems like it has been successful and useful at Cornell, so it stands to reason that VIVOweb could be just as successful and helpful to biomedical scientists. I think anything that is designed to help link researchers together, to make it easier to find current projects, and to speed up research and development can only be a good thing – if it works. I’ve blogged about institutions pouring money into other new, developing cyberinfrastrcuture projects before; this is yet another example of having to wait and see what happens. In the meantime, I suppose scientists can always try to use Facebook to network (there are some science groups using it), if they can find the time to do so, in between tending their Facebook farms and cooking in their Facebook cafes.

Ease of Use Trumps Expertise AND Trustworthiness in 2001 study

Credible or not, and why? Image source.

Any one metric for evaluating credibility probably contains fallacies. A combination of techniques may be the wisest course, but in spite of best efforts, people will do what they do.

A 2001 CHI report found that respondents ranked Ease of use above expertise and trustiworthiness in determining website credibility.
Ease of use is the second in a list of 5 main criteria, coming after Real-world feel, and before Expertise, Trustworthiness, and Tailoring (personalization). This actually makes some sense when considering the Real-world feel includes identifying information such as physical address being available on the site. This openness shows that the institution is not hiding at least this contact information, allowing for investigative possibilities.

The meaning the researchers give for Ease of use includes ability to search an "archive," which is a historically common way of evaluating authority (ie, how long have you known someone?) and professional design which is probably more problematic.

A more recent study from 2003, with one author in common from the first study, found that "Design look" was the TOP credibility issue. Appearing in 48.1% of comments compared to 14% that mentioned "Accuracy of information."
Any analysis is subjective, so perhaps the researchers were lax in their judgment of the comments, but I believe that website browsers really do put a great deal of faith into sites due to their usability.
One user commented, "The design is sloppy and looks like some adolescent boys in a garage threw this together."

Perhaps this low(er) importance of accuracy as a metric stems from the fact that users don't necessarily already know what is accurate, that is, after all, what they are judging.
One user wrote, "Most of the articles on this Web site seem to be headline
news that I have already heard, so they are believable." Seeming to be headline news may or may not be a better a metric for accuracy than the design look.

The authors do admit that due to the nature of their study, design look may have surfaced as more important because participants were not actually invested in their search, and so allowed themselves to rely more heavily on superficialities than they otherwise would have.

However, both studies contain results that point to 5 or 6 main ways that users decide whether or not to believe information on the web, and towards the top of both lists is design/use. Although this might not be the way we wish the world to work, it is wise to take it into consideration. We don't want to "Look childish."

In order of ranking from the 2003 study, these criteria were related to determining credibility either in a negative or positive way (I think which is which will be obvious):

Design Look 46.1%
Information Design/Structure 28.5%
Information Focus 25.1%
Company Motive 15.5%
Usefulness of Information 14.8%
Accuracy of Information 14.3% (!)
Name Recognition & Reputation 14.1% (!)
Advertising 13.8%
Bias of Information 11.6%
Tone of the Writing 9.0%
Identity of Site Sponsor 8.8%
Functionality of Site 8.6%
Customer Service 6.4%
Past Experience with Site 4.6%
Information Clarity 3.7%
Performance on a Test 3.6%
Readability 3.6%
Affiliations 3.4%

Sunday, October 25, 2009

They Did What?!?: Trust & Scholarly Publishing

Fister, Barbara. "This Journal Brought to You By..." ACRLog. Blog posted May 9, 2009. Available at http://acrlog.org/2009/05/09/this-journal-brought-to-you-by/(Accessed 10/25/2009).

Grant, Bob. "Merck Published Fake Journal" The Scientist.com. Posted April 30, 2009. Available at http://www.the-scientist.com/templates/trackable/display/blog.jsp?type=blog&o_url=blog/display/55671&id=55671 (version not requiring free registration: http://www.the-scientist.com/blog/print/55671/). (Accessed 10/25/2009).

Grant, Bob. "Elsevier Published 6 Fake Journals." The Scientist.com. Posted May 8, 2009. Available at http://www.the-scientist.com/templates/trackable/display/blog.jsp?type=blog&o_url=blog/display/55679&id=55679. (Accessed 10/25/2009)




As I was looking around on the ACRLog this week, I came across a blog entry from May of this year popped out at me given this week's readings on trust and authenticity. In a post titled "This Journal Brought to You By...", Barbara Fister writes of how journal publisher Elsevier received money from the pharmaceutical company Merck to publish a journal, the Australasian Journal of Bone and Joint Medicine, that was designed to appear like a scholarly, peer-review style journal. Fister's post in turn led me to the two articles written by Bob Grant for The Scientist.com that first exposed the ruse to the blogosphere.

The "journal" appears to have been disseminated to doctors in Australia and many of its contents were summaries or reprints of articles that were favorable toward Merck drugs.
(The existence of the fake journal was first reported in the Australian in connection with a lawsuit by a man who suffered a heart attack after using Vioxx). Though it is unlikely that the publication would have fooled researchers, it may have hoodwinked some doctors. According to Grant, Elsevier has since apologized for the publication and has stated that they regret that the journal's sponsorship by Merck had not been made clear (nowhere in the issues Grant reviewed was this sponsorship acknowledged). A spokesman for Merck, Sharp & Dohme Australia (MDSA) explained to Grant that "MSDA understood that Elsevier envisaged the complimentary publication would draw on the vast resources of Elsevier, publishers of many leading peer-reviewed journals including Lancet, Bone, Joint Bone Spine and others, to deliver novel and timely full text articles and abstracts to physicians."

In the second of his articles, Grant uncovered that the Australasian Journal of Bone and Joint Medicine was not an aberration, but that Elsevier had in fact between 2000 and 2005 published five other such fake journals--the Australasian Journal of Neurology, the Australasian Journal of Cardiology, the Australasian Journal of Clinical Pharmacy, the Australasian Journal of Cardiovascular Medicine, and the Australasian Journal of Bone & Joint [Medicine]. (Other bloggers have suggested that the number may be even higher). These "journals" were published by Elsevier's Australian office under the Excerpta Medica imprint. At the time of Bob Grant's second article, Elsevier had declined to reveal which companies had sponsored these other publications.

From what I can tell--though none of the many blog posts I examined were perfectly clear on this point--the series of "sponsored article publications" (Elsevier's term) appeared either primarily or exclusively in print (analog) form. They were never indexed in Medline and they had no journal websites (though I'm unclear as to whether they ever appeared in Elsevier's online listing of publications). Grant's original article post includes a link to PDFs for Vol. 2, Issue 1 and Issue 2 of the Australasian Journal of Bone and Joint Medicine. The reason why I decided to write about what appears to have been a print phenomenon is that it also raises interesting questions regarding trust and authenticity in the digital world. I suspect that if these fake journals had had a real web presence, their status as crypto-advertising would have been exposed much more quickly. That said, both print and digital scholarly communication are affected by several shared trust and authenticity issues. The questions that,
in this week's readings, Nancy Van House states are important for establishing a digital library's provenance--"Who designed it, and for what purpose? What design choices were made? Including, what is not visible to the user?"--are precisely those questions whose answers were intentionally obscured in the Elsevier fiasco.

Moreover, it has been interesting to see how this revelation has impacted the debates surrounding Open Access and publisher gate-keeping. As a blogger at Small Gray Matters writes: "The bitter irony is that Elsevier, along with the other major academic publishers, have spent the last few years ceaselessly lobbying against the open access movement, on the grounds that open access journals can’t be trusted to maintain the high quality of peer review that the commercial publishers provide." Is scholarly communication (and, by extension, digital curation) too important a function to entrust to commercial entities,
like Elsevier or Google, whose bottom line will always be financial?

The Ocean Observatories Initiative

The Ocean Observatories Initiative (OOI) is an infrastructure project that will collect and make available data from a network of ocean-based sensors that will measure various physical, chemical, geological, and biological variables in the ocean and sea floor. The data collected from hundreds of such sensors will be made available in near-real time, and openly accessible. As stated by Tim Cowles, the program director of Ocean Observing in the Consortium for Ocean Leadership, the data collected will belong not to the OOI, but to the people, whether they be professors, researchers, or the average citizen. The goal of the project is to "address a multitude of important science and societal question, including those centering around climate change, ecosystem health, ocean acidification and carbon cycling."

The Ocean Observatories Initiative is a joint project of the National Science Foundation and the Consortium for Ocean Leadership. The OOI is the National Science Foundation's contribution to the U.S. Integrated Ocean Observing System (IOOS). The OOI will focus on making scientific discoveries via new technologies, while the role of the IOOS will be to apply these discoveries to social needs. The University of California, San Diego will be responsible for implementing the cyberinfrastructure necessary to make the sensor data available to the public. The cyberinfrastructure is intended to allow everyone to interact with the ocean, not just those who have access to special equipment.

The project has recently moved from the planning to initialization phase, the first official project year having begun this September. Initial data flow is scheduled for early 2013, with full capability being established by 2015. Grants for the first year alone total more than $110 million, with $106 million coming from the 2009 American Recovery and Reinvestment Act, and $5.91 million from the NSF to cover construction costs. The first workshop for introducing the OOI to the science community is scheduled for November 11-12 of this year in Baltimore.

This program is an example of the way in which science is capitalizing on the advances in information technology, as well as an example of the move toward data-driven science. The project is touted by the creators as a revolution in the field of ocean science, and it seems that the open accessibility of data will be a key feature of this revolution. As the project is only in the initial construction phase, one cannot yet determine whether the project will be as successful as, for example, the National Virtual Observatory. If it is, the OOI will undoubtedly be an invaluable resource for ocean science.

Saturday, October 24, 2009

Spezify

I have died and gone to heaven. Spezify "is a search tool that presents textual, graphic, and photographic results in a visual format. Blogs, videos, microblogs and images, couple with web-based versions of more traditional print media to give comprehensive search results."

For a visual thinker that loves to scour everything from Twitter to YouTube to blogs in search of the latest from galleries, artists, conferences, scholars, designers...Spezify pools and presents information in a visual collage format. It also shows you related search terms (which may or may not be relevant - i.e. an important interntational conference for Contemporary Asian Art is going on right now - so the current search of that term brought up "hotel" as a related term.

Like Twitter it also shows you recent "hot" search terms. Spezify displays the information in a giant visual wall that you can scroll in all dimensions. Via Twitter, Cooliris told me that there is a way that I can d/l their program even though I have a PowerPC Mac mini. This weekend I will do so and see if Spezify's search results can be navigated on the visual wall of Cooliris. This would make for faster and improved scanning functionality.

Wednesday, October 21, 2009

Technology in the Arts compares Virtual Gallery software

New Report Compares Virtual-Gallery Software.

A free report put out by Technology in the Arts takes a comparative look at three popular pieces of Virtual Gallery software: Virtual Galleries, Image Armada, and SceneCaster/3D Scenes. Such software can be used by curators and artists wishing to create a 3D gallery environment that is either online or otherwise digitally portable. Uses include serving as a design tool for curators or students planning installations - to serving as an access tool for those unable to physically explore the gallery space.

The report graded each application according to the following criteria: "ease of creating an exhibition; quality of images; flexibility of creating an exhibition; ease of navigation in the three-dimensional space; and ease of publishing or distributing the gallery." Comparative information was also provided re: the following technical information: "Mac and PC compatibility, hardware requirements, internet-based viewing, and image and lighting manipulation."

One of the most useful aspects of the report were the recommendations - whether your organization has a limited budget, requires Mac compatibility, and whether the gallery needs to be shared with people online. These concerns seemed to be the make or break features, and these recommendations were included on the next to the last page.

In addition to introducing me to some fascinating programs for the other form of "digital curation" the report gave a succinct presentation of software comparative review. My favorite, though expensive ($3750-$6750) would have to be Virtual Gallerie as it combines walk-through 3-D functionality with workability in a web browser.

The possibilities for having portable and interactive online exhibitions of artistic, historic or curated archival collections - in a 3-D manner, with perhaps abilities to zoom in and read documents and rotate/analyze 3-D objects - is fascinating. At a glance it appears similar to Second Life. I posit that if museum and gallery can create exhibits (which communicate narratives and encourage innovative thinking and evaluation of art and history) that are digitally portable
and preservable - I find the prospects exciting.

More Google. Wave (sorry)


Did you get a Google Wave invite? Yeah well, who cares. Some people say it will change everything; some people say it's all hype. Whatever you think, it's kinda fun to think about what Wave could do for science and research. Cameron Neylon wrote a short article about Wave and science for Nature and wrote quite a bit more about it on his blog "Science in the open." He suggests things like a Creative Commons bot that stamps a document and asks contributors if they're ok with the license. Or how about a bot that registers a document with a repository and automatically sends metadata to the repository? Data can be linked from within a Wave to a live spreadsheet that is being constantly updated. Neylon also makes some interesting suggestions for using the playback feature in the peer review process. Neylon is even sponsoring a Google Wave Science Hack Day where a bunch of people will get together, in person and in Wave, to discuss scientific applications of Wave. I've noticed people talking about the implications for interactions between doctors and patients. So what do you think? Are Wave's implications for research and scholarly communication interesting or not?

Netherlands Coalition for Digital Preservation

In September 2009 the Netherlands Coalition for Digital Preservation released an English summary of the Dutch National Digital Preservation Survey, held in 2009. During the course of the survey they spoke with many different groups from three different sectors (government, research community, and the private sector). Many of the issues that they identified we have already seen so far in this class, but I appreciated the insight in steps towards solutions they came up with.

One problem identified was a shortage of knowledge on just how to perform digital curation. The approach of one group in the study to solve this problem was to set up a forum for exchange of knowledge and expertise. Another approach (and one that has seen some development in the USA government) is the appointment of chief information officers (CIOs) for government institutions.

Another problem was the shortage of tools: on the international tools projects such as PLANETS, Caspar Preserves and PrestoPRIME were all being developed with EU funding. Although not finished, the idea here is to create toolboxes for smaller institutions that don’t have preservation as a main concern, so they can still maintain custody of their files.

As we have seen before, another common problem is the hangover from pre-online publication. Permanent access, or cradle-to-grave custodianship, is called the “records continuum” in the Dutch model. Since the producers of digital information don’t recognize the long-term concern of curation as their own and are instead concerned with publishing to the public, project plans must provide provisions in how to provide long-term access, beyond just simple funding requirements. This would include dedicated selection mechanisms spelled out in the project plans with support in the data creation phase from data management professionals. This is a good way to avoid placing the burden strictly upon the researchers and if done correctly could even enhance research.

One side concern that I found interesting was the possibility of outsourcing file storage to commercial groups to use scale advantages. This has a lot of associated concerns we haven’t looked at yet. There are a lot of fascinating projects going on in the Netherlands, including “Images for the Future”, an audio-visual digitization project with collaboration from six groups that has a budget of 154 million € and will be tackling tens of thousands of hours of video, audio, and millions of photos over the next six years.

Google Books and the Digital Curation as Preservation Myth

On October 8th, Sergey Brin, co-founder of Google, wrote an op-ed piece in the New York Times in which he defended the Google Books project presumably as a response to the Department of Justice's antitrust probe into the Google Books settlement. In the piece, Brin lays out all of the reasons why Google Books is an accessibility revolution. The piece is strewn, however, with inaccuracies and outright lies. Brin claims, for example, that "if you want to access a typical out-of-print book, you have only one choice — fly to one of a handful of leading libraries in the country and hope to find it in the stacks." OK, certainly there are books for which there are only one or two copies in non-circulating rare books libraries, but the "typical out-of-print book?" What?! Ever heard of inter-library loan?

This is far from the only egregious inaccuracy in the article. Brin makes the (now common) argument for digitization as preservation by referring to incidents of flood and fire after which hundreds of books were lost including a 1978 flood at the Stanford University Library (Brin's alma mater). Brin cleverly claims that one could have read about the reports of loss due to this flood in the Stanford-Lockheed Meyer Library Flood Report, but this report is also no longer available. A NYT commentator, however, managed to find four copies of this very report on WorldCat.

Brin really hammers-in the argument for digitization as preservation, even implicitly comparing Google Books to the Library of Alexandria which was deliberately burned down on three separate occasions. But what is keeping Google's serves from being destroyed in similar catastrophes? What happens to all of this heavily proprietary content when Google's stock prices are in the gutter? What happens when an electronic terrorist hacks into those servers and wipes everything out? The myth of digital curation as preservation in the case of Google Books rests on two pretty precarious presumptions: 1) that the digital content is invulnerable, and 2) that Google (which has been around for 15 or so years) is immortal. Are we really that naive? It's not like Google is putting everything into the Granite Mountain Records Vault.

Rather than focus on the fairy tale of digitization as preservation, I think Google has a better chance of selling the accessibility argument. That argument will be much more convincing, of course, if Google Books improves its metadata.

Digital Collection Building: Not Just for Digital Libraries!

This week I read an article, written by an academic librarian in Scotland (feel free to read the rest of this post in a brogue), that seeks to weigh the respective values of digital and traditional collection building in the context of lean economic times. (I wonder where that idea came from.)

Intrinsic to the argument is the postulate that to build a digital collection effectively and to build a traditional print collection effectively one must make a significant investment of roughly equal size. Furthermore, Nicholas Joint (the author) assumes that libraries will not necessarily possess the budget to engage fully in both activities. Thus, Joint presents the dilemma facing librarians as a choice between embracing a digital future and focusing on digital collection building or maintaining current practices and continuing to purchase paper books.

Joint's arguments come in two phases. First, he cites some usage statistics from Scotland. Apparently, the circulation rate in Scotland is only about fifty percent of the collection size, which is actually really good as far as academic libraries go. (As someone who has traveled to Edinburgh in February, I can attest that there is indeed probably a lot of reading go on indoors.) However, the same libraries found that each digital item in the collection was downloaded twenty times within a school year. While there are some validity issues with the statistics, namely that the university only possessed three thousand presumably popular digital titles, usage statistics seem to suggest that even conventional libraries may wish to devote greater resources to digital collection building than print acquisitions.

The second phase of Joint's argument is theoretical. In the article, he compares to works on digital publishing and the emergence of the ebook as a medium. One, Luke Tredinnick's Digital Information Culture: the Individual and Society in the Digital Age essentially argues that a new digital culture has either already emerged or is emerging and that not only is resistance futile, but that it will be doomed to appear to future generations as a socially-conservative backlash akin to fears of erosion of values wrought by tv. Conversely, in Never Mind the Web, Here Comes the Book, Miha Kovač argues that everyone is misinterpreting the term ebook. All books nowadays start in electronic or digital form and yet get transformed to print objects because that is how users prefer to engage them. As such, print collections will continue to be supreme according to Kovac.

Ultimately, Joint decides that Tredinnick possesses the stronger argument. Joint views print-outs as little more than useful transitions of essentially digital books. He points out that no one donates printouts of ejournals to libraries. For this reason, as well as the statistical one, Joint recommends that libraries devote more of their resources to building digital collections than traditional print collections.

I found the argument interesting primarily as an example that digital curation is not limited solely to digital libraries such as Valley of the Shadow, nor even to electronic publishing conglomerates such as JSTOR, but is also of increasingly crucial importance to traditional librarians as well.

Polymath Project

The Polymath Project began in February 2009 as an experiment to collectively work on an unsolved mathematics problem using a blog. Timothy Gowers and Michael Nelson started this project, and wrote this article in Nature, to test the success of an open-source platform to group problem solving. They wanted to prove the density Hales–Jewett theorem (DHJ), which while is known to be true still needs proofs.

The project was successful and in just over a month, the problem was solved. There were many contributors from all aspects of the mathematical community, such as an award winning researcher and a high school teacher. Gowers and Nielson liked the way their experiment differed from other team projects in the sciences, since in those teams "work is usually divided up in a static, hierarchical way. In the Polymath Project, everything was out in the open, so anybody could potentially contribute to any aspect."

Farsightedly, Nielson and Gowers realize this situation raises authorship and preservation issues. Who gets credit for solving the DHJ theorem? The authors of this article will provide a link to the blog as a list of collaborators, so that someone who didn't provide insightful comments won't get the same credit as another who made major advances. Of course, this brings preservation into it, since the blog might not always be a permanent link. Nielson an dGowers mention how the Library of Congress saves certain legal blogs and hope to emulate that somehow. I think the issue is more complex than that, and more investigation should be done on how other web sites are preserved.

The articles notes that "outside mathematics, open-source approaches have only slowly been adopted by scientists" but the authors give several examples for how other sciences could benefit from this kind of collaboration and online data sharing. Almost summing up what we've learned in class, Nielson and Gowers realize "the widespread adoption of such open-source techniques will require significant cultural changes in science." I'm not sure how a blog can be considered that revolutionary as 'open-source' , but opening up a process that is usually confined to classrooms and chalkboards is a step in the right direction. It's nice to see people within the community noting the problem and challenging their colleagues.

Furthering the Open Access Movement

In recognition of open access week, Jason Baird Jackson wrote a blog post titled Getting Yourself Out of the Business in Five Easy Steps. His five steps are as follows:
• Choose not to submit scholarly journal articles or other works to publications owned by for-profit firms.
• Say no, when asked to undertake peer-review work on a book or article manuscript that has been submitted for publication by a for-profit publisher or a journal under the control of a commercial publisher.
• Do not seek or accept the editorship of a journal owned or under the control of a commercial publisher.
• Do not take on the role of series editor for a book series being published by a for-profit publisher.
• Turn down invitations to join the editorial boards of commercially published journals or book series.

In effect, he is advocating a boycott against for-profit journals. In theory this sounds good, but I have trouble seeing it play out well in practice. AS we have discussed many times, many of the people writing articles are doing so in part to fulfill tenure requirements. Those requirements include publishing to the top journals in the field which are almost exclusively for profit. While the idea behind a boycott is noble, it seems hard to believe that researchers would be willing to give up publishing in the top journals.

Stevan Harnad wrote a blog post in response to Jackson saying just that. But more interestingly, he explained that just such a boycott was threatened in 2000 among biological researchers. As expected, the journals ignored the threat and the biologists never followed through with the boycott. That's when they decided to try something new. They launched PLoS, an alternative open-access journal for the field.

PLoS is now one of the great success stories of open access, becoming known for its quality and becoming a leading journal.
But may other open-access journals have not been so successful. Harnad suggests the solution to open-access lies in the institutional repository mandating of all published works of the researchers at an institution being placed into the repository.

Again, I have my doubts that this will work as institutional repositories do not have a good track record. Even when researchers are required to place articles in their repository, they don't always do it. And many for-profit journals will not allow a final version of the article to be placed in the repository, leaving the researcher forced to place preprints into the repository.

Clearly the system we have for journals is not working, but with few exceptions, institutional repositories are not the answer. Currently, it's pretty rare for an open-access journal to be considered as good as the top for-profit journals. We still haven't found an appropriate strategy to further the open access movement, only strategies that have had mixed results in the past. It seems that it is time for something new, but what form would it take?

Tuesday, October 20, 2009

How do policies help?

In chapter two of the NSF publication we read this week, The Elements of the Digital Data Collections, the NSF says that policy needs to be developed in order to support the life of digital data. This chapter does not exactly explain why policies for data collections are needed or what exactly they hope to accomplish with policies. Within the document the NSF mentions that "a “one-size-fits-all” approach to policy development is inadequate". So why are they still asking for people to work on creating policies?

Also this week I read the short blog entry William Uricchio and the object/subject in participatory media, which is about a talk from William Uricchio last week. This a cursory look at participatory culture and how our culture has developed systems based on the participation of millions of people. Two of the examples Uricchio provides are SETI and iPhone apps. Systems of participation exist without models for how they should operate, instead these tools rely simply on what was possible when they were being created and based on what the needs of the users and developers were. This is very much the case for the cyberinfrastructure tools we've seen this semester.

One difference between the NSF document and the Uricchio talk is that the NSF is continually looking for a monolithic answer to cyberinfrastructure while Uricchio provides examples of working systems that could inspire the cyberinfrastructure. In particular, there's Photosynth, a tool from Microsoft that creates 3D "experiences" based on collective digital photography. It is definitely worth downloading Silverlight to see how the photos have been crowdsourced and curated in a meaninful way. About this example, the blog author says "Rather than building a space by building a model, the model emerges from the production of thousands of amateur photographers". This is certainly something I've been thinking about in terms of the cyberinfrastructure - before deciding what it will be and what it should be called, the cyberinfrastrcuture should emerge as a production of collaborative effort. After all, there is no working model for the cyberinfrastructure.

After all the discussion in class about building the cyberinfrastructure and reading about the alledged need for policies that will shape the cyberinfrastructure, I'm a bit skeptical that focusing on the end result of policy is helpful. Any work that's done on developing policy may be premature until there are more applicable cyberinfrastructure tools. Also, I'm afraid of what kinds of developments could potentially be stifled by inadequate policies. And again, there is no "one-size-fits-all" answer to all digital curation projects, these projects arise based on very specific needs of the people who use the data.

The one policy issue that I think will be the most difficult to tame long-term is copyright. I have been thinking lately that the best way to shape these policies is to simply participate in fair use activities. My hope is that if enough people are asserting their right to fair use, policies will be developed and strengthened that reflect those behaviors.

Ultimately, I wonder what policies will actually do for cyberinfrastructure. Why do the people involved want these policies? Do they simply desire a map to follow because that is the model they are familiar with?

History of the Old and New

First thing I should say is happy Open Access week everyone. I am sure you have discovered in your perusals of the interwebs that this week is set aside for the celebration of free information for all. I recommend you take time to explain to your loved ones the value that open access has in their lives.

In the non-open access side someone posted to the insider about a historical document site called Footnote, I did not immediately investigate but I put it on the back burner to check out later. Then a few minuets later when I was reading a blog Footnote came up yet again. I felt that I must look into it.

Footnote is a semi-free, (read mostly pay) historical document image dataset. There are all sorts of interesting things to explore, News Papers, Texas Death records, Records from WWII, the founding of Liberia and more. Some 50 million images or so the site claims.

The viewer is really nice. The images are clear (better than Google books), the zoom is good and the interface is pretty easy. The ability to annotate images is very interesting but I am not sure how that works; if the annotations are just for you or for everyone. Users can create connections between documents and images that they upload. Did I mention that users can up load images? You can upload images and create pages about people or places or events. Interesting not all of these items are “historic” in the proper term. These are very odd. You can also print copies of the documents. Which I did… There is yet more functionality that I have not had the time to explore. This has the potential to be a very powerful cultural heritage tool.

Now here are my caveats. The free version is highly restrictive. It seems like ever document I tried to query on came up with a premium page notice. This is kind of unfortunate because I really wanted a chance to explore the scope of documentation. Many of the documents come from NARA or the LOC and I understand that what the site is selling is the “service” but still I would like the opportunity to play a bit more before I buy.

Another caveat is provenance. And I guess this is the age old question, how do you trust what you find on-line. I assume that there is some sort of community in place here that cleans the documents and ensures that the information is valid and accurate. Another question I have is how to ensure the contents move into the future.

This site is very interesting to me but I am not quite sure why. There is a mix here of the personal and the historical that you don’t usually see elsewhere. I really find it fascinating, it is an experiment with meaning and historical documentation that does not usually happen.

Research Remix

Open Access is one of the major issues that we keep revisiting in this class. In honor of International Open Access Week (which spans from October 19-23, and is an “international movement that pushes for broad and free access to research findings and publicly funded studies”), I’m blogging about the Issue Lab and its Research Remix, which I discovered at the Open Access Week website.

IssueLab is an open access archive and online publishing forum that gathers and shares nonprofit research about social issues. The mission of IssueLab is to “more effectively archive, distribute and promote the extensive and diverse body of work being produced by the third sector.” IssueLab aims to make it easier for nonprofits to manage and share their research with others.

On Monday, coinciding perfectly with the beginning of Open Access Week, the Research Remix was announced. This is a contest issued by IssueLab that calls for contestants to creatively reuse nonprofit research data to make videos that provide commentary on social issues. The goal of the IssueLab’s Research Remix is to get interested people to “apply their skills for social good, to base their social commentary in fact, and to learn about legally remixing and re-purposing existing content.”

At the IssueLab website there are over 300 (I believe the current count to be 335) Creative Commons openly licensed research reports archived that contestants can use when making their videos; these reports all have openly licensed video footage, images and music for video-makers to repurpose. This video contest “aims to engage working artists and digital media students” while at the same time “encouraging nonprofits to make their research more broadly available and usable through open licensing.”

The deadline for this competition is December 31, and one of the requirements is that the submitted contest videos must be openly licensed, which I think is a great requirement because fair is fair; after all, it wouldn't be right to use someone else's openly available materials, and then not let anyone else remix your own. All final videos will be made available to the organizations whose research was used in remixing.

I love the idea of the Research Remix and I think it will be something fun that inventive video-making types can really get involved with; as evidenced by the growing numbers of fan-made videos at websites like YouTube, people love to remix and repurpose things. By using openly licensed material, the Research Remix is a great way for people to make a statement about social issues while creatively expressing themselves legally. Legally is the key word in that sentence; much of the fan video remixing that is found online is illegal because users repurpose copyrighted material. I’m not sure how much this contest is really going to be encouraging for nonprofits to make their research more openly accessible, but I like that the thought is there and that an effort is being made.

That being said, I hope everyone is having a happy International Open Access Week.

Digital Repositories and Sustainability

In two blog posts from earlier this year, Dorothea Salo mused on the issue of sustainability with regard to digital repositories.

In the first post, entitled (aptly) "Sustainability," Salo was inspired to think about the sustainability of digital repositories when she learned that arXiv is looking for different funding sources because Cornell is no longer willing to foot the bill. The question that this brings to Salo's mind has to do with role of libraries in the support of digital archives, particularly open access digital archives. She asks, "Are we [libraries] really memory organizations, or are we only memory organizations for print?" Salo seems to lean toward the belief that the library's role in the maintenance of digital archives should be larger. This sentiment reminded me of the article by Sayeed Choudhury and Timothy Stinson that I posted on a few weeks ago; they, too, envisioned the library as having a prominent role in the future of digital archives.

Salo also spends a bit of time in her post venting her feelings about irresponsible digital curation. She cites an instances in which searching for a long-term home for an item that originated outside of her institution she came across a repository that had been abandoned. This repository, Salo relates, seems to have been "hacked up" by two people in their spare time, and abandoned after the founders "got a couple of publications out of the attempt." Understandably, Salo finds this sort of behavior inexcusable, particularly because other won't attempt to establish repositories for this discipline while this one is still in existence. Such irresponsibility is, she says, a betrayal of trust.

The second post, called "Object lesson: when researchers run repositories," relates the story of the Mana'o anthropology repository. Dr. Alex Golub, founder of this repository, is no longer able to maintain it and is attempting to find an institution that will pick up it his stead. Although she considers the Mana'o repository a worthy project with the best of intentions (unlike the one cited in the previous post), Salo believes that any one-person effort of this type will eventually have to be rescued, and that it is irresponsible not to anticipate this from the inception. Salo again muses on the role of the library in rescuing such digital archives, naming librarianship's "continuing error" as the fact that "they have no infrastructure or plan in place for accomplishing these rescues."

Salo present us with two scenarios in which the sustainability of digital repositories is at issue. In the first case, the repositories failure was due to the negligence of its founders. In the second scenario, the well-intentioned creator of a repository came to a point at which he was no longer able to handle it, and had not planned for such an eventuality. The larger issue in both of these cases is who will pick up the slack. Salo, being an academic librarian, questions the library's role in the maintenance of digital archives and repositories. However, one could also put things in a larger context and ask what type of institution is best suited to be the custodian of the kinds of data repositories and archives that are emerging. Should the burden fall on libraries? Or are agencies like the Nation Science Foundation better suited to handle such large projects? Or should new entities be founded whose sole purpose is to create and curate digital repositories? Finding the best home for a particular digital curation project is one way to ensure its continued sustainability.

Computer program proves Shakespeare didn't work alone, researchers claim - Times Online

Computer program proves Shakespeare didn't work alone, researchers claim - Times Online

and

http://news.yahoo.com/s/time/20091020/us_time/08599193097100

Following a blurb on Yahoo! news, I found this neat article from Times Online, in conjunction with a Time.com article, discussing how a distinguished Shakespearean authority, Sir Brian Vickers, used plagiarism detection software to determine authorship of the history play Edward III.

Some background concerning this literary mystery: there's been a mystery for over 400 years as to whether Shakespeare co-authored Edward III with Thomas Kyd. The earliest copies of the play are unattributed and was published anonymously in 1596. The possibility of Edward III being penned, in part, by Shakespeare has been debated back and forth and was completely ignored by mainstream publication until 1997.

The plagiarism software, called Pl@giarism and developed at the University of Maastricht, was created to help educators catch cases of plagiarism. Pl@giarism operates by comparing all documents in a specified directory or by scanning the internet for files or works that match a phrase from the document being scrutinized.

Sir Brian Vickers, in his study of Shakespeare and other period works, came up with the concept of a "lingusitic fingerprint," in other words, playwrights often use the same patterns of speech. Vickers argues that, while it is established that a group or society can share common words three word phrases, it is pretty unlikely for a match of metaphors and unusual parts of speech to come up unless it was written by the same person due to the concept of the linguistic fingerprint.


Vickers came up with 200 matches when he tested Edward III with Shakespeare's work published before 1596. The average match rate is 20 matches. The 200 matches that Vickers came up with, in terms of relevancy, to use a database term, is a pretty strong indication that the provenance of four scenes in Edward III can be given to Shakespeare. Vickers tested a series of "recycled" phrases such as "come in person hither," or "thou art thyself" and whole phrases like "lilies that fester smell worse than weeds," and sure enough, there were some closely related matches in Edward III!

The article ends with Vickers stating his hope that further computer developments can be made in order to help prove provenance of major anonymous works and that the English literature field will come to utilize such handy materials in their study of text.

In the light of what we've been talking about in class and how this ties in with digital curation, I thought that these two articles and the project it discussed, proved to be a compelling example of how the cyberinfrastructure can transform and add a new dimension to research in the humanities and essentially put to rest some stubborn mysteries that have plagued the English academia.

visualizing data on demand!


The National Geophysical Data Center makes available large amounts of data, divided into sub-sections based on terrain and source: land, marine, satellite, snow & ice, and solar-terrestial. The NGDC is part of the US Department of Commerce, the National Oceanic & Atmospheric Administration, the National Environmental Satellite, and the Data and Information Service. The goal of NGDC is “is to be the world's leading provider of geophysical and environmental data, information, and products.”

The NGDC seems to be a good model for digital curation, and for a repository that makes accessible the datasets in their holdings. NGDC currently contains over 400 databases (digital and analog) from many sources, and users of these data sets include researchers, private industry, government agencies, universities, as well as the public.

Perhaps the most interesting aspect of the NGDC is their “Visualizing Data” section. This area of the site makes available a selection of images based on datasets in the collection, and furthermore, “can produce custom images, on request, for many of our databases.” This really amazes me! That the NGDC is not only providing datasets (for free!) but also offering to work with users of the sets to create useful images derived from these sets makes it an excellent example of providing data, and the tools to visualize that data.