Wednesday, October 28, 2009
Improved Twitter feed curation
1) I can create lists for subjects that interest me, and move those that I follow into that list (functionally this is akin to creating labels in Gmail). Doing this will create a link on the right of the home-page Twitter feed that I can click and read only those in that list.
2) others can subscribe to this list - which creates a kind of sub-follower rating. Let's say for example that I compile a list of those who tweet about Contemporary Asian art: this would be its own list that interested people could subscribe to (essentially bypassing me and my tweets but reading those that I have aggregated - or perhaps just going to my list to harvest them for one's own list). This is an idea similar to reading filtered blog subscriptions or friends groups in Facebook as well as customized blog-rolls.
3) There are great implications for use for larger or more widely trusted entities whose lists could be very popular to subscribe to. The introduction of this to Twitter's networking capabilities is a tremendous improvement for ease of use and both content and contact searching.
Now, if only YouTube would similarly jump on board with this concept.
Get the gamers involved
I had intended to write about this topic earlier in the semester and had forgotten all about it until I saw a reference in this week's reading by Amy Friedlander, The Triple Helix: Cyberinfrastructure, Scholarly communication, and Trust. Friedlander discusses how problems that can be cleanly parsed into discrete tasks are well suited for a distributed capacity model which allows open, but still structured, participation from the public. She provides the example of protein folding, which is the topic of an April 2009 Wired Magazine article: Gamers Unravel the Secret Life of Protein.
The protein chemistry world has a biennial World Series competition to see who can predict the shape of a protein only knowing the sequence of its constitute parts (Community-Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction, or CASP). CASP surveys labs around the world to find proteins that are about to be solved, and compile a list of puzzles online.
David Baker, whose team had dominated the competition since 1998, had been using Rosetta@home, similar to SETI@home, which farmed out computations to volunteer PCs distributed globally - providing Baker with the equivalent of a supercomputer. However, the computers were unable to complete certain puzzles, which humans should be able to solve, having better spatial reasoning. Baker's friend David Salesin, a computer scientist, brought him together with Zoran Popovic, another computer scientist and graphics expert, and the three developed what turned into a massively multiplayer competition. Gamers are given a multicolored knot of spirals and clumps, which they fold and wiggle into its optimum shape.
Baker then entered potentially accurate CASP protein structures into the biennial competition. Of 15 submissions, 7 finished "in the money" and one took first place. The gamer team, led by a 13-year old, beat the best biochemists. Baker was also hoping to find prodigies... when "Cheese" (the 13-year old) was asked how he did it, he said, "it just looks right."
Baker has given the players a new challenge to design a new protein drug with the right size and binding properties. Baker will synthesize and test the most promising structures and if any have value in the real world, the gamers will share in the credit.
The article doesn't specifically address authenticity or trust, but the gamer submissions are not automatically deemed correct, even though the game is based on laws of physics. Submissions are reviewed by CASP, and/or tested in the lab. But, given the fact that there are more ways to fold protein than atoms in the universe, and they arrange in a fraction of a second, collective efforts are crucial.
Old Problem, New Scale
Bearman and Trant see a number of possible solutions to the problem without really recommending one single method. They divide their methods into public methods, secret methods, and functionally dependent methods. Public methods may either be social (e.g. creating a collecting institution of record along the lines of a certified website which raises issues of who performs the certification) or technological, such as using public key encryption to create digital signatures for documents. Secret methods consist mainly of technological clues buried in empty space within files such as stegonography. Functionally dependent methods probably do the best job in authenticating because the very working of the file is tied to the end user's ability to authenticate it. However, this is much trickier to set up and probably a bit overkill. (After all, there is a certain amount of responsibility traditionally borne by the scholar...)
I found this article interesting largely because of my history background. When evaluating physical reprints or translations of primary sources, a historical scholar will check the publisher (kind of an early form of public authentication, I suppose.) But in digitizing, nearly anyone can be a publisher, so count me as convinced, scholarly responsibility notwithstanding, that some sort of authenticity determination scheme will be necessary as more and more primary documents are digitized
Searching Twitter
The New York Times had a good article this week about Twitter's approach to innovation. At this point, the concept of the "democratization of innovation" is probably familiar to most people who know a thing or two about the interwebs, but Twitter was pretty exceptional in its adherence to a credo of bottom-up innovation. At the outset, the service afforded only the ability to write "micro-blogs" of 140 characters. Methods of searching these micro-blogs developed organically by twitterers themselves. Nearly every search convention:
1. The '@' before the screen name of another user
2. The letters 'RT' before a reproduction of something that already been tweeted (or a "re-tweet")
3. The '#' used before a certain topic so that topics with more than one word were easily searchable (e.g., #iranelection)
4. Even the term 'tweet' to refer to a single micro-blog
...all of these were developed by twitter users and later adopted as standard form by Twitter developers. Now that Microsoft and Google are in on the game....sigh. So much for democracy.
Admittedly, it's too early to tell what kind of affect this new way to search Twitter is going to have on things like the use of the hash tag, but it's pretty easy to see why Microsoft and Google want to get in on the Twitter action. Back in March, on the Tech Crunch blog, Michael Arrington wrote a post about just how useful Twitter is to businesses. Because tweets are so short, many people (and I mean MANY) use Twitter to do nothing more than gripe about things, especially their experiences with products. One can only imagine how useful this information is to businesses. In fact this information is SO useful, that Twitter, itself, has been valued at nearly $1 Billion, even though it has around 35 million users (compared to Facebook's 175 million). With that kind of valuation, it's surprising that Twitter hasn't decided to cash in and (like Facebook) allow more ad space. Since Twitter is so amenable to being searched, however, they may never have to get into the ad game in order to make serious money. Here's to hoping that Twitter holds-on to its original democratic, bottom-up approach to innovation despite all of the corporate interest.
Averting a Digital Katrina: Sustaining Trust in the Research Infrastructure (EDUCAUSE Review) | EDUCAUSE
Amy Friedlander, Director of Programs at the Council on Library and Information Resources, wrote a very helpful and insightful article on the issuing of sustaining trust in the research infrastructure. The article was published in the Educause Review website.
Friedlander starts off the article by explaining her Katrina anaology. When the Hurricane Katrina stormed over New Orleans and the levees broke, the victims placed a huge trust on the infrastructures to step in and help them out. However, the infrastructure themselves had broken down: social, engineering, and political, thus creating this huge Catch-22 with huge repercussions that are still felt today.
She states that trust in an infrastructure stems from familiarity of a system; when things are running along as expected we build trust on its reliability, which explains in part why traditional scholarly publication remains robust even with the development of other methods of communication are evolving. Friedlander quotes Christine Borgman in identifying the three major functions that traditional scholarly communication achieves:
1) legitimizes scholarly work
2) disseminates that work to an audience (or several)
3) provides access, preservation and citation
The writer and reader share the same expectations and the repeated "success" of such a system causes scholars to cling to the older methods of scholarly communication. Plus, the long tradition of building prestige from having your work published in a famous or prestigious journal is still in effect.
However, the electronic publishing format is going happen regardless and Friedlander points out some of inconsistencies and examples that display a lack of trust in such systems and then proceeds to explain why this is rightly so. Scientists rarely participate in social-networking or create blogs or wikis of their research, namely data, partly because there is a huge confusion of how to correctly cite e-works in a paper. Print is seen as more authentic and is easier to cite, to boot.
Friedlander spends some time discussing the issue of citation: print articles and journals are a form of efficiency in citation and also uphold the core values of research:
1) attribution and credit
2) reliability
3) persistance
4) validation of sources
5) integrity of audience
4) replication of results
She states that very interactivity and the dynamism of the digital medium undermine the core values of trust in scholarly research. For example, the ability to download data in order to verify the experiment means inconsistency in how the data is displayed, which introduces unexpected results in the infrastructure which is hardly conducive for one to relax in the cyberinfrastructure.
Friedlander ends the article by saying that "retreating to analog is hardly the answer," nor is giving up the standard methods of interpretation, however in building up trust in the cyberinfrastructure requires long-term management and preservation of the digital data along with the policies and methods that encourage discovery, community, and discussion is the key.
Friedlander's article echoed many of the sentiments that were discussed in the class readings in struggling to define where trust comes from and how the traditional methods of scholarly publication are still widely preferred because of familiarity and a well-established expectation of be able to gain scholarly prestige. She did not really come up with any solutions to the issue at hand, similar to the other articles. This issue is a philosophical one and will remain under discussion for quite some time because I feel that scholarly communication is under a huge transistion and will happen in spite of itself with issues more or less resolving itself in response to problems that may or may not occur.
The Information Bottleneck

In his article Institutional Repositories and Research Data Curation in a Distributed Environment, Michael Witt discusses the lack of a framework for organizing information digitally and the effect that has had on datasets.
Witt begins by talking about scientific research of the past and how carefully lab notebooks were kept by scientists and later preserved in archives as part of the scientific record. While I have trouble believing this was always the case, I would agree that changes in technology have changed the way the records are being kept by scientists.
The bulk of his article is spent talking about the existing repository infrastructure of the Purdue Libraries, but I found his ideas about an information bottleneck more interesting. Witt believes that the original data is narrowed for use within the scope of a particular article. This new data is all that most people ever see of the original data. Assuming that the long tail holds true, then that data is valuable for its many different future uses. Unfortunately, the nature of the information bottleneck has left the data stripped of much of its original information and perhaps also of its future value.
Witt insists that datasets need to be presented in context to remain meaningful and useful. And this seems to me a fairly obvious idea. So why aren't we getting the raw data into repositories along with the narrow-focus versions of the data? Based on some of the other readings that I have done, I can only suggest it is because we are still having trouble getting even the narrowly focused data into repositories.
Witt does not address how this is to be done, he just suggests that "...at some point in the future, the process and units of scholarly communication may be reconsidered to fully recognize and include research datasets. In some cases, such as the Human Genome Project, the value of a genome dataset itself is generally recognized to be greater than any single, published finding resulting from its analysis."
I agree with his sentiment, and can imagine many uses for such a collection of datasets. But while we are still struggling to have truly useful repositories, it seems like an idea that will have to remain in the future.
Current Open Access Income Models from the Scholarly Publishing and Academic Resources Coalition (SPARC)
This is a lengthy document which describes in detail every tactic currently in use by open-access journals for financial support. SPARC compiled this overview to encourage existing journals who are thinking of becoming open-access as well as anybody who is thinking of starting an open-access journal from scratch. SPARC's definition of open-access is: "free and immediate online access to peer-reviewed journal literature".
Aside from wanting to participate in open scholarship, there are a few reasons why a journal might consider an open-access model for business purposes, these are those reasons according to SPARC:
• to increase access to its published research by lowering or eliminating market barriers to the content;
• to maximize market reach and support a new journal launch when the market will not support a traditional subscription model; or
• to implement a supply-side model (discussed below) in response to funder-mandated content deposit policies.
Legitimate advertising and/or sponsorship is one model for income for journals, these streams of revenue vary from Google Adsense (Open Government) to long-term corporate partnership (CERN Courier). I was particularly interested in this revenue source after reading about Elsevier's sheisty "sponsorship" so I looked at the examples that SPARC provides to see where the money is coming from. It turns out that for all of Elsevier's talk about open-access journals not being practical, they do sponsor at least one open-access journal. That journal is the The Journal of Electronic Publishing. JEP is interesting, these are the rest of their sponsors:
LexisNexis| O'Reilly Media| Newsbank Readex| Aptara| Wiley-Blackwell
JEP has a bunch of articles about open-access, one of them is Two Scenarios for How Scholarly Publishers Could Change Their Business Model to Open Access. How do they keep their corporate sponsors?
Corporate funding is probably the most problematic approach to revenue and SPARC provides only one example (Evidence-based Complementary and Alternative Medicine). In order for this model to work, the journal has to have an underwriting policy in order to prevent something like the Elsevier/MERCK situation from happening.
10% of open-access journals use the supply-side model of article processing fees for revenue, this is where authors pay the fee to publish in a journal. Most of these fees are subsidized by the author's institution or through a grant, only 5% of the fees are paid for by authors out-of-pocket. Some journals give authors the option to pay the fee to make their article open (BMJ Journals Unlocked).
PLoS is one journal that uses the article processing fee for income, but they do it slightly differently. Their model is based on an institutional membership fee which covers all of the papers submitted through an institution. PLoS even encourages consortial memberships to offset costs.
So far I've focused on some of the supply-side models for income. In terms of demand-side revenue, these are the models currently in place:
Versioning - Either selling an aggregate print copy of all issues at the end of the year or a print copy of each issue with additional non-research content throughout the year, it may also be considered the "special print" model.
Use-trigger Fees - In this model, the income is from institutions who use the journal heavily. There is a threshold of use, so most people will have access but when use is heavy from a single institution the fees are triggered. SPARC provides only one example of a journal who uses this model - Anthropological Index Online.
Convenience-Format License - This is what's currently happening with law journals. While the journals may be open access (Open Access Law Journals), they have an additional license with LexisNexis or WestLaw.
Value-Added Fee-Based Services - This is something I assume we're all familiar with, the web is full of sites like this.
Contextual E-Commerce - When a journal also has an e-commerce site, PLoS does this to a certain extent.
Tuesday, October 27, 2009
Digitizing Books: A Multi-Step Process
Digitization Primer
Sutherland provides a quick breakdown of mass digitization for books:
- Page Pictures
- Raw OCR
- Corrected OCR
- Semantic Coding
I'm surprised that Sutherland doesn't mention digitization of other parts of the book like the spine, headband, jacket, binding, endpapers and so on that would fall into this section. There must be some issues regarding the best way to proceed digitizing these, but I suppose the interested audience is not substantial for this article.
Shortcoming to page pictures are fairly obvious; there's nothing new discussed here, but Sutherland does point the wide range in quality page pictures encompass, Google marks the lower end of that spectrum.
Raw OCR is typically included as a next step after actual photographing of a page. Sutherland states that "commercial OCR of modem [sic] printed materials is now virtually perfect." She uses that typo to point out some of the shortcomings of OCR with older texts with their odd fonts, partial punctuation spacing, deteriorated paper and so on.
Corrected OCR fixes these problems by ensuring fidelity to the printed text. Humans can do this, but the labor is expensive. It's even expensive when it's cheap: double-key entry from workers in foreign countries is still costly for the amount text and books corrected.
Alternatives exist, and as Meg blogged a few weeks ago Recaptcha is one such solution. This service uses the ambiguous words an OCR device encounters as human-verification criteria for various sites. As users type the words they see in the image Recaptcha provides, they also provide a +1 for their reading of that word. Ultimately, this resolves the OCR's ambiguity. Recaptcha has important shortcomings though. It can't present words with old orthography or spacing, it will not present words the OCR device did not tag as ambiguous or words which it simply skipped on account of illegibility, and word context is stripped.
Sutherland's own Distributed Proofreaders is another solution. As expected accuracy is high, output is relatively low.
Semantic coding is another topic we've covered in class, particularly in regard to XML's and JSON's appropriateness for such marking. The semantic encoding mentioned here is fairly pedestrian as such things go: chapter headings or is Washington a city, state or person?, etc., but Sutherland leaves space for more interesting semantic markup in the future.
Sutherland concludes that digitization should be seen as a multistep process. I agree 100%, and I think it's clear Google similarly conceives of digitization as iterative, with it's "Getting there!" disposition toward accurate data and metadata (compared to OCA's metadata which is on the whole much much better). The slippery slope to this approach is that every step is a little less accomplished than it could be; raw OCR quality is contingent on page picture quality, corrected OCR needs somewhat accurate raw OCR for context. With that in mind I wonder if it would not be better to more closely consider first steps rather than applying so much functionality and completeness to later steps.
Collaboration and Trust in Digital Preservation
▪ Bureaucratic collaborations
▪ Leaderless collaborations
▪ Non-specialized collaborations
▪ Participatory collaborations
He suggests that data curation facilities are more likely to grow out of participatory collaboration efforts, which tend to be egalitarian in nature. His work also underscores one of the main themes we have been discussing all semester: that data curation projects are often created within disciplines or sub-disciplines and are confined to small communities that share data standards. Day acknowledged the importance of dealing with the social and legal implications of developing infrastructures for sharing data and collaborating. He takes up the subject of institutional repositories, but encourages collaboration across institutions and with other parties to increase preservation efforts. Day presents international initiatives to study and implement shared preservation environments.
Collaboration for digital preservation of scientific data requires a network of institutions that share trust. Day defines the relationship by stating that, “trust is at least partly about participants accepting a level of vulnerability in exchange for certain perceived benefits.” Trust often develops over time as institutions work together over a period of years and reach a level of expectation about the other’s reliability. Day relates trust to notions of control, describing the relationship as interdependent. Formal control mechanisms include policies, rules, and procedures, while informal controls can be simple shared institutional cultures. It is important to develop relationships of trust between institutions participating in digital curation initiatives. The trustworthiness of digital repositories is based on the strength of the institutions contributing data and participating in preservation efforts.
Finally, Day reviews various certification measures that may help establish a level of trust. He mentions the responsibilities set forth as the OAIS Reference model, discusses the RLG & OCLC published criteria for “trusted digital repositories,” and outlines the purpose of the RLG-NARA Digital Repository Certification Task Force.
The underlying theme of these initiatives is self-evaluation and definition. Repositories are encouraged to document policies, understand their own strengths and responsibilities, and consider more formal certification. The Digital Curation Centre has created a self-assessment toolkit, DCC Digital Repository Audit Method Based on Risk Assessment (DRAMBORA). The tool is designed to help institutions examine their strengths and weaknesses with an eye toward digital collaboration and risk management. The overarching theme of the article is that while collaboration is important and trust is necessary to adequately collaborate, an institution must thoroughly understand their contributions and roles in collaborating and must be aware of risks.
Renting Scholarship
Basically DeepDyve's plan is to offer any article in their database for rent, at various prices for various time limits. $1 can get you 1 day, but no printing allowed! $9.99 to $19.99 per month will get you more articles and longer time limits, and a special yearly package allows printing. Read Write Web says that printing isn't an issue "because most users are just looking at these articles for a few facts or a bibliography and don't need them for extended periods of time". DeepDyve calls these users "knowledge workers" who are people that "use the web for research for their education, health and careers". Hm. Firstly, I would think that if someone is doing research, she would actually want some article she downloaded for a little while, possibly printed out in order for later reference. Secondly, the term "knowledge worker" is silly, although I don't doubt that someday we could see the iSchool turn into the kwSchool. DeepDyve sees "knowledge workers" as professionals needing up-to-date research but not actually doing the research, I think. An untapped consumer base.
I understand that for the Average Joe or Jane, accessing articles must be a real pain, especially if one isn't associated with an institution with a library of any type. But does the Average Joe or Jane actually look for scientific articles about how "High Nucleotide Divergence in Developmental Regulatory Genes Contrasts With the Structural Elements of Olfactory Pathways in Caenorhabditis" (an article listed on DeepDyve's homepage)? I might guess not, unless a serious medical situation is at hand. Perhaps a researcher's institution is private without interlibrary loan and does not have a particular journal, then I could see a person needing this service. I'm trying to envision a nurse or doctor trying to figure out a medical problem or wants the latest information on a topic not in JAMA (or maybe I'm just envisioning House). Then DeepDyve could be quite useful.
The part of this new service that I find particularly intriguing, however, is the strange middle-ground this takes between traditional journal access and open access. I don't think this model flies in the face of the current paradigm of scholarly publishing, since the company is only offering the product of the publishing, not changing the process. But I do like how extremely technical scholarship can now leave the hallowed halls of academia and its libraries and can slum it with knowledge workers for just a dollar. Even if for only a day.
VIVOweb - Facebook for Scientists
According to the AP press release, presently scientists can use search engines and educated guessing to connect with others needed for their research. The new Facebook-style professional network would use "emerging" Semantic Web technology (hasn't the Semantic Web been supposedly emerging for years now?) to make information from institutions, academic journals, and researchers themselves available to scientists.
VIVOweb will be based on an open-source software developed by Cornell in 2003 called VIVO (which sounds like it should be an acronym, but actually isn't). At Cornell, VIVO is a tool that aids research and discovery by creating a single access point to find scholarly information. VIVO was created to "bring together in one site publicly available information on the people, departments, graduate fields, facilities, and other resources that collectively make up the research and scholarship environment in all disciplines." This technology uses the Vitro integrated ontology editor and semantic web application software, and was created in part to ease the mounting frustrations from Cornell faculty members who had a difficult time finding collaborators in different disciplines across campus because there was no simple, single place to go to look. The VIVO website boasts that users can search the VIVO website to find information about “anything related to academic and research pursuits at Cornell.”
VIVO has since branched out to other institutions.
The VIVOweb project that the NIH is funding for the sciences will use Cornell’s VIVO as a basis to "build a multi-institutional platform for the biomedical community." By creating this multi-institutional Facebook for scientists, VIVOweb will make it easier for scientists to extend their research networks, and to find existing projects or fellow scientists with similar research interests. Cornell will be in charge of developing the multi-dimensional aspect of the project, while the University of Florida will be focusing on creating and developing technology to keep the VIVOweb data current, and the University of Indiana will be responsible for developing the social networking tools that the scientists will need to make VIVOweb useful.
Trying to emulate something as successful and popular as Facebook, but for a more specific purpose, sounds like a pretty good idea. Based on what little I’ve read, VIVO seems like it has been successful and useful at Cornell, so it stands to reason that VIVOweb could be just as successful and helpful to biomedical scientists. I think anything that is designed to help link researchers together, to make it easier to find current projects, and to speed up research and development can only be a good thing – if it works. I’ve blogged about institutions pouring money into other new, developing cyberinfrastrcuture projects before; this is yet another example of having to wait and see what happens. In the meantime, I suppose scientists can always try to use Facebook to network (there are some science groups using it), if they can find the time to do so, in between tending their Facebook farms and cooking in their Facebook cafes.
Ease of Use Trumps Expertise AND Trustworthiness in 2001 study
Credible or not, and why? Image source.Any one metric for evaluating credibility probably contains fallacies. A combination of techniques may be the wisest course, but in spite of best efforts, people will do what they do.
A 2001 CHI report found that respondents ranked Ease of use above expertise and trustiworthiness in determining website credibility.
Ease of use is the second in a list of 5 main criteria, coming after Real-world feel, and before Expertise, Trustworthiness, and Tailoring (personalization). This actually makes some sense when considering the Real-world feel includes identifying information such as physical address being available on the site. This openness shows that the institution is not hiding at least this contact information, allowing for investigative possibilities.
The meaning the researchers give for Ease of use includes ability to search an "archive," which is a historically common way of evaluating authority (ie, how long have you known someone?) and professional design which is probably more problematic.
A more recent study from 2003, with one author in common from the first study, found that "Design look" was the TOP credibility issue. Appearing in 48.1% of comments compared to 14% that mentioned "Accuracy of information."
Any analysis is subjective, so perhaps the researchers were lax in their judgment of the comments, but I believe that website browsers really do put a great deal of faith into sites due to their usability.
One user commented, "The design is sloppy and looks like some adolescent boys in a garage threw this together."
Perhaps this low(er) importance of accuracy as a metric stems from the fact that users don't necessarily already know what is accurate, that is, after all, what they are judging.
One user wrote, "Most of the articles on this Web site seem to be headline
news that I have already heard, so they are believable." Seeming to be headline news may or may not be a better a metric for accuracy than the design look.
The authors do admit that due to the nature of their study, design look may have surfaced as more important because participants were not actually invested in their search, and so allowed themselves to rely more heavily on superficialities than they otherwise would have.
However, both studies contain results that point to 5 or 6 main ways that users decide whether or not to believe information on the web, and towards the top of both lists is design/use. Although this might not be the way we wish the world to work, it is wise to take it into consideration. We don't want to "Look childish."
In order of ranking from the 2003 study, these criteria were related to determining credibility either in a negative or positive way (I think which is which will be obvious):
Design Look 46.1%
Information Design/Structure 28.5%
Information Focus 25.1%
Company Motive 15.5%
Usefulness of Information 14.8%
Accuracy of Information 14.3% (!)
Name Recognition & Reputation 14.1% (!)
Advertising 13.8%
Bias of Information 11.6%
Tone of the Writing 9.0%
Identity of Site Sponsor 8.8%
Functionality of Site 8.6%
Customer Service 6.4%
Past Experience with Site 4.6%
Information Clarity 3.7%
Performance on a Test 3.6%
Readability 3.6%
Affiliations 3.4%
Sunday, October 25, 2009
They Did What?!?: Trust & Scholarly Publishing
Grant, Bob. "Merck Published Fake Journal" The Scientist.com. Posted April 30, 2009. Available at http://www.the-scientist.com/templates/trackable/display/blog.jsp?type=blog&o_url=blog/display/55671&id=55671 (version not requiring free registration: http://www.the-scientist.com/blog/print/55671/). (Accessed 10/25/2009).
Grant, Bob. "Elsevier Published 6 Fake Journals." The Scientist.com. Posted May 8, 2009. Available at http://www.the-scientist.com/templates/trackable/display/blog.jsp?type=blog&o_url=blog/display/55679&id=55679. (Accessed 10/25/2009)

As I was looking around on the ACRLog this week, I came across a blog entry from May of this year popped out at me given this week's readings on trust and authenticity. In a post titled "This Journal Brought to You By...", Barbara Fister writes of how journal publisher Elsevier received money from the pharmaceutical company Merck to publish a journal, the Australasian Journal of Bone and Joint Medicine, that was designed to appear like a scholarly, peer-review style journal. Fister's post in turn led me to the two articles written by Bob Grant for The Scientist.com that first exposed the ruse to the blogosphere.
The "journal" appears to have been disseminated to doctors in Australia and many of its contents were summaries or reprints of articles that were favorable toward Merck drugs. (The existence of the fake journal was first reported in the Australian in connection with a lawsuit by a man who suffered a heart attack after using Vioxx). Though it is unlikely that the publication would have fooled researchers, it may have hoodwinked some doctors. According to Grant, Elsevier has since apologized for the publication and has stated that they regret that the journal's sponsorship by Merck had not been made clear (nowhere in the issues Grant reviewed was this sponsorship acknowledged). A spokesman for Merck, Sharp & Dohme Australia (MDSA) explained to Grant that "MSDA understood that Elsevier envisaged the complimentary publication would draw on the vast resources of Elsevier, publishers of many leading peer-reviewed journals including Lancet, Bone, Joint Bone Spine and others, to deliver novel and timely full text articles and abstracts to physicians."
In the second of his articles, Grant uncovered that the Australasian Journal of Bone and Joint Medicine was not an aberration, but that Elsevier had in fact between 2000 and 2005 published five other such fake journals--the Australasian Journal of Neurology, the Australasian Journal of Cardiology, the Australasian Journal of Clinical Pharmacy, the Australasian Journal of Cardiovascular Medicine, and the Australasian Journal of Bone & Joint [Medicine]. (Other bloggers have suggested that the number may be even higher). These "journals" were published by Elsevier's Australian office under the Excerpta Medica imprint. At the time of Bob Grant's second article, Elsevier had declined to reveal which companies had sponsored these other publications.
From what I can tell--though none of the many blog posts I examined were perfectly clear on this point--the series of "sponsored article publications" (Elsevier's term) appeared either primarily or exclusively in print (analog) form. They were never indexed in Medline and they had no journal websites (though I'm unclear as to whether they ever appeared in Elsevier's online listing of publications). Grant's original article post includes a link to PDFs for Vol. 2, Issue 1 and Issue 2 of the Australasian Journal of Bone and Joint Medicine. The reason why I decided to write about what appears to have been a print phenomenon is that it also raises interesting questions regarding trust and authenticity in the digital world. I suspect that if these fake journals had had a real web presence, their status as crypto-advertising would have been exposed much more quickly. That said, both print and digital scholarly communication are affected by several shared trust and authenticity issues. The questions that, in this week's readings, Nancy Van House states are important for establishing a digital library's provenance--"Who designed it, and for what purpose? What design choices were made? Including, what is not visible to the user?"--are precisely those questions whose answers were intentionally obscured in the Elsevier fiasco.
Moreover, it has been interesting to see how this revelation has impacted the debates surrounding Open Access and publisher gate-keeping. As a blogger at Small Gray Matters writes: "The bitter irony is that Elsevier, along with the other major academic publishers, have spent the last few years ceaselessly lobbying against the open access movement, on the grounds that open access journals can’t be trusted to maintain the high quality of peer review that the commercial publishers provide." Is scholarly communication (and, by extension, digital curation) too important a function to entrust to commercial entities, like Elsevier or Google, whose bottom line will always be financial?
The Ocean Observatories Initiative
The Ocean Observatories Initiative (OOI) is an infrastructure project that will collect and make available data from a network of ocean-based sensors that will measure various physical, chemical, geological, and biological variables in the ocean and sea floor. The data collected from hundreds of such sensors will be made available in near-real time, and openly accessible. As stated by Tim Cowles, the program director of Ocean Observing in the Consortium for Ocean Leadership, the data collected will belong not to the OOI, but to the people, whether they be professors, researchers, or the average citizen. The goal of the project is to "address a multitude of important science and societal question, including those centering around climate change, ecosystem health, ocean acidification and carbon cycling." Saturday, October 24, 2009
Spezify
For a visual thinker that loves to scour everything from Twitter to YouTube to blogs in search of the latest from galleries, artists, conferences, scholars, designers...Spezify pools and presents information in a visual collage format. It also shows you related search terms (which may or may not be relevant - i.e. an important interntational conference for Contemporary Asian Art is going on right now - so the current search of that term brought up "hotel" as a related term.
Like Twitter it also shows you recent "hot" search terms. Spezify displays the information in a giant visual wall that you can scroll in all dimensions. Via Twitter, Cooliris told me that there is a way that I can d/l their program even though I have a PowerPC Mac mini. This weekend I will do so and see if Spezify's search results can be navigated on the visual wall of Cooliris. This would make for faster and improved scanning functionality.
Wednesday, October 21, 2009
Technology in the Arts compares Virtual Gallery software
A free report put out by Technology in the Arts takes a comparative look at three popular pieces of Virtual Gallery software: Virtual Galleries, Image Armada, and SceneCaster/3D Scenes. Such software can be used by curators and artists wishing to create a 3D gallery environment that is either online or otherwise digitally portable. Uses include serving as a design tool for curators or students planning installations - to serving as an access tool for those unable to physically explore the gallery space.
The report graded each application according to the following criteria: "ease of creating an exhibition; quality of images; flexibility of creating an exhibition; ease of navigation in the three-dimensional space; and ease of publishing or distributing the gallery." Comparative information was also provided re: the following technical information: "Mac and PC compatibility, hardware requirements, internet-based viewing, and image and lighting manipulation."
One of the most useful aspects of the report were the recommendations - whether your organization has a limited budget, requires Mac compatibility, and whether the gallery needs to be shared with people online. These concerns seemed to be the make or break features, and these recommendations were included on the next to the last page.
In addition to introducing me to some fascinating programs for the other form of "digital curation" the report gave a succinct presentation of software comparative review. My favorite, though expensive ($3750-$6750) would have to be Virtual Gallerie as it combines walk-through 3-D functionality with workability in a web browser.
The possibilities for having portable and interactive online exhibitions of artistic, historic or curated archival collections - in a 3-D manner, with perhaps abilities to zoom in and read documents and rotate/analyze 3-D objects - is fascinating. At a glance it appears similar to Second Life. I posit that if museum and gallery can create exhibits (which communicate narratives and encourage innovative thinking and evaluation of art and history) that are digitally portable
and preservable - I find the prospects exciting.
More Google. Wave (sorry)
Did you get a Google Wave invite? Yeah well, who cares. Some people say it will change everything; some people say it's all hype. Whatever you think, it's kinda fun to think about what Wave could do for science and research. Cameron Neylon wrote a short article about Wave and science for Nature and wrote quite a bit more about it on his blog "Science in the open." He suggests things like a Creative Commons bot that stamps a document and asks contributors if they're ok with the license. Or how about a bot that registers a document with a repository and automatically sends metadata to the repository? Data can be linked from within a Wave to a live spreadsheet that is being constantly updated. Neylon also makes some interesting suggestions for using the playback feature in the peer review process. Neylon is even sponsoring a Google Wave Science Hack Day where a bunch of people will get together, in person and in Wave, to discuss scientific applications of Wave. I've noticed people talking about the implications for interactions between doctors and patients. So what do you think? Are Wave's implications for research and scholarly communication interesting or not?
Netherlands Coalition for Digital Preservation
Google Books and the Digital Curation as Preservation Myth
This is far from the only egregious inaccuracy in the article. Brin makes the (now common) argument for digitization as preservation by referring to incidents of flood and fire after which hundreds of books were lost including a 1978 flood at the Stanford University Library (Brin's alma mater). Brin cleverly claims that one could have read about the reports of loss due to this flood in the Stanford-Lockheed Meyer Library Flood Report, but this report is also no longer available. A NYT commentator, however, managed to find four copies of this very report on WorldCat.
Brin really hammers-in the argument for digitization as preservation, even implicitly comparing Google Books to the Library of Alexandria which was deliberately burned down on three separate occasions. But what is keeping Google's serves from being destroyed in similar catastrophes? What happens to all of this heavily proprietary content when Google's stock prices are in the gutter? What happens when an electronic terrorist hacks into those servers and wipes everything out? The myth of digital curation as preservation in the case of Google Books rests on two pretty precarious presumptions: 1) that the digital content is invulnerable, and 2) that Google (which has been around for 15 or so years) is immortal. Are we really that naive? It's not like Google is putting everything into the Granite Mountain Records Vault.
Rather than focus on the fairy tale of digitization as preservation, I think Google has a better chance of selling the accessibility argument. That argument will be much more convincing, of course, if Google Books improves its metadata.
Digital Collection Building: Not Just for Digital Libraries!
Intrinsic to the argument is the postulate that to build a digital collection effectively and to build a traditional print collection effectively one must make a significant investment of roughly equal size. Furthermore, Nicholas Joint (the author) assumes that libraries will not necessarily possess the budget to engage fully in both activities. Thus, Joint presents the dilemma facing librarians as a choice between embracing a digital future and focusing on digital collection building or maintaining current practices and continuing to purchase paper books.
Joint's arguments come in two phases. First, he cites some usage statistics from Scotland. Apparently, the circulation rate in Scotland is only about fifty percent of the collection size, which is actually really good as far as academic libraries go. (As someone who has traveled to Edinburgh in February, I can attest that there is indeed probably a lot of reading go on indoors.) However, the same libraries found that each digital item in the collection was downloaded twenty times within a school year. While there are some validity issues with the statistics, namely that the university only possessed three thousand presumably popular digital titles, usage statistics seem to suggest that even conventional libraries may wish to devote greater resources to digital collection building than print acquisitions.
The second phase of Joint's argument is theoretical. In the article, he compares to works on digital publishing and the emergence of the ebook as a medium. One, Luke Tredinnick's Digital Information Culture: the Individual and Society in the Digital Age essentially argues that a new digital culture has either already emerged or is emerging and that not only is resistance futile, but that it will be doomed to appear to future generations as a socially-conservative backlash akin to fears of erosion of values wrought by tv. Conversely, in Never Mind the Web, Here Comes the Book, Miha Kovač argues that everyone is misinterpreting the term ebook. All books nowadays start in electronic or digital form and yet get transformed to print objects because that is how users prefer to engage them. As such, print collections will continue to be supreme according to Kovac.
Ultimately, Joint decides that Tredinnick possesses the stronger argument. Joint views print-outs as little more than useful transitions of essentially digital books. He points out that no one donates printouts of ejournals to libraries. For this reason, as well as the statistical one, Joint recommends that libraries devote more of their resources to building digital collections than traditional print collections.
I found the argument interesting primarily as an example that digital curation is not limited solely to digital libraries such as Valley of the Shadow, nor even to electronic publishing conglomerates such as JSTOR, but is also of increasingly crucial importance to traditional librarians as well.
Polymath Project
The project was successful and in just over a month, the problem was solved. There were many contributors from all aspects of the mathematical community, such as an award winning researcher and a high school teacher. Gowers and Nielson liked the way their experiment differed from other team projects in the sciences, since in those teams "work is usually divided up in a static, hierarchical way. In the Polymath Project, everything was out in the open, so anybody could potentially contribute to any aspect."
Farsightedly, Nielson and Gowers realize this situation raises authorship and preservation issues. Who gets credit for solving the DHJ theorem? The authors of this article will provide a link to the blog as a list of collaborators, so that someone who didn't provide insightful comments won't get the same credit as another who made major advances. Of course, this brings preservation into it, since the blog might not always be a permanent link. Nielson an dGowers mention how the Library of Congress saves certain legal blogs and hope to emulate that somehow. I think the issue is more complex than that, and more investigation should be done on how other web sites are preserved.
The articles notes that "outside mathematics, open-source approaches have only slowly been adopted by scientists" but the authors give several examples for how other sciences could benefit from this kind of collaboration and online data sharing. Almost summing up what we've learned in class, Nielson and Gowers realize "the widespread adoption of such open-source techniques will require significant cultural changes in science." I'm not sure how a blog can be considered that revolutionary as 'open-source' , but opening up a process that is usually confined to classrooms and chalkboards is a step in the right direction. It's nice to see people within the community noting the problem and challenging their colleagues.
Furthering the Open Access Movement
• Choose not to submit scholarly journal articles or other works to publications owned by for-profit firms.
• Say no, when asked to undertake peer-review work on a book or article manuscript that has been submitted for publication by a for-profit publisher or a journal under the control of a commercial publisher.
• Do not seek or accept the editorship of a journal owned or under the control of a commercial publisher.
• Do not take on the role of series editor for a book series being published by a for-profit publisher.
• Turn down invitations to join the editorial boards of commercially published journals or book series.
In effect, he is advocating a boycott against for-profit journals. In theory this sounds good, but I have trouble seeing it play out well in practice. AS we have discussed many times, many of the people writing articles are doing so in part to fulfill tenure requirements. Those requirements include publishing to the top journals in the field which are almost exclusively for profit. While the idea behind a boycott is noble, it seems hard to believe that researchers would be willing to give up publishing in the top journals.
Stevan Harnad wrote a blog post in response to Jackson saying just that. But more interestingly, he explained that just such a boycott was threatened in 2000 among biological researchers. As expected, the journals ignored the threat and the biologists never followed through with the boycott. That's when they decided to try something new. They launched PLoS, an alternative open-access journal for the field.
PLoS is now one of the great success stories of open access, becoming known for its quality and becoming a leading journal.
But may other open-access journals have not been so successful. Harnad suggests the solution to open-access lies in the institutional repository mandating of all published works of the researchers at an institution being placed into the repository.
Again, I have my doubts that this will work as institutional repositories do not have a good track record. Even when researchers are required to place articles in their repository, they don't always do it. And many for-profit journals will not allow a final version of the article to be placed in the repository, leaving the researcher forced to place preprints into the repository.
Clearly the system we have for journals is not working, but with few exceptions, institutional repositories are not the answer. Currently, it's pretty rare for an open-access journal to be considered as good as the top for-profit journals. We still haven't found an appropriate strategy to further the open access movement, only strategies that have had mixed results in the past. It seems that it is time for something new, but what form would it take?
Tuesday, October 20, 2009
How do policies help?
Also this week I read the short blog entry William Uricchio and the object/subject in participatory media, which is about a talk from William Uricchio last week. This a cursory look at participatory culture and how our culture has developed systems based on the participation of millions of people. Two of the examples Uricchio provides are SETI and iPhone apps. Systems of participation exist without models for how they should operate, instead these tools rely simply on what was possible when they were being created and based on what the needs of the users and developers were. This is very much the case for the cyberinfrastructure tools we've seen this semester.
One difference between the NSF document and the Uricchio talk is that the NSF is continually looking for a monolithic answer to cyberinfrastructure while Uricchio provides examples of working systems that could inspire the cyberinfrastructure. In particular, there's Photosynth, a tool from Microsoft that creates 3D "experiences" based on collective digital photography. It is definitely worth downloading Silverlight to see how the photos have been crowdsourced and curated in a meaninful way. About this example, the blog author says "Rather than building a space by building a model, the model emerges from the production of thousands of amateur photographers". This is certainly something I've been thinking about in terms of the cyberinfrastructure - before deciding what it will be and what it should be called, the cyberinfrastrcuture should emerge as a production of collaborative effort. After all, there is no working model for the cyberinfrastructure.
After all the discussion in class about building the cyberinfrastructure and reading about the alledged need for policies that will shape the cyberinfrastructure, I'm a bit skeptical that focusing on the end result of policy is helpful. Any work that's done on developing policy may be premature until there are more applicable cyberinfrastructure tools. Also, I'm afraid of what kinds of developments could potentially be stifled by inadequate policies. And again, there is no "one-size-fits-all" answer to all digital curation projects, these projects arise based on very specific needs of the people who use the data.
The one policy issue that I think will be the most difficult to tame long-term is copyright. I have been thinking lately that the best way to shape these policies is to simply participate in fair use activities. My hope is that if enough people are asserting their right to fair use, policies will be developed and strengthened that reflect those behaviors.
Ultimately, I wonder what policies will actually do for cyberinfrastructure. Why do the people involved want these policies? Do they simply desire a map to follow because that is the model they are familiar with?
History of the Old and New
In the non-open access side someone posted to the insider about a historical document site called Footnote, I did not immediately investigate but I put it on the back burner to check out later. Then a few minuets later when I was reading a blog Footnote came up yet again. I felt that I must look into it.
Footnote is a semi-free, (read mostly pay) historical document image dataset. There are all sorts of interesting things to explore, News Papers, Texas Death records, Records from WWII, the founding of Liberia and more. Some 50 million images or so the site claims.
The viewer is really nice. The images are clear (better than Google books), the zoom is good and the interface is pretty easy. The ability to annotate images is very interesting but I am not sure how that works; if the annotations are just for you or for everyone. Users can create connections between documents and images that they upload. Did I mention that users can up load images? You can upload images and create pages about people or places or events. Interesting not all of these items are “historic” in the proper term. These are very odd. You can also print copies of the documents. Which I did… There is yet more functionality that I have not had the time to explore. This has the potential to be a very powerful cultural heritage tool.
Now here are my caveats. The free version is highly restrictive. It seems like ever document I tried to query on came up with a premium page notice. This is kind of unfortunate because I really wanted a chance to explore the scope of documentation. Many of the documents come from NARA or the LOC and I understand that what the site is selling is the “service” but still I would like the opportunity to play a bit more before I buy.
Another caveat is provenance. And I guess this is the age old question, how do you trust what you find on-line. I assume that there is some sort of community in place here that cleans the documents and ensures that the information is valid and accurate. Another question I have is how to ensure the contents move into the future.
This site is very interesting to me but I am not quite sure why. There is a mix here of the personal and the historical that you don’t usually see elsewhere. I really find it fascinating, it is an experiment with meaning and historical documentation that does not usually happen.
Research Remix
IssueLab is an open access archive and online publishing forum that gathers and shares nonprofit research about social issues. The mission of IssueLab is to “more effectively archive, distribute and promote the extensive and diverse body of work being produced by the third sector.” IssueLab aims to make it easier for nonprofits to manage and share their research with others.
On Monday, coinciding perfectly with the beginning of Open Access Week, the Research Remix was announced. This is a contest issued by IssueLab that calls for contestants to creatively reuse nonprofit research data to make videos that provide commentary on social issues. The goal of the IssueLab’s Research Remix is to get interested people to “apply their skills for social good, to base their social commentary in fact, and to learn about legally remixing and re-purposing existing content.”
At the IssueLab website there are over 300 (I believe the current count to be 335) Creative Commons openly licensed research reports archived that contestants can use when making their videos; these reports all have openly licensed video footage, images and music for video-makers to repurpose. This video contest “aims to engage working artists and digital media students” while at the same time “encouraging nonprofits to make their research more broadly available and usable through open licensing.”
The deadline for this competition is December 31, and one of the requirements is that the submitted contest videos must be openly licensed, which I think is a great requirement because fair is fair; after all, it wouldn't be right to use someone else's openly available materials, and then not let anyone else remix your own. All final videos will be made available to the organizations whose research was used in remixing.
I love the idea of the Research Remix and I think it will be something fun that inventive video-making types can really get involved with; as evidenced by the growing numbers of fan-made videos at websites like YouTube, people love to remix and repurpose things. By using openly licensed material, the Research Remix is a great way for people to make a statement about social issues while creatively expressing themselves legally. Legally is the key word in that sentence; much of the fan video remixing that is found online is illegal because users repurpose copyrighted material. I’m not sure how much this contest is really going to be encouraging for nonprofits to make their research more openly accessible, but I like that the thought is there and that an effort is being made.
That being said, I hope everyone is having a happy International Open Access Week.
Digital Repositories and Sustainability
Computer program proves Shakespeare didn't work alone, researchers claim - Times Online
and
http://news.yahoo.com/s/time/20091020/us_time/08599193097100
Following a blurb on Yahoo! news, I found this neat article from Times Online, in conjunction with a Time.com article, discussing how a distinguished Shakespearean authority, Sir Brian Vickers, used plagiarism detection software to determine authorship of the history play Edward III.
Some background concerning this literary mystery: there's been a mystery for over 400 years as to whether Shakespeare co-authored Edward III with Thomas Kyd. The earliest copies of the play are unattributed and was published anonymously in 1596. The possibility of Edward III being penned, in part, by Shakespeare has been debated back and forth and was completely ignored by mainstream publication until 1997.
The plagiarism software, called Pl@giarism and developed at the University of Maastricht, was created to help educators catch cases of plagiarism. Pl@giarism operates by comparing all documents in a specified directory or by scanning the internet for files or works that match a phrase from the document being scrutinized.
Sir Brian Vickers, in his study of Shakespeare and other period works, came up with the concept of a "lingusitic fingerprint," in other words, playwrights often use the same patterns of speech. Vickers argues that, while it is established that a group or society can share common words three word phrases, it is pretty unlikely for a match of metaphors and unusual parts of speech to come up unless it was written by the same person due to the concept of the linguistic fingerprint.
Vickers came up with 200 matches when he tested Edward III with Shakespeare's work published before 1596. The average match rate is 20 matches. The 200 matches that Vickers came up with, in terms of relevancy, to use a database term, is a pretty strong indication that the provenance of four scenes in Edward III can be given to Shakespeare. Vickers tested a series of "recycled" phrases such as "come in person hither," or "thou art thyself" and whole phrases like "lilies that fester smell worse than weeds," and sure enough, there were some closely related matches in Edward III!
The article ends with Vickers stating his hope that further computer developments can be made in order to help prove provenance of major anonymous works and that the English literature field will come to utilize such handy materials in their study of text.
In the light of what we've been talking about in class and how this ties in with digital curation, I thought that these two articles and the project it discussed, proved to be a compelling example of how the cyberinfrastructure can transform and add a new dimension to research in the humanities and essentially put to rest some stubborn mysteries that have plagued the English academia.
visualizing data on demand!
The National Geophysical Data Center makes available large amounts of data, divided into sub-sections based on terrain and source: land, marine, satellite, snow & ice, and solar-terrestial. The NGDC is part of the US Department of Commerce, the National Oceanic & Atmospheric Administration, the National Environmental Satellite, and the Data and Information Service. The goal of NGDC is “is to be the world's leading provider of geophysical and environmental data, information, and products.”
The NGDC seems to be a good model for digital curation, and for a repository that makes accessible the datasets in their holdings. NGDC currently contains over 400 databases (digital and analog) from many sources, and users of these data sets include researchers, private industry, government agencies, universities, as well as the public.
Perhaps the most interesting aspect of the NGDC is their “Visualizing Data” section. This area of the site makes available a selection of images based on datasets in the collection, and furthermore, “can produce custom images, on request, for many of our databases.” This really amazes me! That the NGDC is not only providing datasets (for free!) but also offering to work with users of the sets to create useful images derived from these sets makes it an excellent example of providing data, and the tools to visualize that data.