Showing posts with label metadata. Show all posts
Showing posts with label metadata. Show all posts

Tuesday, September 29, 2009

Who's Afraid of Google Scholar?

Jasco, Peter. "Newswire Analysis: Google Scholar's Ghost Authors, Lost Authors, and Other Problems." Library Journal (9/24/2009), available at http://www.libraryjournal.com/article/CA6698580.html (accessed 9/28/2009).

The latest critique of Google's metadata comes from Peter Jasco writing on Google Scholar for Library Journal. In his analysis, Jasco illustrates how ill-trained parsers and web crawlers are responsible for creating a variety of metadata errors (particularly as pertains to authorship) which in turn produce incorrect publication and citation counts.

Jasco begins by noting that whereas Google Books blamed the embarrassing metadata errors exposed by Geoffrey Nunberg on wonky data from libraries and publishers, they do not have this excuse for errors in Google Scholar since Google chose not to use the metadata offered them by publishers and instead rely on its crawlers and parsers. The result, Jasco asserts, is that Google's algorithms replace real authors with ersatz ones like "P Login" (here Google has misconstrued a login prompt on a publisher's page, "Please Login," as an article's author). Alternately, Google Scholar also inflates publication and citation rates by counting citations as records in author searches and by generating multiple master records for the same article. Jasco notes that a search for his own name as author (e.g. author:jasco) yielded 578 records, but 403 (or 79%) of these records were for articles that cite his papers, not articles he authored. Such problems are compounded when Google Scholar's figures are naively adopted by tools such as the Google Scholar Citation Count gadget or Publish or Perish (PoP) software.

Jasco's figure depicting the now-corrected "Password" as author problem.


I'm torn as to how to respond to Jasco's article. On the one hand, it is clear that Google Scholar's metadata is error-riddled. One can't help but wince at some of Jasco's examples: the Google parser identified as author names things such as headers, section titles, and publication information, including Methods (42,700 author records) Contents (25,200), Limited (234,000) and Ltd (452,000). On the other hand, any university or department that's willing to make tenure decisions based on a cursory search of Google Scholar has such deep-rooted problems that inaccuracies on the part of Google Scholar can only be the least of these. It is to be expected that an unwitting high school or college student might be led astray by mis-attributed authorship or incorrect publication dates in Google Books, but it's less understandable for professionals in the field to outsource their committee work to Google Scholar or (even more perplexing) for an individual to assume that Google Scholar knows more about how many articles s/he has published than s/he does. It is part of being a professional in the field to keep an up-to-date CV and be cognizant of the publications one has authored or co-authored. Citation indexes are admittedly a trickier issue since such information is not as easily compiled by individuals, however shouldn't we expect (demand?) that academics and administrators would vet any results Google Scholar yields?

Jasco is not optimistic about Google Scholar's future. He sees their refusal of publisher/indexer metadata in favor of their own under-trained crawlers and parsers as a "lethal mix of ignorance and arrogance." Though Google Scholar has corrected some widely publicized errors, Jasco asserts "[t]he parsers [themselves] have not improved much in the past five years despite much criticism."

In addition to the actively "malevolent" threats to cyberinfrasctructure posed by hackers, phishers, and other cybertransgressors that are discussed in the Workshop on Cyberinfrastructure for the Social and Behavioral Sciences Final Report, cyberinfrastructure is at risk from the perhaps even more damaging effects of incompetence, laziness, and insouciance--factors which undermine our confidence in cyberinfrastructures and the collections they support. There is plenty of blame to go around. Google needs to be more concerned about inaccuracies in its products. Perhaps the responsible thing would be for Google Scholar either to preface certain functionalities, such as citation counts, with clear "user beware" warnings or to suspend them until it can ensure that its records are more accurate (it would still yield record results for the user to vet and tabulate). However, it is also incumbent upon users to educate themselves better regarding the tools they are adopting for important tasks. After all, Google Scholar announces itself as "Beta" in its logo. Hiring, promotion, and funding decisions should not be based on numbers drawn from tools that are still very much works-in-progress simply because it's easier to get your numbers from Google Scholar than to crunch them yourselves.

Wednesday, September 16, 2009

Flickr launches Flickr Galleries

Sept. 14th, I found out on Mashable that Flickr launched a Galleries application: http://mashable.com/2009/09/14/flickr-galleries/. What interests me the most is that it differs strongly from the already existing "favorites" ability (which is more the accumulating of images that interest one than "curating" of images). Their Galleries limit one to only 18 images (pulled from available images on Flickr), and one may title and write about what one finds meaningful about them.

Please explore the Galleries that people have created here: http://www.flickr.com/galleries. One can comment on the gallery, but not rate it or add it (which I think would be great features). Two of my favorites are Moleskinerie and tea. I would love the ability to enter metadata on gallery collections as a whole - as part of cataloging a curated exhibit goes beyond each individual image to address the relationship the images have to each other and how they interact.

The architect that I code metadata in his image database for - has been explaining to me his vision for a metadata hierarchy or schema. His vision seems to be one that is personal, intuitive and based upon his own use of images as visual inspiration. It is fascinating to work through and help develop this kind of metadata functionality in a workable way. For example: He wants a category called "Activities" - in which would include courtyards, outdoor cafes, public squares, people. He would like a category "Skylines" which would include architectural facades against the sky.

Going back to the Galleries on Flickr - I am inspired by seeing the arrangements and limited selections selected by individuals, named brief and poetic titles. I find these to be creative and mental exercises that not only allow others a glimpse into how others see the world, but allow others the chance to craft and curate one's own expression through the pairing, juxtoposition, arranging and poetic guidance of images that speak to us.

Flickr user "nonac" has created numerous Galleries on Flickr and what I especially love about them is that they contain written text - the voice of the curator. The eyes have it.

I expect to follow great things on Flickr because of this new capability - but would like to see gallery-specific tagging in addition to the ability to favorite these collections. The potential to comment on and discuss these arrangements as well as the possibility for artists to collaborate via this application is very exciting.

Can this creative use of digital/social-tagging and curating be extended to other creative uses of digital content online? I find YouTube ripe with creative usability problems. Amazon.com allows one to create book lists. I am very interested in how one might move beyond the "accumulation" level of online content collection to the "curating" level - and how that might play out in other creative online platforms.