The latest critique of Google's metadata comes from Peter Jasco writing on Google Scholar for Library Journal. In his analysis, Jasco illustrates how ill-trained parsers and web crawlers are responsible for creating a variety of metadata errors (particularly as pertains to authorship) which in turn produce incorrect publication and citation counts.
Jasco begins by noting that whereas Google Books blamed the embarrassing metadata errors exposed by Geoffrey Nunberg on wonky data from libraries and publishers, they do not have this excuse for errors in Google Scholar since Google chose not to use the metadata offered them by publishers and instead rely on its crawlers and parsers. The result, Jasco asserts, is that Google's algorithms replace real authors with ersatz ones like "P Login" (here Google has misconstrued a login prompt on a publisher's page, "Please Login," as an article's author). Alternately, Google Scholar also inflates publication and citation rates by counting citations as records in author searches and by generating multiple master records for the same article. Jasco notes that a search for his own name as author (e.g. author:jasco) yielded 578 records, but 403 (or 79%) of these records were for articles that cite his papers, not articles he authored. Such problems are compounded when Google Scholar's figures are naively adopted by tools such as the Google Scholar Citation Count gadget or Publish or Perish (PoP) software.
I'm torn as to how to respond to Jasco's article. On the one hand, it is clear that Google Scholar's metadata is error-riddled. One can't help but wince at some of Jasco's examples: the Google parser identified as author names things such as headers, section titles, and publication information, including Methods (42,700 author records) Contents (25,200), Limited (234,000) and Ltd (452,000). On the other hand, any university or department that's willing to make tenure decisions based on a cursory search of Google Scholar has such deep-rooted problems that inaccuracies on the part of Google Scholar can only be the least of these. It is to be expected that an unwitting high school or college student might be led astray by mis-attributed authorship or incorrect publication dates in Google Books, but it's less understandable for professionals in the field to outsource their committee work to Google Scholar or (even more perplexing) for an individual to assume that Google Scholar knows more about how many articles s/he has published than s/he does. It is part of being a professional in the field to keep an up-to-date CV and be cognizant of the publications one has authored or co-authored. Citation indexes are admittedly a trickier issue since such information is not as easily compiled by individuals, however shouldn't we expect (demand?) that academics and administrators would vet any results Google Scholar yields?
Jasco is not optimistic about Google Scholar's future. He sees their refusal of publisher/indexer metadata in favor of their own under-trained crawlers and parsers as a "lethal mix of ignorance and arrogance." Though Google Scholar has corrected some widely publicized errors, Jasco asserts "[t]he parsers [themselves] have not improved much in the past five years despite much criticism."
In addition to the actively "malevolent" threats to cyberinfrasctructure posed by hackers, phishers, and other cybertransgressors that are discussed in the Workshop on Cyberinfrastructure for the Social and Behavioral Sciences Final Report, cyberinfrastructure is at risk from the perhaps even more damaging effects of incompetence, laziness, and insouciance--factors which undermine our confidence in cyberinfrastructures and the collections they support. There is plenty of blame to go around. Google needs to be more concerned about inaccuracies in its products. Perhaps the responsible thing would be for Google Scholar either to preface certain functionalities, such as citation counts, with clear "user beware" warnings or to suspend them until it can ensure that its records are more accurate (it would still yield record results for the user to vet and tabulate). However, it is also incumbent upon users to educate themselves better regarding the tools they are adopting for important tasks. After all, Google Scholar announces itself as "Beta" in its logo. Hiring, promotion, and funding decisions should not be based on numbers drawn from tools that are still very much works-in-progress simply because it's easier to get your numbers from Google Scholar than to crunch them yourselves.
