Showing posts with label Google. Show all posts
Showing posts with label Google. Show all posts

Thursday, November 19, 2009

Google Swirl

I am consistently excited by new developments in visual browsing and searching on the web. Google's new development to come out of Google Labs, Google Image Swirl is very provocative and is very close to something that I have been imagining a need for. It would be fascinating it the application could be employed on pre-curated collections of images (ie. to "Google Swirl" a collection of visual art, such as ARTstor). But first, what is Google Swirl?

As a bridging of Picasa Face Recognition and Similar Images, search results depend upon both image metadata and computer vision research. There are comparisons to Google's Wonder Wheel (which displays search results graphically) and Visual Thesaurus. One enters a search term and 12 groupings of images appear visualized as photo stacks. One chooses a particular image and the Flash experimental interface "swirls" to display your image and branches to numerous other images with varying degrees of relationship to that image.

Wednesday, October 7, 2009

Google ambitions ignore digital ephemera of yesteryear

Wired recently had an article critiquing Googles ambitious reach into digital curation of books with a reminder of the last large digital library it undertook and all but abandoned. Usenet.

First Google rescued post-1995 Usenet (a dial-up message board system founded in 1980) from Dejanews in 2001. Google later was able to add millions of posts to this from the magtape of an old Unix guru, Marc Spencer. This gave Google a digital library from over 2 decades of 700 million articles from 35,000 newsgroups.

Wired interviews the Unix guru whose archive supplies a bulk of Google groups (and can't seem to get his name right - is it Marc or Henry Spencer?). Speaking on behalf of his community of Usenet users, Mr. Spencer is very disappointed with the poor searchibility of the material he provided Google.

I would ask: Where are the finding aids? How are things catalogued? Does Google use any of the methods that have worked for traditional archives spanning hundreds of years? No, Google can't even retrieve the legendary alt.gothic flamewars of 1993.

Well, perhaps the world is a better place for that.

But in all seriousness, if anything required professional curation, Usenet could be that. Many from my generation compiled subject knowledge in the early 90s in the form of F.A.Qs, discographies, videographies, and many other encyclopediac collections of unpublished popular (or not very popular) material. Did we save it all? Probably not. There were news-groups for great swaths of information that probably did not carry over to the world wide web post 1995, or even post 2001.

Granted, an enormous amount of Usenet one would not want to have see the light of day, for reasons legal and sundry that could make Craigslist pale in comparison. Still, there are volumes of music, film, and sub-cultural anthropology I have in my file cabinets printed on dot-matrix printers in 1992 that I have yet to see mirrored online or in zines, magazines or books. As I get older and de-clutter, and deem certain material unnecessary or not keeping in with my current interests or values, that information will most likely be lost.

I think a great idea would be for there to be a way to feed and properly code Usenet material (i.e. Usenet group "such and such" date: 1996) and feed the archive anonymously into Wikipedia.

I think the anonymous Wiki framework would be the best platform to mine and add Usenet knowledge and curate it in a central way that would be searchable. Usenet entries would be date-logged as historical data, and contestation, edits and updates would need to follow the historical Usenet information. This would give Wiki a valuable historical layer that I currently find missing. There are times when Wiki feels to me like urban Las Vegas, where the old is torn down and forgotten to make room for the new. What is Wiki kept historical layers to their entries? I think it would be interesting to see a Wiki article from 7 years ago on a given topic.

Tuesday, September 29, 2009

Who's Afraid of Google Scholar?

Jasco, Peter. "Newswire Analysis: Google Scholar's Ghost Authors, Lost Authors, and Other Problems." Library Journal (9/24/2009), available at http://www.libraryjournal.com/article/CA6698580.html (accessed 9/28/2009).

The latest critique of Google's metadata comes from Peter Jasco writing on Google Scholar for Library Journal. In his analysis, Jasco illustrates how ill-trained parsers and web crawlers are responsible for creating a variety of metadata errors (particularly as pertains to authorship) which in turn produce incorrect publication and citation counts.

Jasco begins by noting that whereas Google Books blamed the embarrassing metadata errors exposed by Geoffrey Nunberg on wonky data from libraries and publishers, they do not have this excuse for errors in Google Scholar since Google chose not to use the metadata offered them by publishers and instead rely on its crawlers and parsers. The result, Jasco asserts, is that Google's algorithms replace real authors with ersatz ones like "P Login" (here Google has misconstrued a login prompt on a publisher's page, "Please Login," as an article's author). Alternately, Google Scholar also inflates publication and citation rates by counting citations as records in author searches and by generating multiple master records for the same article. Jasco notes that a search for his own name as author (e.g. author:jasco) yielded 578 records, but 403 (or 79%) of these records were for articles that cite his papers, not articles he authored. Such problems are compounded when Google Scholar's figures are naively adopted by tools such as the Google Scholar Citation Count gadget or Publish or Perish (PoP) software.

Jasco's figure depicting the now-corrected "Password" as author problem.


I'm torn as to how to respond to Jasco's article. On the one hand, it is clear that Google Scholar's metadata is error-riddled. One can't help but wince at some of Jasco's examples: the Google parser identified as author names things such as headers, section titles, and publication information, including Methods (42,700 author records) Contents (25,200), Limited (234,000) and Ltd (452,000). On the other hand, any university or department that's willing to make tenure decisions based on a cursory search of Google Scholar has such deep-rooted problems that inaccuracies on the part of Google Scholar can only be the least of these. It is to be expected that an unwitting high school or college student might be led astray by mis-attributed authorship or incorrect publication dates in Google Books, but it's less understandable for professionals in the field to outsource their committee work to Google Scholar or (even more perplexing) for an individual to assume that Google Scholar knows more about how many articles s/he has published than s/he does. It is part of being a professional in the field to keep an up-to-date CV and be cognizant of the publications one has authored or co-authored. Citation indexes are admittedly a trickier issue since such information is not as easily compiled by individuals, however shouldn't we expect (demand?) that academics and administrators would vet any results Google Scholar yields?

Jasco is not optimistic about Google Scholar's future. He sees their refusal of publisher/indexer metadata in favor of their own under-trained crawlers and parsers as a "lethal mix of ignorance and arrogance." Though Google Scholar has corrected some widely publicized errors, Jasco asserts "[t]he parsers [themselves] have not improved much in the past five years despite much criticism."

In addition to the actively "malevolent" threats to cyberinfrasctructure posed by hackers, phishers, and other cybertransgressors that are discussed in the Workshop on Cyberinfrastructure for the Social and Behavioral Sciences Final Report, cyberinfrastructure is at risk from the perhaps even more damaging effects of incompetence, laziness, and insouciance--factors which undermine our confidence in cyberinfrastructures and the collections they support. There is plenty of blame to go around. Google needs to be more concerned about inaccuracies in its products. Perhaps the responsible thing would be for Google Scholar either to preface certain functionalities, such as citation counts, with clear "user beware" warnings or to suspend them until it can ensure that its records are more accurate (it would still yield record results for the user to vet and tabulate). However, it is also incumbent upon users to educate themselves better regarding the tools they are adopting for important tasks. After all, Google Scholar announces itself as "Beta" in its logo. Hiring, promotion, and funding decisions should not be based on numbers drawn from tools that are still very much works-in-progress simply because it's easier to get your numbers from Google Scholar than to crunch them yourselves.