I might not be the only one to write about Ethan Zuckerman's blog post this week called Google Books and big data, which is good for discussion, but bad for discussing anything else. Too bad. I got riled up.
This is what started it:
"We’d expect libraries to have information like the place of publication and author birthdates. But it’s harder than you think – some libraries don’t know when books originated and simply used '1899' as a placeholder for any book they didn’t know. You see mentions of 'Internet' in publications dated 1620 – that’s because the mentions were in 1997 from a journal that’s been published since 1620."
Um, libraries are at fault? Not Google, with their secretive scanning practices? I can't tell if he's paraphrasing Google Books' Matthew Gray here or musing himself, and I don't know if that makes a difference. The two were at IBM’s Transparent Text symposium, which is/was a 2-day affair this week with the tagline "Text is Data". Zuckerman and Gray spoke on a panel titled "Analyzing the Written Record" but only Zuckerman has an attached abstract.
Never fear, however - the solution is at hand. A person "can analyze language models from texts published in different years and then make intelligent guesses about when a book with bad or no metadata was published." Excellent! So a stereotypical college freshman, up late working on a paper due the next day, can perform some simple algorithmic tests to make sure the source he/she is citing is correct. Of course.
I think Zuckerman is trying to illustrate how Google's vast data sets can lead to some interesting investigations by comparing information. This could be very advantageous to someone studying language evolution or printing history - until the dates aren't correct. Or will there be enough examples that the amount of incorrect metadata will be at the end of the curve? But do we know how many books vary from their physical counterparts? Unfortunately, the example in the blog is an incomplete link, so I'm not exactly sure how it's supposed to work.
To give this author a break, he seems quite well-rounded, as his blog covers Africa, international topics, and technology AND he's a researcher at Harvard's Berkman Center for Internet and Society. With that pedigree I'm not sure I understand how a basic understanding of library cataloging procedures (WorldCat anyone?) was overlooked. And I'm very interested in seeing new research come from the datasets of text and for now Google is the only one with that set, flawed or not.
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.