Feeds

Google crowdsources card index for 'humanity's last library'

Garbage in, garbage out

Boost IT visibility and business value

Google has responded to criticism of the quality of its books metadata - by inviting anyone to write anything they want. Before you read on, remember that Google Books could become the world's digital library by default - it's been called "the last library" - since nobody is likely to do the scanning ever again.

However, for researchers and scholars, a collection is only as good as its metadata - and the quality of the metadata at Google Books falls far short of any library in history. Last year Stanford linguist and columnist Geoffrey Nunberg writing in the Chronicle of Higher Education described the errors as "disastrous".

Nunberg found that potentially hundreds of thousands of books were misdated, with titles credited to authors before they were born. Google Books showed books from Victorian era discussing Jimi Hendrix, or the microprocessor, for example.

Freud had strong views on web browsers

Attribution errors commonly miscredited authors, with Madame Bovary credited to Henry James. And bizarre classification errors abound. A Mae West biography was filed under Religion, for example. Jane Eyre showed up under Love Stories, Architecture, and Antiques and Collectables. And on top of this mass of errors, was a superstructure of erroneous links. Google's "related books" rarely point to anything related.

In short, if this is humanity's last ever library, humanity's last ever scholars won't get very far with their research.

"Our reputation precedes us" - The Victorians discuss Jimi Hendrix

(We've also highlighted problems due to lack of care and attention at Google Books here.)

When Salon revisited Google Books earlier this month, things hadn't improved. And worse, the answer to 'garbage out' is 'more garbage in' - crowdsourcing.

A Google engineer called "SofiaF" now invites us to nominate books that are out of print. They're only suggestions, but given that none of us are as dumb as all of us, can we expect the quality of the metadata to improve? As with classification, knowing the copyright status of a work requires expertise, particularly the intricacies of territorial copyright. It's not something a helpful amateur with time on their hands can usefully do.

For Nunberg, Google's haste to complete the project is the problem - it prefers to get it finished, for competitive reasons, rather than devote expert resources to getting it right.

"People at Google are also saying, 'Let's crowdsource this,' but that is a stupid idea. You and I are both smart, knowledgeable people, but I wouldn't trust either of us to do the skilled work of cataloging a 1890 edition of Madame Bovary," Nunberg told Salon.

He suggests that Google devote more expert resources to the problem - which is expensive - and that librarians, who have up until now trusted Google Books to get it right, become more feisty and pro-active. ®

Related link

Google Book errors, illustrated [PDF, 1.6MB]

Build a business case: developing custom apps

More from The Register

next story
BBC goes offline in MASSIVE COCKUP: Stephen Fry partly muzzled
Auntie tight-lipped as major outage rolls on
iPad? More like iFAD: We reveal why Apple fell into IBM's arms
But never fear fanbois, you're still lapping up iPhones, Macs
Nadella: Apps must run on ALL WINDOWS – PCs, slabs and mobes
Phone egg, meet desktop chicken - your mother
White? Male? You work in tech? Let us guess ... Twitter? We KNEW it!
Grim diversity numbers dumped alongside Facebook earnings
Microsoft: We're making ONE TRUE WINDOWS to rule us all
Enterprise, Windows still power firm's shaky money-maker
HP, Microsoft prove it again: Big Business doesn't create jobs
SMEs get lip service - what they need is dinner at the Club
ITC: Seagate and LSI can infringe Realtek patents because Realtek isn't in the US
Land of the (get off scot) free, when it's a foreign owner
Dude, you're getting a Dell – with BITCOIN: IT giant slurps cryptocash
1. Buy PC with Bitcoin. 2. Mine more coins. 3. Goto step 1
There's NOTHING on TV in Europe – American video DOMINATES
Even France's mega subsidies don't stop US content onslaught
prev story

Whitepapers

Top three mobile application threats
Prevent sensitive data leakage over insecure channels or stolen mobile devices.
Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Top 8 considerations to enable and simplify mobility
In this whitepaper learn how to successfully add mobile capabilities simply and cost effectively.
Application security programs and practises
Follow a few strategies and your organization can gain the full benefits of open source and the cloud without compromising the security of your applications.
The Essential Guide to IT Transformation
ServiceNow discusses three IT transformations that can help CIO's automate IT services to transform IT and the enterprise.