Feeds

Google eyes filing cabinets

Paper files next for the great data hoover

Gartner critical capabilities for enterprise endpoint backup

Google has revealed plans to help convert the world's paper filing cabinets, in Tron-like fashion, into mere nodes in the great hive mind.

The firm will be using an optical character recognition program called Tesseract that was found gathering dust in Hewlett Packard's garage.

"In a nutshell, we are all about making information available to users, and when this information is in a paper document, OCR is the process by which we can convert the pages of this document into text that can then be used for indexing," Google uber techie Luc Vincent said on the firm's code blog today.

Once recognised as one of the three most accurate OCRs on the market, Tesseract had been out of action since 1995.

HP decided it was better out than in if it wasn't making any money and punted it to the Information Science Research Institute at the University of Las Vegas to have it restored for an open source release. The uni gave it to Google, where it was quickly assimilated.

The software has some limitations, Vincent said. Comparatively speaking, it's not that accurate any more, it will only read English, does not like multiple columns or fancy layouts, and baulks at greyscale and colour documents. But, he said it was better than any other open source OCR software.

"Google currently "reads" almost every web page in the world. Come help us read all the printed material as well!" the firm said in an advertisement for OCR engineers. ®

Secure remote control for conventional and virtual desktops

More from The Register

next story
Why has the web gone to hell? Market chaos and HUMAN NATURE
Tim Berners-Lee isn't happy, but we should be
Mozilla's 'Tiles' ads debut in new Firefox nightlies
You can try turning them off and on again
Microsoft boots 1,500 dodgy apps from the Windows Store
DEVELOPERS! DEVELOPERS! DEVELOPERS! Naughty, misleading developers!
'Stop dissing Google or quit': OK, I quit, says Code Club co-founder
And now a message from our sponsors: 'STFU or else'
Apple promises to lift Curse of the Drained iPhone 5 Battery
Have you tried turning it off and...? Never mind, here's a replacement
Uber, Lyft and cutting corners: The true face of the Sharing Economy
Casual labour and tired ideas = not really web-tastic
Linux turns 23 and Linus Torvalds celebrates as only he can
No, not with swearing, but by controlling the release cycle
prev story

Whitepapers

Gartner critical capabilities for enterprise endpoint backup
Learn why inSync received the highest overall rating from Druva and is the top choice for the mobile workforce.
Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Rethinking backup and recovery in the modern data center
Combining intelligence, operational analytics, and automation to enable efficient, data-driven IT organizations using the HP ABR approach.
Consolidation: The Foundation for IT Business Transformation
In this whitepaper learn how effective consolidation of IT and business resources can enable multiple, meaningful business benefits.
Next gen security for virtualised datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.