Feeds

Google admits 'garbage in, garbage out' translation problem

Research supremo cops to Google Translate loop the loop

Build a business case: developing custom apps

Google's ever-so-clever Google Translate service may be falling foul of a problem known to grizzled engineers across the globe: garbage in, garbage out.

The problem was discussed by Google's director of research, Peter Norvig at the Nasa Innovative Advanced Concepts conference at Stanford, California on Wednesday, in response to a question by an audience member.

Norvig admitted that Google was "aware" of a problem caused by some sites using Google's services to translate their body copy into another language to create a localized version of their site.

The problem with this cut-rate method (bare cupboards of out-of-work translators aside) is that if Google indexes this site, it may then factor the translation into the models it itself uses to train its own machine-translation engine.

This post-modern problem means that Google's machines may be training themselves on data generated by Google's machines, which means that rather than getting incrementally better with each new model, they just stagnate.

"It's not a big problem yet – it could get worse," Norvig said. "We mostly address it by judging the quality of a site. If you look good, we'll keep your examples; if you look sketchy we'll toss them out."

Google has already sought to make it difficult for spammers to pollute the web with poorly translated text by shutting down its Google Translate API. Norvig also disclosed an approach in which Google tried to fingerprint each translation through precise word and syntax choices that wouldn't be obvious to the reader, but would be obvious to Google's bots, but said the company had retired the scheme as it was not effective. ®

5 things you didn’t know about cloud backup

More from The Register

next story
PEAK LANDFILL: Why tablet gloom is good news for Windows users
Sinofsky's hybrid strategy looks dafter than ever
Leaked Windows Phone 8.1 Update specs tease details of Nokia's next mobes
New screen sizes, dual SIMs, voice over LTE, and more
Fiendishly complex password app extension ships for iOS 8
Just slip it in, won't hurt a bit, 1Password makers urge devs
Mozilla keeps its Beard, hopes anti-gay marriage troubles are now over
Plenty on new CEO's todo list – starting with Firefox's slipping grasp
Apple: We'll unleash OS X Yosemite beta on the MASSES on 24 July
Starting today, regular fanbois will be guinea pigs, it tells Reg
Another day, another Firefox: Version 31 is upon us ALREADY
Web devs, Mozilla really wants you to like this one
Secure microkernel that uses maths to be 'bug free' goes open source
Hacker-repelling, drone-protecting code will soon be yours to tweak as you see fit
Cloudy CoreOS Linux distro declares itself production-ready
Lightweight, container-happy Linux gets first Stable release
prev story

Whitepapers

7 Elements of Radically Simple OS Migration
Avoid the typical headaches of OS migration during your next project by learning about 7 elements of radically simple OS migration.
Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Consolidation: The Foundation for IT Business Transformation
In this whitepaper learn how effective consolidation of IT and business resources can enable multiple, meaningful business benefits.
Solving today's distributed Big Data backup challenges
Enable IT efficiency and allow a firm to access and reuse corporate information for competitive advantage, ultimately changing business outcomes.
A new approach to endpoint data protection
What is the best way to ensure comprehensive visibility, management, and control of information on both company-owned and employee-owned devices?