Feeds

Google admits 'garbage in, garbage out' translation problem

Research supremo cops to Google Translate loop the loop

Internet Security Threat Report 2014

Google's ever-so-clever Google Translate service may be falling foul of a problem known to grizzled engineers across the globe: garbage in, garbage out.

The problem was discussed by Google's director of research, Peter Norvig at the Nasa Innovative Advanced Concepts conference at Stanford, California on Wednesday, in response to a question by an audience member.

Norvig admitted that Google was "aware" of a problem caused by some sites using Google's services to translate their body copy into another language to create a localized version of their site.

The problem with this cut-rate method (bare cupboards of out-of-work translators aside) is that if Google indexes this site, it may then factor the translation into the models it itself uses to train its own machine-translation engine.

This post-modern problem means that Google's machines may be training themselves on data generated by Google's machines, which means that rather than getting incrementally better with each new model, they just stagnate.

"It's not a big problem yet – it could get worse," Norvig said. "We mostly address it by judging the quality of a site. If you look good, we'll keep your examples; if you look sketchy we'll toss them out."

Google has already sought to make it difficult for spammers to pollute the web with poorly translated text by shutting down its Google Translate API. Norvig also disclosed an approach in which Google tried to fingerprint each translation through precise word and syntax choices that wouldn't be obvious to the reader, but would be obvious to Google's bots, but said the company had retired the scheme as it was not effective. ®

Top 5 reasons to deploy VMware with Tegile

More from The Register

next story
Netscape Navigator - the browser that started it all - turns 20
It was 20 years ago today, Marc Andreeesen taught the band to play
Sway: Microsoft's new Office app doesn't have an Undo function
Content aggregation, meet the workplace ... oh
Sign off my IT project or I’ll PHONE your MUM
Honestly, it’s a piece of piss
Return of the Jedi – Apache reclaims web server crown
.london, .hamburg and .公司 - that's .com in Chinese - storm the web server charts
NetWare sales revive in China thanks to that man Snowden
If it ain't Microsoft, it's in fashion behind the Great Firewall
Chrome 38's new HTML tag support makes fatties FIT and SKINNIER
First browser to protect networks' bandwith using official spec
prev story

Whitepapers

Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Why cloud backup?
Combining the latest advancements in disk-based backup with secure, integrated, cloud technologies offer organizations fast and assured recovery of their critical enterprise data.
Win a year’s supply of chocolate
There is no techie angle to this competition so we're not going to pretend there is, but everyone loves chocolate so who cares.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?
Intelligent flash storage arrays
Tegile Intelligent Storage Arrays with IntelliFlash helps IT boost storage utilization and effciency while delivering unmatched storage savings and performance.