Feeds

An open source Google - without the ads

But is it legal?

  • alert
  • submit to reddit

Intelligent flash storage arrays

With the hope of returning at least one corner of the web to its non-commercial roots, Google watcher Daniel Brandt, who curates the NameBase archive, has released the source code to a Google scraper. Brandt has been making an ad-free proxy available for two years using Google's little known minimal "ie" interface. By using this proxy, users bypass both Google's notorious "2038" cookie (that's when it expires) and the text ads.

Brandt fully expects Google to throw legal and technical resources at him, but says he welcomes the challenge if only to clarify copyright issues. Google took people's free stuff and made a $50 billion business from it, he argues.

"The commercialization of the web became possible only because tens of thousands of noncommercial sites made the web interesting in the first place," he writes. "All search engines should make a stable, bare-bones, ad-free, easy-to-scrape version of their results available for those who want to set up nonprofit repeaters. Even if it cuts into their ad profits slightly, there's no easier way to give back some of what they stole from us."

He explains in more detail in the source code: "Legally, Google probably has the right to block anyone they want. And legally, we believe that as a tiny nonprofit with an interest in Google's violations of privacy, we have the right to access Google's publicly-available data any way we want. If you want to argue about copyright, then let's start with the fact that Google scrapes billions of web pages and doesn't ask permission before making the cache copies available. Thiss craping is used as a carrier for the ads that make Google stinkin' rich.

"Now that, in our opinion, is an interesting copyright issue. As this is written, Google has a market cap of $55bn. This exceeds the market cap of General Motors and Ford combined. Google is probably the single largest information resource on the planet, and they're getting rich off of us. It's time for Google to give something back to the public sector."

The source code, which runs on Linux, asks the users only to use the program for non-commercial purposes.

"We think it would be splendid if scraping Google for nonprofit purposes, and stripping out their wretched advertising, was established someday as an acceptable, legal practice."

In the week since it launched, the source code has been downloaded about a hundred times a day says Brandt.

Google would rather you licensed its beta Web API. However, as Charles Ferguson writing in MIT Technology Review noted recently, the service is "laughably limited" to 1,000 queries a day, and offers little functionality; Google has let the offering languish.

You can find the code here [ZIP archive, 16kb], an explanation here and try out the proxy here. ®

Related stories

Google exposes web surveillance cams
Major flaw found in Google Desktop
Google News' chief robot speaks out
Gates: PC will replace TV, TV will become a giant Google
Google Desktop privacy branded 'unacceptable'

Internet Security Threat Report 2014

More from The Register

next story
The 'fun-nification' of computer education – good idea?
Compulsory code schools, luvvies love it, but what about Maths and Physics?
Facebook, Apple: LADIES! Why not FREEZE your EGGS? It's on the company!
No biological clockwatching when you work in Silicon Valley
Happiness economics is bollocks. Oh, UK.gov just adopted it? Er ...
Opportunity doesn't knock; it costs us instead
Ex-US Navy fighter pilot MIT prof: Drones beat humans - I should know
'Missy' Cummings on UAVs, smartcars and dying from boredom
Yes, yes, Steve Jobs. Look what I'VE done for you lately – Tim Cook
New iPhone biz baron points to Apple's (his) greatest successes
Lords take revenge on REVENGE PORN publishers
Jilted Johns and Jennies with busy fingers face two years inside
Sysadmin with EBOLA? Gartner's issued advice to debug your biz
Start hoarding cleaning supplies, analyst firm says, and assume your team will scatter
Edward who? GCHQ boss dodges Snowden topic during last speech
UK spies would rather 'walk' than do 'mass surveillance'
Doctor Who's Flatline: Cool monsters, yes, but utterly limp subplots
We know what the Doctor does, stop going on about it already
prev story

Whitepapers

Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Why and how to choose the right cloud vendor
The benefits of cloud-based storage in your processes. Eliminate onsite, disk-based backup and archiving in favor of cloud-based data protection.
Three 1TB solid state scorchers up for grabs
Big SSDs can be expensive but think big and think free because you could be the lucky winner of one of three 1TB Samsung SSD 840 EVO drives that we’re giving away worth over £300 apiece.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.
Security for virtualized datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.