Feeds

AOL publishes database of users' intentions

Your search history, right here

Internet Security Threat Report 2014

AOL Labs prompted a weekend of hyperventilation in the 'blogosphere' by publishing the search queries from 650,000 users. This mini-scandal may yet prove valuable, however, as it reveals an intriguing psychological study of the boundaries of what is considered acceptable privacy.

In his turgid book on Google - one so obsequious and unchallenging that Google bought thousands of copies to give away to its staff - former dot.bust publisher John Battelle enthused about something he called the "database of intentions". The information collected by search engines, he trumpeted, would be a marketer's dream, and tell us more about ourselves than we ever realized we could know. AOL's publication is the first general release of such a database to the public.

But hold on a minute. Is it, really?

AOL's data was anonymized, with user identification removed. The search logs contained 10.8m normalized queries from 658,086 unique users, collected between March 1 and 31 May this year, amounting to around a third of all queries made by its US users. The data has since been removed, but an AOL research paper which was built on the data can still be found, here [PDF, 228kb]. You may find it about as enlightening as similar studies we've covered before (see People more drunk at weekends, researchers discover).

Although the user's IDs were hidden, and didn't contain information on what the user actually clicked on, some argued that the data permitted personally identifiable information to be inferred from the query logs.

Something called TechCrunch, a weblog devoted to hyping its publisher's personal investments and companies created by his friends, explained how:

"The most serious problem is the fact that many people often search on their own name, or those of their friends and family, to see what information is available about them on the net. Combine these ego searches with porn queries and you have a serious embarrassment. Combine them with 'buy ecstasy' and you have evidence of a crime. Combine it with an address, social security number, etc., and you have an identity theft waiting to happen. The possibilities are endless."

Now, you may be thinking - that only serves people right for conducting vanity searches. But more seriously, there are dangers in following this line of reasoning.

It's not only individuals who "ego surf", it could be the individual's spouse, a member of their family, a colleague, or even their web stalker. (I've had a few).

Similarly, is the query "buy ecstasy" necessarily the intention of a raver, or tweaker? It might be a parent, a neighborhood watch scheme, or a promoter, keen to stamp out drug dealing at his venue before an event.

So the "database of intentions", then, turns out to be more more of "a database of inferences" - as reflective as it is of the inferrer as the web surfer.

And if, as TechCrunch weakly suggests, the act of typing "buy ecstasy" into a search is itself "evidence of a crime", then there will be a lot of happy policeman out there this evening, for whom the business of catching criminals has just been made a lot easier.

The" precogs" of Phillip K Dick's story Minority Report - who are able to predict crimes before they take place, thus allowing them to be prevented - will no longer be necessary. Plod will simply be able issue a pre-emptive warrant for a crime that never took place, on the basis of a user's Google results, no?

So that's one line of sloppy thinking dealt with. It ignores another, however.

Providing a secure and efficient Helpdesk

Next page: Leave No Trace

More from The Register

next story
Doctor Who's Flatline: Cool monsters, yes, but utterly limp subplots
We know what the Doctor does, stop going on about it already
Facebook, Apple: LADIES! Why not FREEZE your EGGS? It's on the company!
No biological clockwatching when you work in Silicon Valley
'Cowardly, venomous trolls' threatened with TWO-YEAR sentences for menacing posts
UK government: 'Taking a stand against a baying cyber-mob'
Happiness economics is bollocks. Oh, UK.gov just adopted it? Er ...
Opportunity doesn't knock; it costs us instead
The 'fun-nification' of computer education – good idea?
Compulsory code schools, luvvies love it, but what about Maths and Physics?
Ex-US Navy fighter pilot MIT prof: Drones beat humans - I should know
'Missy' Cummings on UAVs, smartcars and dying from boredom
Sysadmin with EBOLA? Gartner's issued advice to debug your biz
Start hoarding cleaning supplies, analyst firm says, and assume your team will scatter
Don't bother telling people if you lose their data, say Euro bods
You read that right – with the proviso that it's encrypted
prev story

Whitepapers

Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Why cloud backup?
Combining the latest advancements in disk-based backup with secure, integrated, cloud technologies offer organizations fast and assured recovery of their critical enterprise data.
Win a year’s supply of chocolate
There is no techie angle to this competition so we're not going to pretend there is, but everyone loves chocolate so who cares.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?
Intelligent flash storage arrays
Tegile Intelligent Storage Arrays with IntelliFlash helps IT boost storage utilization and effciency while delivering unmatched storage savings and performance.