Feeds

IBM Boffins KNOW WHERE YOU LIVE, thanks to Twitter

"Woohoo I'm in Sydney" tells people you're in Sydney, it seems

Secure remote control for conventional and virtual desktops

If you thought refraining from geotagging your Tweets or photos was enough to keep your secrets from the world at large, think again: IBM researchers say a Twitter user's primary location can be inferred from their behaviour, with accuracy as high as 68 per cent.

In this paper at Arxiv, Jalal Mahmud, Jeffrey Nichols and Clemens Drews of IBM Research at Almaden say they can at least get city-level predictions of Twitter users' “home” locations (by which they mean the primary location from which an individual usually Tweets), even though the user isn't using Twitter's location features.

To do this, the researchers produced two algorithms. The first uses behaviours such as volume of Tweets from a user, and external information (a dictionary of location names and services such as Foursquare). They say that while this algorithm works best when users make “explicit references” of locations in Tweets, it “still works with reduced accuracy when no explicit references are available”.

The second algorithm predicts locations “hierarchically using time zone, state or geographic region as the first level and city at the second level”.

With a dataset of around 1.5 million Tweets from 9,551 users, the researchers then extracted classifiers including:

  • All words in the Tweets;
  • All hashtags in the Tweets; and
  • All city and state location names in the Tweets.

Armed with this data, the researchers then note, they can also make some assumptions about location – for example, given America's timezones, a user in New York is more likely to be at home at 7:00PM eastern time, while at the same time, a Californian user is probably still at work. That means a user's volume of Tweets helps become a hint to their location.

The paper notes that “geo-tags are not used in any of our prediction algorithms, although around 65 per cent of the tweets in our dataset are geo-tagged”.

But don't worry, the researchers only intend their work to be used for good: “a journalist tracking an event on Twitter may want to know which tweets are coming from users who are likely to be in a location of that event, vs. tweets coming from users who are likely to be far away. As another example, a retailer or a consumer products vendor may track trending opinions about their products and services and analyse differences across geographies.

“Second, our examination of the discriminative features used by our algorithms suggests strategies for users to employ if they wish to micro-blog publicly but not inadvertently reveal their location”, the study notes. ®

Secure remote control for conventional and virtual desktops

More from The Register

next story
Brit telcos warn Scots that voting Yes could lead to HEFTY bills
BT and Co: Independence vote likely to mean 'increased costs'
Phones 4u slips into administration after EE cuts ties with Brit mobe retailer
More than 5,500 jobs could be axed if rescue mission fails
New 'Cosmos' browser surfs the net by TXT alone
No data plan? No WiFi? No worries ... except sluggish download speed
Radio hams can encrypt, in emergencies, says Ofcom
Consultation promises new spectrum and hints at relaxed licence conditions
Blockbuster book lays out the first 20 years of the Smartphone Wars
Symbian's David Wood bares all. Not for the faint hearted
Bonking with Apple has POUNDED mobe operators' wallets
... into submission. Weve squeals, ditches payment plans
'Serious flaws in the Vertigan report' says broadband boffin
Report 'fails reality test' , is 'simply wrong' and offers ''convenient' justification for FTTN says Rod Tucker
This flashlight app requires: Your contacts list, identity, access to your camera...
Who us, dodgy? Vast majority of mobile apps fail privacy test
prev story

Whitepapers

Secure remote control for conventional and virtual desktops
Balancing user privacy and privileged access, in accordance with compliance frameworks and legislation. Evaluating any potential remote control choice.
Saudi Petroleum chooses Tegile storage solution
A storage solution that addresses company growth and performance for business-critical applications of caseware archive and search along with other key operational systems.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?
Security for virtualized datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.
Providing a secure and efficient Helpdesk
A single remote control platform for user support is be key to providing an efficient helpdesk. Retain full control over the way in which screen and keystroke data is transmitted.