Feeds

There's a tide of unstructured data coming - start swimming

Or you could just work out a plan...

Combat fraud and increase customer satisfaction

Whether you prefer to define the known size of our planet’s total digital universe in petabytes or even zettabytes, we can all agree the collective weight of data production is spiralling ever upwards.

While we focus on the relative merits of transactional versus analytical databases, the unstructured data that fails to fall within the general purview of either these systems is the rising tide beneath.

We are not just talking about non-textual audio, video and graphical data here. Unstructured data must also be thought of in its textual form of Word documents, emails, social media messages and other as yet undefined data shapes.

Different stakeholders view structured and unstructured differently. After all, in the world of video production it does not necessarily follow that all video data will be structured to those companies working with it.

Equally, textual information held in Word or other word processing applications may be regarded as unstructured if it does not align with the structure or access method of the database in which it is housed.

Unstructured data is defined by a combination of the data’s structure, the database or container structure holding the data, and the access method used to reach the data.

Without some form of reference, data value plummets like a stone

Love me, love my data

So how do we build procedures and policies for managing unstructured data? Just how swollen is the rising tide and where are the undertows that can suck us under?

How do we learn to love the new world of structured and unstructured data and live with both?

Do we need to exercise some almost chaos-theory like aptitude for data agility to get through? Would it be wise to hold unstructured data in a structured database but access it via unstructured methods?

The fact is that context will always rank as ace high, says Rob Bamforth, principal analyst at research firm Quocirca. He argues that without some form of reference, data value plummets like a stone.

“This context has to be applied to the data as stored (in the form of metadata, tags, or anything to provide some context that can be built upon), otherwise it is applied when accessed, even by unstructured methods,” he says.

“For example, a Google search may appear as a complete open search of all the unstructured data on the internet, but in reality it is the specific product of the search and ranking algorithms used.

"Plus we need to factor in how far and fast the web spiders have trawled the data that appears to be available at any moment.”

Staying with the Google example, we need to remember that Google typically determines the value of the data rather more often than the content provider, such as the journalist writing this, for example.

Define the context

Whoever defines the context adds the value to the data – and it could also come from how different forms of data are combined.

As another example, if a government agency were to combine sufficient quantities of essentially public and shared data in such a way that its value increases dramatically and becomes secret intelligence, then once again we have brought structure to bear upon chaos.

So is the unstructured data tsunami is out of control?

In a recent survey carried out by Unisphere and MarkLogic, 86 per cent of respondents said unstructured data is important to their organisation, but only 11 per cent had clear procedures and policies for managing it.

Andrew Anderson, CEO of information stream company Celaton, suggests that forward-thinking organisations are starting to use artificial intelligence and automation in making sense of unstructured data.

“Those who are still relying on human interpretation will be trying to stay afloat on the unstructured data tsunami with one hand tied behind their back,” he says.

Combat fraud and increase customer satisfaction

More from The Register

next story
This time it's 'Personal': new Office 365 sub covers just two devices
Redmond also brings Office into Google's back yard
Batten down the hatches, Ubuntu 14.04 LTS due in TWO DAYS
Admins dab straining server brows in advance of Trusty Tahr's long-term support landing
Inside the Hekaton: SQL Server 2014's database engine deconstructed
Nadella's database sqares the circle of cheap memory vs speed
Microsoft lobs pre-release Windows Phone 8.1 at devs who dare
App makers can load it before anyone else, but if they do they're stuck with it
Half of Twitter's 'active users' are SILENT STALKERS
Nearly 50% have NEVER tweeted a word
Oh no, Joe: WinPhone users already griping over 8.1 mega-update
Hang on. Which bit of Developer Preview don't you understand?
Internet-of-stuff startup dumps NoSQL for ... SQL?
NoSQL taste great at first but lacks proper nutrients, says startup cloud whiz
Windows 8.1, which you probably haven't upgraded to yet, ALREADY OBSOLETE
Pre-Update versions of new Windows version will no longer support patches
IRS boss on XP migration: 'Classic fix the airplane while you're flying it attempt'
Plus: Condoleezza Rice at Dropbox 'maybe she can find ... weapons of mass destruction'
prev story

Whitepapers

Designing a defence for mobile apps
In this whitepaper learn the various considerations for defending mobile applications; from the mobile application architecture itself to the myriad testing technologies needed to properly assess mobile applications risk.
3 Big data security analytics techniques
Applying these Big Data security analytics techniques can help you make your business safer by detecting attacks early, before significant damage is done.
Five 3D headsets to be won!
We were so impressed by the Durovis Dive headset we’ve asked the company to give some away to Reg readers.
The benefits of software based PBX
Why you should break free from your proprietary PBX and how to leverage your existing server hardware.
Securing web applications made simple and scalable
In this whitepaper learn how automated security testing can provide a simple and scalable way to protect your web applications.