Feeds

There's a tide of unstructured data coming - start swimming

Or you could just work out a plan...

Remote control for virtualized desktops

Whether you prefer to define the known size of our planet’s total digital universe in petabytes or even zettabytes, we can all agree the collective weight of data production is spiralling ever upwards.

While we focus on the relative merits of transactional versus analytical databases, the unstructured data that fails to fall within the general purview of either these systems is the rising tide beneath.

We are not just talking about non-textual audio, video and graphical data here. Unstructured data must also be thought of in its textual form of Word documents, emails, social media messages and other as yet undefined data shapes.

Different stakeholders view structured and unstructured differently. After all, in the world of video production it does not necessarily follow that all video data will be structured to those companies working with it.

Equally, textual information held in Word or other word processing applications may be regarded as unstructured if it does not align with the structure or access method of the database in which it is housed.

Unstructured data is defined by a combination of the data’s structure, the database or container structure holding the data, and the access method used to reach the data.

Without some form of reference, data value plummets like a stone

Love me, love my data

So how do we build procedures and policies for managing unstructured data? Just how swollen is the rising tide and where are the undertows that can suck us under?

How do we learn to love the new world of structured and unstructured data and live with both?

Do we need to exercise some almost chaos-theory like aptitude for data agility to get through? Would it be wise to hold unstructured data in a structured database but access it via unstructured methods?

The fact is that context will always rank as ace high, says Rob Bamforth, principal analyst at research firm Quocirca. He argues that without some form of reference, data value plummets like a stone.

“This context has to be applied to the data as stored (in the form of metadata, tags, or anything to provide some context that can be built upon), otherwise it is applied when accessed, even by unstructured methods,” he says.

“For example, a Google search may appear as a complete open search of all the unstructured data on the internet, but in reality it is the specific product of the search and ranking algorithms used.

"Plus we need to factor in how far and fast the web spiders have trawled the data that appears to be available at any moment.”

Staying with the Google example, we need to remember that Google typically determines the value of the data rather more often than the content provider, such as the journalist writing this, for example.

Define the context

Whoever defines the context adds the value to the data – and it could also come from how different forms of data are combined.

As another example, if a government agency were to combine sufficient quantities of essentially public and shared data in such a way that its value increases dramatically and becomes secret intelligence, then once again we have brought structure to bear upon chaos.

So is the unstructured data tsunami is out of control?

In a recent survey carried out by Unisphere and MarkLogic, 86 per cent of respondents said unstructured data is important to their organisation, but only 11 per cent had clear procedures and policies for managing it.

Andrew Anderson, CEO of information stream company Celaton, suggests that forward-thinking organisations are starting to use artificial intelligence and automation in making sense of unstructured data.

“Those who are still relying on human interpretation will be trying to stay afloat on the unstructured data tsunami with one hand tied behind their back,” he says.

Beginner's guide to SSL certificates

More from The Register

next story
Microsoft to bake Skype into IE, without plugins
Redmond thinks the Object Real-Time Communications API for WebRTC is ready to roll
Mozilla: Spidermonkey ATE Apple's JavaScriptCore, THRASHED Google V8
Moz man claims the win on rivals' own benchmarks
Microsoft promises Windows 10 will mean two-factor auth for all
Sneak peek at security features Redmond's baking into new OS
FTDI yanks chip-bricking driver from Windows Update, vows to fight on
Next driver to battle fake chips with 'non-invasive' methods
DEATH by PowerPoint: Microsoft warns of 0-day attack hidden in slides
Might put out patch in update, might chuck it out sooner
Ubuntu 14.10 tries pulling a Steve Ballmer on cloudy offerings
Oi, Windows, centOS and openSUSE – behave, we're all friends here
Apple's OS X Yosemite slurps UNSAVED docs into iCloud
Docs, email contacts... shhhlooop, up it goes
Was ist das? Eine neue Suse Linux Enterprise? Ausgezeichnet!
Version 12 first major-number Suse release since 2009
prev story

Whitepapers

Choosing cloud Backup services
Demystify how you can address your data protection needs in your small- to medium-sized business and select the best online backup service to meet your needs.
A strategic approach to identity relationship management
ForgeRock commissioned Forrester to evaluate companies’ IAM practices and requirements when it comes to customer-facing scenarios versus employee-facing ones.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.
New hybrid storage solutions
Tackling data challenges through emerging hybrid storage solutions that enable optimum database performance whilst managing costs and increasingly large data stores.
Business security measures using SSL
Examines the major types of threats to information security that businesses face today and the techniques for mitigating those threats.