Feeds

IBM gets handle on unstructured data

Almaden incarnates search and BI

Providing a secure and efficient Helpdesk

“A few years ago IBM got together with a few agencies from the US Government, plus several research institutions, where it was decided to develop a standard framework which would allow all the proprietary algorithms to be plugged in so that others could take advantage of them,” he said. “It was announced last year that we were going to give the implementation of this framework (called Unstructured Information Management Architecture - UIMA) to the open source community. It can now be used to use search arguments to identify documents where a connection has been inferred, even though none of the keywords of the search argument are in them.”

Project Avatar puts a layer of intelligence on top of the primitive interface layer. This performs semantic analyses of search requests in order to surface more comprehensive results. To demonstrate, Mattos used the simple example of searching for the words 'John' and 'Phone'.

"This will normally lead to all documents containing those words. But Avatar will infer the likelihood that the required answer is actually John's phone number. The system will then search for a document containing John's phone number, even if it does not contain those two words together. I may even use some description that I know I can associate with 'John' and the system will find a string that it associates with 'phone number'."

According to Mattos, the objective here is to use intelligence in the search interface to allow users to start looking for concepts rather than keywords. From here it is then possible to start extracting the concepts out of documents, which in turn will allow users to start generating information even before anyone has asked for specific facts. An example would be examining records from a call centre and being able to determine the percentage of calls complaining about quality and/or new maintenance contract terms, and from that rapidly pinpoint areas in customer and product support that need to be addressed to improve the customers’ experience.

This is, arguably, the essence of BI, surfacing possible answers to questions that can impact the ongoing performance of a business, before the business user has even formulated them.

One of the most important types of unstructured data is the voice, and Almaden has already developed speech recognition that can record the voice of the caller and do speech to text conversion in any of the major languages. "We can even do language to language translation, so could take worldwide records, translate them all into English and do analysis on the results. This information can then be incorporated with typed records.

An interesting side issue here is the added ability to analyse customers' spoken interactions with company staff to assess factors such as customer satisfaction. "This is done using sentiment analysis, which can only be obtained from the voice, whether someone was angry, upset or whatever, not the text,” Mattos said. "That can't be done in real time, however."

From a users' point of view, Project Avatar will make it possible to have a single, customer-defined `company-standard’ UI to make inquiries of both structured and unstructured data, where users will be able to search for concepts.

As this is now out in the open source community, it is likely that DIYBI tools will start appearing that are a mixture of search engine and BI tool. It is unlikely to come from IBM of course, though the company has delivered a search engine for enterprises called Omnifind which is built on the technology. ®

Internet Security Threat Report 2014

More from The Register

next story
Google+ goes TITSUP. But WHO knew? How long? Anyone ... Hello ...
Wobbly Gmail, Contacts, Calendar on the other hand ...
UNIX greybeards threaten Debian fork over systemd plan
'Veteran Unix Admins' fear desktop emphasis is betraying open source
Preview redux: Microsoft ships new Windows 10 build with 7,000 changes
Latest bleeding-edge bits borrow Action Center from Windows Phone
Microsoft promises Windows 10 will mean two-factor auth for all
Sneak peek at security features Redmond's baking into new OS
Netscape Navigator - the browser that started it all - turns 20
It was 20 years ago today, Marc Andreeesen taught the band to play
Redmond top man Satya Nadella: 'Microsoft LOVES Linux'
Open-source 'love' fairly runneth over at cloud event
Chrome 38's new HTML tag support makes fatties FIT and SKINNIER
First browser to protect networks' bandwith using official spec
prev story

Whitepapers

Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Why and how to choose the right cloud vendor
The benefits of cloud-based storage in your processes. Eliminate onsite, disk-based backup and archiving in favor of cloud-based data protection.
Three 1TB solid state scorchers up for grabs
Big SSDs can be expensive but think big and think free because you could be the lucky winner of one of three 1TB Samsung SSD 840 EVO drives that we’re giving away worth over £300 apiece.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.
Security for virtualized datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.