Feeds

Two centuries of Hansard to move online

Web no longer 'somewhere data goes to die'

Intelligent flash storage arrays

Parliament hopes to place all Hansard reports - from 1804 to 2004 - online by the end of this year.

Its information management department is using optical character recognition (OCR) technology to turn three million printed pages of the record of Parliamentary proceedings into digitised text. Some is already online, although the project has not yet been officially approved as a version of Hansard.

Edward Wood, Parliament's director of information management, said the department has sliced up original bound copies of Hansard to obtain the pages for scanning – adding that such books are commonly available, as many libraries are selling them.

"For me, it symbolised opening up the data," he told Kable's Electronic Document and Records Management conference.

Wood said the main aim was to avoid expensive conservation work on printed versions of Hansard used by Parliament's members and staff, but also to allow better searching and reduce storage costs.

The process compares the results of three OCR scans with 100 per cent of the results proof read by a contractor. Parliament also proof reads one per cent to check the quality of the work. Wood said although the likes of Google and Microsoft have digitised some of Hansard as part of other projects, their work "is not particularly good, on the whole – there's very little metadata".

Robert Brook, a developer working on the project, said the system aims to provide excellent metadata, with material linked by bill, MP, constituency and even monarch. "Previously, we've treated the web as somewhere data goes to die," he said, but the aim of this project is to open it to numerous uses.

This has been evident in the eclecticism of searches made by users so far, Brook added. "I expected them to look for Tony Blair and Iraq," he said, but instead popular searches have included Telic, the code name for Britain's operations in Iraq, asbestos use in playgrounds, and Corsham's military communications centre.

Around 95 per cent of searches come through Google. "No one uses our search engine, which is really galling," said Brook. But he added that this means people are finding the nascent system as part of general search, rather than specifically looking for Hansard.

Brook said the system, which relies entirely on open source software and uses open data standards to allow reuse and mash-ups on other websites, will add another decade's worth of material in the next month. If it wins approval, it will eventually "get a portcullis on top", he said, and be adopted as an official archive of Hansard.

This article was originally published at Kablenet.

Kablenet's GC weekly is a free email newsletter covering the latest news and analysis of public sector technology. To register click here.

Internet Security Threat Report 2014

More from The Register

next story
The 'fun-nification' of computer education – good idea?
Compulsory code schools, luvvies love it, but what about Maths and Physics?
Facebook, Apple: LADIES! Why not FREEZE your EGGS? It's on the company!
No biological clockwatching when you work in Silicon Valley
Happiness economics is bollocks. Oh, UK.gov just adopted it? Er ...
Opportunity doesn't knock; it costs us instead
Ex-US Navy fighter pilot MIT prof: Drones beat humans - I should know
'Missy' Cummings on UAVs, smartcars and dying from boredom
Yes, yes, Steve Jobs. Look what I'VE done for you lately – Tim Cook
New iPhone biz baron points to Apple's (his) greatest successes
Lords take revenge on REVENGE PORN publishers
Jilted Johns and Jennies with busy fingers face two years inside
Sysadmin with EBOLA? Gartner's issued advice to debug your biz
Start hoarding cleaning supplies, analyst firm says, and assume your team will scatter
Edward who? GCHQ boss dodges Snowden topic during last speech
UK spies would rather 'walk' than do 'mass surveillance'
Doctor Who's Flatline: Cool monsters, yes, but utterly limp subplots
We know what the Doctor does, stop going on about it already
prev story

Whitepapers

Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Why and how to choose the right cloud vendor
The benefits of cloud-based storage in your processes. Eliminate onsite, disk-based backup and archiving in favor of cloud-based data protection.
Three 1TB solid state scorchers up for grabs
Big SSDs can be expensive but think big and think free because you could be the lucky winner of one of three 1TB Samsung SSD 840 EVO drives that we’re giving away worth over £300 apiece.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.
Security for virtualized datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.