Feeds

British Library tracks rise and fall of file formats

Analysis of 2.5 billion online files suggests software obsolescence slowing

  • alert
  • submit to reddit

The essential guide to IT transformation

File formats and the software capable of reading them are living longer than previously thought, according to a British Library and UK Web Archive study.

Formats over Time: Exploring UK Web History (PDF, slides as PDF) considers 2.5 billion files author Andrew N Jackson retrieved with the help of the Internet Archive and the Joint Information Systems Committee (JISC). All the files come from “the UK web domain” and come from the period between 1996 and 2010.

Jackson used Apache Tika and PRONOM's DROID tool to inspect the files and determine the format they use. Central to the research was Jeff Rothenberg's 1997 prediction that “Digital Information Lasts Forever – Or Five Years, Whichever Comes First.” Jackson is also keen on a rebuttal from David Rosenthal, who he quotes as saying: “When challenged, proponents of [format migration strategies] have failed to identify even one format in wide use when Rothenberg [made that assertion] that has gone obsolete in the intervening decade and a half.”

Jackson's take is that file formats seem to last rather longer than five years even if they don't survive forever.

“While there were just two active versions of HTML in 1996 (2.0 and 3.2), all six were still active in 2010,” he writes. “Similarly, there were three active versions of PDF in 1996 (1.0-1.2) and eleven different versions in 2010 (1.0-1.7, 1.7 Extension Level 3, A-1a and A-1b, with 1.2-1.6 dominant). In general, it appears that format versions, like formats, are quick to arise but slow to fade away.

HTML versions found online in the UK between 1996 and 2010

Jackson attributes formats' longevity to the Network Effect, but also writes that he is uncomfortable drawing firm conclusions about software obsolescence given the sample is UK-centric and the tools used to analyse data identify files imperfectly.

He nonetheless concludes:

Our initial analysis supports Rosenthal's position; that most formats last much longer than five years, that network effects to appear to stabilise formats, and that new formats appear at a modest, manageable rate.

But he also warns that “a number of formats and versions that are fading from use, and these should be studied closely in order to understand the process of obsolescence.” ®

Boost IT visibility and business value

More from The Register

next story
The Return of BSOD: Does ANYONE trust Microsoft patches?
Sysadmins, you're either fighting fires or seen as incompetents now
Munich considers dumping Linux for ... GULP ... Windows!
Give a penguinista a hug, the Outlook's not good for open source's poster child
Intel's Raspberry Pi rival Galileo can now run Windows
Behold the Internet of Things. Wintel Things
Microsoft cries UNINSTALL in the wake of Blue Screens of Death™
Cache crash causes contained choloric calamity
Eat up Martha! Microsoft slings handwriting recog into OneNote on Android
Freehand input on non-Windows kit for the first time
Linux kernel devs made to finger their dongles before contributing code
Two-factor auth enabled for Kernel.org repositories
prev story

Whitepapers

5 things you didn’t know about cloud backup
IT departments are embracing cloud backup, but there’s a lot you need to know before choosing a service provider. Learn all the critical things you need to know.
Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Build a business case: developing custom apps
Learn how to maximize the value of custom applications by accelerating and simplifying their development.
Rethinking backup and recovery in the modern data center
Combining intelligence, operational analytics, and automation to enable efficient, data-driven IT organizations using the HP ABR approach.
Next gen security for virtualised datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.