Feeds

British Library tracks rise and fall of file formats

Analysis of 2.5 billion online files suggests software obsolescence slowing

  • alert
  • submit to reddit

Security for virtualized datacentres

File formats and the software capable of reading them are living longer than previously thought, according to a British Library and UK Web Archive study.

Formats over Time: Exploring UK Web History (PDF, slides as PDF) considers 2.5 billion files author Andrew N Jackson retrieved with the help of the Internet Archive and the Joint Information Systems Committee (JISC). All the files come from “the UK web domain” and come from the period between 1996 and 2010.

Jackson used Apache Tika and PRONOM's DROID tool to inspect the files and determine the format they use. Central to the research was Jeff Rothenberg's 1997 prediction that “Digital Information Lasts Forever – Or Five Years, Whichever Comes First.” Jackson is also keen on a rebuttal from David Rosenthal, who he quotes as saying: “When challenged, proponents of [format migration strategies] have failed to identify even one format in wide use when Rothenberg [made that assertion] that has gone obsolete in the intervening decade and a half.”

Jackson's take is that file formats seem to last rather longer than five years even if they don't survive forever.

“While there were just two active versions of HTML in 1996 (2.0 and 3.2), all six were still active in 2010,” he writes. “Similarly, there were three active versions of PDF in 1996 (1.0-1.2) and eleven different versions in 2010 (1.0-1.7, 1.7 Extension Level 3, A-1a and A-1b, with 1.2-1.6 dominant). In general, it appears that format versions, like formats, are quick to arise but slow to fade away.

HTML versions found online in the UK between 1996 and 2010

Jackson attributes formats' longevity to the Network Effect, but also writes that he is uncomfortable drawing firm conclusions about software obsolescence given the sample is UK-centric and the tools used to analyse data identify files imperfectly.

He nonetheless concludes:

Our initial analysis supports Rosenthal's position; that most formats last much longer than five years, that network effects to appear to stabilise formats, and that new formats appear at a modest, manageable rate.

But he also warns that “a number of formats and versions that are fading from use, and these should be studied closely in order to understand the process of obsolescence.” ®

Website security in corporate America

More from The Register

next story
New 'Cosmos' browser surfs the net by TXT alone
No data plan? No WiFi? No worries ... except sluggish download speed
'Windows 9' LEAK: Microsoft's playing catchup with Linux
Multiple desktops and live tiles in restored Start button star in new vids
iOS 8 release: WebGL now runs everywhere. Hurrah for 3D graphics!
HTML 5's pretty neat ... when your browser supports it
'People have forgotten just how late the first iPhone arrived ...'
Plus: 'Google's IDEALISM is an injudicious justification for inappropriate biz practices'
Mathematica hits the Web
Wolfram embraces the cloud, promies private cloud cut of its number-cruncher
Mozilla shutters Labs, tells nobody it's been dead for five months
Staffer's blog reveals all as projects languish on GitHub
SUSE Linux owner Attachmate gobbled by Micro Focus for $2.3bn
Merger will lead to mainframe and COBOL powerhouse
iOS 8 Healthkit gets a bug SO Apple KILLS it. That's real healthcare!
Not fit for purpose on day of launch, says Cupertino
prev story

Whitepapers

Secure remote control for conventional and virtual desktops
Balancing user privacy and privileged access, in accordance with compliance frameworks and legislation. Evaluating any potential remote control choice.
WIN a very cool portable ZX Spectrum
Win a one-off portable Spectrum built by legendary hardware hacker Ben Heck
Storage capacity and performance optimization at Mizuno USA
Mizuno USA turn to Tegile storage technology to solve both their SAN and backup issues.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?
The next step in data security
With recent increased privacy concerns and computers becoming more powerful, the chance of hackers being able to crack smaller-sized RSA keys increases.