Feeds

NSA open sources Google database mimic

Crypto masters incubate Son of BigTable

Build a business case: developing custom apps

The US National Security Agency is open sourcing a distributed "NoSQL" database based on Google's proprietary BigTable platform.

Known as Accumulo, the platform has been in development at the NSA for over three years, and it's built atop Hadoop, the open source distributed file system and distributed number-crunching platform that mimics Google's internal infrastructure.

Unlike existing BigTable mimics such as HBase, Accumolo has "fine-grained" access controls and a new server-side programming mechanism that can modify data that's written to disk, or returned to the user. Using the cell-level access labels, you can provide external servers with access to some cells in the Accumolo data store but not others.

The NSA believes this may be of interest to government and health care operations and other outfits concerned with privacy. It acknowledges, however, that the access labels do not constitute a "complete security solution".

As noticed by H Online, the NSA has officially proposed Accumulo as an incubator project at Apache. Though the agency has little experience with public open source work, it says that the project has been "treated internally" as an open source project since its inception in 2008.

"We intend to strongly encourage the community to help with and contribute to the code. We will actively seek potential committers and help them become familiar with the codebase," the agency says. "We do not anticipate difficulty in operating under Apache's development process."

It does acknowledges, however, that the project overlaps with HBase.

"Accumulo and HBase are both based on the design of Google's BigTable, so there is a danger that potential users will have difficulty distinguishing the two or that they will not see an incentive in adopting Accumulo. There are a few key areas in which Accumulo differs from HBase. Some of the desired features of Accumulo could be incorporated into HBase, however the most important of these may be unlikely to be adopted," the agency continues, referring to both the cell-level access labels and the server-side programming mechanism.

"It is a possibility that the codebases will ultimately converge, but the number of differences at the current time warrants a separate project for Accumulo."

Google does not open source the software platforms underpinning its internal infrastructure. But in 2004, it published papers describing its GFS distributed file system and its MapReduce distributed number crunching platform, and these gave rise to the independent Hadoop, which resides at Apache. Google's BigTable paper followed in 2006, and this served as the basis for HBase as well as Acumulo.

According to the NSA, its project now spans over 200,000 lines of (mostly Java) code and hundreds of pages of documentation. It's built atop not only the core Hadoop platforms, but Apache Zookeeper (a means of managing distributed services) and Thift (a framework for developing services across multiple languages).

At this point, there is no indication that the database platform is part of a top secret NSA mission to plant a trojan horse on the machines of innocent open source mavens across the globe. But we'll keep you updated. ®

The essential guide to IT transformation

More from The Register

next story
Munich considers dumping Linux for ... GULP ... Windows!
Give a penguinista a hug, the Outlook's not good for open source's poster child
The Return of BSOD: Does ANYONE trust Microsoft patches?
Sysadmins, you're either fighting fires or seen as incompetents now
Microsoft cries UNINSTALL in the wake of Blue Screens of Death™
Cache crash causes contained choloric calamity
Time to move away from Windows 7 ... whoa, whoa, who said anything about Windows 8?
Start migrating now to avoid another XPocalypse – Gartner
You'll find Yoda at the back of every IT conference
The piss always taking is he. Bastard the.
HANA has SAP cuddling up to 'smaller partners'
Wanted: algorithm wranglers, not systems giants
prev story

Whitepapers

Endpoint data privacy in the cloud is easier than you think
Innovations in encryption and storage resolve issues of data privacy and key requirements for companies to look for in a solution.
Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Top 8 considerations to enable and simplify mobility
In this whitepaper learn how to successfully add mobile capabilities simply and cost effectively.
Solving today's distributed Big Data backup challenges
Enable IT efficiency and allow a firm to access and reuse corporate information for competitive advantage, ultimately changing business outcomes.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.