Feeds

NSA open sources Google database mimic

Crypto masters incubate Son of BigTable

Beginner's guide to SSL certificates

The US National Security Agency is open sourcing a distributed "NoSQL" database based on Google's proprietary BigTable platform.

Known as Accumulo, the platform has been in development at the NSA for over three years, and it's built atop Hadoop, the open source distributed file system and distributed number-crunching platform that mimics Google's internal infrastructure.

Unlike existing BigTable mimics such as HBase, Accumolo has "fine-grained" access controls and a new server-side programming mechanism that can modify data that's written to disk, or returned to the user. Using the cell-level access labels, you can provide external servers with access to some cells in the Accumolo data store but not others.

The NSA believes this may be of interest to government and health care operations and other outfits concerned with privacy. It acknowledges, however, that the access labels do not constitute a "complete security solution".

As noticed by H Online, the NSA has officially proposed Accumulo as an incubator project at Apache. Though the agency has little experience with public open source work, it says that the project has been "treated internally" as an open source project since its inception in 2008.

"We intend to strongly encourage the community to help with and contribute to the code. We will actively seek potential committers and help them become familiar with the codebase," the agency says. "We do not anticipate difficulty in operating under Apache's development process."

It does acknowledges, however, that the project overlaps with HBase.

"Accumulo and HBase are both based on the design of Google's BigTable, so there is a danger that potential users will have difficulty distinguishing the two or that they will not see an incentive in adopting Accumulo. There are a few key areas in which Accumulo differs from HBase. Some of the desired features of Accumulo could be incorporated into HBase, however the most important of these may be unlikely to be adopted," the agency continues, referring to both the cell-level access labels and the server-side programming mechanism.

"It is a possibility that the codebases will ultimately converge, but the number of differences at the current time warrants a separate project for Accumulo."

Google does not open source the software platforms underpinning its internal infrastructure. But in 2004, it published papers describing its GFS distributed file system and its MapReduce distributed number crunching platform, and these gave rise to the independent Hadoop, which resides at Apache. Google's BigTable paper followed in 2006, and this served as the basis for HBase as well as Acumulo.

According to the NSA, its project now spans over 200,000 lines of (mostly Java) code and hundreds of pages of documentation. It's built atop not only the core Hadoop platforms, but Apache Zookeeper (a means of managing distributed services) and Thift (a framework for developing services across multiple languages).

At this point, there is no indication that the database platform is part of a top secret NSA mission to plant a trojan horse on the machines of innocent open source mavens across the globe. But we'll keep you updated. ®

Remote control for virtualized desktops

More from The Register

next story
Nexus 7 fandroids tell of salty taste after sucking on Google's Lollipop
Web giant looking into why version 5.0 of Android is crippling older slabs
Be real, Apple: In-app goodie grab games AREN'T FREE – EU
Cupertino stands down after Euro legal threats
Download alert: Nearly ALL top 100 Android, iOS paid apps hacked
Attack of the Clones? Yeah, but much, much scarier – report
SLURP! Flick your TONGUE around our LOLLIPOP – Google
Android 5 is coming – IF you're lucky enough to have the right gadget
Microsoft: Your Linux Docker containers are now OURS to command
New tool lets admins wrangle Linux apps from Windows
Bada-Bing! Mozilla flips Firefox to YAHOO! for search
Microsoft system will be the default for browser in US until 2020
prev story

Whitepapers

Choosing cloud Backup services
Demystify how you can address your data protection needs in your small- to medium-sized business and select the best online backup service to meet your needs.
Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.
The hidden costs of self-signed SSL certificates
Exploring the true TCO for self-signed SSL certificates, including a side-by-side comparison of a self-signed architecture versus working with a third-party SSL vendor.
Top 5 reasons to deploy VMware with Tegile
Data demand and the rise of virtualization is challenging IT teams to deliver storage performance, scalability and capacity that can keep up, while maximizing efficiency.