Feeds

IBM's tools give Big Data a good seeing to

Company shares nothing but Hadoop and GPFS

Choosing a cloud hosting partner with confidence

IBM is using Hadoop to make its General Parallel File System capable of dealing with Big Data - extremely large data sets - for cloud-based analytic computing.

Announced at the Supercomputing 2010 conference, the General Parallel File System-Shared Nothing Cluster (GPFS-SNC) project at IBM Research Almaden involves an architecture designed to provide higher availability through clustering technologies, dynamic file system management and replication.

GPFS is the basis for IBM's High Performance Computing Systems, Information Archive, Scale-Out NAS (SONAS), and Smart Business Compute Cloud. GPFS-SNC is a distributed, shared-nothing, computing architecture in which each node is self-sufficient; tasks are divided up between these independent computers and no one node waits on any other.

Hadoop, which is used by Yahoo!, has evolved from Google's MapReduce technology for computations involving petabyte-level data sets distributed across thousands of commodity hsrdware-based computational nodes. The Hadoop Distributed File System (HDFS) is a distributed, scalable and portable file system, written in Java, involving a cluster of data nodes.

HDFS is aware of the location, in a network switch sense, of servers (worker nodes) in the cluster and the system uses this to ensure they compute data local to them and thus reduce data traffic across the network. Different copies of data are kept on different sets of worker nodes, with data being replicated across nodes this way to avoid unnecessary redundancy and high availability, without RAID, should a worker node rack or network switch fail.

HDFS is not POSIX-compliant and one aspect of the GPFS-SNC project is to provide POSIX-compliance. GPFS on its own is POSIX-compliant.

IBM says running data analytics applications in the cloud on extremely large data sets is gaining traction because it is affordable and the underlying infrastructure can store and compute the immense amount of data involved. A POSIX interface means traditional applications using POSIX interfaces can use the cloud resources.

The end-user apps IBM has in mind are things like business intelligence, digital media processing and surveillance video searches. GPFS-SNC technology decomposes the large computation involved into a set of smaller parallelisable computations. IBM reckons GPFS-SNC can work around the frequent failures expected in large-scale commodity server and storage deployments, while being an efficient user of compute, storage and network resources.

IBM's announcement statement says GPFS-SNC "will convert terabytes of pure information into actionable insights twice as fast as previously possible... the design provides a common file system and namespace across disparate computing platforms, streamlining the process and reducing disk space."

The GPFS-SNC project is likely to be used in the EU-funded, IBM-led VISION cloud project announced in the beginning of November. ®

Secure remote control for conventional and virtual desktops

More from The Register

next story
Download alert: Nearly ALL top 100 Android, iOS paid apps hacked
Attack of the Clones? Yeah, but much, much scarier – report
You stupid BRICK! PCs running Avast AV can't handle Windows fixes
Fix issued, fingers pointed, forums in flames
Microsoft: Your Linux Docker containers are now OURS to command
New tool lets admins wrangle Linux apps from Windows
Facebook, working on Facebook at Work, works on Facebook. At Work
You don't want your cat or drunk pics at the office
Soz, web devs: Google snatches its Wallet off the table
Killing off web service in 3 months... but app-happy bonkers are fine
First in line to order a Nexus 6? AT&T has a BRICK for you
Black Screen of Death plagues early Google-mobe batch
prev story

Whitepapers

Why and how to choose the right cloud vendor
The benefits of cloud-based storage in your processes. Eliminate onsite, disk-based backup and archiving in favor of cloud-based data protection.
Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Driving business with continuous operational intelligence
Introducing an innovative approach offered by ExtraHop for producing continuous operational intelligence.
10 threats to successful enterprise endpoint backup
10 threats to a successful backup including issues with BYOD, slow backups and ineffective security.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?