Feeds

SKA precursor starts streaming firehosing astrodata to the world

Cramming the universe into a fibre at 1.5 terabytes an hour

Internet Security Threat Report 2014

Hard on the heels of yesterday's discussion of high-performance computing with the International Centre for Radio Astronomy Research and the National Computational Infrastructure comes the announcement that real data has started to stream out of Western Australia's Murchison Widefield Array.

In fact, to anybody but the biggest big data enthusiast, “stream” seems an inadequate word for what's leaving this one source. That's because the MWA's 2,048 dual-polarisation dipole antennas, arranged as 128 “tiles” (mostly in a relatively small 1.5 km core region, with some beyond that to yield a 3 km baseline) yield a veritable firehose of data.

The iVEC-managed Pawsey Centre, 800 km distant in Perth is receiving 400 megabytes per second from the telescope – and as discussed here and here, that's after on-site correlators, the first step in processing the telescope data, reduce the amount of data leaving the site to a manageable level.

A dedicated 10 Gbps fibre runs from Murchison to Geraldton, after which dark fibre on the Nextgen Networks-operated Regional Backhaul Blackspots Project link to Perth provides connectivity to the Pawsey Centre.

Murchison iVEC - ICRAR team

The Murchison Widefield Array Data Archive team from ICRAR: Dave Pallot (left) Professor Andreas Wicenec (centre) and Associate Professor Chen Wu (right) at the Pawsey Centre. Image: ICRAR

Just how much work is needed to create a manageable data set is demonstrated by the fact that the MWA's correlators perform half of the entire computation needed by the facility, iVEC says.

That gets the raw data – about 1.5 TB per hour – down to a much more manageable requirement that the Pawsey Centre store a mere 3 PB annually.

And, since in science data becomes more useful the more available it is, there are other institutions in on the act. MIT in the US and the Victoria University in New Zealand already have links to the Pawsey Centre, with another planned to connect India, another partner in the MWA project.

MIT's research is the early Universe, which means there's further filtering carried out. A mere 150 TB has been sent to the USA of astronomy data already collected in the MWA's pilot operations and shipped to Perth on an earlier 1 Gbps link. The new data streams will add 4 TB per day to that.

As previously discussed by The Register, the Pawsey Centre will also host an advanced hierarchical storage management facility, including a 20 Petabyte Spectra Logic tape library, a 6 Petabyte SGI disk-based storage system, and a significant visualisation and post-processing capability provided by SGI in the form of a 6 Terabyte UV2000 and 34 Data Analysis Engines. All of these systems will be interconnected via a high-speed (FDR) Infiniband network.

Managing the distribution of this data is the open-source Next Generation Archive System (NGAS) developed by professor Wicenec while at the European Southern Observatory and modified by an ICRAR for operation at the Pawsey Centre. ®

Beginner's guide to SSL certificates

More from The Register

next story
Docker's app containers are coming to Windows Server, says Microsoft
MS chases app deployment speeds already enjoyed by Linux devs
'Hmm, why CAN'T I run a water pipe through that rack of media servers?'
Leaving Las Vegas for Armenia kludging and Dubai dune bashing
'Urika': Cray unveils new 1,500-core big data crunching monster
6TB of DRAM, 38TB of SSD flash and 120TB of disk storage
Facebook slurps 'paste sites' for STOLEN passwords, sprinkles on hash and salt
Zuck's ad empire DOESN'T see details in plain text. Phew!
SDI wars: WTF is software defined infrastructure?
This time we play for ALL the marbles
Windows 10: Forget Cloudobile, put Security and Privacy First
But - dammit - It would be insane to say 'don't collect, because NSA'
Oracle hires former SAP exec for cloudy push
'We know Larry said cloud was gibberish, and insane, and idiotic, but...'
Symantec backs out of Backup Exec: Plans to can appliance in Jan
Will still provide support to existing customers
prev story

Whitepapers

Forging a new future with identity relationship management
Learn about ForgeRock's next generation IRM platform and how it is designed to empower CEOS's and enterprises to engage with consumers.
Why cloud backup?
Combining the latest advancements in disk-based backup with secure, integrated, cloud technologies offer organizations fast and assured recovery of their critical enterprise data.
Win a year’s supply of chocolate
There is no techie angle to this competition so we're not going to pretend there is, but everyone loves chocolate so who cares.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?
Intelligent flash storage arrays
Tegile Intelligent Storage Arrays with IntelliFlash helps IT boost storage utilization and effciency while delivering unmatched storage savings and performance.