Feeds

Amazon teaches cloud to speak Pig Latin

Adoophay orfay ethay assesmay

New hybrid storage solutions

Amazon has taught its cloud to speak Pig Latin.

In April, Jeff Bezos and company unveiled a new web service based on Hadoop, the open-source phenomenon that seeks to mimic MapReduce, the distributed data-crunching platform that drives Google's online infrastructure. And today, Amazon announced that its Elastic MapReduce service now includes support for Pig Latin, the Hadoop programming language first developed at Yahoo!.

You can write straight to Hadoop in Java, but Pig puts Hadoop programming on a somewhat higher level. As Amazon puts it: "Pig Latin is a SQL-like data transformation language. You can use Pig Latin to run complex processes on large-scale compute clusters without having to spend time learning the MapReduce paradigm."

But the SQL comparison is a tad misleading. Hive - a Hadoop programming language seeded by Facebook - is closer to SQL, as is a second as-yet-unnamed language under development at Yahoo!. Pig sits somewhere between the SQL-like paradigm and the low-level code of MapReduce.

Amazon offers two means of using Pig: an "Interactive mode," which lets you run Pig queries on an existing MapReduce cluster by setting up an secure shell connection, and a "batch mode," which involves launching multiple MapReduce server instances that reference your Pig Latin.

Based on Google-published research papers, Hadoop mimics the company's MapReduce framework, which maps data-crunching tasks across distributed machines, splitting them into sub-tasks, before reducing the results into one master calculation. Thus Amazon Elastic MapReduce.

The Apache-hosted Hadoop was originally developed by Nutch-crawler founder Doug Cutting. After three and half years at Yahoo! developing the platform, Cutting is now headed for Cloudera, a Silicon Valley startup that has commercialized Hadoop - Red Hat-style. ®

Secure remote control for conventional and virtual desktops

More from The Register

next story
Not appy with your Chromebook? Well now it can run Android apps
Google offers beta of tricky OS-inside-OS tech
Keep that consumer browser tat away from our software says Oracle
Big Red decides it will only support Firefox's Extended Support Releases
Greater dev access to iOS 8 will put us AT RISK from HACKERS
Knocking holes in Apple's walled garden could backfire, says securo-chap
WordPress 4.0 is here, complete with one-click upgrade process
Don't relax yet, sysadmins, there's still a chance for some big messes here
NHS grows a NoSQL backbone and rips out its Oracle Spine
Open source? In the government? Ha ha! What, wait ...?
Google extends app refund window to two hours
You now have 120 minutes to finish that game instead of 15
Intel: Hey, enterprises, drop everything and DO HADOOP
Big Data analytics projected to run on more servers than any other app
prev story

Whitepapers

Secure remote control for conventional and virtual desktops
Balancing user privacy and privileged access, in accordance with compliance frameworks and legislation. Evaluating any potential remote control choice.
Intelligent flash storage arrays
Tegile Intelligent Storage Arrays with IntelliFlash helps IT boost storage utilization and effciency while delivering unmatched storage savings and performance.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.
Security for virtualized datacentres
Legacy security solutions are inefficient due to the architectural differences between physical and virtual environments.
Providing a secure and efficient Helpdesk
A single remote control platform for user support is be key to providing an efficient helpdesk. Retain full control over the way in which screen and keystroke data is transmitted.