Feeds

Hadoop's little buddy Nutch 2.0 gulps down web's big data

Apache projects united

Beginner's guide to SSL certificates

Hadoop daddy Doug Cutting's Nutch, the open-source web-search engine written in Java, has been updated to crawl through piles of big data on the web.

Apache Software Foundation (ASF) has released Nutch 2.0 featuring a data abstraction technique that plugs into big-data stores and frameworks Apache Accumulo, Avro, Cassandra, HBase and, yes, the Hadoop Distributed File System (HDFS).

The abstraction layer that was employed is yet another Apache project, Gora – a framework that provides an in-memory data model and persistence layer for big data.

Gora works with NoSQL column stores, key value stores and document stores, as well as with RDBMSes.

The ASF website where Gora makes its home states its goal as becoming "the standard data representation and persistence framework for big data".

Nutch 2.0 also builds on the Apache open-source search server Soir, which adds a crawler, and a link-graph database with parsing support handled by the Apache Tika project.

Cutting wrote Nutch in 2003 with Mike Cafarella, while the pair were also developing Hadoop – using Google's MapReduce distributed data processing framework to make Hadoop system work at scale. Cutting also wrote Lucene, but it was Hadoop that made his name and he was brought in by Yahoo! to implement the system on its servers.

Nutch has since been somewhat eclipsed by Hadoop, which is used by Amazon.com, Facebook and Yahoo! to name just three web giants. Search engines written using Nutch include Krugle and mozDex. ®

Internet Security Threat Report 2014

More from The Register

next story
Download alert: Nearly ALL top 100 Android, iOS paid apps hacked
Attack of the Clones? Yeah, but much, much scarier – report
NSA SOURCE CODE LEAK: Information slurp tools to appear online
Now you can run your own intelligence agency
Whistling Google: PLEASE! Brussels can only hurt Europe, not us
And Commish is VERY pro-Google. Why should we worry?
Microsoft: Your Linux Docker containers are now OURS to command
New tool lets admins wrangle Linux apps from Windows
Soz, web devs: Google snatches its Wallet off the table
Killing off web service in 3 months... but app-happy bonkers are fine
First in line to order a Nexus 6? AT&T has a BRICK for you
Black Screen of Death plagues early Google-mobe batch
prev story

Whitepapers

10 ways wire data helps conquer IT complexity
IT teams can automatically detect problems across the IT environment, spot data theft, select unique pieces of transaction payloads to send to a data source, and more.
Why CIOs should rethink endpoint data protection in the age of mobility
Assessing trends in data protection, specifically with respect to mobile devices, BYOD, and remote employees.
A strategic approach to identity relationship management
ForgeRock commissioned Forrester to evaluate companies’ IAM practices and requirements when it comes to customer-facing scenarios versus employee-facing ones.
Reg Reader Research: SaaS based Email and Office Productivity Tools
Read this Reg reader report which provides advice and guidance for SMBs towards the use of SaaS based email and Office productivity tools.
Business security measures using SSL
Examines the major types of threats to information security that businesses face today and the techniques for mitigating those threats.