Feeds

Yahoo! open sources uber web server

'400-terabytes-a-day' Traffic Server lands at Apache

Boost IT visibility and business value

Yahoo! has open sourced the back-end software platform that underpins the company's webmail client and countless other applications offered up across its sweeping web portal.

Known as Traffic Server, the platform handles general edge caching, edge processing, and load balancing at Yahoo!, but it's also used to manage traffic on the company's internal storage and server-virtualization services.

"It's used at the edge, but it also gets used almost like an application server," Chuck Neerdaels, Yahoo! vp of content storage, delivery, and edge, tells The Reg. It serves up the latest version of Yahoo! Mail, for instance, which the company calls "Candy Gram."

Acquired with Yahoo!'s purchase of Inktomi, Traffic Server has been in active use at the two companies for the past eight years. According to Yahoo!, it now handles 30,000 requests per second, serving 30 billion Web objects and 400 terabytes of data a day.

"It is a very mature, very reliable piece of technology," says Shelton Shugar, Yahoo!'s senior vp of cloud computing. "In some form, it supports more than half of Yahoo!'s traffic."

The company donated a version of the platform to The Apache Software Foundation last week through the Apache Incubator program. "This is part of our overhaul strategy to open source cloud services that are mature and not laden with Yahoo!-specific stuff that wouldn't make sense for open source," says Shugar.

In June, the company open sourced its internal Hadoop distro, an internet-scale distributed data-crunching platform based the Apache project of the same name.

Neerdaels says Traffic Server was originally designed as a proxy cache. But it's Yahoo!'s go-to tool for http session management, and it includes an API for tweaking content all the way down at the protocol level. "You can poke around with various headers and inject content and direct client requests to different backends, all through a relatively clean API," he says.

Neerdaels and Shuger also describe Traffic Server as an "extensible framework" that lets you tweak the architecture according to the task at hand. And, according to the company, it suits web operations both large and small. "It's simple enough for a small operation to pick it up quickly," says Neerdaels. "But the big players can pick it up and it will scale to meet their needs pretty impressively."

Those 400 terabytes are served up from between 100 and 150 "commodity" machines. ®

Build a business case: developing custom apps

More from The Register

next story
The Return of BSOD: Does ANYONE trust Microsoft patches?
Sysadmins, you're either fighting fires or seen as incompetents now
Linux turns 23 and Linus Torvalds celebrates as only he can
No, not with swearing, but by controlling the release cycle
China hopes home-grown OS will oust Microsoft
Doesn't much like Apple or Google, either
Sin COS to tan Windows? Chinese operating system to debut in autumn – report
Development alliance working on desktop, mobe software
Apple promises to lift Curse of the Drained iPhone 5 Battery
Have you tried turning it off and...? Never mind, here's a replacement
Why has the web gone to hell? Market chaos and HUMAN NATURE
Tim Berners-Lee isn't happy, but we should be
Eat up Martha! Microsoft slings handwriting recog into OneNote on Android
Freehand input on non-Windows kit for the first time
Linux kernel devs made to finger their dongles before contributing code
Two-factor auth enabled for Kernel.org repositories
prev story

Whitepapers

Implementing global e-invoicing with guaranteed legal certainty
Explaining the role local tax compliance plays in successful supply chain management and e-business and how leading global brands are addressing this.
Endpoint data privacy in the cloud is easier than you think
Innovations in encryption and storage resolve issues of data privacy and key requirements for companies to look for in a solution.
Scale data protection with your virtual environment
To scale at the rate of virtualization growth, data protection solutions need to adopt new capabilities and simplify current features.
Boost IT visibility and business value
How building a great service catalog relieves pressure points and demonstrates the value of IT service management.
High Performance for All
While HPC is not new, it has traditionally been seen as a specialist area – is it now geared up to meet more mainstream requirements?