Welcome!

@CloudExpo Authors: Elizabeth White, Yeshim Deniz, Ian Khan, Liz McMillan, Pat Romanski

Related Topics: Symbian, Containers Expo Blog, @CloudExpo

Symbian: Article

Exclusive Q&A with Rob Weltman, Director of Grid Services, Yahoo!

Cloud-Based Tools Like Hadoop Are Booming Says Yahoo Exec

Cloud-based tools, including large-scale data-intensive computing as offered by Hadoop, are key to the rise and rise of cloud computing. In this wide-ranging Exclusive Q&A with SYS-CON's Cloud Computing Journal, the Director of Grid Services at Yahoo! - Rob Weltman - explains to Jeremy Geelan, Conference Chair of SYS-CON's 1st International Cloud Computing Conference & Expo held last week in San Jose, CA, how analyzing and learning from ever-growing volumes of business data is essential to continuously refining and improving service offerings.

Cloud Computing Journal: Yahoo! has been the largest contributor to the Hadoop project and uses Hadoop extensively in its Web search and advertising businesses. Can you explain a little of the background to that?
Rob Weltman: Yahoo! Search (and before it Inktomi) was a pioneer in using large clusters of commodity computers to speed up the crawling and indexing of Web sites. While working on the architecture and design of the next generation of Web Search crawling and indexing, we came in touch with Doug Cutting and the open source Lucene project for text indexing/search. Lucene contained a distributed file system with integrated computation using the map-reduce paradigm. It looked very promising and appropriate for many data-intensive applications. Hadoop was then split out as its own project. Yahoo! supported Hadoop in a big way, both in contributing to its development as an open source project and in applying it to solve many large-scale data/computation problems in the company.

Hadoop has matured at an amazingly fast pace. From a 20-node cluster two years ago, to many 2,000-node clusters today; from a somewhat embarrassing terasort (a benchmark) performance to the terasort leader; from a no-access control to user- and group-owned files and directories. There is now a high-level language - Pig - that allows you to express complex operations on data in an intuitive way and have them translated into Hadoop map-reduce jobs.

In 2007, Hadoop at Yahoo! was used primarily for research - analyzing enormous volumes of data to find the best algorithms and parameters for selecting search results or ads to present to users. Now it is also a central component in many production operations, including Web Search, ad serving, and personalization.

Cloud Computing Journal: Are cloud-based tools like Hadoop the most important kinds of tools for the future, do you think?
RW: Being able to add capacity as needed without major software or infrastructure changes is clearly important for many organizations. Sharing resources and dynamically allocating more or less to various functions on demand is highly attractive as companies strive to control costs while the computing needs grow and shift. Analyzing and learning from ever-growing volumes of business data is essential to continuously refining and improving service offerings. The ability to quickly explore new algorithms and put them into production will be a competitive advantage for those with the resources to apply them. All of these speak to the importance of Cloud Computing,

Cloud Computing Journal: How important a role does Java play in the project? Is that because of the need to scale horizontally (and massively)?
RW: Hadoop supports programming and scripting in many languages. Hadoop, itself, is written in Java. The language provides strong support for the central infrastructure needs of system and network programming. There is a large body of experience in developing robust, performance-optimized, scalable platforms in Java.

Java provides portability to many hardware and software environments however Hadoop's horizontal scalability is not a result of the choice of language but rather of a design that is strongly focused on fault-tolerance and distribution.

Cloud Computing Journal: Is the Yahoo! Search Webmap still the world's largest Hadoop production application so far as you are aware? Can you share some size data about Webmap with us?
RW: Yes, as far as I know, the Yahoo! WebMap is the largest Hadoop application in production. It uses 2,000+ computers and is still continuously growing. It produces 300TB of data per run, including 1.2 trillion links.

Cloud Computing Journal: How important are Hadoop clusters to Yahoo! Overall? Do your Web search queries depend on them?
RW: Hadoop isn't directly involved in responding to queries typed in by users, but it is responsible for much of the backend work that produces the indexes used to service those queries. If the Hadoop clusters were down, the quality of search results would quickly degrade as the indexes became stale.

Cloud Computing Journal: Who else besides Yahoo! uses Hadoop to run large distributed computations?
RW: Many of the major Hadoop users are listed at http://wiki.apache.org/hadoop/PoweredBy. Facebook has several hundred nodes in a cluster for backend processing and analysis. Quantcast has several thousand cores in a very large cluster. Many companies, including AOL, A9 (Amazon), and IBM have deployed somewhat smaller clusters. It's likely that almost all of the uses involve large quantities of data.

Cloud Computing Journal: Can Hadoop be run on Amazon EC2?
RW: Absolutely! There is a ready-to-run AMI (virtual machine definition for EC2) for Hadoop. Among many others, Powerset (now owned by Microsoft) runs on EC2.

Cloud Computing Journal: What about Sun's Grid Engine - can it also be run on that?
RW: Yes, Hadoop works with Sun's Grid Engine but you lose the benefit of data locality (putting the computation of each piece of a distributed job near the data needed by that piece).

Cloud Computing Journal: Does the Hadoop team have any kind of a blog or forum?
RW: We have a blog at http://developer.yahoo.net/blogs/hadoop/. The team is also heavily engaged in the user and developer Hadoop mailing lists at hadoop.apache.org.

Cloud Computing Journal: Doug Cutting named it after his child's stuffed elephant. Is there any downside to an Enterprise IT tool having the name of a stuffed elephant?
RW: I did get some ribbing during the election period when I wore my Hadoop Summit t-shirt with the elephant on it, but I was able to clarify Hadoop's open source and non-partisan nature.

Cloud Computing Journal: What else have you and your team developed at Yahoo!, in terms of data-analytics applications for example?
RW: The Grid Computing development team at Yahoo! works on the Hadoop core software, the Pig high-level language, the ZooKeeper distributed coordination service, and the Chukwa monitoring and metric analysis system. In addition, it provides various Hadoop add-ons and tools to e.g. facilitate joining of very large data sets or to understand and improve the performance and efficiency of Hadoop jobs. We provide consulting to application teams that develop large-scale Hadoop programs (often involving feature extraction, modeling, optimization, and index creation) but do not produce them ourselves. 

More Stories By Jeremy Geelan

Jeremy Geelan is Chairman & CEO of the 21st Century Internet Group, Inc. and an Executive Academy Member of the International Academy of Digital Arts & Sciences. Formerly he was President & COO at Cloud Expo, Inc. and Conference Chair of the worldwide Cloud Expo series. He appears regularly at conferences and trade shows, speaking to technology audiences across six continents. You can follow him on twitter: @jg21.

Comments (0)

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


@CloudExpo Stories
SYS-CON Events announced today that Interface Masters Technologies, a leader in Network Visibility and Uptime Solutions, will exhibit at the 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. Interface Masters Technologies is a leading vendor in the network monitoring and high speed networking markets. Based in the heart of Silicon Valley, Interface Masters' expertise lies in Gigabit, 10 Gigabit and 40 Gigabit Eth...
In this strange new world where more and more power is drawn from business technology, companies are effectively straddling two paths on the road to innovation and transformation into digital enterprises. The first path is the heritage trail – with “legacy” technology forming the background. Here, extant technologies are transformed by core IT teams to provide more API-driven approaches. Legacy systems can restrict companies that are transitioning into digital enterprises. To truly become a lea...
In his session at @ThingsExpo, Kausik Sridharabalan, founder and CTO of Pulzze Systems, Inc., will focus on key challenges in building an Internet of Things solution infrastructure. He will shed light on efficient ways of defining interactions within IoT solutions, leading to cost and time reduction. He will also introduce ways to handle data and how one can develop IoT solutions that are lean, flexible and configurable, thus making IoT infrastructure agile and scalable.
SYS-CON Events announced today that Sheng Liang to Keynote at SYS-CON's 19th Cloud Expo, which will take place on November 1-3, 2016 at the Santa Clara Convention Center in Santa Clara, California.
Just over a week ago I received a long and loud sustained applause for a presentation I delivered at this year’s Cloud Expo in Santa Clara. I was extremely pleased with the turnout and had some very good conversations with many of the attendees. Over the next few days I had many more meaningful conversations and was not only happy with the results but also learned a few new things. Here is everything I learned in those three days distilled into three short points.
Cognitive Computing is becoming the foundation for a new generation of solutions that have the potential to transform business. Unlike traditional approaches to building solutions, a cognitive computing approach allows the data to help determine the way applications are designed. This contrasts with conventional software development that begins with defining logic based on the current way a business operates. In her session at 18th Cloud Expo, Judith S. Hurwitz, President and CEO of Hurwitz & ...
So, you bought into the current machine learning craze and went on to collect millions/billions of records from this promising new data source. Now, what do you do with them? Too often, the abundance of data quickly turns into an abundance of problems. How do you extract that "magic essence" from your data without falling into the common pitfalls? In her session at @ThingsExpo, Natalia Ponomareva, Software Engineer at Google, provided tips on how to be successful in large scale machine learning...
An IoT product’s log files speak volumes about what’s happening with your products in the field, pinpointing current and potential issues, and enabling you to predict failures and save millions of dollars in inventory. But until recently, no one knew how to listen. In his session at @ThingsExpo, Dan Gettens, Chief Research Officer at OnProcess, will discuss recent research by Massachusetts Institute of Technology and OnProcess Technology, where MIT created a new, breakthrough analytics model f...
The Transparent Cloud-computing Consortium (abbreviation: T-Cloud Consortium) will conduct research activities into changes in the computing model as a result of collaboration between "device" and "cloud" and the creation of new value and markets through organic data processing High speed and high quality networks, and dramatic improvements in computer processing capabilities, have greatly changed the nature of applications and made the storing and processing of data on the network commonplace.
Without a clear strategy for cost control and an architecture designed with cloud services in mind, costs and operational performance can quickly get out of control. To avoid multiple architectural redesigns requires extensive thought and planning. Boundary (now part of BMC) launched a new public-facing multi-tenant high resolution monitoring service on Amazon AWS two years ago, facing challenges and learning best practices in the early days of the new service. In his session at 19th Cloud Exp...
Digitization is driving a fundamental change in society that is transforming the way businesses work with their customers, their supply chains and their people. Digital transformation leverages DevOps best practices, such as Agile Parallel Development, Continuous Delivery and Agile Operations to capitalize on opportunities and create competitive differentiation in the application economy. However, information security has been notably absent from the DevOps movement. Speed doesn’t have to negat...
The Internet of Things can drive efficiency for airlines and airports. In their session at @ThingsExpo, Shyam Varan Nath, Principal Architect with GE, and Sudip Majumder, senior director of development at Oracle, will discuss the technical details of the connected airline baggage and related social media solutions. These IoT applications will enhance travelers' journey experience and drive efficiency for the airlines and the airports. The session will include a working demo and a technical d...
While DevOps promises a better and tighter integration among an organization’s development and operation teams and transforms an application life cycle into a continual deployment, Chef and Azure together provides a speedy, cost-effective and highly scalable vehicle for realizing the business values of this transformation. In his session at @DevOpsSummit at 19th Cloud Expo, Yung Chou, a Technology Evangelist at Microsoft, will present a unique opportunity to witness how Chef and Azure work tog...
Your business relies on your applications and your employees to stay in business. Whether you develop apps or manage business critical apps that help fuel your business, what happens when users experience sluggish performance? You and all technical teams across the organization – application, network, operations, among others, as well as, those outside the organization, like ISPs and third-party providers – are called in to solve the problem.
Almost two-thirds of companies either have or soon will have IoT as the backbone of their business in 2016. However, IoT is far more complex than most firms expected. How can you not get trapped in the pitfalls? In his session at @ThingsExpo, Tony Shan, a renowned visionary and thought leader, will introduce a holistic method of IoTification, which is the process of IoTifying the existing technology and business models to adopt and leverage IoT. He will drill down to the components in this fra...
Digital transformation is too big and important for our future success to not understand the rules that apply to it. The first three rules for winning in this age of hyper-digital transformation are: Advantages in speed, analytics and operational tempos must be captured by implementing an optimized information logistics system (OILS) Real-time operational tempos (IT, people and business processes) must be achieved Businesses that can "analyze data and act and with speed" will dominate those t...
If you had a chance to enter on the ground level of the largest e-commerce market in the world – would you? China is the world’s most populated country with the second largest economy and the world’s fastest growing market. It is estimated that by 2018 the Chinese market will be reaching over $30 billion in gaming revenue alone. Admittedly for a foreign company, doing business in China can be challenging. Often changing laws, administrative regulations and the often inscrutable Chinese Interne...
I'm a lonely sensor. I spend all day telling the world how I'm feeling, but none of the other sensors seem to care. I want to be connected. I want to build relationships with other sensors to be more useful for my human. I want my human to understand that when my friends next door are too hot for a while, I'll soon be flaming. And when all my friends go outside without me, I may be left behind. Don't just log my data; use the relationship graph. In his session at @ThingsExpo, Ryan Boyd, Engi...
Internet of @ThingsExpo, taking place November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with the 19th International Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world and ThingsExpo Silicon Valley Call for Papers is now open.
Adobe is changing the world though digital experiences. Adobe helps customers develop and deliver high-impact experiences that differentiate brands, build loyalty, and drive revenue across every screen, including smartphones, computers, tablets and TVs. Adobe content solutions are used daily by millions of companies worldwide-from publishers and broadcasters, to enterprises, marketing agencies and household-name brands. Building on its established design leadership, Adobe enables customers not o...