Welcome!

@CloudExpo Authors: Shelly Palmer, Elizabeth White, Karthick Viswanathan, Pat Romanski, Liz McMillan

Related Topics: @CloudExpo, Cloud Security, @DevOpsSummit

@CloudExpo: Blog Post

Even Netflix Goes Down Once in a While By @Neotys | @CloudExpo [#Cloud]

Thanks to their constant evaluation of cloud operations, the severity of the outage was not extensive

Even Chaos Generating Netflix Goes Down Once in a While

Where were you on February 3, 2015 at 3:40 p.m. PST? Snowed in? Desperately trying to refresh Netflix? If so, you weren't alone. It turns out even best and biggest companies experience failure from time to time. Despite the success of their Chaos Monkey approach to operations, the Internet streaming media provider experienced an outage for a little over an hour. Responses from users ranged from outrage to well, hysteria. You can read more about the media frenzy and see how the masses took to twitter to discuss here.

All was not lost however. Thanks to their constant evaluation of cloud operations, the severity of the outage was not extensive. Given the level of dedication Netflix has to constant self-improvement, you can be sure that in the future, they will be learning from this experience to stay on top of potential issues. That being said, we can still look at what Netflix has done right thus far, and how they maintain a service that so rarely has outages.

Method Behind the Madness: Netflix's Chaos Infrastructure
Netflix definitely subscribes to the saying: "If you fail to plan, you plan to fail." Their strategy integrates quality assurance with the operational infrastructure. This combination provides for the ultimate chaos-proof product. In accordance with the best Testing in Production (TiP) procedures, this type of testing is conducted in a live environment, and familiarizes developers with uncommon situations on a regular, controlled basis. It also unites the QA team and operational team, which reduces miscommunication in the production process.

Netflix's Chaos Monkey randomly simulates the failure of various components within the application stack at random times. Engineers must be alert and ready to find a solution for future recovery scenarios. There's a method behind the madness. The genius is that by embracing failure, they actually reduce its potential to wreak havoc.

Netflix has found ways to function despite major hiccups. They keep operational processes at the forefront, reducing the amount of angry customers, as well as the potential for lost revenue. Some of their techniques include displaying popular picks instead of the entire personalized menu and making streaming the utmost priority. If Netflix can hold on to their ability to suggest and stream movies during a major outage, they've won.

Getting to the Root of the Problem
There's more logic beyond the chaos system than one might think. Netflix puts into practice the exact techniques for root cause analysis. To put this strategy to work, you must:

  • Define the issue at hand.
  • Discover why it happened.
  • Develop strategies to reduce the likelihood that the problem will happen again and assess potential risk factors.

Of course, of all of these steps it's the discovery process that is the most difficult. Here, Netflix instills the "5 Why's" principle, which means that they ask themselves five times why the problem occurred. Asking yourself "why" as many times as possible will encourage you to review the angle from all sides and get deep enough to see the roots of the problem. They ask themselves what circumstances were present which allowed or caused its occurrence? How can they change and manipulate those circumstances to prevent it from happening again?

In the problem solving phase, it is imperative to use data collection which provides information about the duration, direct impact and potential effects on other instances. By implementing this technique, Netflix spots weak points in their system and develops the automatic recovery system needed to strengthen them.

Gorilla Warfare and the Simian Army
Netflix's chaos-causing monkey is actually pretty serious business. The entire Simian army consists of several different monkeys who each have different tasks. As a unit the Simian army is responsible for causing a variety of failures within the cloud. The security of the cloud is tested by identifying any abnormal conditions and assessing Netflix's ability to function in spite of them.

Each service, or monkey, has its own specific role. Some of the most vital members of the troop include:

  • Chaos Monkey disables systems within groups. Despite its troublemaking ways, it only runs during business hours so that engineers are available to solve the problems as quickly as possible.
  • Conformity Monkey helps maintain status quo. He looks for instances within groups and disables them. The idea behind this practice is that the instances can be re-launched at a later date appropriately.
  • The Security Monkey is Conformity's right hand primate. His tasks are primarily two-fold: looking for vulnerabilities or security violations while checking SSL and DRM certifications for validity.
  • Janitor Monkey keeps the place looking neat. He trashes unused resources.
  • The Chaos Gorilla lives up to his name. This big guy simulates a full-scale outage of a complete Amazon availability zone.

The monkey team continues to grow to provide a complete cloud testing strategy. Best of all, the Simian Army is open-source and allows them to test the cloud's operation and security at any time.

Incorporate Simulated Users for Maximum Results
You may not be ready to deploy the entire Simian army, but there are lots of ways you can put in place your own early warning systems to keep cloud operations running as smoothly as possible. One great way of doing this is with products like NeoSense. NeoSense generates simulated users, which highlight weak spots in performance happening within business transactions in the production environment. The critical data these simulated users generate allows teams to spot problems before a catastrophic situation occurs. It also gives operations teams the ability to get working quicker, with more informed data so they can get systems back up and running.

Beyond Monkeying Around
Jokes aside, we can see that Netflix's strategies for cloud operational testing are right on target. They do more than allow developers to identify weak areas - they maximize testing for failure in real, everyday environments. By doing so, the emphasis is moved from dealing with paralyzing fear to building necessary solutions. This technique is applicable to a multitude of tech enterprises that need to provide the highest level of service for their demanding users.

More Stories By Tim Hinds

Tim Hinds is the Product Marketing Manager for NeoLoad at Neotys. He has a background in Agile software development, Scrum, Kanban, Continuous Integration, Continuous Delivery, and Continuous Testing practices.

Previously, Tim was Product Marketing Manager at AccuRev, a company acquired by Micro Focus, where he worked with software configuration management, issue tracking, Agile project management, continuous integration, workflow automation, and distributed version control systems.

@CloudExpo Stories
The Internet giants are fully embracing AI. All the services they offer to their customers are aimed at drawing a map of the world with the data they get. The AIs from these companies are used to build disruptive approaches that cannot be used by established enterprises, which are threatened by these disruptions. However, most leaders underestimate the effect this will have on their businesses. In his session at 21st Cloud Expo, Rene Buest, Director Market Research & Technology Evangelism at Ara...
WebRTC is great technology to build your own communication tools. It will be even more exciting experience it with advanced devices, such as a 360 Camera, 360 microphone, and a depth sensor camera. In his session at @ThingsExpo, Masashi Ganeko, a manager at INFOCOM Corporation, will introduce two experimental projects from his team and what they learned from them. "Shotoku Tamago" uses the robot audition software HARK to track speakers in 360 video of a remote party. "Virtual Teleport" uses a mu...
Trying to improve density, lower costs and run applications faster than before? Today, enterprises looking for a secure cloud strategy are increasingly turning to container-based Platform as a Service solutions for on-premises hosted DevOps. In her session at 21st Cloud Expo, Alise Cashman Spence, Offering Manager, Power Systems Cloud Solutions at IBM, will discuss the driving factors behind these cloud trends and how IBM customers are realizing exceptional performance, security and control for ...
Internet of @ThingsExpo, taking place October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 21st Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound change in personal and enterprise IT since the creation of the Worldwide Web more than 20 years ago. All major researchers estimate there will be tens of billions devic...
Mobile device usage has increased exponentially during the past several years, as consumers rely on handhelds for everything from news and weather to banking and purchases. What can we expect in the next few years? The way in which we interact with our devices will fundamentally change, as businesses leverage Artificial Intelligence. We already see this taking shape as businesses leverage AI for cost savings and customer responsiveness. This trend will continue, as AI is used for more sophistica...
"When we talk about cloud without compromise what we're talking about is that when people think about 'I need the flexibility of the cloud' - it's the ability to create applications and run them in a cloud environment that's far more flexible,” explained Matthew Finnie, CTO of Interoute, in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
SYS-CON Events announced today that SourceForge has been named “Media Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. SourceForge is the largest, most trusted destination for Open Source Software development, collaboration, discovery and download on the web serving over 32 million viewers, 150 million downloads and over 460,000 active development projects each and every month.
"NetApp's vision is how we help organizations manage data - delivering the right data in the right place, in the right time, to the people who need it, and doing it agnostic to what the platform is," explained Josh Atwell, Developer Advocate for NetApp, in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
What You Need to Know You know you need the cloud, but you’re hesitant to simply dump everything at Amazon since you know that not all workloads are suitable for cloud. You know that you want the kind of ease of use and scalability that you get with public cloud, but your applications are architected in a way that makes the public cloud a non-starter. You’re looking at private cloud solutions based on hyperconverged infrastructure, but you’re concerned with the limits inherent in those technolog...
SYS-CON Events announced today that DXWorldExpo has been named “Global Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Digital Transformation is the key issue driving the global enterprise IT business. Digital Transformation is most prominent among Global 2000 enterprises and government institutions.
One of the biggest challenges with adopting a DevOps mentality is: new applications are easily adapted to cloud-native, microservice-based, or containerized architectures - they can be built for them - but old applications need complex refactoring. On the other hand, these new technologies can require relearning or adapting new, oftentimes more complex, methodologies and tools to be ready for production. In his general session at @DevOpsSummit at 20th Cloud Expo, Chris Brown, Solutions Marketi...
SYS-CON Events announced today that Nihon Micron will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Nihon Micron Co., Ltd. strives for technological innovation to establish high-density, high-precision processing technology for providing printed circuit board and metal mount RFID tags used for communication devices. For more inf...
SYS-CON Events announced today that Ryobi Systems will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Ryobi Systems Co., Ltd., as an information service company, specialized in business support for local governments and medical industry. We are challenging to achive the precision farming with AI. For more information, visit http:...
SYS-CON Events announced today that mruby Forum will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. mruby is the lightweight implementation of the Ruby language. We introduce mruby and the mruby IoT framework that enhances development productivity. For more information, visit http://forum.mruby.org/.
SYS-CON Events announced today that Mobile Create USA will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Mobile Create USA Inc. is an MVNO-based business model that uses portable communication devices and cellular-based infrastructure in the development, sales, operation and mobile communications systems incorporating GPS capabi...
SYS-CON Events announced today that Daiya Industry will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Daiya Industry specializes in orthotic support systems and assistive devices with pneumatic artificial muscles in order to contribute to an extended healthy life expectancy. For more information, please visit https://www.daiyak...
SYS-CON Events announced today that Fusic will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Fusic Co. provides mocks as virtual IoT devices. You can customize mocks, and get any amount of data at any time in your test. For more information, visit https://fusic.co.jp/english/.
SYS-CON Events announced today that Keisoku Research Consultant Co. will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Keisoku Research Consultant, Co. offers research and consulting in a wide range of civil engineering-related fields from information construction to preservation of cultural properties. For more information, vi...
SYS-CON Events announced today that MIRAI Inc. will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. MIRAI Inc. are IT consultants from the public sector whose mission is to solve social issues by technology and innovation and to create a meaningful future for people.
SYS-CON Events announced today that Interface Corporation will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Interface Corporation is a company developing, manufacturing and marketing high quality and wide variety of industrial computers and interface modules such as PCIs and PCI express. For more information, visit http://www.i...