{"id":737,"date":"2016-10-19T08:44:56","date_gmt":"2016-10-19T08:44:56","guid":{"rendered":"http:\/\/adriangrigoras.com\/blog\/?p=737"},"modified":"2016-10-19T08:44:56","modified_gmt":"2016-10-19T08:44:56","slug":"7-years-youtube-scalability-lessons-30-minutes","status":"publish","type":"post","link":"https:\/\/adriangrigoras.com\/blog\/7-years-youtube-scalability-lessons-30-minutes\/","title":{"rendered":"7 Years Of YouTube Scalability Lessons In 30 Minutes"},"content":{"rendered":"<p>If you started out building a dating site and instead ended up \u00a0building a video sharing site (YouTube) that handles 4 billion views a day, then it\u2019s just possible you learned something along the way. And indeed, Mike Solomon, one of the original engineers at YouTube, did learn a lot and he has given a talk about it at <a href=\"https:\/\/us.pycon.org\/2012\/\">PyCon<\/a>: <a href=\"http:\/\/www.youtube.com\/watch?v=G-lGCC4KKok\">Scalability at YouTube<\/a>.<\/p>\n<p>This isn\u2019t an architecture driven talk where we are led through a description of how a lot of boxes connect to each other. Mike could give that sort of talk. He has worked on building YouTube\u2019s servlet infrastructure, video indexing feature, video transcoding system, their full text search, a CDN, and much more. But instead, he\u2019s taken a step back, took a long look around at what time has wrought, and shared some deep lessons, obviously hard won from experience.<\/p>\n<p>The key takeaway away of the talk for me was doing <strong>a lot with really simple tools<\/strong>. While many teams are moving on to more complex ecosystems, YouTube really does keep it simple. They program primarily in Python, use MySQL as their database, they\u2019ve stuck with Apache, and even new features for such a massive site start as a very simple Python program.<\/p>\n<p>That doesn\u2019t mean YouTube doesn\u2019t do cool stuff, they do, but what makes everything work together is more a philosophy or a way of doing things than technological hocus pocus. What made YouTube into one of the world\u2019s largest websites? Read on and see&#8230;<\/p>\n<h2 dir=\"ltr\">Stats<\/h2>\n<ul>\n<li>4 billion Views a day<\/li>\n<li>60 hours of video is uploaded every minute<\/li>\n<li>350+ million devices are YouTube enabled<\/li>\n<li>Revenue double in 2010<\/li>\n<li>The number of videos has gone up 9 orders of magnitude and the number of developers has only gone up two orders of magnitude.<\/li>\n<li>1 million lines of Python code<\/li>\n<\/ul>\n<h2 dir=\"ltr\">Stack<\/h2>\n<ul>\n<li><strong>Python &#8211;<\/strong>\u00a0most of the lines of code for YouTube are still in Python. Everytime you watch a YouTube video you are executing a bunch of Python code.<\/li>\n<li><strong>Apache<\/strong> &#8211; when you think you need to get rid of it, you don\u2019t. Apache is a real rockstar technology at YouTube because they keep it simple. Every request goes through Apache.<\/li>\n<li><strong>Linux<\/strong> &#8211; the benefit of Linux is there\u2019s always a way to get in and see how your system is behaving. No matter how bad your app is behaving, you can take a look at it with Linux tools like strace and tcpdump.<\/li>\n<li><strong>MySQL<\/strong> &#8211; is used a lot. When you watch a video you are getting data from MySQL. Sometime it\u2019s used a relational database or a blob store. It\u2019s about tuning and making choices about how you organize your data.<\/li>\n<li><a href=\"http:\/\/code.google.com\/p\/vitess\/\"><strong>Vitess<\/strong><\/a> &#8211; a \u00a0new project released by YouTube, written in Go, it\u2019s a frontend to MySQL. It does a lot of optimization on the fly, it rewrites queries and acts as a proxy. Currently it serves every YouTube database request. It\u2019s RPC based.<\/li>\n<li><strong>Zookeeper<\/strong> &#8211; a distributed lock server. It\u2019s used for configuration. Really interesting piece of technology. Hard to use correctly so read the manual<\/li>\n<li><strong>Wiseguy<\/strong> &#8211; a CGI servlet container.<\/li>\n<li><strong>Spitfire<\/strong> &#8211; a templating system. It has an abstract syntax tree that let\u2019s them do transformations to make things go faster.<\/li>\n<li><strong>Serialization formats<\/strong> &#8211; no matter which one you use, they are all expensive. Measure. Don\u2019t use pickle. Not a good choice. Found protocol buffers slow. They wrote their own BSON implementation which is 10-15 time faster than the one you can download.<\/li>\n<\/ul>\n<h2 dir=\"ltr\">General Lessons<\/h2>\n<ul>\n<li><strong>Tao of YouTube<\/strong>: choose the simplest solution possible with the loosest guarantees that are practical. The reason you want all these things is you need flexibility to solve problems. The minute you over specify something you paint yourself into a corner. You aren\u2019t going to make those guarantees. Your problem becomes automatically more complex when you try and make all those guarantees. You leave yourself no way out.<\/li>\n<li><strong>That whole process is what scalability is about<\/strong>. A scalable system is one that\u2019s not in your way. That you are unaware of. It\u2019s not buzz words. It\u2019s a general problem solving ethos.<\/li>\n<li><strong>Hallmark of big system design<\/strong>: Every system is tailored to its specific requirements. Everything depends on the specifics of what you are building.<\/li>\n<li><strong>YouTube is not asynchronous<\/strong>, everything is blocking.<\/li>\n<li><strong>Believes more in philosophy than doctrine<\/strong>. Make it simple. What does that mean? You\u2019ll know when you see it. If you do code review that changes thousands of lines of code and many files then there was probably a simpler way. Your first demo should be simple, then iterate.<\/li>\n<li><strong>To solve a problem: One word &#8211; simple<\/strong>. Look for the most simple thing that will address the problem space. There are lots of complex problems, but the first solution doesn\u2019t need to be complicated. The complexity will come naturally over time.<\/li>\n<li><strong>A lot of YouTube systems start as one Python file<\/strong> and become large ecosystems after many many years. All their prototype were written in Python and survived for a surprising amount of time.<\/li>\n<li><strong>In a design review<\/strong>:\n<ul>\n<li>What\u2019s the first solution?<\/li>\n<li>How are you going to iterate?<\/li>\n<li>What do we know about how this data is going to be used?<\/li>\n<\/ul>\n<\/li>\n<li><strong>Things change over time<\/strong>. How YouTube started out has no bearing on what happens later. YouTube started out as a dating site. If they had designed for that they would have different conversation. Stay flexible.<\/li>\n<li><strong>YouTube CDN<\/strong>. Originally contracted it out. Was very expensive so they did it themselves. You can build a pretty good video CDN if you have a good hardware dude. You build a very large rack, stick machines in, then take lighttpd, and then override the 404 handler to find the video that you didn\u2019t find. That took two weeks and it\u2019s first day served 60 gigabits. You can do a lot with really simple tools.<\/li>\n<li><strong>You have to measure<\/strong>.\u00a0Vitess swapped out one its protocols for an HTTP implementation. Even though it was in C it was slow. So they ripped out HTTP and did a direct socket call using python and that was 8% cheaper on global CPU. The enveloping for HTTP is really expensive.<\/li>\n<\/ul>\n<h2 dir=\"ltr\">Scalability Techniques<\/h2>\n<ul>\n<li>These are not new ideas, but it\u2019s amazing how a few core ideas can apply in a lot different dimensions.<\/li>\n<li><strong>Divide and Conquer &#8211; The Scalability Technique<\/strong>\n<ul>\n<li>This is the scalability technique. Everything is about partitioning out work. Deciding how to execute it. Applies to many things, from web tier, you have a lot of web servers that are more or less identically and independently and you grow them horizontally. That\u2019s divide and conquer.<\/li>\n<li>This is the crux of database sharding. How do you partitions things out and communicate between the parts that you\u2019ve subdivided. These are things you want to figure out early on because they influence how you grow.<\/li>\n<li>Simple and loose connections are really valuable.<\/li>\n<li>The dynamic nature of Python is a win here. No matter how bad your API is you can stub or modify or decorate your way out of a lot of problems.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Approximate Correctness &#8211; Cheat a Little<\/strong>\n<ul>\n<li>Another favorite technique. The state of the system is that which it is reported to be. If a user can\u2019t tell a part of the system is skewing and inconsistent, then it\u2019s not.<\/li>\n<li>A real world example. If you write a comment and someone loads the page at the same time, they might not get it for 300-400ms, the user who is reading won\u2019t care. The writer of the comment will care, so you make sure the user who wrote the comment will see it. So you cheat a little bit. Your system doesn\u2019t have to have globally consistent transactions. That would be super expensive and overkill. Not every comment is a financial transaction. So know when you can cheat.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Expert Knob Twiddling<\/strong>\n<ul>\n<li>Ask, what do you know about your consistency model? For comments is eventually consistent good enough? Renting a movie is different. When renting there\u2019s money so we\u2019ll do the best we can to never lose that. Different consistency models are needed depending on the data.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Jitter &#8211; Add Entropy Back into Your System<\/strong>\n<ul>\n<li>Hot word in their group all of the time. If your system doesn\u2019t jitter then you get<a href=\"http:\/\/highscalability.com\/blog\/2008\/3\/14\/problem-mobbing-the-least-used-resource-error.html\">thundering herds<\/a>. Distributed applications are really weather systems. Debugging them is as deterministic as predicting the weather. Jitter introduces more randomness because surprisingly, things tend to stack up.<\/li>\n<li>For example, cache expirations. For a popular video they cache things as best they can. The most popular video they might cache for 24 hours. If everything expires at one time then every machine will calculate the expiration at the same time. This creates a thundering herd.<\/li>\n<li>By jittering you are saying \u00a0randomly expire between 18-30 hours. That prevents things from stacking up. They use this all over the place. Systems have a tendency to self synchronize as operations line up and try to destroy themselves. Fascinating to watch. You get slow disk system on one machine and everybody is waiting on a request so all of a sudden all these other requests on all these other machines are completely synchronized. This happens when you have many machines and you have many events. Each one actually removes entropy from the system so you have to add some back in.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Cheating &#8211; Know How to Fake Data<\/strong>\n<ul>\n<li>Awesome technique. The fastest function call is the one that doesn\u2019t happen. When you have a monotonically increasing counter, like movie view counts or profile view counts, you could do a transaction every update. Or you could do a transaction every once in awhile and update by a random amount and as long as it changes from odd to even people would probably believe it\u2019s real. Know how to fake data.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Scalable Components &#8211; Make Your own Luck<\/strong>\n<ul>\n<li><strong>You can look at an API and get a good feel<\/strong>. Are the inputs well defined? Do you know what you are getting out? A lot of this ends up being about data. Have a tight specification of what data comes out every function and how it flows actually helps you understand the application without documentation. You can tell what\u2019s happening before and after a function is called.<\/li>\n<li><strong>In Python things tend to move towards RPCs<\/strong>. The structure of your code is based on the discipline of your programmers. So establish good conventions, when all else fails there\u2019s an RPC wall so you know what goes in and what comes out.<\/li>\n<li><strong>Your components will not be perfect<\/strong>. A component might last a month or six months, who knows. By drawing these lines you are making some of your own luck. When things go south you can swap it out and do something different. Sometimes that rewriting someing in python and C and sometimes that means getting rid of it entirely. You don\u2019t know until you are able to observe.<\/li>\n<li><strong>With so many people on a team nobody can know the whole system<\/strong>, so you need to define components. This is video transcode it\u2019s distinct from video search. You want well defined subcomponents. It\u2019s good software design. These things end up talking to each other so having a good data specification is helpful. The greatest sin he made was communication between the servlet layer and the template layer to be a dictionary. Very bad idea. Should have added a WatchPage and said a watch page had a video and some comments and some related videos. This caused a lot of problems because the dictionary can have a few hundred attributes. They don\u2019t always make the right choice.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Efficiency &#8211; Traded Off for Scalability<\/strong>\n<ul>\n<li><strong>Efficiency is traded off for scalability<\/strong>. The most efficient thing is to write it in C and cram it into one process, but that\u2019s not scalable.<\/li>\n<li><strong>Focus on the macro level<\/strong><strong>, your components, and how they break out<\/strong>. Does it makes sense to do this an RPC or do it inline? Break it into a subpackage and just someday this may be different.<\/li>\n<li><strong>Focus on algorithms<\/strong>. In Python the effort to implement a good algorithm is low. There\u2019s the bisect module, for example, where you can take a list, do something meaningful, and serialize it to disk and read it back again. There\u2019s a penalty versus C, but it\u2019s very easy.<\/li>\n<li><strong>Measurement<\/strong>. In Python measurement is like reading tea leaves. There\u2019s a lot of things in Python that are counter intuitive, like the cost of grabage colleciton. Most of chunks of their apps spend their time serializing. Profiling serialization is very depending on what you are putting in. Serializing ints is very different than serializing big blobs.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Efficiency in Python &#8211; Knowing What Not to Do<\/strong>\n<ul>\n<li><strong>More about knowing what not to do<\/strong>. How dynamic you make things correlates to how expensive it is to run your Python app.<\/li>\n<li><strong>Dummer code is easier to grep for and easier to maintain<\/strong>. The more magical the code is the harder is to figure out how it works.<\/li>\n<li><strong>They don\u2019t do a lot of OO<\/strong>. They use a lot of namespaces. Use classes to organize data, but rarely for OO.<\/li>\n<li><strong>What is your code tree going to look like?<\/strong> He wants these words to describe it: simple, pragmatic, elegant, orthogonal, composable. This is an ideal, reality is a bit different.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<p>source:\u00a0http:\/\/highscalability.com\/blog\/2012\/3\/26\/7-years-of-youtube-scalability-lessons-in-30-minutes.html<\/p>\n","protected":false},"excerpt":{"rendered":"<p>If you started out building a dating site and instead ended up \u00a0building a video sharing site (YouTube) that handles 4 billion views a day, then it\u2019s just possible you learned something along the way. And indeed, Mike Solomon, one of the original engineers at YouTube, did learn a lot and he has given a\u2026 <span class=\"read-more\"><a href=\"https:\/\/adriangrigoras.com\/blog\/7-years-youtube-scalability-lessons-30-minutes\/\">Read More &raquo;<\/a><\/span><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[30],"tags":[],"class_list":["post-737","post","type-post","status-publish","format-standard","hentry","category-architecture"],"_links":{"self":[{"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/posts\/737","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/comments?post=737"}],"version-history":[{"count":1,"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/posts\/737\/revisions"}],"predecessor-version":[{"id":738,"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/posts\/737\/revisions\/738"}],"wp:attachment":[{"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/media?parent=737"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/categories?post=737"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/adriangrigoras.com\/blog\/wp-json\/wp\/v2\/tags?post=737"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}