{"id":25176,"date":"2019-01-03T11:58:07","date_gmt":"2019-01-03T11:58:07","guid":{"rendered":"http:\/\/wealinternational.com.br\/?p=25176"},"modified":"2019-01-03T11:58:07","modified_gmt":"2019-01-03T11:58:07","slug":"how-the-artificial-intelligence-program-alphazero-mastered-its-games-by-james-somers","status":"publish","type":"post","link":"https:\/\/wealinternational.com.br\/en\/how-the-artificial-intelligence-program-alphazero-mastered-its-games-by-james-somers\/","title":{"rendered":"How the Artificial-Intelligence Program AlphaZero Mastered Its Games (By James Somers)"},"content":{"rendered":"<p><\/p>\n<div class=\"SectionBreak SectionBreak__sectionBreak___1ppA7\">\n<p>&#8220;A few weeks ago, a group of researchers from Google\u2019s artificial-intelligence subsidiary, DeepMind, published a <a class=\"ArticleBody__link___1FS03\" href=\"http:\/\/science.sciencemag.org\/content\/362\/6419\/1140\" target=\"_blank\" rel=\"noopener\">paper<\/a> in the journal <em class=\"\">Science<\/em> that described an A.I. for playing games. While their system is general-purpose enough to work for many two-person games, the researchers had adapted it specifically for Go, chess, and shogi (\u201cJapanese chess\u201d); it was given no knowledge beyond the rules of each game. At first it made random moves. Then it started learning through self-play. Over the course of nine hours, the chess version of the program played forty-four million games against itself on a massive cluster of specialized Google hardware. After two hours, it began performing better than human players; after four, it was beating the best chess engine in the world.<\/p>\n<div class=\"Callout__inset-right___3Etg2\" data-type=\"callout\" data-callout=\"inset-right\">\n<div class=\"CuratedEmbed__container___1VdYz\"><img decoding=\"async\" title=\"\" src=\"https:\/\/media.newyorker.com\/photos\/5c24f4778822322ea4b3befe\/master\/w_727,c_limit\/Somers-AlphaZero.jpg\" alt=\"\" \/><\/div>\n<div>\n<div>\n<p>In 2016, a Google program soundly defeated Lee Sedol, the world\u2019s best Go player, in a match viewed by more than a hundred million people.\u00a0<small class=\"ImageCaption__credit___rg3mC \">Photograph by Ahn Young-joon \/ AP<\/small><\/p>\n<\/div>\n<\/div>\n<\/div>\n<p>The program, called AlphaZero, descends from AlphaGo, an A.I. that became known for defeating Lee Sedol, the world\u2019s best Go player, in March of 2016. Sedol\u2019s defeat was a stunning upset. In \u201cAlphaGo,\u201d a documentary released earlier this year on Netflix, the filmmakers follow both the team that developed the A.I. and its human opponents, who have devoted their lives to the game. We watch as these humans experience the stages of a new kind of grief. At first, they don\u2019t see how they can lose to a machine: \u201cI believe that human intuition is still too advanced for A.I. to have caught up,\u201d Sedol says, the day before his five-game match with AlphaGo. Then, when the machine starts winning, a kind of panic sets in. In one particularly poignant moment, Sedol, under pressure after having lost his first game, gets up from the table and, leaving his clock running, walks outside for a cigarette. He looks out over the rooftops of Seoul. (On the Internet, more than fifty million people were watching the match.) Meanwhile, the A.I., unaware that its opponent has gone anywhere, plays a move that commentators called creative, surprising, and beautiful. In the end, Sedol lost, 1-4. Before there could be acceptance, there was depression. \u201cI want to apologize for being so powerless,\u201d he said in a press conference. Eventually, Sedol, along with the rest of the Go community, came to appreciate the machine. \u201cI think this will bring a new paradigm to Go,\u201d he said. Fan Hui, the European champion, agreed. \u201cMaybe it can show humans something we\u2019ve never discovered. Maybe it\u2019s beautiful.\u201d<\/p>\n<p>AlphaGo was a triumph for its creators, but still unsatisfying, because it depended so much on human Go expertise. The A.I. learned which moves it should make, in part, by trying to mimic world-class players. It also used a set of hand-coded heuristics to avoid the worst blunders when looking ahead in games. To the researchers building AlphaGo, this knowledge felt like a crutch. They set out to build a new version of the A.I. that learned on its own, as a \u201ctabula rasa.\u201d<\/p>\n<p>The result, AlphaGo Zero, detailed in a <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/www.nature.com\/articles\/nature24270\" target=\"_blank\" rel=\"noopener\">paper<\/a> published in October, 2017, was so called because it had zero knowledge of Go beyond the rules. This new program was much less well-known; perhaps you can ask for the world\u2019s attention only so many times. But in a way it was the more remarkable achievement, one that no longer had much to do with Go at all. In fact, less than two months later, DeepMind published a <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/arxiv.org\/abs\/1712.01815\" target=\"_blank\" rel=\"noopener\">preprint<\/a> of a third paper, showing that the algorithm behind AlphaGo Zero could be generalized to any two-person, zero-sum game of <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/en.wikipedia.org\/wiki\/Perfect_information\" target=\"_blank\" rel=\"noopener\">perfect information<\/a> (that is, a game in which there are no hidden elements, such as face-down cards in poker). DeepMind dropped the \u201cGo\u201d from the name and christened its new system AlphaZero. At its core was an algorithm so powerful that you could give it the rules of humanity\u2019s richest and most studied games and, later that day, it would become the best player there has ever been. Perhaps more surprising, this iteration of the system was also by far the simplest.<\/p>\n<\/div>\n<div class=\"SectionBreak SectionBreak__sectionBreak___1ppA7\">\n<p>A typical chess engine is a hodgepodge of tweaks and shims made over decades of trial and error. The best engine in the world, Stockfish, is open source, and it gets better by a kind of Darwinian selection: someone suggests an idea; tens of thousands of games are played between the version with the idea and the version without it; the best version wins. As a result, it is not a particularly elegant program, and it can be hard for coders to understand. Many of the changes programmers make to Stockfish are best formulated in terms of chess, not computer science, and concern how to evaluate a given situation on the board: Should a knight be worth 2.1 points or 2.2? What if it\u2019s on the third rank, and the opponent has an opposite-colored bishop? To illustrate this point, David Silver, the head of research at DeepMind, once listed the moving parts in Stockfish. There are more than fifty of them, each requiring a significant amount of code, each a bit of hard-won chess arcana: the Counter Move Heuristic; databases of known endgames; evaluation modules for Doubled Pawns, Trapped Pieces, Rooks on (Semi) Open Files, and so on; strategies for searching the tree of possible moves, like \u201caspiration windows\u201d and \u201citerative deepening.\u201d<\/p>\n<p>AlphaZero, by contrast, has only two parts: a neural network and an algorithm called Monte Carlo Tree Search. (In a nod to the gaming mecca, mathematicians refer to approaches that involve some randomness as \u201cMonte Carlo methods.\u201d) The idea behind M.C.T.S., as it\u2019s often known, is that a game like chess is really a tree of possibilities. If I move my rook to d8, you could capture it or let it be, at which point I could push a pawn or move my bishop or protect my queen. . . . The trouble is that this tree gets incredibly large incredibly quickly. No amount of computing power would be enough to search it exhaustively. An expert human player is an expert precisely because her mind automatically identifies the essential parts of the tree and focusses its attention there. Computers, if they are to compete, must somehow do the same.<\/p>\n<p>This is where the neural network comes in. AlphaZero\u2019s neural network receives, as input, the layout of the board for the last few moves of the game. As output, it estimates how likely the current player is to win and predicts which of the currently available moves are likely to work best. The M.C.T.S. algorithm uses these predictions to decide where to focus in the tree. If the network guesses that \u2018knight-takes-bishop\u2019 is likely to be a good move, for example, then the M.C.T.S. will devote more of its time to exploring the consequences of that move. But it balances this \u201cexploitation\u201d of promising moves with a little \u201cexploration\u201d: it sometimes picks moves it thinks are unlikely to bear fruit, just in case they do.<\/p>\n<p>At first, the neural network guiding this search is fairly stupid: it makes its predictions more or less at random. As a result, the Monte Carlo Tree Search starts out doing a pretty bad job of focussing on the important parts of the tree. But the genius of AlphaZero is in how it learns. It takes these two half-working parts and has them hone each other. Even when a dumb neural network does a bad job of predicting which moves will work, it\u2019s still useful to look ahead in the game tree: toward the end of the game, for instance, the M.C.T.S. can still learn which positions actually lead to victory, at least some of the time. This knowledge can then be used to improve the neural network. When a game is done, and you know the outcome, you look at what the neural network predicted for each position (say, that there\u2019s an 80.2 per cent chance that castling is the best move) and compare that to what actually happened (say, that the percentage is more like 60.5); you can then \u201ccorrect\u201d your neural network by tuning its synaptic connections until it prefers winning moves. In essence, all of the M.C.T.S.\u2019s searching is distilled into new weights for the neural network.<\/p>\n<div class=\"SectionBreak SectionBreak__sectionBreak___1ppA7\">\n<p>With a slightly better network, of course, the search gets slightly less misguided\u2014and this allows it to search better, thereby extracting better information for training the network. On and on it goes, in a feedback loop that ratchets up, very quickly, toward the plateau of known ability.<\/p>\n<\/div>\n<div class=\"SectionBreak SectionBreak__sectionBreak___1ppA7\">\n<p>When the AlphaGo Zero and AlphaZero papers were published, a small army of enthusiasts began describing the systems in <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/medium.com\/applied-data-science\/how-to-build-your-own-alphazero-ai-using-python-and-keras-7f664945c188\" target=\"_blank\" rel=\"noopener\">blog posts<\/a> and <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/www.youtube.com\/watch?v=Fbs4lnGLS8M\" target=\"_blank\" rel=\"noopener\">YouTube videos<\/a> and building their own <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/github.com\/suragnair\/alpha-zero-general\" target=\"_blank\" rel=\"noopener\">copycat versions<\/a>. Most of this work was explanatory\u2014it flowed from the amateur urge to learn and share that gave rise to the Web in the first place. But a couple of efforts also sprung up to replicate the work at a large scale. The DeepMind papers, after all, had merely described the greatest Go- and chess-playing programs in the world\u2014they hadn\u2019t contained the source code, and the company hadn\u2019t made the programs themselves available to players. Having declared victory, its engineers had departed the field.<\/p>\n<p>Gian-Carlo Pascutto, a computer programmer who works at the Mozilla Corporation, had a track record of building competitive game engines, first in chess, then in Go. He followed the latest research. As the combination of Monte Carlo Tree Search and a neural network became the state of the art in Go A.I.s, Pascutto built the world\u2019s most successful open-source Go engines\u2014first <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/www.sjeng.org\/leela.html\" target=\"_blank\" rel=\"noopener\">Leela<\/a>, then <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/github.com\/gcp\/leela-zero\" target=\"_blank\" rel=\"noopener\">LeelaZero<\/a>\u2014which mirrored the advances made by DeepMind. The trouble was that DeepMind had access to Google\u2019s vast cloud and Pascutto didn\u2019t. To train its Go engine, DeepMind used five thousand of Google\u2019s \u201cTensor Processing Units\u201d\u2014chips specifically designed for neural-network calculations\u2014for thirteen days. To do the same work on his desktop system, Pascutto would have to run it for seventeen hundred years.<\/p>\n<p>To compensate for his lack of computing power, Pascutto distributed the effort. LeelaZero is a federated system: anyone who wants to participate can download the latest version, donate whatever computing power he has to it, and upload the data he generates so that the system can be slightly improved. The distributed LeelaZero community has had their system play more than ten million games against itself\u2014a little more than AlphaGo Zero. It is now one of the strongest existing Go engines.<\/p>\n<div class=\"\" data-type=\"callout\" data-callout=\"inline-recirc\">\n<div id=\"RecircCarousel\" class=\"recircCarouselUnit RecirculationCarousel__carousel___3oV_x\">\n<div class=\"SectionBreak SectionBreak__sectionBreak___1ppA7\">\n<p>It wasn\u2019t long before the idea was extended to chess. In December of last year, when the AlphaZero preprint was published, \u201cit was like a bomb hit the community,\u201d Gary Linscott said. Linscott, a computer scientist who had worked on Stockfish, used the existing LeelaZero code base, and the new ideas in the AlphaZero paper, to create <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/github.com\/LeelaChessZero\/lczero\" target=\"_blank\" rel=\"noopener\">Leela Chess Zero<\/a>. (For Stockfish, he had developed a testing framework so that new ideas for the engine could be distributed to a fleet of volunteers, and thus vetted more quickly; distributing the training for a neural network was a natural next step.) There were kinks to sort out, and educated guesses to make about details that the DeepMind team had left out of their papers, but within a few months the neural network began improving. The chess world was already obsessed with AlphaZero: <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/www.chess.com\/survey\/is-google-s-alphazero-the-best-chess-player-on-the-planet\" target=\"_blank\" rel=\"noopener\">posts on<\/a> chess.com celebrated the engine; commentators and grandmasters <a class=\"ArticleBody__link___1FS03\" href=\"https:\/\/www.youtube.com\/watch?v=lFXJWPhDsSY\" target=\"_blank\" rel=\"noopener\">pored over<\/a> the handful of AlphaZero games that DeepMind had released with their paper, declaring that this was \u201chow chess ought to be played,\u201d that the engine \u201cplays like a human on fire.\u201d Quickly, Lc0, as Leela Chess Zero became known, attracted hundreds of volunteers. As they contributed their computer power and improvements to the source code, the engine got even better. Today, one core contributor suspects that it is just a few months away from overtaking Stockfish. Not long after, it may become better than AlphaZero itself.<\/p>\n<p>When we spoke over the phone, Linscott marvelled that a project like his, which would once have taken a talented doctoral student several years, could now be done by an interested amateur in a couple of months. Software libraries for neural networks allow for the replication of a world-beating design using only a few dozen lines of code; the tools already exist for distributing computation among a set of volunteers, and chipmakers such as Nvidia have put cheap and powerful G.P.U.s\u2014graphics-processing chips, which are perfect for training neural networks\u2014into the hands of millions of ordinary computer users. An algorithm like M.C.T.S. is simple enough to be implemented in an afternoon or two. You don\u2019t even need to be an expert in the game for which you\u2019re building an engine. When he built LeelaZero, Pascutto hadn\u2019t played Go for about twenty years.<\/p>\n<\/div>\n<div class=\"SectionBreak SectionBreak__sectionBreak___1ppA7\">\n<p>David Silver, the head of research at DeepMind, has pointed out a seeming paradox at the heart of his company\u2019s recent work with games: the simpler its programs got\u2014from AlphaGo to AlphaGo Zero to AlphaZero\u2014the better they performed. \u201cMaybe one of the principles that we\u2019re after,\u201d he said, in a talk in December of 2017, \u201cis this idea that by doing less, by removing complexity from the algorithm, it enables us to become more general.\u201d By removing the Go knowledge from their Go engine, they made a better Go engine\u2014and, at the same time, an engine that could play shogi and chess.<\/p>\n<p>It was never obvious that things would turn out this way. In 1953, Alan Turing, who helped create modern computing, wrote a short paper titled, \u201cDigital Computers Applied to Games.\u201d In it, he developed a chess program \u201cbased on an introspective analysis of my thought processes while playing.\u201d The program was simple, but in its case simplicity was no virtue: like Turing, who wasn\u2019t a gifted chess player, it missed much of the depth of the game and didn&#8217;t play very well. Even so, Turing conjectured that the idea that \u201cone cannot programme a machine to play a better game than one plays oneself\u201d was a \u201crather glib view.\u201d Although it sounds right to say that \u201cno animal can swallow an animal heavier than itself,\u201d plenty of animals can. Similarly, Turing suggested, there might be no contradiction in a bad chess player making a chess program that plays brilliantly. One tantalizing way to do it would be to have the program learn for itself.<\/p>\n<p>The success of AlphaZero seems to bear this out. It has a simple structure, but it\u2019s capable of learning surprisingly deep features of the games it plays. In one section of the AlphaGo Zero paper, the DeepMind team illustrates how their A.I., after a certain number of training cycles, discovers strategies well-known to master players, only to discard them just a few cycles later. It is odd and a little unsettling to see humanity\u2019s best ideas trundled over on the way to something better; it hits close to home in a way that seeing a physical machine exceed us\u2014a bulldozer shifting a load of earth, say\u2014doesn\u2019t. In a recent editorial in <em class=\"\">Science<\/em>, Garry Kasparov, the former chess champion who lost to I.B.M.\u2019s Deep Blue in 1997, argues that AlphaZero doesn\u2019t play chess in a way that reflects the presumably systematic \u201cpriorities and prejudices of programmers\u201d; instead\u2014even though it searches far fewer positions per move than a traditional engine\u2014it plays in an open, aggressive style and seems to think in terms of strategy rather than tactics, like a human with uncanny vision. \u201cBecause AlphaZero programs itself,\u201d Kasparov writes, \u201cI would say that its style reflects the truth.\u201d<\/p>\n<p>Playing chess like a human, of course, isn&#8217;t the same thing as thinking about chess like a human, or learning like one. There is an old saying that game-playing is the <em class=\"\">Drosophila<\/em> of A.I.: as the fruit fly is to biologists, so games like Go and chess are to computer scientists studying the mechanisms of intelligence. It\u2019s an evocative analogy. And yet it could be that the task of playing chess, once it\u2019s converted into the task of searching tens of thousands of nodes per second in a game tree, exercises a different kind of intelligence than the one we care about most. Played in this way, chess might be more like earth-moving than we thought: an activity that, in the end, isn\u2019t our fort\u00e9, and so shouldn\u2019t be all that dear to our souls. To learn, AlphaZero needs to play millions more games than a human does\u2014 but, when it\u2019s done, it plays like a genius. It relies on churning faster than a person ever could through a deep search tree, then uses a neural network to process what it finds into something that resembles intuition. Surely the program teaches us something new about intelligence. But its success also underscores just how much the world\u2019s best human players can see by means of a very different process\u2014one based on reading, talking, and feeling, in addition to playing. What may be most surprising is that we humans have done as well as we have in games that seem, now, to have been made for machines.&#8221;<\/p>\n<p>Author:\u00a0James Somers,\u00a0a writer and a programmer based in New York.<\/p>\n<p>Source:\u00a0https:\/\/www.newyorker.com\/science\/elements\/how-the-artificial-intelligence-program-alphazero-mastered-its-games<\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<p><\/p>","protected":false},"excerpt":{"rendered":"<p>&#8220;A few weeks ago, a group of researchers from Google\u2019s artificial-intelligence subsidiary, DeepMind, published a paper in the journal Science that described an A.I. for playing games. While their system is general-purpose enough to work for many two-person games, the researchers had adapted it specifically for Go, chess, and shogi (\u201cJapanese chess\u201d); it was given [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_et_pb_use_builder":"","_et_pb_old_content":"","_et_gb_content_width":"","footnotes":""},"categories":[17,16],"tags":[38,22,24],"_links":{"self":[{"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/posts\/25176"}],"collection":[{"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/comments?post=25176"}],"version-history":[{"count":2,"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/posts\/25176\/revisions"}],"predecessor-version":[{"id":25178,"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/posts\/25176\/revisions\/25178"}],"wp:attachment":[{"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/media?parent=25176"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/categories?post=25176"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wealinternational.com.br\/en\/wp-json\/wp\/v2\/tags?post=25176"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}