AlphaGo's Reasoning vs. LLMs: The Debate Over AI's True Intelligence
In March 2016, a program named AlphaGo, developed with contributions from the author, made a pivotal move on a Go board in Seoul that initially appeared to be a mistake. Move 37 in game two of a five-game match against Lee Sedol, one of the greatest professional Go players, seemed so absurd that commentators speculated it was a programming glitch. AlphaGo, however, went on to win that game and ultimately triumphed 4-1 over Sedol, prompting the player to remark, “I thought AlphaGo was based on probability calculation and that it was merely a machine. But when I saw this move, I changed my mind. Surely, AlphaGo is creative.”

This victory contrasted sharply with Deep Blue's defeat of Garry Kasparov in 1997, which relied on evaluating 200 million chess positions per second using hard-coded rules. Go, a vastly more complex game, demands a different approach. AlphaGo succeeded by sensing who was ahead and inventing moves no human had conceived, demonstrating a capability beyond brute-force calculation that would take supercomputers billions of years to compute for even a fraction of possibilities.
Many accounts mistakenly attribute Move 37 to pure machine intuition. However, the author asserts that it was AlphaGo’s powers of reasoning that facilitated this creative choice—a capability that today’s Large Language Models (LLMs) currently lack. For future AI systems to deliver trustworthy results and genuinely novel insights in critical fields like science and medicine, equipping them with such reasoning capabilities is deemed essential.
AlphaGo's architecture comprised two distinct systems: a policy network and search machinery. The policy network, trained to predict strong human moves, regarded Move 37 as having a mere one in 10,000 chance of being played by an expert. It was the search machinery that enabled AlphaGo to look beyond immediate plausibility, construct a game tree with thousands of branches, and weigh the future consequences of proposed moves, ultimately leading to the selection of Move 37.
This dual system mirrored Daniel Kahneman's theory of human thought: System 1 (fast, intuitive) and System 2 (slow, deliberative). AlphaGo's networks provided the 'hunches,' while its search mechanism supplied the 'deliberation,' testing these hunches against subsequent moves. Neither component functioned effectively in isolation; intuition alone would not have chosen Move 37, and brute-force search would have struggled to filter the vast number of possibilities.
In contrast, contemporary AI models, such as LLMs like ChatGPT, operate differently. They primarily function by predicting the next token, which aligns with System 1—fast, associative, and adept at pattern completion. While the field recognized that language fluency alone was insufficient for true usefulness, the development of 'chain of thought' processes aimed to introduce deliberation. These processes generate intermediate steps to decompose problems and influence subsequent reasoning, leading to real gains, particularly in mathematics and coding.
However, this 'chain of thought' mechanism, unlike AlphaGo's search, does not introduce a genuinely separate reasoning process. The intermediate reasoning is still generated by the same iterated next-token prediction. Chatbots exhibit three key shortcomings preventing their processes from qualifying as genuine reasoning: they lack an explicit, persistent epistemic state; they do not cleanly separate knowledge from how it is manipulated; and research indicates they often concoct chains of thought after reaching an answer, rather than using them for genuine deliberation.
This distinction is crucial for high-stakes applications in medicine, engineering, and scientific research, where understanding 'how' a system reaches its conclusion is as important as the conclusion itself. When errors occur, pinpointing whether the reasoning was flawed, evidence invalid, or assumptions incorrect is vital. The author recently departed Google DeepMind to pursue a fresh approach to machine reasoning, drawing inspiration from AlphaGo's architecture, specifically its ability to maintain and update a game tree as a record of its knowledge and considered possibilities.
What to watch: The development of AI architectures that integrate explicit reasoning mechanisms akin to AlphaGo.
Editor's note: The draft excellently synthesizes the author's technical argument regarding AlphaGo, System 2 reasoning, and the limitations of current LLMs.
AI-generated and fact-checked against the original report; claims the gate cannot verify are held back.