Few scientific debates are as loud today as the one about large language models. Do they understand anything, or merely imitate? Are they reasoning? Do they need symbols, world models or even bodies before we can call them intelligent? While these questions are general and have implications that will affect everyone’s life and jobs, these arguments are usually presented as technical; to an outsider (like I was a few years ago) they may even seem as recent as the current technology. In reality, almost all of them go a long way back to the early days of AI. In this essay, I will argue that to fully understand the current controversy, one has to understand where it comes from. And in that light, those questions take a rather different, more human shape.
A quarrel with a long history
The history of artificial intelligence and machine learning has always been a history of conceptual disagreement between people who held different views of how a mind works and of what it means to learn and to solve problems. This is easy to forget now that AI looks like a branch of computer science but in its first decades it was nothing of the sort: the field of AI predates the first departments of computer science and its founders were neurophysiologists, psychologists, linguists, mathematicians and at least one political scientist, who brought with them their disciplines’ oldest quarrels. Is learning the strengthening of associations or the discovery of rules? Is knowledge built from experience or largely innate? Is thinking the manipulation of symbols, or something that happens in the graded activity of many simple units?
Broadly, the answers split the founders into two tribes. The symbolists held that thinking is the manipulation of symbols according to explicit rules, much as a mathematician manipulates an equation or a grammarian parses a sentence; the content lands on pre-written rules and moves within limits set by those rules. To learn how to play chess, one needs to learn the rules that move every piece first and if we deviate from those rules, we are no longer playing chess. According to symbolists, an intelligent machine is a program into which we write knowledge and above all the rules for using it. Their roots were in logic, mathematics and linguistics, to the point that Chomsky could be considered the most influential intellectual ally of that approach.
The connectionists held instead that intelligence emerges from large networks of simple, neuron-like units whose connections are strengthened or weakened by experience: nobody writes the knowledge in, because the machine learns it from examples. We are not born knowing things and, argues a connectionist, we are not born knowing rules either. Their roots were in neurophysiology and psychology. Rules against learning, symbols against networks: almost every argument in this essay is a variation on that split, with a few hybrids in between and, as we will see, a third party sceptical of both.
As is often the case in academia, some of these quarrels were also personal. They were carried on for decades, sometimes between people who had known each other since school, and they shaped who was funded, who was published and who was heard. Many of us in academia can say by experience that more often than not, scientists fight about attention and ego more than they fight about truth.
It is now generally agreed that the field of AI went through two long periods of stagnation, known as the AI winters. In both cases, it was the symbolists who had the louder voice whenever the field stagnated. And in both cases, the way out came not from winning an argument, but from a technical advance made by the connectionists: backpropagation ended the first winter, and graphics processors the second. Large language models have taken us into what I would call the first spring.
A map of the quarrel
I want to emphasise that behind this argument there are people, and their very human interactions. The figure below, therefore, is the backbone of the first part of this essay as it places the main actors and their intellectual contribution. The map is best read by position, from left to right, and by season, from top to bottom, with arrows for who trained whom, who influenced whom and who attacked whom.
Read over time, the map shows that history has followed the same pattern twice:
- the connectionists come up with some amazingly surprising finding or observation, often announced with excessive and annoying fanfare;
- the symbolists sceptically point out and describe the limitations of the system;
- the field stagnates into a winter of discontent in which activity and funding decline;
- a new technology or intuition rescues the connectionists and moves the field to a new age of progress, and the cycle restarts from (1).
We are now on the third iteration and have survived two long winters.
The first summer: the founders (1943–1969)
The first model of a neural network was proposed in 1943 by Warren McCulloch, a neurophysiologist and psychiatrist, and Walter Pitts, a self-taught logician (McCulloch and Pitts, 1943). Their neurons were logic gates, which makes the founding document of connectionism also a founding document of symbolic computation: at the start, the two traditions were one. In 1949 the psychologist Donald Hebb proposed that learning consists in strengthening the connections between cells that are active together. His intuition is often summarised as “neurons that fire together, wire together” and it is still a founding concept of Neuroscience. More or less at the same time, Alan Turing had argued that machines could think and a few years later, in 1958, Frank Rosenblatt, a research psychologist at Cornell, finally joined the dots and turned that idea into a learning machine: the perceptron.
The perceptron was not just a concept but a gigantic electromechanical machine inspired by the biological understanding of how a nervous system works in minimal terms. Rosenblatt started with a software simulator but eventually he wanted to create a robotic brain, composed of input neurons receiving visual stimuli (a retina), a group of processing neurons hidden in the middle, and a group of output neurons translating stimuli into action1. Its retina was a twenty-by-twenty grid of photocells, wired through a plugboard to a bank of “association units”. The initial wiring was random because Rosenblatt believed that is the state in which animals are born: random connections (or weights). The weight of each association unit’s vote on the output was held in a physical potentiometer, and when the machine gave a wrong answer, small electric motors physically turned the dials.
Rosenblatt proved that this simple rule, adjusting the weights only when the machine was wrong, would always find a solution whenever one existed, and the machine learned to recognise simple shapes such as triangles. The press, for its part, reported that the Navy expected the perceptron’s descendants to walk, talk, see and even become aware of themselves (the Cornell Chronicle tells the story well). The Mark I Perceptron machine now sits in the Smithsonian.
The other tribe formed more or less at the same time. At the Dartmouth workshop of 1956, the mathematicians John McCarthy and Marvin Minsky, with Claude Shannon and Nathaniel Rochester, launched “artificial intelligence” on the conjecture that every feature of intelligence could be described precisely enough for a machine to simulate it. One interesting aspect of that proposal (which is two pages long and worth reading) is that it immediately identified what I think is still the elephant in the symbolists’ room: the role of creativity and imagination which, by definition, should span beyond the containment of rules. This is worth a separate essay in which I would argue that you can either have a machine that creatively solves new science or a machine that is intrinsically based on rules and never makes mistakes.
At any rate, at the Dartmouth workshop Allen Newell and Herbert Simon, the latter a political scientist by training, showed a program that proved theorems by searching through symbols, a line of work they would later codify as the physical symbol system hypothesis. In 1959 the linguist Noam Chomsky published a review of B. F. Skinner’s Verbal Behavior that later became the manifesto against associationism: language, he argued, could not be learned from associations alone, so the mind must come equipped with rules. Again, these were considered problems not of computer science (computer science departments did not exist until 1962) but of cognition, language and neuroscience. Marvin Minsky himself liked to repeat the quip that “any field that feels the urge to put the word science in its name, is not a science“.
The first and most consequential feud was the one between Minsky and Rosenblatt and, to a certain extent, the two very similar machines they had built: Rosenblatt’s perceptron and Minsky’s SNARC. In 1951, as a Princeton graduate student, Minsky conceived and built SNARC, one of the very first learning machines, made of forty artificial neurons built from vacuum tubes and clutches. The machine was meant to emulate a rat moving in a maze, and each decision of the virtual animal could be rewarded by turning knobs, each associated with a neuron. The knobs were operated automatically by a motor salvaged from a B-24 bomber’s autopilot, cranking a chain, thus nudging the neurons and their weights. SNARC was impressive but Minsky soon realised it had properties that were hard to explain. For instance, it was built to simulate one rat, but somehow it could follow two or three rats in the maze at once. Minsky found himself stuck and followed one of the guiding principles that drove all his career: “when you find yourself stuck on a problem, it means it’s too difficult for you. Move on“. And so he did. He donated the machine to graduate students at Dartmouth who cannibalised it for parts, and only one bone remains.
Both SNARC and the perceptron were funded by the US Navy’s Office of Naval Research, but only Rosenblatt’s machine was presented to the press as the embryo of a thinking computer. Partly because of the arms race of the time, the Navy was eager to sell it as the first step towards the conscious machines described above. Minsky was, by most accounts, irritated by those claims, which he considered wildly inflated. Through the 1960s the two clashed loudly at conferences, although they are said to have remained on friendly terms. I find the video below to give a particularly well documented summary of that time.
There is also an important human factor in the first quarrel between Rosenblatt and Minsky. The two had actually been schoolmates at the Bronx High School of Science, a school for young intellectual prodigies (nine of its graduates went on to win Nobel Prizes). It is the kind of environment that probably inflates self-confidence and intellectual rivalry in equal parts. Minsky had grown up in such places: before Bronx Science he attended the Ethical Culture School in Manhattan, where Oppenheimer had studied a generation earlier.
Before I move to the next season, I want to take a digression and highlight one aspect that was common to many of the incredible people I mentioned so far. These were all multifaceted, multidisciplinary geniuses who gave immense contributions to multiple fields. Between building SNARC and founding the MIT AI Lab, Marvin Minsky invented the confocal microscope, now a workhorse of biology laboratories everywhere. Between computers and codebreaking, Turing wrote one of the founding papers of mathematical biology, showing how chemical reactions could lay down the spots and stripes of a developing embryo. McCulloch and Pitts went on, with Jerome Lettvin and Humberto Maturana, to describe what the frog’s eye tells the frog’s brain, a classic of neurophysiology. Simon, the political scientist, collected a Nobel Prize in economics as well as a Turing Award. And after his AI work, Rosenblatt moved to the Section of Neurobiology and Behavior at Cornell, where he spent his last years trying to transfer learned behaviour between rats by injecting brain extracts. He also built a small observatory and proposed a method for detecting planets around other stars. This pattern will continue in the future, as we shall see: Hopfield contributed to both AI and molecular biology. Hinton was an experimental psychologist; Hassabis a neuroscientist and former video-game designer. This should be a lesson for how we fund and hire today. Grant panels, hiring committees and journals reward depth within a single discipline. The history of this field and many others suggests that we should also reward people who come to its problems carrying different baggage: different training, different instruments, different ideas.
The first winter (1969–1986)
The success of the perceptron (and the “failure” of SNARC) moved Minsky to change approach. He was determined to show that the field was inflated by propaganda and worked to mathematically prove that neural networks could never achieve the wonders they were credited with. In 1969, together with Seymour Papert, he published Perceptrons. The book mathematically proved that a single layer of trainable connections cannot compute certain simple functions (XOR being the most famous example) and that single-handedly killed the entire field.
Minsky and Papert’s formalisations did not really surprise Rosenblatt, who knew the limits of networks with a single layer. Rosenblatt also knew that networks with more layers could in principle compute such functions and had worked on them himself. The problem was that he did not know how to train those networks. When the machine made a mistake, the perceptron rule said how to adjust the connections feeding the output, but nothing said which of the hidden connections, deeper inside, were to blame. This is known as the credit assignment problem.
In the book’s closing pages, Minsky and Papert offered their “intuitive judgment” that extending the analysis to many layers would still prove sterile, and it was this (wrong) guess, more than the (right) theorems, that was read as a verdict on learning machines in general. That verdict was read by funding agencies as a mathematical proof that machines cannot learn, and money and prestige drained away from neural networks for more than a decade.
In the folklore of the field, this made Minsky the villain of the story: the man who killed neural networks and brought on the first winter, perhaps out of a rivalry that went back to high school. Some have tried to rehabilitate him. The sociologist Mikel Olazaran, who made the controversy the subject of his doctorate at Edinburgh, argued that the theorems were correct but narrow, that researchers at the time read them with considerable flexibility, and that the verdict was really delivered by funding agencies, the military in particular, which had decided to move their money to the symbolists (Olazaran, 1996). Olazaran may be right about how the verdict was delivered, but I am less sure about the intent, and so, it seems, was Papert himself. In 1988 he admitted that “there was some hostility in the energy behind the research reported in Perceptrons“, and that part of their drive came from seeing research money spent on connectionist projects they considered misleading. He insisted that most of their motivation was more fundamental, yet he also told the whole story as a fairy tale of two sisters competing for DARPA’s coffers, with Minsky and himself cast as the agents of the “artificial” one.
The winter spread outside of the USA and the field froze in the UK too. The Lighthill report of 1973, which hit symbolic AI as hard as anyone, judged that AI had failed to deliver on its promises, and the televised debate in which Lighthill faced McCarthy, Donald Michie and Richard Gregory at the Royal Institution is still worth watching because it replays many of the arguments we are witnessing today.
Rosenblatt moved to a new field (animal neuroscience) and, sadly, died in a boating accident on his forty-third birthday in 1971. He never saw the end of that winter.
A few people kept working despite the winter. Geoffrey Hinton, who had studied experimental psychology at Cambridge, did his doctorate in Edinburgh under Christopher Longuet-Higgins, who by then had turned from neural networks to symbolic AI and was one of the invited respondents to the Lighthill report. Hinton joined his lab but not his ideas and the two disagreed throughout. Hinton kept working stubbornly on a connectionist approach and he was massively vindicated in 1986. In a short paper in Nature, together with Rumelhart and Williams he showed a way to teach the deeper layers of the network. Using a technique called backpropagation they could train exactly the hidden layers that had defeated everyone in the 1960s. Backpropagation solved the credit assignment problem: it tells every connection, however deep, how much it contributed to the error. The missing ingredient had never been a new kind of network, only a way of sending the error signal backwards through it. Backpropagation almost single-handedly got the field unstuck and resuscitated connectionism, and Hinton would later share the 2024 Nobel Prize in Physics. At the same time Rumelhart, James McClelland and the PDP Research Group, most of them psychologists, published the two volumes of Parallel Distributed Processing, laying the conceptual foundations for an idea that would pay off decades later on parallel hardware.
The second winter (1988–2012)
The counter-attack from symbolists was almost immediate. In 1988 a special issue of Cognition carried two long critiques of connectionism. Jerry Fodor, who had spent nearly three decades at MIT, and Zenon Pylyshyn argued that networks could not explain the systematicity of thought, while Steven Pinker and Alan Prince took apart the celebrated PDP model of English past-tense learning. Minsky and Papert reissued Perceptrons the same year, with an epilogue arguing that the new connectionists had not escaped the old limits. Paul Smolensky’s replies, culminating in his tensor product representations, tried to show that networks could implement structured representations without becoming symbol systems.
The attacks were not as rigorous or important, though, and this time the cold came more slowly. Backpropagation simply worked and the results were impressive: in 1989 Yann LeCun, who had been Hinton’s postdoc in Toronto, used it to train convolutional networks that read handwritten zip codes, and descendants of his system went on to read a substantial fraction of the cheques written in the United States (the IEEE now lists the work as an engineering milestone). Yet deeper networks proved very hard to train, because the error signal faded as it travelled backwards through many layers, a problem analysed most clearly for recurrent networks, whose layers are steps in time (Bengio et al., 1994), and the computers and datasets of the day were far too small. In cognitive science and in public debate, the symbolist critique remained the louder voice throughout the 1990s, from Pinker’s bestsellers to The Algebraic Mind (2001) by Gary Marcus, who had done his doctorate with Pinker at MIT. Within machine learning, though, networks were displaced less by symbols than by statistics. Support vector machines and probabilistic graphical models offered convex optimisation and mathematical guarantees, and by their own accounts the neural network researchers of the early 2000s struggled to get their papers accepted. Symbolic AI, meanwhile, suffered its own collapse: the market for specialised Lisp machines crashed in 1987, expert systems disappointed, and Japan’s Fifth Generation project ended in 1992 short of its goals.
The second winter’s thaw came from several directions at once, and most of them had been prepared by the handful of researchers the Canadian Institute for Advanced Research had kept funding through the lean years, Hinton, LeCun and Yoshua Bengio among them; Richard Sutton and Andrew Barto, meanwhile, had been building reinforcement learning into a field of its own. Hinton’s deep belief networks (2006) showed that deep architectures could be trained; Nvidia’s CUDA platform (2007) made graphics processors programmable for general computation, and Raina, Madhavan and Ng (2009) showed how much faster they could train networks; ImageNet (2009) supplied labelled images at scale. In 2012 AlexNet, built by Hinton’s students Alex Krizhevsky and Ilya Sutskever and trained on two gaming graphics cards, won the ImageNet competition by a wide margin. The second winter ended, once again, with engineering. As Sutton later summarised, general methods that scale with computation always end up beating knowledge built in by hand. He called this The Bitter Lesson.
The first spring (2012–)
What GPUs unleashed was unlike anything before. Being able to finally train dozens, and soon hundreds, of layers, deep networks took over vision and speech. Backpropagation was then married to reinforcement learning and that changed the game once more. DeepMind’s DQN learned to play Atari games from raw pixels (2015), and AlphaGo beat Lee Sedol at Go in 2016. The transformer (2017) made it possible to train language models on a scale nobody had attempted, and with ChatGPT in late 2022 they reached hundreds of millions of people. The old learning camp collected the field’s highest honours: the 2018 Turing Award for Hinton, LeCun and Bengio, the 2024 award for Barto and Sutton, and in 2024 the Nobel Prize in Physics for Hopfield and Hinton, and a share of the Nobel Prize in Chemistry for Demis Hassabis and John Jumper of DeepMind, for AlphaFold.
By any reading, all this should have marked the victory of connectionism over symbolism.
Yet it did not and the quarrel is back. Interestingly, the map shows that its participants have not changed sides. Gary Marcus, heir to the Chomsky–Pinker line, argues that language models need explicit symbols and rules; Chomsky himself dismissed ChatGPT in a 2023 op-ed. Bengio (who debated Marcus in 2019 in a two-hour exchange that is a good primer on the whole controversy) wants deliberate, “System 2” reasoning built inside networks rather than bolted on. François Chollet, whose ARC benchmark was designed to expose the limits of pattern-matching, argues for combining deep learning with program search.
Emily Bender stands in a different lineage altogether. Following Searle’s Chinese room (1980) and Harnad’s symbol grounding problem (1990), she argues that a system trained on linguistic form alone cannot recover meaning. The twist is that Harnad’s problem was originally aimed at symbolic AI: if the argument works, it sinks a symbol system trained on text just as surely as a language model. These are not so much technical opinions about transformers as inherited views of what a mind is.
LeCun, from the Hinton line, takes yet another stance. Like the symbolists, he argued that language models, which learn by predicting the next word, will never reach human-level intelligence as they lack a model of how the physical world works, a persistent memory and the ability to plan. Unlike them, his remedy is more learning, not symbols.
All these positions have intellectual merit to an extent. The problem is that the LLM field is growing at a pace never seen before and proving them obsolete week after week. This creates some confusion within the public: on a Monday we are told LLMs are not to be trusted because, by design, they can’t understand things. The following Tuesday, we are told an LLM has solved 400 new problems of mathematics. When is the quarrel becoming a farce?
What does it mean to be symbolic anyway?
Today’s symbolic systems come in two flavours. In the first, a learning system uses symbolic tools when it needs them, as a person reaches for a calculator or a tape measure: AlphaGo searches with a classical tree search, AlphaGeometry hands its constructions to a deduction engine, and AlphaProof‘s proofs are checked by a proof assistant. In the second, the rules are built into the network itself. The figure below sorts the landmark systems of the past forty years.
Perhaps surprisingly, it is worth asking once more what the word symbolic still means. In 1988 the distinction was sharp and testable: either the mind computes by applying explicit rules to structured symbols, as Fodor and Pylyshyn held, or it computes with distributed patterns of activity, as the PDP group held, and the two views predicted different things. Today “symbolic” has stretched to cover both flavours above, which some call hard (or tight) and soft (or loose) symbolism, and a good deal more.
Marcus defines a neurosymbolic system as any AI that combines neural networks with symbolic operations such as conditionals, operations over variables or code interpreters, a definition under which, as the computer scientist Florian Tramèr first pointed out, every language model since roughly GPT-3.5 would qualify. Some go as far as to call reinforcement learning with verifiable rewards (RLVR), in which a program checks the model’s answers during training, a form of symbolism. But then why not reinforcement learning from human feedback (RLHF), in which the judge is a neural network trained on people’s preferences? The two differ in who writes the reward, not in how the model computes. And if training against an external judge is what makes a system symbolic, language models were symbolic almost from the beginning, because they only became usable conversational partners once RLHF was applied to them (InstructGPT was the direct precursor of ChatGPT). Symbolic became a word that applies to everything and no longer distinguishes anything and, frankly, many of these definitions, once layered onto the map of the actors, read more like a post-hoc attempt to say “I was right!” than a useful description of the systems.
This, in my view, is what makes the present discontent different from the old ones, and less productive. Minsky and Papert proved theorems. Fodor and Pylyshyn set a challenge precise enough to be answered (and in 2023 it finally was, by Lake and Baroni, with a standard neural network). Today the field is progressing faster than at any point in its history, and each advance tends to be met not by a revised assessment but by a new criterion. The old critics gave the field problems to solve. Current critiques give definitions that change every passing week.
One reason why the word “symbolic” has lost part of its meaning is that it has become less and less clear what problem symbols are supposed to solve. In the early 2020s the case for neurosymbolic AI was mostly about reliability: language models hallucinated, and explicit symbols would keep them honest. Then LLMs became more knowledgeable and more reliable than expert humans and it was about reasoning and planning. Then they started winning high school mathematics competitions, and it was about genuine novelty, the capacity to find new approaches rather than recombine old ones. Now they are solving Millenium prizes and it is increasingly about alignment and verifiable safety. These are all real problems. But each time language models made visible progress on one of them, largely through scale and learning, the emphasis moved to something else.
All of which raises a question that is oddly absent from the debate: is the human brain neurosymbolic? As far as neuroscience can tell, there is no module in it that applies explicit rules to discrete symbols; there are neurons, synapses and patterns of activity, shaped by learning. The brain – not just the human one – is incredibly resilient and generalist and we simply do not understand where this adaptability is coming from. We do know that some people can live normal lives even sporting brains that are a fraction in size of what most humans have. We know that we can take the connectome of a fruit fly and, through reinforcement learning, have it play video games and even chess, showing the versatility of a complex-enough network.
“Neurosymbolic” is now often used as a synonym for grounded: a system that does not claim things that go against experience, the opposite of one that hallucinates. That is a reasonable thing to want from a tool. I would not buy a tape measure that hallucinates. As a criterion for intelligence, however, it is a strange one, because human minds fail it all the time. People give confident explanations of their own behaviour that have little to do with its real causes (Nisbett and Wilson, 1977). Shown a face they did not choose, they explain why they chose it, without noticing the swap (Johansson et al., 2005). Split-brain patients fluently invent reasons for actions started by the hemisphere that cannot speak, a habit of mind that Michael Gazzaniga called the interpreter and that has since been found at work in all of us. And most of the explanations humanity has offered for the natural world, from the four humours to miasmas, were coherent, rule-governed and wrong. Being grounded in reality is something minds achieve slowly, with effort and with external help, from experiment to peer review. It is not something they come equipped with, symbols or no symbols. Human beings acquire most of their theoretical knowledge through books and tales, exactly like an LLM does.
The need for grounding perhaps comes from the fact that we hold language models to a standard of infallibility, as if an intelligent system should never err. Turing anticipated why this is wrong in a 1947 lecture to the London Mathematical Society, observing that a machine expected to be infallible cannot also be intelligent. Brains are fallible; that is the price of generalising beyond what they were given. What we want to be infallible are our tools, which is why humans invented notation, calculators, formal proof and peer review, and why the spring’s best systems reach for exactly such tools.
It is fair to say, of course, that until the 2010s, the networks that worked in practice were classifiers: LeNet looked at a handwritten digit and returned “7”; AlexNet looked at a photograph and returned “dog”. Generative networks existed, from Boltzmann machines to Hinton’s deep belief nets, but they were not what anyone used. A classifier in that sense is no more than a tool. On that ground the symbolists were at home, because a system whose job is to produce the correct symbol is naturally measured by the standards of rules and proofs. Language models changed this. They do not end in a label but in open-ended text, images or code; there is rarely a single right answer, and what we ask of them looks less like the work of a tool and more like the work of a mind. This, I think, is one reason the old critique feels at once so familiar and so out of place: it is a toolmaker’s critique, applied to something that is no longer quite a tool and does not want to be one.
How will this end?
In The Society of Mind, published in 1986, the same year as the PDP volumes, Minsky argued that intelligence does not come from any single principle or trick. It comes from the interaction of many small, specialised agents, each of them mindless on its own. When he taught the theory at MIT, he presented it explicitly as a way out of the old oppositions, symbolic against connectionist among them, and his lectures are still online.
As a neuroscientist, I find this the most persuasive picture of where artificial minds are heading. Brains are not one kind of computation. They are assemblies of specialised, sometimes hyper-specialised, structures, each good at a narrow job and none of them in charge, which together behave as a single mind. Artificial systems are already moving the same way. Mixture-of-experts models route each input to a few specialised sub-networks, and the most capable systems now coordinate teams of agents, some of which write code, search the literature or check proofs2.
Seen this way, the quarrel between symbols and networks dissolves into a question about organisation. Symbolic tools become one more kind of specialist in the society, consulted when the job calls for it, as a calculator is consulted by a mind that knows its own limits. And the whole society, seen from outside, will look like what we have always called a mind. It would be a fitting irony if the man blamed for the first winter turned out to have described the shape of the spring.
- A few years later, Eric Kandel would show that a handful of neurons in the sea slug Aplysia californica were enough to produce behaviour and to learn, much as Rosenblatt’s architecture assumed. ↩︎
- Interestingly enough, the MoE modes do not route by topic but by token suggesting the connectivity that emerges from MoE training is acting on dimensions we do not fully understand. ↩︎