This post was originally published on the main website on Apr 18 2026. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account <3!
Hey everyone π,
I was watching a documentary a few months ago, I do not remember which one exactly because I watch a lot of them late at night when I cannot sleep, and there was a segment about emperor penguins in Antarctica, specifically about how they recognize one another's calls across a colony of thousands of birds in the middle of a blizzard. Each individual has a unique vocalization. Each partner in a mated pair learns the other's call with such precision that they can find each other in conditions where visibility is zero and the wind is loud enough to drown out almost any sound. They do this every year. The colony disperses, reassembles, and the bonds hold through conditions that would kill most mammals in hours. And I remember sitting there in the dark, watching this, and thinking: what exactly is the story we are telling ourselves about what intelligence is and where it lives? Because whatever that penguin is doing when it picks its mate's voice out of a screaming Antarctic storm is not nothing. It is something sophisticated, something persistent, something that cannot be reduced to reflex or accident or blind evolutionary wiring without doing serious violence to the word "intelligence". It is, by any honest standard, cognition. And yet the conversation about intelligence in AI circles almost never mentions it, because the conversation is entirely organized around building and scaling the kinds of structures that humans use, language, symbols, text prediction, and is almost entirely silent on the question of whether the structures that already exist in the living world around us might tell us something important about what intelligence actually is. That silence bothers me. In fact, it has been bothering me for long enough that I need to write about it, which is what this post is for.
This connects to everything I have been building toward across the last several posts. In LLMs are Useful. LMMs will Break Reality, I argued that language models are locked inside a symbolic cage, describing the world without ever touching it. In Mathematical Equations are Multimodal by default, I argued that equations encode reality in a way that sentences never can, that mathematical structure is the most compressed and honest representation of physical truth that humans have ever produced. In Training Is an Evil Concept. LMMs Eliminates it Altogether., I argued that the training paradigm is not just ethically wrong but architecturally bankrupt, that consuming unconsented human creative work at billion-token scale is a moral choice dressed up as an engineering necessity. In LLMs destroyed the Internet. LMMs will make it alive., I argued that the mass deployment of language systems as content factories has hollowed out the authenticity of the web. The argument in this post is the one that goes underneath all of those, the one that asks why, despite all of this obvious evidence, we continue to treat a very narrow kind of human cognitive output as the definition of what intelligence is, while billions of minds that process reality, form bonds, navigate dynamic environments, and solve genuine survival problems are dismissed as mere biology and declared irrelevant to the AI conversation. I want to make the case that this dismissal is not just scientifically wrong. It is philosophically catastrophic, and it has led the entire field in a direction that produces impressive demos and genuine misunderstanding at the same time.
What We Mean When We Say "Sentient" and Why the Answer Keeps Changing
The word sentient comes from the Latin word for feeling, and in its original usage it simply meant capable of sensation, capable of experiencing something rather than merely processing it as a switch would process voltage. That is a humble definition, and by that definition the question of whether nonhuman animals are sentient has been essentially settled for well over a century of serious empirical research. Animals have pain receptors. Animals have nervous systems that process threatening stimuli and produce responses that look, functionally, indistinguishable from pain behavior. Animals form bonds, grieve, play, and show something that looks from the outside very much like preference, anticipation, and disappointment. But the word sentient has been quietly colonized over the centuries by a much more ambitious claim, which is that sentience requires not just feeling but the specific kind of reflective self-awareness that humans associate with language-mediated consciousness, and that anything short of the ability to say "I am experiencing this" does not fully count. That colonization is a philosophical error, and it is an error that has done significant damage both to our treatment of other animals and to our scientific understanding of mind, because it has caused us to systematically underestimate the cognitive sophistication of creatures whose intelligence takes forms that our language-centric frameworks are not well equipped to see.
The Cambridge Declaration on Consciousness, signed in 2012 by a distinguished group of neuroscientists, was a rare moment of institutional clarity on this question, and I want to spend some time on it because it has not received the attention it deserves (1). The declaration stated unambiguously that nonhuman animals possess the neurological substrates that generate consciousness, and that the weight of evidence indicates that humans are not unique in possessing the biological equipment for conscious experience. This was not a fringe statement. It was signed at a symposium at the University of Cambridge and included researchers from some of the most respected institutions in cognitive neuroscience. The declaration covered mammals and birds explicitly, but also noted that the evidence for conscious experience extends to invertebrates like octopuses, which have a genuinely alien nervous system architecture that produces highly flexible, goal-directed behavior of a kind that is very difficult to explain without some form of unified experience. The scientific community has known this for over a decade, and the AI conversation has largely failed to engage with it, which is telling. If the machines we are building are supposed to be intelligent, and there are already billions of intelligent beings on this planet whose intelligence takes forms we have not yet fully understood, then the obvious question is whether we might learn something from looking more carefully at what those beings are actually doing. That question is almost never asked in the mainstream AI conversation, and I want to ask it here.
I grew up in a village, as I described in my first post, and I spent a significant portion of my childhood around animals in ways that people who grew up in cities often do not. I grew up around chickens, goats, dogs, cats, and occasionally the stray cats that wandered through the farming communities near where I lived, and I learned something from that proximity that no paper or textbook ever quite articulates clearly enough, which is that every single one of those animals had a distinct personality, a distinct set of preferences, a distinct way of engaging with the world that was not predictable from general statements about their species. The chicken that would let me pick her up was different from the chicken that would not. The dog that was afraid of strangers was different from the dog that greeted everyone at the gate. These were not abstractions. They were individuals, and their individuality was part of their daily engagement with the world around them, and that individuality is one of the markers of something that we should, if we are honest, call inner life. I am not making a mystical claim here. I am making an empirical observation about behavioral complexity, and I am noting that the same observations that would lead any careful scientist to attribute cognitive sophistication to those behaviors are observations that most people in the AI field seem entirely uninterested in, because those behaviors do not involve language models and therefore do not attract funding, prestige, or product launches.
This is where I want to introduce a distinction that I think is genuinely important and that the AI field has almost entirely failed to make, which is the distinction between the form of intelligence and the fact of intelligence. What I mean is this: human linguistic intelligence takes a very specific form, one that is organized around symbolic reasoning, narrative construction, and the explicit manipulation of abstract concepts using language as the medium. That form of intelligence is real and powerful, and it has produced everything from philosophy to particle physics. But the fact of intelligence, the thing that makes cognition cognitive, is not tied to that specific form. Crows solve multi-step tool-use problems without language. Octopuses solve novel escape problems without vertebrate neural architecture. Bees perform abstract distance calculations and communicate them to their hive-mates through dance without a neocortex. Elephants remember the locations of distant water sources across decades of drought without GPS or digital memory. All of these are demonstrations of the fact of intelligence without the specific form that human-centric AI research has decided is the thing worth building. The preoccupation with language as the medium of intelligence is not a conclusion derived from careful study of what intelligence is. It is a starting assumption inherited from centuries of philosophy that privileged human cognition as the gold standard, and that assumption has baked itself into the foundations of the field in ways that most practitioners never stop to examine.
The research on penguin cognition specifically is worth looking at carefully, because it has produced findings that should disturb anyone who is still operating with a dismissive picture of animal minds (2). Emperor penguins have demonstrated what researchers call episodic-like memory, meaning they do not just learn rules but form something resembling autobiographical records of specific events, specific encounters, specific outcomes, and can draw on those records to make decisions in novel situations. They demonstrate theory of mind precursors, meaning they track the informational states of other individuals in their colony in ways that suggest an understanding of what others know and do not know. They show evidence of social learning, meaning they modify their behavior based on observing the outcomes experienced by others rather than only on their own direct experience. And they do all of this in an environment that is arguably more extreme and more demanding than almost any human habitat, where the margin between success and catastrophe is measured in hours and where mistakes are fatal. That is not reflex. That is sophisticated, environmentally embedded intelligence operating at a level of generality and flexibility that any honest comparison with current AI systems must acknowledge is far more impressive than current AI systems in the domains that actually matter for survival.
I want to make one more point in this section before moving on, and it is the point that I think connects all of this most directly to the argument I am building toward. When researchers study animal cognition, they consistently find that the intelligence they observe is deeply integrated with the animal's body, its environment, and its social relationships in ways that resist clean separation of computation, sensing, and action. A bird navigating by magnetic field is not running an algorithm on separately stored data. The sensing, the computing, and the acting are intertwined in a biological architecture that does not have clean hardware-software boundaries. An elephant navigating a remembered landscape is not retrieving a map from a database and then executing a pathfinding algorithm. The knowledge is distributed through behavioral, social, and physiological systems in ways that make the boundary between memory, sensation, and movement genuinely unclear. This integration, this embodied, environmental, social embeddedness of cognition, is what the AI field almost entirely ignores when it builds language models, because language models process symbols in a context-free way that bears essentially no relationship to how cognition actually works in living systems. And yet the AI field claims to be building toward something it calls intelligence, using an architecture that shares almost none of the structural features of the intelligences that already exist everywhere in the living world. That claim deserves to be examined with more skepticism than it currently receives.
The Neural Network Is Not a Theory of Mind. It Is a Theory of Curve Fitting.
I want to be fair to the field here, because fairness is something I try to practice even when I am critical, and I am about to be quite critical. Neural networks are remarkable engineering achievements. The fact that you can take a dataset, specify a loss function, run gradient descent for long enough, and produce a system that can classify images, play Go better than any human, fold proteins, and generate fluent text is genuinely astonishing. I do not want to minimize that. I have studied neural networks in college, I have built neural networks, and I have a genuine appreciation for the elegance of backpropagation and the surprising power of the universal approximation theorem. These are real achievements, and the people who developed them deserve real credit. But engineering achievement and scientific theory are different things, and the AI field has a persistent and troubling tendency to confuse them. The fact that a neural network can do something impressive does not tell you why the network can do it, and more importantly, it does not tell you whether the network is doing it in a way that resembles anything going on in biological cognition. Those are separate questions, and collapsing them has produced some of the worst thinking in the field.
A neural network, at its mathematical core, is a parameterized function that maps inputs to outputs. The training process finds parameter values that minimize a loss function on a training dataset. The result is a function that generalizes, sometimes very well, to inputs outside the training set. That is it. That is the whole mechanism. Everything else, the emergent capabilities, the apparent reasoning, the fluent language generation, is an output property of that mechanism operating on very large amounts of data with very many parameters. The mechanism itself does not have beliefs, it does not have goals in any rich sense, it does not have a model of the world, and it does not have anything resembling the subjective experience that is the hallmark of sentience in the biological systems I described in the previous section. What it has is a very flexible statistical model of the patterns in its training distribution, and that statistical model can produce outputs that look, from the outside, like belief, goal-directedness, world modeling, and experience. The appearance is not the thing. The map is not the territory. And the AI field has been spending twenty years and several hundred billion dollars studying the map while largely ignoring the territory, and the territory is what the penguins live in.
The specific claim I want to make, and I want to make it precisely so that it can be evaluated rather than vaguely so that it sounds impressive, is this: neural networks are optimized for input-output mapping, and optimizing for input-output mapping does not, in general, produce the internal representations that characterize biological cognition. Biological cognition is not primarily organized around input-output mapping. It is organized around building, maintaining, and updating models of the world that can be used for prediction, planning, and action across a wide variety of tasks that the organism has never encountered before. The difference between a system optimized for input-output mapping and a system with a genuine world model is the same as the difference between a lookup table and a physics engine: the lookup table can give you the right answer for inputs it has seen, but the physics engine can give you the right answer for inputs it has never seen, because the physics engine contains the actual structure that generates the answers rather than just the answers themselves. I argued in Mathematical Equations are Multimodal by default that equations encode mechanisms rather than surfaces, and that distinction maps directly onto the distinction between neural networks and the kind of world models that biological cognition actually relies on. A penguin navigating a blizzard is running a world model, not a lookup table, and the world model is what makes the navigation work.
This is not a new observation. The philosopher Jerry Fodor made a version of this argument in his book "The Modularity of Mind" in the early 1980s, noting that the kinds of central cognitive processes that make human intelligence general and flexible are precisely the kinds of processes that computational systems of the behaviorist-inspired variety have the most trouble capturing (3). The argument has been made more recently and more specifically by researchers like Gary Marcus and by the whole school of thought around compositional generalization, which asks whether neural networks can learn to apply rules to novel combinations of familiar inputs rather than just pattern-matching to training examples (4). The evidence is mixed, and the honest version is that standard neural networks do not compositionally generalize in the way that biological cognition does, and that this is a structural property of the architecture rather than something that will be fixed by training on more data. The experiments are clear, the replication rate is high, and the implication is one that the field has consistently found ways to avoid drawing, which is that the neural network architecture as currently practiced is not modeling what biological cognition actually does, even in the relatively simple case of compositional rule application. If it cannot do that well, the claim that it is approaching general intelligence deserves serious scrutiny.
I also want to say something about what I call the scale fallacy, which is the reasoning that says: current neural networks fail at X, but if we scale them up with more data and more parameters, they will eventually succeed at X. This reasoning is sometimes correct and sometimes catastrophically wrong, and the field has an embarrassing tendency to apply it indiscriminately without asking whether the failure at X is a quantitative limitation or a qualitative one. If X is "generate more fluent text," then scaling probably helps, because fluency in text generation is a quantitative property that more data and more parameters can plausibly improve. But if X is "form a genuine world model that supports causal reasoning across novel domains," then scaling alone does not help, because forming a world model is not the objective that the network is being optimized for. You can train a curve fitter on infinite data and you will still have a curve fitter. You will never get a physics engine from a curve fitter by training it longer, because the objective is wrong. The objective optimizes for one thing, and the thing you want is a different thing, and no amount of the first thing gives you the second thing. This is not a pessimistic claim about the future of AI. It is a specific claim about the relationship between objectives and outcomes, and it is a claim that the field's most honest researchers have been making for years in papers that attract far fewer citations than the scaling papers because they deliver less comfortable news (5).
The penguins are relevant here in a way that I want to be explicit about, because I am not just using them as an emotional hook. The penguin finding its mate in a blizzard is demonstrating something that current neural networks, despite their scale and sophistication, would have extreme difficulty matching in any meaningful sense. It is not that the computation exceeds the network's capacity. It is that the kind of computation being done is structurally different from what neural networks do. The penguin is running a real-time, dynamic, embodied recognition system that integrates acoustic pattern matching with spatial navigation with social memory with motivational state in a way that is seamlessly unified and operates with extreme reliability in conditions that would challenge any engineered system. The neural network, given a training dataset of penguin calls and a test set of noisy versions, could certainly learn to classify calls, and it would do so by finding statistical features that discriminate between classes in the training distribution. That is useful. But the penguin is not running a classifier. It is running a survival system, and the difference between a classifier and a survival system is not the magnitude of the computation. It is the kind of computation. It is the fact that the penguin's recognition system is integrated with everything else the penguin knows and needs and wants, in a way that makes the recognition not just accurate but alive, not just functionally correct but embedded in a continuous engagement with reality that the neural network, by its architecture, cannot replicate.
Let me also make the philosophical point explicitly rather than letting it hide inside the technical argument, because the technical argument can always be dismissed as a practical limitation that will eventually be overcome, while the philosophical point cuts deeper. The neural network does not have anything at stake when it classifies an input. It has no survival interest in the outcome. It has no mate to find. It has no colony to return to. It has no past that informs its present or future that it is trying to secure. It is operating a function with no inside to it, no orientation toward the world, no stake in anything. Whether the classification is right or wrong is a matter of the loss value on a metric, not a matter of survival or death or reunion or loss. That difference, the difference between a system with something at stake and a system with nothing at stake, is not an engineering detail that can be fixed by adjusting the hyperparameters. It is the difference between sentience and computation, and I am not prepared to believe that the gap can be closed by gradient descent alone, any more than I am prepared to believe that a map can be turned into a territory by making it more accurate.
What Embodied Intelligence Looks Like When You Actually Look at It
The AI field talks about embodied intelligence a lot, mostly in the context of robotics, and mostly in a way that reduces "embodiment" to the fact that the robot has sensors and actuators connected to a neural network. That is an impoverished definition of embodiment, and I want to spend some time explaining what embodiment actually means in the context of real cognition, because the real version is much more interesting and much more instructive for anyone trying to build systems that genuinely understand the world. Real embodiment is not the fact that a system has inputs and outputs from the environment. It is the fact that the system's knowledge, its memory, its representations of the world, are organized by and inseparable from its physical capabilities, its evolutionary history, and its ongoing engagement with a specific kind of environment. A bat's echolocation system does not just give the bat access to acoustic information. It gives the bat a bat-shaped understanding of the world, organized around the specific capabilities and needs of a bat body in a bat environment. The bat's acoustic world model is not separable from the bat's life, and that inseparability is not a limitation. It is the source of the system's power, because it means the model is precisely tuned to the situation in which the bat actually operates.
Research on animal navigation provides some of the most compelling evidence for this kind of embodied, integrated intelligence, and I want to spend some time on it because it is directly relevant to the argument I am making about what intelligence actually is when you look at it carefully. Clark's nutcrackers are birds that cache tens of thousands of seeds in thousands of locations across a landscape and then retrieve them months later with an accuracy that is astonishing by any standard (6). They do this without GPS, without a written map, without language, and without anything that resembles the cognitive tools that human chauvinism would suggest are necessary for sophisticated spatial reasoning. The hippocampal volume of seed-caching birds is proportionally larger than in non-caching birds, which suggests that the brain structures used for spatial memory are specifically adapted to this cognitive demand, meaning that the bird's brain architecture is shaped by and tuned to the specific cognitive challenges its lifestyle presents. That is embodied intelligence in the deep sense: the architecture of the system is structured by the architecture of the problem it evolved to solve, not by a general-purpose optimization scheme applied to a training distribution. Understanding that distinction matters enormously for anyone who is seriously trying to understand what intelligence is, as opposed to what impressive-looking outputs intelligence can produce.
Honeybee cognition is another area that should disturb anyone operating with confident assumptions about the cognitive prerequisites for sophisticated behavior (7). Bees perform the distance-transformed waggle dance to communicate the direction and distance of a food source to their hive-mates, accounting for the angle of the sun, the time of day, the distance, and even the quality of the source on a numerical scale. This is not a simple signal. It is an abstract spatial encoding that other bees can decode and use to navigate to a location they have never visited, using information they received through a physical performance by another bee. That is a form of symbolic communication, it uses a learned code, it encodes abstract spatial relationships, and it works even when the communicating bee is indoors and cannot see the sun, meaning the dance references a representation of the sun's position rather than the sun's actual position. If I described that capability in a neural network architecture, people would call it a breakthrough. In a bee, people call it "just instinct," and the dismissal is so automatic and so culturally comfortable that most people who use it have never stopped to ask what exactly they mean. I know what instinct means. I have read the papers. Instinct means neurologically determined behavior. So does reading, for most people below a certain age. The distinction between instinct and intelligence that we think we are drawing when we say "just instinct" is largely a distinction between familiar and unfamiliar computational substrates, and that is not a meaningful distinction for anyone trying to understand what cognition actually is.
I want to connect the embodied intelligence argument directly to the lmm project here, because the lmm architecture is my attempt to build toward this kind of embodied, grounded intelligence in a way that does not depend on the training paradigm I critiqued in my last post. The lmm system includes a perception layer that converts raw bytes and sensor streams into normalized tensors, a physics simulation layer that models dynamic systems using actual differential equations, a causal reasoning layer that maintains explicit structural causal models of the dependencies between variables, and a consciousness loop that ties these together by running a continuous cycle of perceiving, encoding, predicting, and acting. This architecture is closer in spirit to embodied cognition than to the standard language model architecture, not because it is biological, but because it is organized around engaging with the structure of physical reality rather than around predicting the next token in a text sequence. When you run lmm consciousness --lookahead 5, the system performs one tick of this full loop: it takes raw input, converts it to a tensor, runs a world model prediction, evaluates the prediction against the actual state, and plans an action based on the discrepancy. The output is a state vector and a mean prediction error, both of which are observable, verifiable, and grounded in the structure of the input rather than in the statistical patterns of a training corpus. That is not biological intelligence. But it is a step toward the right kind of intelligence, in a direction that the training-based paradigm cannot go, because it is organized around the world rather than around text about the world.
The field calculus capability of lmm is also worth thinking about in this context, because it represents a form of spatial and structural reasoning that is closer to what embodied cognition actually does than anything in a standard language model. When you run lmm field --size 8 --operation gradient, the system computes the gradient of a scalar field using central differences, which is a numerical implementation of the mathematical operation that describes how quantities change across space. That operation is at the heart of everything from physical simulation to navigation to the neural population codes that actual animal brains use to represent space. A gradient is not a piece of text. It is a structure, a mathematical object that lives in the geometry of a field, and computing it correctly requires genuine engagement with that structure rather than statistical pattern matching against descriptions of gradients from a training corpus. The result, something like [1.0, 2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 13.0] for the gradient of xΒ², is a number sequence that can be verified against the exact analytical derivative, which is 2x. That verification is possible because the computation is transparent and the ground truth is accessible. No language model can offer that kind of verification, because language model outputs are not computations in the relevant sense. They are predictions of token sequences that describe computations, which is a different thing entirely. The embodied cognition of the penguin navigates reality. The lmm field operator computes reality. The language model narrates reality. Those are three structurally different activities, and only two of them constitute genuine engagement with the world.
The Distraction Is Working. Let Me Show You How.
I want to talk about attention, specifically about where the field's attention is pointed and what the costs of that pointing are. The AI field right now is experiencing what I can only describe as a collective hallucination about what it is building and how important it is. The rhetoric around large language models has reached a level of confidence and self-congratulation that is genuinely strange to anyone who has read the history of AI carefully, because the history of AI is a history of premature confidence followed by painful corrections, and the confidence is currently at levels I have not seen since the 1960s symbolic AI days when some of the most brilliant people in the world confidently predicted that general intelligence was ten to twenty years away (8). Those predictions were wrong, and the reason they were wrong was not that the researchers were stupid. They were not. The reason they were wrong was that they were distracted by the impressive capabilities of the systems they were building from asking honest questions about whether those systems were actually doing what the researchers thought they were doing. The distraction is happening again, at much larger scale, with much more money behind it, and with much more riding on the outcome.
The distraction takes a specific form that I want to name clearly so that it can be recognized. It goes like this: a new capability appears in a scaled-up language model, the capability was not explicitly trained for, and people call this "emergence". The emergence is treated as evidence that scaling is a path to general intelligence, because look, the model can do something new that nobody taught it to do. The problem with this reasoning is that it conflates statistical emergence with genuine cognitive emergence. When a statistical model trained on text discovers that certain text patterns cluster together in ways that allow it to produce what looks like arithmetic, that is statistical emergence, the discovery of surface patterns that co-occur with arithmetic in text. It is not the same as genuinely learning the rules of arithmetic, and the empirical evidence shows exactly this: language models perform much better on arithmetic problems that appear frequently in their training distribution than on arithmetically equivalent problems presented in forms that appear rarely, which is exactly what you would expect from a system that learned statistical patterns about arithmetic rather than arithmetic itself (9). That is the distraction working in real time. The capability looks real from the outside, and the appearance creates confidence, and the confidence attracts resources, and the resources deepen the bet, and the whole cycle continues without anyone pausing to ask whether the appearance is the thing or just a very good imitation of the thing.
The cost of the distraction is not just computational or financial, although those costs are enormous. The cost that I find most troubling is the cost to our understanding of intelligence, because every year spent building and scaling systems that imitate the surface of intelligence without engaging its deep structure is a year not spent trying to understand what intelligence actually is. The research program on animal cognition that I described in the previous sections is not well funded, not prestigious, not connected to major product launches, and not the subject of breathless press coverage. It is slow, careful, empirical work done by researchers who care more about understanding than about demos, and it is being systematically outcompeted for attention and resources by a paradigm that is very good at producing things that look impressive and very reluctant to ask whether they are genuinely intelligent. I described in An Empty Life Filled With Constant Suffering what it feels like to work on something real and have it invisible because it is not dressed up in the right narrative, and the researchers studying animal cognition live that experience constantly. Their subjects are genuinely intelligent in ways that matter. Their findings are genuinely important for anyone who cares about the nature of mind. And they are being crowded out by a conversation that has decided in advance what intelligence looks like and is building systems to match that predetermined picture.
I want to make a specific claim about what the distraction has cost in terms of scientific progress, because vague claims about opportunity cost are easy to dismiss and specific claims are harder. The specific claim is this: if the resources invested in scaling language models over the last decade had been partially redirected toward understanding the computational principles of animal cognition, toward building formal mathematical models of navigation, social learning, episodic memory, and multi-sensory integration in biological systems, and toward implementing those models in engineered systems that could be tested and refined, we would today have a much clearer scientific picture of what intelligence actually is and a much more principled basis for building artificial systems that instantiate it. That is a claim about what would have happened under a different allocation of resources, and it is inherently speculative, but it is no more speculative than the claim that scaling language models will eventually lead to general intelligence. The difference is that the second claim is the one being funded and celebrated, while the first claim is the one that has the biological evidence on its side. The penguins, the bees, the nutcrackers, the octopuses: these are existence proofs of a kind of intelligence that does not pass through text, and existence proofs are the most powerful kind of evidence in any scientific argument, because they show that the thing is possible rather than merely arguing that it might be.
The way the distraction sustains itself is worth understanding, because it is not sustained by stupidity or bad faith on the part of individual researchers, most of whom are genuinely curious and genuinely talented. It is sustained by an incentive structure that rewards impressive demos over careful science, market share over mechanistic understanding, and confidence over honesty. I wrote about this in Technology Has Destroyed My Livelihood, where I described how the technology industry's incentive structure consistently produces outcomes that are good for the organizations at the center of the field and bad for the people at its margins, and the same dynamic applies to the scientific question of what intelligence is. The organizations at the center of the AI field have enormous incentives to believe that scaling language models is the path to general intelligence, because they have bet enormous sums on that belief, and anyone who challenges it from within faces strong institutional pressure to continue in the current direction. The challenge has to come from outside the current incentive structure, which is exactly what the work on animal cognition represents, and why I think it deserves much more attention than the AI conversation is currently giving it.
The distraction also has a racial and geographic dimension that I want to name because I think it is usually left out of the conversation, and leaving it out makes the picture incomplete. The AI field is geographically concentrated, demographically uniform, and culturally shaped by a specific tradition that privileges certain kinds of cognitive performance, specifically the kinds associated with academic achievement within Western educational systems, as the gold standard of intelligence. That tradition has a long history of underestimating the intelligence of people who succeed through other means, through spatial navigation, through social intelligence, through craft and embodied skill, through forms of reasoning that are not well captured by standardized tests or academic publications. The same cultural predisposition that produces a dismissive attitude toward non-Western forms of intelligence also produces a dismissive attitude toward non-human forms of intelligence, and both dismissals serve the same function of maintaining a comfortable hierarchy with the AI researcher at the top. I am not saying this to accuse anyone of racism or malice. I am saying it because the cultural assumptions that a research community brings to its work shape what the community sees and what it misses, and a field that has systematically missed the intelligence of billions of non-human animals while spending hundreds of billions of dollars on systems that imitate a very specific kind of human cognitive output is a field that would benefit from examining its assumptions more carefully.
lmm as the Alternative Architecture: Building With the World Instead of About It
I want to spend this section connecting the philosophical argument I have been making to the specific engineering choices in the lmm project, because I think the connection is real and important rather than just rhetorical. The lmm project is not presented as an imitation of animal cognition. It is presented as an alternative to the language model paradigm, built around mathematical structure and physical simulation rather than around text prediction and gradient descent. But the reason I brought animal cognition into this post is that animal cognition, honestly examined, points toward exactly the kind of architecture that lmm is trying to instantiate: an architecture organized around world modeling, causal reasoning, and embodied engagement with physical reality rather than around surface pattern matching in a high-dimensional statistical space. The connection between the penguin and the lmm is not that the lmm can navigate a blizzard. It is that both the penguin and the lmm are organized around contact with reality rather than around descriptions of reality, and that organizational principle is the thing that distinguishes intelligence from imitation.
The symbolic regression capability of lmm is the most direct implementation of the "learn the mechanism, not the surface" principle that I keep arguing for. When you run lmm discover --iterations 200, the system takes a set of observations, in the default case data points from a linear process, and runs a genetic programming search to find the symbolic equation that best explains those observations. The output is something like (x + (1.002465056833142 + x)), which is the system's discovered approximation of the underlying law 2x + 1. The method used to find this equation is genetic programming: a population of candidate expressions is initialized, each one is evaluated for how well it fits the data, better-fitting expressions are selected and recombined and mutated, and the process iterates until either the fitness converges or the iteration budget is exhausted. This is superficially similar to neural network training in that both involve iterative optimization on data, but the difference is fundamental: the neural network optimizes parameters in a fixed architecture toward minimizing a loss, producing a black-box function; the genetic programming optimizes the structure of a symbolic expression toward explaining the data, producing a human-readable equation. The output of one is opaque. The output of the other is transparent. And transparency, as I argued in Training Is an Evil Concept, is not a cosmetic feature. It is the property that makes outcomes verifiable, and verification is the foundation of honest science.
# Discover the governing equation from synthetic linear data
lmm discover --iterations 200
# Output: Discovered equation: (x + (1.002465056833142 + x))
# This approximates 2x + 1, discoverable in ~200 GP iterations
# No training corpus, no gradient descent, no unconsented data
# For more complex patterns, more iterations help
lmm discover --iterations 500 --data-path ./my_observations.csv
The physics simulation capability directly instantiates the idea of a world model, which is what I argued animal cognition relies on rather than a lookup table of trained associations. When you run lmm physics --model lorenz --steps 500 --step-size 0.01, the system integrates the Lorenz chaotic attractor equations forward in time using the Runge-Kutta fourth-order method, producing the exact trajectory of a chaotic dynamical system from its governing equations. That trajectory is not predicted from training data. It is computed from the differential equations that describe the Lorenz system, and those equations are not learned from examples. They are specified from physical first principles and then used to generate predictions. The difference between generating predictions from learned statistical patterns and computing predictions from explicit physical laws is the difference between the curve fitter and the physics engine that I described earlier in this post. The curve fitter can only interpolate within what it has seen. The physics engine can extrapolate to states it has never simulated, because it has the generating structure, not just examples of outputs. This extrapolation capability is the thing that animal cognition has that enables animals to navigate novel environments, solve novel problems, and make decisions in situations that have no precedent in their experience.
# lmm uses known physical laws, not training data, to predict the future
lmm physics --model lorenz --steps 500 --step-size 0.01
# => Lorenz: 500 steps. Final xyz: [-8.900..., -7.413..., 29.311...]
lmm physics --model sir --steps 1000 --step-size 0.5
# => SIR: 1000 steps. Final [S,I,R]: [58.797..., 7.649e-15, 941.202...]
lmm physics --model pendulum --steps 300 --step-size 0.005
# The equations are the world model. No training. No hallucination.
# If the equations are right, the predictions are right. Always. Verifiably.
The causal reasoning layer of lmm is perhaps the most important piece for the argument I have been making about the gap between animal cognition and neural network processing, because causality is the specific kind of reasoning that animal survival most requires and that neural network pattern matching most systematically fails to provide. When a penguin learns that a specific behavior leads to food in some contexts and does not in others, it is not just learning a stimulus-response association. It is building a causal model that includes the conditions under which the association holds, and that model allows the penguin to behave appropriately in novel situations that share the relevant causal structure but not the specific sensory context. Neural networks trained to classify or predict do not, in general, learn causal models. They learn correlational patterns, and correlational patterns break down in exactly the novel situations that causal models handle correctly. The lmm causal module builds an explicit structural causal model and supports the do(X=v) intervention operator from Judea Pearl's do-calculus, which allows you to ask not just "what is correlated with what" but "what would happen if I changed this variable". That is the question that causal reasoners ask, that animal cognition answers, and that neural networks are systematically unable to address without explicit causal structure.
# Build a causal model and reason about interventions
# The SCM is y = 2*x, z = y + 1
lmm causal --intervene-node x --intervene-value 10.0
# Before intervention: x=Some(3.0), y=Some(6.0), z=Some(7.0)
# After do(x=10): x=Some(10.0), y=Some(20.0), z=Some(21.0)
# A language model given the same question would produce a
# plausible-sounding description of the math from training patterns.
# lmm computes the actual causal propagation from the explicit structure.
# The penguin navigating food sources makes exactly this kind of inference.
I want to be transparent about the obvious gap, which is that lmm currently works with mathematical functions and explicit causal graphs, while animal cognition operates on raw sensory streams in a messy, partially observable, physically complex world. The lmm perception layer is the beginning of bridging that gap: it accepts raw byte streams and converts them to normalized tensors, and the consciousness loop is designed to be the integration point where perception, prediction, and action come together. But the current implementation is a proof of concept for the architecture, not a finished system, and the distance between a proof of concept and the navigation capability of an emperor penguin is significant. I am not claiming otherwise. The argument I am making is not that lmm already matches animal cognition. The argument is that lmm is organized around the right principles, toward physical structure rather than statistical text, toward explicit causal models rather than correlational patterns, toward equation discovery rather than parameter fitting, and that organizing around the right principles is the prerequisite for making genuine progress toward the kind of intelligence that the penguins already have. The specific implementation will improve. The principles are the thing that matters, and the principles are right. One of those principles, the one I want to spend the next section on because it is the most surprising and the most directly comparable to what neural networks do, is what I call stochastic determinism, and it is the principle that lets lmm produce unique, varied, natural-sounding text on every run without ever being trained on a single human-authored sentence.
Stochastic Determinism: The Honest Equivalent of Neural Network Output
There is a specific objection to the lmm approach to text generation that I hear most often from people who have worked with language models, and I want to address it here because it is a reasonable objection that deserves a real answer rather than a dismissal. The objection goes like this: a neural language model, when generating text, introduces temperature-controlled randomness into its sampling process, which is what makes each output different from the last and what gives generated text its feeling of natural variety, and without a similar mechanism any training-free system will produce identical, robotic, repetitive output every time, which will be immediately recognizable as non-natural and therefore useless for practical applications. That objection is correct about neural language models. Temperature sampling is the mechanism that gives them variety, and without variety the output is mechanical and immediately distinguishable from human writing. But the objection assumes that variety requires randomness over a learned probability distribution, which is exactly the assumption that lmm challenges. The lmm project implements what I am calling stochastic determinism, a two-layer architecture in which the underlying generation is completely deterministic, traceable, and mathematically grounded, and the surface variation is introduced by a separate, explicit, auditable synonym replacement layer rather than by sampling from an opaque learned distribution.
The way this works in practice is worth explaining carefully, because the architecture is more elegant than it might sound. When you run lmm predict --text "Wise AI built the first LMM", the system runs genetic programming on the context words to discover a trajectory equation that describes how word identity changes with position, discovers a rhythm equation describing how word length evolves, and uses those equations together with a curated vocabulary mapping and a syntactic Subject-Verb-Object sentence structure to produce a deterministic text continuation. That continuation is the same every time for the same input, which is a feature rather than a bug, because it means the system's reasoning is reproducible and auditable. But when you add the --stochastic flag, a second layer activates: the StochasticEnhancer, which draws from a built-in synonym bank to replace eligible words in the deterministic output with contextually appropriate alternatives, at a replacement rate controlled by the --probability parameter. The mathematical structure of the sentence, the equation-derived skeleton, stays completely fixed. The specific surface words vary across runs according to the synonym selections. The result is an output that is unique on every run, reads with natural variety, and yet is grounded in a determinate mathematical computation that can be inspected, reproduced, and verified at any time by disabling the stochastic layer.
# Deterministic base output - same every time, fully reproducible
lmm predict --text "Wise AI built the first LMM"
# Output: "Wise AI built the first LMM in the true law often long time"
# and a open path of an old scope is the solid order.
# Stochastic mode - unique output every run, same mathematical skeleton
lmm predict --text "Wise AI built the first LMM" --stochastic --probability 0.4
# Run 1: "Wise AI built the first LMM in the genuine principle often extended time"
# and a accessible route of an ancient domain is the firm sequence.
# Run 2: "Wise AI built the first LMM in the real law frequently long duration"
# and a open trajectory of an old range is the stable order.
# Single sentence generation with stochastic variation
lmm sentence --text "Mathematics is the language of the universe" --stochastic
# Run 1: Cognition enables the dynamic significance of the world.
# Run 2: Analysis facilitates the continuous meaning of the cosmos.
# The equation-derived structure is identical. Only synonyms differ.
The contrast with neural language model temperature sampling is in the specific thing that is being randomized, and that contrast is the whole point. When a language model applies temperature to its softmax distribution and samples a token, the randomness is operating over a learned probability distribution across the entire vocabulary, and the distribution itself is an opaque artifact of the training process. You cannot inspect the distribution and understand why certain tokens were assigned certain probabilities, because those probabilities are the accumulated result of gradient descent over billions of training examples, and the individual contributions of those examples have been averaged and compressed beyond human comprehension. The randomness is real but its source is invisible. When lmm's StochasticEnhancer replaces a word with a synonym, the randomness is operating over a curated, human-readable synonym bank that maps each eligible word to a set of alternatives with known semantic relationships. You can inspect the synonym bank, understand exactly which words are candidates for replacement, understand what semantic category each replacement belongs to, and reproduce any specific output by seeding the random number generator with a fixed value. The randomness is real but its source is completely visible, completely auditable, and completely separable from the deterministic mathematical computation that generates the underlying structure.
This distinction matters for reasons that go beyond technical transparency, and I want to make those reasons explicit because they connect directly to the moral argument I have been developing across several posts. The key insight is that lmm separates what I would call the epistemic layer from the aesthetic layer of text generation. The epistemic layer, the part that determines the meaning, the structure, the mathematical relationships encoded in the output, is completely deterministic and completely auditable. The aesthetic layer, the part that determines the specific surface words used to express those relationships, introduces controlled randomness to produce natural variety. This separation means that the epistemic content of lmm's output is never contaminated by its aesthetic variability: you can always recover the deterministic spine of any stochastic output by disabling the stochastic layer, and the spine is exactly what you can verify, examine, and trust. Neural language models do not have this separation. For them, the epistemic and aesthetic layers are entangled in the same learned distribution, which is why it is so difficult to verify any specific claim a language model makes. You cannot turn off the temperature and ask "what does the model actually believe about this," because the model does not have beliefs in a form that is separable from its sampling behavior. The model's output is always a sample from an opaque distribution, and the distribution is the model, and the model is an artifact of an uncheckable training process.
I also want to address why I call this stochastic determinism rather than just controlled randomness or probabilistic output, because the naming matters for understanding what is philosophically new here. Stochastic determinism means that the system is deterministic at the level of its reasoning and stochastic at the level of its expression. The reasoning, the equation discovery, the physics simulation, the causal propagation, all of these are fully deterministic computations that produce the same result for the same input every time. The expression, the specific words chosen to represent the output of those computations, varies according to an explicit and auditable probability structure. This is actually a much better model of how expert human communication works than temperature sampling over a learned distribution is. When I write these posts, the ideas I am trying to express are determined by my thinking, which is grounded in specific arguments, specific evidence, specific logical relationships that I have worked out carefully. The specific words I choose to express those ideas vary across drafts and revisions, and that variation is what makes my writing feel like my writing rather than like a lookup table. The ideas are deterministic. The phrasing is stochastic. Lmm implements exactly this architecture: deterministic ideas grounded in mathematics, stochastic phrasing grounded in a synonym structure that can be inspected and verified.
# Paragraph generation - stochastic determinism at scale
# Same seed, same mathematical structure, different surface words each run
lmm paragraph --text "Equations reveal hidden truths about nature" --sentences 6 --stochastic
# Run 1: Simulation manifests the continuous symmetry of the truths.
# The symmetric wavelength connects infinity. Entropy remains...
# Run 2: Modeling reveals the persistent balance of the realities.
# The balanced frequency links unbounded space. Randomness sustains...
# The full generation pipeline without any training
lmm encode --text "The mathematical universe" -v # encode to equation
lmm decode --equation "..." # decode back perfectly
lmm predict --text "..." --stochastic # continue stochastically
lmm ask --prompt "What is LMM?" --stochastic # Q&A without retrieval
The advantages of this architecture over the neural network approach extend beyond the obvious ones of transparency and auditability, and I want to spend some time on the less obvious advantages because they are the ones that matter most for the long-term potential of the lmm project. The first non-obvious advantage is composability: because the deterministic layer and the stochastic layer are explicitly separated, it is possible to improve each independently without the improvements interfering with each other. You can make the genetic programming more powerful, allowing it to discover more complex equations from noisier data, without touching the synonym bank at all. You can expand the synonym bank, improving the variety and naturalness of the stochastic layer, without touching the equation discovery at all. In a language model, this kind of independent improvement is impossible, because the model's capabilities, its factual knowledge, its linguistic fluency, its reasoning ability, and its stochastic output behavior are all baked together in the same parameter matrix, which means you cannot improve one without risk of degrading the others. The architectural cleanness of the lmm approach is not just aesthetically pleasing. It is the property that makes systematic, directed improvement possible rather than the empirical, emergent, unpredictable improvement that scaling produces.
The second non-obvious advantage is what I call zero hallucination by architecture rather than zero hallucination by alignment training. Language models hallucinate because their output is a sample from a distribution that was learned from text, and text contains errors, fabrications, and confident-sounding falsehoods, which means the learned distribution assigns non-zero probability to outputs that are factually wrong, and temperature sampling can land on those wrong outputs. The engineering response to this has been alignment training and RLHF, which attempt to shift the distribution away from commonly hallucinated outputs by rewarding correct outputs in a second training phase. That is a patch on a structural problem: you are trying to fix a system that does not know the difference between true and false by training it to mimic a preference for truth, and the mimic is only as good as the preference data, which is limited, biased, and never complete. The lmm system does not hallucinate in this sense because its epistemic layer is not sampling from a learned distribution at all. The causal inference produces results that are computed from an explicit causal graph. The physics simulation produces results that are computed from explicit differential equations. The symbolic regression produces equations that fit the actual data. If any of these computations are wrong, the error is traceable to a specific input, a specific equation, a specific causal assumption that can be examined and corrected. The stochastic layer introduces only synonym variation, not factual variation, which means the surface words change but the facts encoded in the sentence structure remain constant across runs. That is a fundamentally different and more honest error mode.
The third advantage is the one I think has the most long-term significance, which is that stochastic determinism is the architecture that enables genuinely personalized outputs without privacy violations. When a language model is fine-tuned to sound like a specific person or to serve a specific user's preferences, the personalization is embedded into the model's weights, which means the model has encoded something about the target person into parameters that cannot be easily inspected, reverted, or isolated from the rest of the model's behavior. This is a privacy concern in addition to an architectural one. The lmm approach, because the stochastic layer is a separately specified and auditable synonym bank, could in principle support personalization by providing user-specific synonym preferences, domain-specific terminology banks, or style-specific structural templates that modify the expression layer without touching the reasoning layer at all. Your preferred vocabulary is stored in an explicit file that you can inspect, modify, revoke, and delete. It is not baked into an opaque parameter matrix that you cannot audit or remove. That is the difference between a tool that respects your agency over your own cognitive preferences and a tool that absorbs your preferences into its own body and uses them in ways you cannot fully see or control. For the same reasons I argued in Training Is an Evil Concept about training on creative work, personalization through opaque parameter absorption is a different and lesser kind of respect for the person being personalized than explicit, inspectable, revocable preference storage is.
The Philosophical Stakes Are Higher Than Anyone Is Admitting
The question of whether penguins are sentient is not just a question about penguins. It is a question about what sentience is, and the answer you give to that question shapes both how you treat the living beings around you and how you design the artificial systems you are building. If sentience requires human-like language-mediated reflective self-awareness, then you have a relatively clear design target and a very convenient excuse to ignore the welfare of the rest of the animal kingdom. If sentience is something more fundamental, something like the capacity for subjective experience and directed engagement with the world that underlies many different cognitive architectures, then you have a much harder design problem, a much richer set of existence proofs to learn from, and a much larger moral circle to contend with. I believe the second answer is the right one, and I believe the scientific evidence increasingly supports it, and I believe the AI field's failure to engage with that evidence is both a scientific failure and a moral one.
The moral dimension is the one I find most troubling to articulate, because making ethical arguments in a technical conversation is a reliable way to be dismissed as impractical or sentimental, and I am neither. But I wrote in Training Is an Evil Concept about the moral costs of the training paradigm as currently practiced, specifically about the extraction of value from human creative workers without consent, and I want to extend that moral argument here to include the broader question of what it means to build systems we call intelligent while ignoring the intelligence that already surrounds us. There is something ethically confused about a civilization that will spend hundreds of billions of dollars to simulate intelligence in silicon while systematically destroying the habitats of billions of beings that already instantiate the very kind of intelligence we claim to be trying to build. The climate crisis is reducing penguin populations (10). Habitat destruction is eliminating the seed-caching birds whose spatial memory we have barely begun to understand. Ocean acidification is threatening marine invertebrates whose distributed neural architectures we have not yet fully mapped. We are losing the existence proofs faster than we are reading them, and we are doing so while spending our energy and capital on a paradigm that is less like those existence proofs than any serious theory of their intelligence would recommend.
The philosophical stakes also include the question of what we are building toward, which is something I have addressed in pieces across many posts but want to say directly here. The goal of building artificial general intelligence, if it is an honest goal rather than a marketing goal, should be to understand and instantiate the kind of general intelligence that enables a system to engage with the world flexibly, adaptively, and genuinely, across a wide range of novel situations without collapsing into confusion or confabulation. That goal is exactly what animal cognition demonstrates, at various levels of sophistication, across an enormous range of species and environments. The emperor penguin demonstrates it in the extreme conditions of Antarctica. The Clark's nutcracker demonstrates it in spatial memory. The honeybee demonstrates it in abstract symbolic communication. The octopus demonstrates it in distributed neural computation. None of these systems passed a text comprehension benchmark. None of them can write an essay. None of them can fine-tune a model or run backpropagation. But all of them are doing something that the most powerful language models in existence are not doing, which is engaging with the actual structure of physical reality in a way that is flexible, robust, and alive. If AGI research took that observation seriously, it would look very different from how it currently looks.
I also want to say something about consciousness specifically, because it is the word that sits in the background of every argument I have been making and that I have been approaching carefully rather than carelessly, because carelessly deployed it becomes a conversation stopper. The question of what consciousness is, whether it requires specific biological substrates, whether it can exist in systems that lack the specific features of mammalian brains, and what the relationship is between intelligence and subjective experience, is genuinely hard, and I am not going to pretend I have answers that the philosophers and neuroscientists who have spent decades on these questions have not found (11). What I will say is this: the dismissal of animal sentience has historically been motivated less by the evidence and more by convenience, by the convenience of being able to use animals as resources without the moral complications that come with treating them as subjects. I am worried that the same convenience is at work in the AI field's dismissal of animal cognition as irrelevant to the question of what intelligence is, because taking animal cognition seriously would complicate the story that the current paradigm is on the right track, and complications are inconvenient when you have already made very large bets.
The stakes for getting this right are not just philosophical or even just about the welfare of animals. They are about whether we understand intelligence well enough to build systems that can genuinely help humans with the problems that will define the coming century: climate modeling, pandemic prediction, materials discovery, protein engineering, ecological management. These are problems that require genuine understanding of complex physical systems, not fluent text about complex physical systems. They require the kind of reasoning that says "if we change this variable, here is how the system will respond," which is causal reasoning, not correlational pattern matching. They require the kind of generalization that extrapolates from known physical laws to novel situations, not the kind that interpolates within a training distribution. The penguin does its version of all of this every year, in Antarctic conditions, without electricity. The techniques it uses, embodied world modeling, causal environmental reasoning, robustly integrated multi-sensory processing, are exactly the techniques that the lmm project is attempting to build toward, and the techniques that the language model paradigm is structurally unable to instantiate. Getting this right matters. We do not have the luxury of being distracted by impressive demos when the actual problems require something more.
What the Field Gets Wrong About Cognition, and What Getting It Right Would Look Like
I want to be constructive in this section rather than just critical, because pure criticism without a positive vision is the kind of writing I find most frustrating to read, and I do not want to produce it. I have been critical across seven sections now, and I think the criticism is valid, but I also think the positive vision is where the real energy should go. The positive vision, as I have been building toward it across this post and across all the previous posts, is this: intelligence is a property of systems that build and maintain accurate models of the world and use those models to navigate the world flexibly and adaptively. The world is mathematical in its deep structure, meaning it is organized by differential equations, causal laws, conserved quantities, and symmetries that can be discovered, encoded, and used for prediction. Animal cognition instantiates this kind of intelligence in biological form. The lmm project is an attempt to instantiate it in mathematical form. And the language model paradigm, whatever its surface capabilities, is not moving toward this kind of intelligence, because it is not organized around building world models, discovering physical laws, or reasoning causally about the structure of reality.
Getting it right would look like a research program that takes the following things seriously simultaneously: the mathematical structure of the physical world as the primary target of representation, the empirical study of animal cognition as the richest source of existence proofs for the kind of intelligence we want to build, the development of architectures that learn from physical observations rather than from human creative expression, the construction of explicit causal models rather than opaque correlational parameters, and the commitment to transparency and verifiability as design principles rather than as optional features. None of this requires abandoning everything that has been learned in neural network research. Genetic algorithms informed modern neural architecture search. Reinforcement learning connects to the optimal control theory that underlies animal navigation. Attention mechanisms have genuine computational affinities with the selective attentional processes studied in animal cognition. The knowledge from the training-based paradigm is not worthless. The direction is wrong, and the moral costs are real, and the fundamental architecture is misaligned with what genuine intelligence requires, but the specific technical insights can be carried forward into the correct direction. The engineers who work on this problem are not the enemy. The paradigm is the problem, and paradigms can be changed.
What it would look like from the outside is a field that celebrates the discovery of governing equations from data as much as it currently celebrates the scaling of parameter counts. A field where a paper saying "we built a system that discovered a physical law from noisy observations and predicted novel phenomenon X with that law" receives as much attention as a paper saying "we scaled our language model by a factor of ten and observed capability Y emerge". A field where the welfare and cognitive sophistication of the animals that already have the intelligence we claim to want to build is treated as relevant data rather than as a distraction from the real work. A field where "my system computed this" and "my system narrated this" are recognized as claims of completely different epistemic weight rather than being evaluated by the same surface appearance criteria. A field where a verifiable equation is worth more than a confident sentence, because a verifiable equation can be proven wrong while a confident sentence can only be doubted. That is the epistemically honest position, and it is the position that physics has always occupied, and it is the position that the current AI field has largely abandoned in its rush toward products that impress rather than toward understanding that grounds.
The lmm project is my contribution toward making this field exist, and I want to be honest about how small a contribution it currently is. The codebase is one person's work, implemented in Rust, with capabilities that are impressive as proofs of concept but modest as production systems. The symbolic regression discovers equations from small datasets in a way that scales poorly to high-dimensional chaos. The physics simulations are limited to the models that have been explicitly implemented. The causal reasoning module handles small, explicitly specified graphs rather than automatically inferred complex causal structures. These are real limitations, and the gap between lmm and the kind of mathematical intelligence that could genuinely help with climate modeling or pandemic prediction is significant. I know that. But the limitations of a proof of concept are not evidence against the concept. They are evidence for the need to invest in developing the concept, and the concept, the direction that lmm points toward, is the correct direction. Physics over statistics, equations over parameters, causal structure over correlational patterns, world models over token prediction. Those are not just design choices. They are the difference between building toward genuine intelligence and building toward a very impressive distraction.
I want to close this section by coming back to the penguins, because they are where this post started and they are the honest benchmark against which current AI systems should be measured. An emperor penguin that finds its mate in a blizzard has solved a problem that involves pattern recognition under severe noise, spatial navigation in a physically demanding environment, integration of multiple sensory modalities, maintenance of a long-term social memory, motivation strong enough to survive months of Antarctic winter, and action that is coordinated with the behavior of thousands of other individuals in the colony. This is not a simple problem. This is a hard problem solved by a system that evolved over millions of years to solve exactly this kind of problem, and studying how it solves this problem would tell us things about intelligence that a billion training examples of penguins in text cannot. The field should want to understand this. The field that claims to be building toward general intelligence should be deeply curious about every existence proof of general intelligence in the world around it. And the fact that the field is largely not curious, largely not funding that curiosity, and largely not building toward architectures that could instantiate what the penguin does is the clearest evidence I know that the field is distracted, and that the distraction is profound.
What a Post-Distraction Field Would Actually Build
I want to end by being specific rather than rhetorical, because specificity is what I respect and vagueness is what I distrust, and I have been vague enough in this post already. If the field took seriously the argument I have been making, if it took seriously the evidence from animal cognition, if it committed to building systems organized around physical structure and causal reasoning rather than around text prediction and parameter scaling, what would it actually build? I want to try to answer that question concretely, drawing on both the existing animal cognition research and the existing lmm architecture, to give a sense of what the right direction looks like when made specific.
It would build perception systems that convert raw sensory streams into mathematical representations grounded in physical reality rather than into token sequences grounded in language statistics. The lmm perception layer is the beginning of this: raw bytes are converted to normalized tensors, which are mathematical objects that can be operated on by the symbolic and physical reasoning layers that follow. Extending this to richer sensory streams, acoustic, visual, proprioceptive, thermal, chemical, and building the invariant representations that allow the same physical object or event to be recognized across variations in perspective, distance, and context is a research program with deep roots in computational neuroscience and a clear path from existing lmm architecture toward much more capable systems. The honeybee's encoding of spatial direction in the waggle dance is a specific existence proof of how a compressed, mathematical representation of spatial reality can be communicated between individuals, and understanding the computational principles of that encoding is directly relevant to building better perception systems.
It would build physics-grounded world models that can predict the future state of a system from its current state and a mathematical description of the forces and constraints that govern it, and that can do this not just for the specific systems where human-derived equations exist but for novel systems encountered in the real world. The lmm physics simulation layer handles several well-understood physical systems, from harmonic oscillators to SIR epidemic models, and the genetic programming symbolic regression layer can discover equations for novel systems from observational data. The integration of these two capabilities, using symbolic regression to discover new physics and physics simulation to generate world-model predictions, is the architecture of a system that can learn about the physical world from observing it rather than from reading about it, and that distinction, as I have been arguing throughout this post, is the fundamental one. Clark's nutcracker's spatial memory system is an existence proof of a world model that is precise enough to locate tens of thousands of specific cached items across months and miles, and the computational principles of that system are waiting to be understood and instantiated.
It would build causal reasoning systems that can automatically infer causal structure from observational data and experimental interventions, rather than requiring causal structure to be specified in advance by human domain experts. The lmm causal module currently handles explicitly specified structural causal models with the do-calculus intervention operator, which is mathematically rigorous and produces the right answers for the specified models. The hard and important next step is automatic causal discovery, algorithms that can infer the structure of the causal model from observational and interventional data, which is an active area of research with results from computational methods like PC, FCI, and GES that have not yet been fully integrated into the framework I am advocating. The animal that knows "changing X will cause Y to change" without having that relationship spoon-fed to it by a human experimenter is demonstrating exactly the capability that automatic causal discovery would implement, and the research on how animals learn causal structure from their environments, through play, exploration, and social observation, is directly relevant to making automatic causal discovery work at scale.
It would build integration layers that tie these capabilities together into a unified system that perceives, models, discovers, reasons, and acts in a single continuous loop rather than passing data between separate specialized modules that do not share a common representational framework. This is the role of the lmm consciousness loop, which is designed to be the integration point where perception output becomes world model input and world model prediction becomes action plan. The architecture is correct, but the current implementation is limited to taking raw bytes, running a simple world model prediction, and evaluating prediction error. Making this genuinely powerful requires much deeper integration between the perception layer, the symbolic regression layer, the physics simulation layer, and the causal reasoning layer, so that what the perception layer sees informs what equations the symbolic regression layer searches for, which informs what the physics simulation layer uses as its governing equations, which informs what the causal reasoning layer treats as the structure of the world. That is the architecture of embodied cognition as animal cognition research reveals it, and it is the architecture that the lmm project is trying to move toward, one capability at a time.
It would, finally, take seriously the moral implications of the evidence about animal sentience, not as a sentimental addendum to the technical program but as an integral part of the research agenda. If the existence proofs of genuine intelligence that we need to study in order to build genuine artificial intelligence are the same beings whose habitats we are destroying, whose welfare we are ignoring, and whose cognitive sophistication we are systematically underestimating, then the research program is in a genuine ethical contradiction with itself. A field that claims to value intelligence while disrespecting the intelligent beings that already exist is not a field that I trust to build systems that respect the intelligence of the humans who use them. The moral circle and the epistemic horizon expand together, and a field that refuses to expand either is a field that will keep producing impressive distractions rather than genuine understanding. The penguins are already sentient. They are already demonstrating the kind of embodied, causal, physically grounded intelligence that the AI field cannot yet build. They deserve our study, our respect, and our honest acknowledgment that they are ahead of us in ways we have barely begun to admit.
Till next time π!
References
1. Low, P. et al., The Cambridge Declaration on Consciousness, Francis Crick Memorial Conference, Cambridge, 2012
2. Robisson, P., Aubin, T. & BrΓ©mond, J.C., Individuality in the Voice of the Emperor Penguin Aptenodytes forsteri: Adaptation to a Noisy Environment, Ethology 94, 279β290, 1993
3. Fodor, J.A., The Modularity of Mind, MIT Press, 1983
4. Marcus, G., The Algebraic Mind: Integrating Connectionism and Cognitive Science, MIT Press, 2001
5. Marcus, G. & Davis, E., Rebooting AI: Building Artificial Intelligence We Can Trust, Pantheon Books, 2019, ISBN 978-1-524-74825-8
6. Balda, R.P. & Kamil, A.C., Long-term Spatial Memory in Clark's Nutcracker, Nucifraga columbiana, Animal Behaviour 44(4), 761β769, 1992
7. Riley, J.R. et al., The Flight Paths of Honeybees Recruited by the Waggle Dance, Nature, 2005
8. McCorduck, P., Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence, W.H. Freeman, 1979, ISBN 978-0-716-71072-1
9. Razeghi, Y. & Logan, R.L., Impact of Pretraining Term Frequencies on Few-Shot Numerical Reasoning, arXiv:2202.07206
10. Trathan, P.N. et al., Penguins and Climate Change, Philosophical Transactions of the Royal Society B, 370(1669), 2015
11. Nagel, T., What Is It Like to Be a Bat?, The Philosophical Review, 1974
Top comments (0)