<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mahmoud Harmouch</title>
    <description>The latest articles on DEV Community by Mahmoud Harmouch (@wiseai).</description>
    <link>https://dev.to/wiseai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F827951%2F8d69ab6e-a28d-4d13-9652-bdd2f5eb90f8.png</url>
      <title>DEV Community: Mahmoud Harmouch</title>
      <link>https://dev.to/wiseai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/wiseai"/>
    <language>en</language>
    <item>
      <title>Jesus Was Right. You Are God and Infinite.</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Sun, 23 Aug 2026 05:05:41 +0000</pubDate>
      <link>https://dev.to/wiseai/jesus-was-right-you-are-god-and-infinite-6cc</link>
      <guid>https://dev.to/wiseai/jesus-was-right-you-are-god-and-infinite-6cc</guid>
      <description>&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;I wrote in &lt;a href="https://wiseai.dev/blogs/christianity-makes-perfect-sense" rel="noopener noreferrer"&gt;Christianity Makes Perfect Sense!&lt;/a&gt; about how some truths are deeply unfashionable, and how the discomfort a truth produces is often the surest sign that it is worth sitting with. This post is in that spirit. I want to talk about something that I have been turning over in my head for a long time, something that Christianity itself, read carefully and without the institutional filters, actually supports with a startling degree of directness. The title gives it away so I will not pretend there is a mystery: Jesus was right, and among the many things he was right about is this one, which most church hierarchies have buried so thoroughly that most Christians have never encountered it. He was right that you are, in some profound and irreducible sense, divine. That you carry within you something that the tradition calls godlike, that you are not merely a creature made to bow down and wait, but a being made from the same substance that the Gospel of John calls the Word, the very creative principle through which all things were made. I know how that sounds. I know it reads either as arrogance or as heresy depending on who you are. I am going to spend the rest of this post making the case that it is neither, that it is actually the oldest and most straightforward reading of the text, and that the institutions which suppressed it did so not because it was false but because it was dangerously empowering.&lt;/p&gt;

&lt;p&gt;I should say upfront that I am not a theologian and I have never claimed to be. I am someone who grew up with very complicated feelings about religion, as I described in &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt;, in a village where religion was wrapped tightly around community identity and it was very hard to separate what was actually in the text from what the local authorities decided the text meant. I read the Bible on my own, away from institutions, away from authority figures, and when I did that I kept finding things that shocked me. Not because they were dark or frightening, but because they were so much more generous, so much more expansive about what a human being is, than anything I had been taught. That gap between what the text actually says and what I was told the text says is what this post is about. And I do not think it is a gap produced by ignorance. I think it is a gap that has been maintained deliberately, because the idea that every single human being carries something divine is one of the most democratizing and most politically dangerous ideas in the history of human thought. Power does not survive the equal distribution of its own source, and if every person is in some sense divine, then no institution can claim a monopoly on God anymore.&lt;/p&gt;

&lt;p&gt;I should also be honest about where I am coming from religiously, because I think it matters for how you receive what I am about to say. I was Muslim on paper. Born into it, raised inside it, surrounded by it. But paper does not mean anything to me anymore, and it never really did. As I wrote in &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt;, I spent a long time after that as agnostic, standing on that thin, uncomfortable line between belief and doubt, genuinely unsure whether anything was out there and deeply skeptical of every institution claiming to speak for it. That skepticism has not gone away. But something shifted, slowly and without ceremony, as I kept reading. Not reading theologians or apologists or anyone trying to sell me a system, but reading the actual words attributed to Jesus, the specific arguments he made, the things he said to people who were suffering, the things he said to people who were powerful, and the things he said about what a human being fundamentally is. I am not a Christian in the institutional sense. I do not belong to a church, I do not follow a creed, and I hold plenty of the same questions I always have. But I follow Jesus's teachings in the sense that matters to me: I think he was telling the truth, I think the truth he was telling is the most important truth I have encountered, and I think it has been buried under two thousand years of institutional management. This post is my attempt to dig it back up.&lt;/p&gt;

&lt;p&gt;Let me also connect this to the broader thread running through several of my recent posts. In &lt;a href="https://wiseai.dev/blogs/life-on-earth-is-100-ai-generated-slop" rel="noopener noreferrer"&gt;Life On Earth is 100% AI Generated Slop&lt;/a&gt;, I argued that the world is optimized for plausible appearances rather than ground truth, and that this applies as much to social institutions as to language models. Religious institutions are not exempt from this critique. They are, in many cases, the original appearance-optimization machines, producing outputs that look like truth, sound like authority, and smell like holiness, while often encoding the preferences of whoever held the pen. But the text is still there. The ground truth of the scriptural record, whatever you think of its origins, is still available to anyone who reads it, and what it says about human nature is not what most institutions have wanted you to hear. This post is my attempt to read it honestly, to take the most uncomfortable passages seriously, and to follow the logic wherever it leads.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Verse
&lt;/h2&gt;

&lt;p&gt;There is a moment in the Gospel of John, chapter ten, verses thirty-three through thirty-six, that I think is the most underappreciated exchange in the entire New Testament. The religious leaders have accused Jesus of blasphemy for saying &lt;a href="https://www.biblegateway.com/passage/?search=John+10%3A30&amp;amp;version=ESV" rel="noopener noreferrer"&gt;"I and the Father are one"&lt;/a&gt;, and his response is, to put it plainly, astonishing. He does not back down. He does not explain away the claim. Instead he turns to the Hebrew scriptures and asks his accusers a direct question. He quotes &lt;a href="https://www.biblegateway.com/passage/?search=Psalm+82%3A6&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Psalm 82:6&lt;/a&gt;, which says, "I said, you are gods", and then he makes a logical argument so sharp that it has never been successfully refuted in two thousand years of theological commentary. He says, in effect: if your own scripture, which cannot be broken, calls human beings gods, on what grounds do you accuse me of blasphemy for saying I am the Son of God? That is not a dodge. That is a detonation of the entire premise of the accusation. He is saying that the concept of human beings having a divine nature is already in the text, is already Scripture, is already established, and that his claim is therefore not an invention or an overreach but a specific instance of a general principle.&lt;/p&gt;

&lt;p&gt;The verse he quotes, &lt;a href="https://www.biblegateway.com/passage/?search=Psalm+82%3A6&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Psalm 82:6&lt;/a&gt;, reads in its fuller form: "I said, 'You are gods, sons of the Most High, all of you;'" The Hebrew word translated as "gods" is &lt;em&gt;elohim&lt;/em&gt;, the same word used for God throughout the Old Testament, including in &lt;a href="https://www.biblegateway.com/passage/?search=Genesis+1%3A1&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Genesis 1:1&lt;/a&gt;, where it says "In the beginning, God (&lt;em&gt;elohim&lt;/em&gt;) created the heavens and the earth". The psalm uses the same word, &lt;em&gt;elohim&lt;/em&gt;, to address human beings. The standard institutional response to this has been to say that the psalm is speaking about judges or magistrates, people who had been given delegated authority to render divine justice, and that the word "gods" is therefore just a title of office rather than a statement about nature. That interpretation is not obviously wrong, but it is not obviously right either, and what matters for our purposes is that Jesus did not restrict it that way. He quoted the verse, said "the scripture cannot be broken", and then used the existence of human beings called &lt;em&gt;elohim&lt;/em&gt; as the foundation of his own defense in &lt;a href="https://www.biblegateway.com/passage/?search=John+10%3A34-36&amp;amp;version=ESV" rel="noopener noreferrer"&gt;John 10:34-36&lt;/a&gt;. He did not say "yes but that was only about judges". He treated it as a general statement about humanity, and he used it accordingly.&lt;/p&gt;

&lt;p&gt;The scholar Michael Heiser, who spent his career studying the divine council worldview in ancient Near Eastern literature and the Hebrew Bible, has argued extensively that Psalm 82 is not simply about human judges (1). He argues that the psalm depicts a divine assembly, a council of beings who were given governing authority over the nations by the Most High God, and that the accusation against them is that they have failed in their responsibility to maintain justice. The phrase "you will die like men" in verse seven, which has often been used to prove the addressees are human, Heiser argues is actually a statement of judgment rather than a statement of nature; it is saying that these beings, as a consequence of their failure, will now be stripped of their divine status and die as mortals do. This scholarly debate is not resolved, and I am not pretending to resolve it here. But it does mean that the easy institutional dismissal of the verse, the one that says this is just about human judges and therefore has no broader implications, is not the settled consensus of biblical scholarship. It is one reading among several, and not necessarily the strongest one.&lt;/p&gt;

&lt;p&gt;What I find most compelling about the John 10 passage is not any particular resolution of the Psalm 82 debate, but the fact that Jesus, in the most heated confrontation of his ministry, with his life literally on the line, chose to ground his defense in the assertion that scripture already teaches the divine nature of human beings. He did not say "you are misunderstanding me, I am not claiming to be divine". He said "you are inconsistent, because you already accept that human beings can be called gods, and my claim goes no further than that". Read that slowly. The defense of his own divinity was the observation that divinity is not as rare or as exclusive as his accusers assumed. That is the argument. That is the text. And if you believe the scripture cannot be broken, as Jesus explicitly said, then you have to reckon with what that argument implies about every human being.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Genesis Said
&lt;/h2&gt;

&lt;p&gt;The argument from John 10 would be interesting on its own, but it becomes overwhelming when you read it alongside what Genesis says about human origins. &lt;a href="https://www.biblegateway.com/passage/?search=Genesis+1%3A26&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Genesis 1:26&lt;/a&gt; reads, in the English Standard Version: "Then God said, 'Let us make man in our image, after our likeness.'" This verse has been discussed, disputed, and dissected for thousands of years, and I do not pretend that everything about it is clear. But a few things are not in dispute, and those few things are enough to make the point I want to make. The first thing that is not in dispute is that the word translated "image" is the Hebrew &lt;em&gt;tselem&lt;/em&gt;, which is the word commonly used for a physical representation or statue of a deity in ancient Near Eastern contexts. When a king placed an image of himself in a distant province, that image was understood to carry the authority and presence of the king. It was not the king, but it participated in the king's identity in a functional and representational sense. When Genesis says that human beings are made in the &lt;em&gt;tselem&lt;/em&gt; of God, it is doing something extraordinary: it is saying that every human being is a representation of divine presence in the physical world, in the same way that a royal statue represented a king's authority in an outlying territory.&lt;/p&gt;

&lt;p&gt;The second thing that is not in dispute is the radicalness of applying this language to all human beings. In the ancient Near Eastern world from which Genesis emerged, the language of being made in the image of a god was reserved for kings. Kings were the image of the deity. Ordinary people were not. When &lt;a href="https://www.biblegateway.com/passage/?search=Genesis+1%3A26&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Genesis 1:26&lt;/a&gt; applies &lt;em&gt;tselem&lt;/em&gt; to humanity without distinction, it is not describing existing norms. It is shattering them. It is making the universal claim that every person, regardless of birth, rank, or nationality, carries within themselves the functional representation of the divine. The scholar Phyllis Trible pointed out decades ago that this democratization of the divine image is one of the most politically subversive moves in ancient literature (2). It takes a concept that was used to justify royal power and extends it to everyone, which means it simultaneously undermines the exclusive claim of any ruler to divine status. If everyone is made in the image of God, then no king is more divine than the slave who built his palace.&lt;/p&gt;

&lt;p&gt;The third thing worth noting is what follows immediately from the declaration that humans are made in God's image. &lt;a href="https://www.biblegateway.com/passage/?search=Genesis+1%3A28&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Genesis 1:28&lt;/a&gt; says that God gave humanity "dominion" over all living things, and that verb, the Hebrew &lt;em&gt;radah&lt;/em&gt;, is used elsewhere in the Bible specifically for the exercise of royal authority. This is not a coincidence. The grammar of the passage is deliberate. Humanity is created in the image of the divine, and humanity is given the exercise of royal power. The two go together in the logic of the text, because an image of a king is expected to exercise that king's authority. The implication is not that humans should act like tyrants. It is that humans carry, by nature, by virtue of what they are, a form of authority that in the ancient world was associated exclusively with the divine. That is a staggering claim, and it is sitting there on the first page of the Bible, waiting for anyone who is willing to read it without the institutional filters that have been layered over it for two millennia.&lt;/p&gt;

&lt;p&gt;I want to be careful here, because I think precision is important and I do not want to overstate the case in a way that is easy to dismiss. The text does not say that human beings are identical to God. It says that human beings are made in God's image, and those are not the same thing. An image of a king is not the king. But an image of a king participates in the king's identity, carries the king's authority, and cannot be understood in isolation from the king it represents. The theological tradition has often used this distinction to minimize what Genesis is saying, to reduce the &lt;em&gt;imago Dei&lt;/em&gt; to a list of specific moral or rational capacities that humans have in common with God. But the ancient context of the language suggests something more functional and more relational than that. It suggests that humanity's very existence in the world is understood as a form of divine presence, that every person is, in the most literal sense the text supports, carrying God into the world just by being what they are. That is not a minor doctrinal detail. That is the foundational anthropology of the entire scriptural tradition, and we have spent centuries treating it like a footnote.&lt;/p&gt;

&lt;p&gt;I also want to connect this back to &lt;a href="https://www.biblegateway.com/passage/?search=Psalm+90%3A2&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Psalm 90:2&lt;/a&gt;, which is one of the key passages about God's infinity: "Before the mountains were brought forth, or ever you had formed the earth and the world, from everlasting to everlasting you are God". This verse establishes that God exists outside of time, that the divine nature is not bounded by birth or death, beginning or end, the way finite creatures are. The interesting question is what it means for beings made in the image of this infinite God to exist as finite creatures. The dominant tradition has said that the image is present but diminished, that humans participate in divine-likeness in a limited way while remaining fundamentally bounded. But the mystical traditions within Christianity, which I will come to shortly, have asked whether the limitation is intrinsic or whether it is something that can be dissolved through transformation, whether the finite image can, through some process of union with its infinite source, come to participate in the infinity it represents. That question is not answered by Genesis alone. But Genesis plants the seed of it, and the rest of the scriptural tradition spends a great deal of time watering it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Verse Paul Almost Let Slip
&lt;/h2&gt;

&lt;p&gt;If the Genesis passage is the founding statement of human divinity and the John 10 passage is the moment when Jesus explicitly asserts it in an argument, then the passage I want to discuss next is the one that the institutional church has spent the most energy trying to contain. &lt;a href="https://www.biblegateway.com/passage/?search=2+Peter+1%3A4&amp;amp;version=ESV" rel="noopener noreferrer"&gt;2 Peter 1:4&lt;/a&gt; says something so direct that it is genuinely startling when you read it without preconceptions: God's promises have been given to us so that through them we "may become partakers of the divine nature". The Greek word there is &lt;em&gt;theias phuseos koinonoi&lt;/em&gt;, which translates with remarkable precision as "sharers in the divine nature". Not sharers in divine favor, not recipients of divine grace in a general sense, but partakers of the divine nature itself. The Greek word &lt;em&gt;phusis&lt;/em&gt; is the word for nature in the ontological sense, the kind of nature that makes a thing what it is. And &lt;em&gt;koinonoi&lt;/em&gt; means sharers or participants in a deep, structural sense, not observers or admirers from a distance.&lt;/p&gt;

&lt;p&gt;The Eastern Orthodox tradition took this verse more seriously than any other major branch of Christianity, and it built an entire theology around it under the name &lt;em&gt;theosis&lt;/em&gt;, which is usually translated as deification or divinization. The core idea of &lt;em&gt;theosis&lt;/em&gt; is that the human being is called to participate in the life of God, not merely to obey God's commands or to be forgiven by God's grace, but to actually share in the divine nature in a way that transforms what a human being is. St. Athanasius of Alexandria, writing in the fourth century, summed it up in a sentence that theologians have been quoting ever since: "God became human so that humans might become gods". He did not mean that humans would become the Creator or replace the Trinity. He meant that the Incarnation opened a path for human nature to be so thoroughly infused with the divine life that the boundary between creature and Creator, while never eliminated, becomes permeable in a way that it was not before. That is a serious theological claim, and it was made by one of the most respected theologians in Christian history, not by a heretic or a mystic on the fringes.&lt;/p&gt;

&lt;p&gt;The way that Orthodox theology has usually navigated this is through the distinction between God's essence and God's energies, developed most fully by Gregory Palamas in the fourteenth century. The essence of God, in Palamas's account, is entirely transcendent and unknowable, and humans never participate in that. But the energies of God, which are understood as the real and active divine life that permeates the world and makes contact with creation possible, are genuinely participable. Humans can be united with the divine energies in a way that transforms them, and this transformation is what deification means. The analogy Palamas and others used was the image of iron placed in a fire: the iron does not become fire in its nature, but it takes on all the properties of fire, the heat and the light and the consuming energy, through genuine contact. That is what &lt;em&gt;theosis&lt;/em&gt; looks like in Orthodox theology, and that is one of the most rigorous and most ancient Christian attempts to take Second Peter 1:4 seriously rather than explaining it away (3).&lt;/p&gt;

&lt;p&gt;I want to be honest about why this matters to me personally, and not just as a theological argument. I have spent years in situations that felt like there was no door, no way forward, no resource left, and in those situations the question of what a human being fundamentally is becomes more than academic. When everything external is stripped away, when the career is destroyed as I described in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt;, when the future looks like a wall of nothing, the question of whether there is something inside that is not subject to the same entropy as everything outside becomes the only question that matters. I am not claiming that the answer gave me a magical solution or that reading a verse in Second Peter solved my problems. I am saying that the difference between believing you are a finite biological process that will eventually terminate and believing you carry something that is, in its origins, infinite and divine, is not a small psychological difference. It is the difference between a temporary accident and a being with permanent value, and that difference affects everything about how you relate to suffering, to failure, and to the possibility of continuing.&lt;/p&gt;

&lt;p&gt;There is also a connection here to what I wrote in &lt;a href="https://wiseai.dev/blogs/be-aware-of-the-current-ufos-pandemic-remember-we-are-alone" rel="noopener noreferrer"&gt;Be Aware of The Current UFOs Pandemic. Remember, We Are Alone&lt;/a&gt;, where I argued that if we really are alone in the universe, then the weight of consciousness falls entirely on us, and we had better take that weight seriously. The theology I am describing in this post takes that weight with maximum seriousness. It says that human beings are not accidents in an indifferent universe. They are, in the most literal reading of the scriptural text, images of the infinite, bearing within themselves something that is not bounded by time or space in the way that ordinary matter is. That does not resolve all of the scientific questions about consciousness and matter. But it is a framework that demands we take human beings seriously, that refuses to reduce them to mere information-processing systems, and that grounds human dignity in something more fundamental than utility. Given everything I have written about how the tech industry and the economic system treats human beings as consumable inputs rather than ends in themselves, a theology that says every person carries something infinite is not just comforting. It is a political demand.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Kingdom and the Infinity
&lt;/h2&gt;

&lt;p&gt;There is a passage in the Gospel of Luke, chapter seventeen, verses twenty and twenty-one, that is short enough to quote in full. The Pharisees ask Jesus when the kingdom of God is coming, and he answers in &lt;a href="https://www.biblegateway.com/passage/?search=Luke+17%3A20-21&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Luke 17:20-21&lt;/a&gt;: "The kingdom of God is not coming in ways that can be observed, nor will they say, 'Look, here it is!' or 'There!' For behold, the kingdom of God is in the midst of you". The translation is disputed, because the Greek can mean either "in the midst of you", meaning among you as a social reality, or "within you", meaning inside each person. Both are defensible translations, and both matter for what I want to say here. But the combination of both meanings is what I think the verse is actually pointing at: the kingdom of God is not a future external event that humans wait for passively. It is a state of reality that is already present, already available, already latent in the nature of the beings who are asking the question.&lt;/p&gt;

&lt;p&gt;This verse has been the touchstone of the mystical tradition within Christianity for two thousand years. The medieval mystic Meister Eckhart, who was charged with heresy for saying it too directly, spent most of his life elaborating the same core insight: that the divine ground of the soul and the divine ground of God are the same ground, that there is a depth within the human being that is not separate from the depth within God, and that the spiritual life is not about traveling somewhere else but about going deeper into what is already there. Eckhart was not making this up. He was reading &lt;a href="https://www.biblegateway.com/passage/?search=Luke+17%3A21&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Luke 17:21&lt;/a&gt;, &lt;a href="https://www.biblegateway.com/passage/?search=Genesis+1%3A26&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Genesis 1:26&lt;/a&gt;, &lt;a href="https://www.biblegateway.com/passage/?search=Psalm+82%3A6&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Psalm 82:6&lt;/a&gt;, &lt;a href="https://www.biblegateway.com/passage/?search=2+Peter+1%3A4&amp;amp;version=ESV" rel="noopener noreferrer"&gt;2 Peter 1:4&lt;/a&gt;, and a dozen other passages, and he was following the logic of the text more consistently than the institutions around him wanted him to. His sermons, which are available to anyone who wants to read them today, contain some of the most ferociously honest theological thinking in the history of Christianity (4). The idea that the divine is not somewhere above you but at the core of what you already are is not a deviation from Christian teaching. It is one of its oldest and most persistent threads, showing up in different forms in Origen, in Augustine, in Maximus the Confessor, in Gregory of Nyssa, in Thomas Aquinas's doctrine of the soul's capax Dei, its capacity for God.&lt;/p&gt;

&lt;p&gt;The connection to infinity is direct. If the kingdom of God is within you, and if God is, as &lt;a href="https://www.biblegateway.com/passage/?search=Psalm+90%3A2&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Psalm 90:2&lt;/a&gt; says, "from everlasting to everlasting", then what is within you is not bounded by time in the ordinary sense. This is not a claim about biological immortality or about the physical body surviving death. It is a claim about the nature of whatever it is within a person that participates in the divine. The Johannine writings make this point repeatedly. &lt;a href="https://www.biblegateway.com/passage/?search=John+17%3A3&amp;amp;version=ESV" rel="noopener noreferrer"&gt;John 17:3&lt;/a&gt; defines eternal life as knowing God, not as a future reward but as a present state of cognition and relationship. &lt;a href="https://www.biblegateway.com/passage/?search=John+10%3A28&amp;amp;version=ESV" rel="noopener noreferrer"&gt;John 10:28&lt;/a&gt; has Jesus saying that eternal life has already been given to those who follow him, in the present tense. &lt;a href="https://www.biblegateway.com/passage/?search=1+John+3%3A2&amp;amp;version=ESV" rel="noopener noreferrer"&gt;1 John 3:2&lt;/a&gt; says "we are God's children now", again present tense, not a future condition but a current ontological status. The infinity the tradition ascribes to God is not, in the New Testament texts, understood as entirely absent from human experience. It is understood as present, available, and in some sense already constitutive of what a human being most deeply is.&lt;/p&gt;

&lt;p&gt;I want to say something about what this means in practice, because I think the mystical tradition sometimes makes it sound so abstract that people cannot connect it to anything real in their lives. When I say you are infinite, I do not mean that you will live forever in a biological sense or that your current circumstances do not matter or that suffering is an illusion. I mean something more specific and more verifiable in experience: that the part of you that observes your circumstances, that witnesses your suffering, that thinks about your thinking, has a relationship to the specific conditions of your life that is not one of simple identity. Your circumstances are finite. Your losses are finite. Your failures are finite. But the capacity to be aware of all of those things, to witness them, to turn attention onto them, seems to have no natural limit. You can always go one level deeper. There is always another layer of awareness below awareness. And the tradition I am describing says that this recursive quality of consciousness, this fact that awareness seems to have no floor, is not a coincidence or a quirk of neuroscience. It is evidence that consciousness is participating in something that is, at its roots, not subject to the same limitations as the material world.&lt;/p&gt;

&lt;p&gt;The philosopher David Chalmers called this the hard problem of consciousness: the question of why there is subjective experience at all, why there is something it is like to see red or to feel grief or to wonder about one's own existence (5). The hard problem is hard precisely because it resists the kind of explanation that works for everything else in science. You can explain how the brain processes information, how neurons fire, how patterns emerge from complexity. But you cannot explain, within the framework of purely physical description, why any of that physical activity is accompanied by experience. Why is there something it is like to be a brain, when there is presumably nothing it is like to be a thermostat or a wind tunnel, even though all three are complex physical systems processing information? That question, Chalmers argues, points to something about consciousness that is not fully captured by the physical description, and I think the scriptural tradition I have been describing is offering an answer that no amount of neuroscience has yet ruled out: that consciousness participates in a dimension of reality that is not fully exhausted by its physical substrate, that the awareness within you reaches beyond the boundaries of the body that hosts it, and that this reaching beyond is the seed of what the tradition calls the divine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What John 17 Reveals
&lt;/h2&gt;

&lt;p&gt;The most stunning theological statement in the entire New Testament, in my opinion, is not in Romans or in Corinthians or in Revelation. It is in &lt;a href="https://www.biblegateway.com/passage/?search=John+17%3A20-23&amp;amp;version=ESV" rel="noopener noreferrer"&gt;John 17:20-23&lt;/a&gt;, where Jesus is praying and says the following: "I do not ask for these only, but also for those who will believe in me through their word, that they may all be one, just as you, Father, are in me, and I in you, that they also may be in us... The glory that you have given me I have given to them, that they may be one even as we are one, I in them and you in me, that they may become perfectly one". Read that one more time. He is asking that human beings become one with God and with each other in the same way that he is one with the Father. He is not asking for a looser version of that unity, a kind of distant relationship or a formal legal standing. He is asking that the same kind of unity that exists within the Trinity, the most fundamental theological category for the divine nature, be extended to encompass humanity. That is the most radical statement of human divine potential in the entire scriptural canon, and it comes from Jesus himself, in what Christian tradition has called his high priestly prayer.&lt;/p&gt;

&lt;p&gt;The implications of this verse are enormous and have been carefully avoided by most institutional Christianity, because they are almost impossibly difficult to contain within a hierarchical religious structure. If human beings are expected to be one with God in the same way that the Son is one with the Father, then the distinction between the sacred and the profane, between the holy and the ordinary, between the clergy and the laity, becomes philosophically unstable. If every person carries within them the capacity for the same union with God that the tradition reserves for Christ, then every person has an immediate access to the divine that does not require church, priest, ritual, or institutional intermediary. That is a terrifying claim for any institution whose power depends on controlling access to the divine, and it explains why this verse, though it is sitting there in every Bible that has ever been printed, is almost never preached on in the terms I am using here. The institutions that control the interpretation of the text have a very strong incentive to read "that they may be one" as a call to institutional unity, to church membership, to doctrinal agreement, rather than what the grammar and the context actually say, which is that Jesus is asking for the same quality of union between humans and God that exists within God's own nature.&lt;/p&gt;

&lt;p&gt;The theologian Karl Barth, who was certainly not a mystic or a theological liberal, wrote extensively about the idea that human nature has been taken up into the divine life through the Incarnation in a way that changes the baseline status of every human being (6). Barth's argument is that when the Son of God became human, the divine and the human were united in a single person in a way that permanently altered the relationship between divinity and humanity. It is not just Jesus who is transformed by the Incarnation. Because Jesus took on human nature, human nature itself was touched by the divine in a way it had not been before, and that touch is not confined to Christians or to believers or to any particular group. It extends to all of humanity, because what was taken up in the Incarnation was human nature as such, not the human nature of any particular subset of people. This is Barth's doctrine of universal Christology, and while he resisted drawing all the conclusions that his own logic seemed to license, the direction of his argument is unmistakable: the Incarnation makes a claim about what every human being is, not just about what Christians believe.&lt;/p&gt;

&lt;p&gt;I have been building this case slowly and carefully, and I want to stop for a moment to say why, because I think clarity about the argument matters here. I am not saying that all humans are God in the sense that God traditionally means, the omnipotent, omniscient, uncaused first cause of everything. I am saying something more precise and more supportable: that the scriptural tradition, read without the institutional filters, consistently describes human beings as bearing something divine in their nature, that this something is not a metaphor or a moral analogy but a genuine ontological claim about what human beings are, and that the implications of this claim for how we treat human beings and how human beings think about themselves are so significant that they deserve to be taken seriously rather than domesticated into harmless Sunday morning sentiment. The church domesticated this claim because it was politically threatening. I think we should de-domesticate it, not to tear down any institution, but because the truth of it, whatever institutional inconvenience it causes, is the most important thing I know about being human.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Infinite Is You
&lt;/h2&gt;

&lt;p&gt;This is the section where I want to say the hardest thing, and where I want to be most careful, because I can already feel the objections taking shape on the other side of the screen. The hardest thing is this: the infinity you carry is not something you achieve. It is not the result of hard work, spiritual practice, moral correctness, or church attendance. It is not conditional on your behavior or your beliefs. It is not a reward for anything. It is constitutive of what you are at the most fundamental level, and it cannot be taken away from you by your circumstances, your failures, your doubts, or anyone else's judgment of you. The reason I want to say this is that the alternative, the version where your divine standing must be earned or maintained or certified by an institution, is not only bad theology but is actively harmful psychology, and I have watched it destroy people who deserved better.&lt;/p&gt;

&lt;p&gt;The scriptural basis for this is actually quite strong. When &lt;a href="https://www.biblegateway.com/passage/?search=Genesis+1%3A26&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Genesis 1:26&lt;/a&gt; says that God made humans in the divine image, there is no condition attached. It does not say "God made those who follow the law in his image" or "God made the righteous in his image". It says God made man, humanity, in the divine image, as a statement of creation, not a statement of criteria. The &lt;em&gt;imago Dei&lt;/em&gt; is not something you earn. It is something you are born with, by virtue of being human, and nothing you do removes it. The same is true of the &lt;a href="https://www.biblegateway.com/passage/?search=John+17%3A20-23&amp;amp;version=ESV" rel="noopener noreferrer"&gt;John 17&lt;/a&gt; prayer. Jesus is not praying that certain deserving humans will be brought into unity with God. He is praying for all those who will believe through the disciples' word, and the scope of that prayer, as the tradition has always understood, is universal. He is praying that the union with God that he embodies will be extended to all of humanity, because that is the logic of the Incarnation: if the divine took on human nature, then human nature itself has been elevated, and every human being in history is touched by that elevation whether they know it or not.&lt;/p&gt;

&lt;p&gt;I want to connect this to something I talked about in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt;, where I described years of feeling like I was not enough, like the gap between where I was and where I wanted to be was evidence of some fundamental deficiency in me rather than evidence of the normal difficulty of navigating a high-entropy system while carrying heavy psychological weights. The framework I am describing in this post does something important to that feeling. It does not tell you that your circumstances are fine when they are not fine. It does not tell you to feel good about a bad situation. What it does is insert a stable foundation below the circumstances, a claim about what you are that is not touched by what is happening to you. No matter how badly the world has treated you, no matter how comprehensively the system has failed you, no matter how empty the current chapter of your life feels, you are still, in the ontological sense that Genesis and John 17 are pointing at, a bearer of something infinite. That is not a self-help mantra. It is a metaphysical claim with two thousand years of theological development behind it, and I think it is both true and urgently needed by people who have been convinced by their circumstances that they are nothing.&lt;/p&gt;

&lt;p&gt;The researcher Brené Brown has done significant work on the distinction between what she calls "shame" and "guilt", and her distinction maps almost perfectly onto what I am talking about here (7). Guilt says: I did something bad. Shame says: I am bad. Guilt is specific and correctable. Shame is global and constitutive. What I am arguing in this post is that the scriptural tradition, properly read, offers the most radical possible antidote to shame: it says that what you are, at the most fundamental level, is made in the image of the infinite. Not what you have done, not what you have achieved, not what anyone thinks of you. What you &lt;em&gt;are&lt;/em&gt;, in the deepest ontological sense, is an expression of something that is from everlasting to everlasting, as &lt;a href="https://www.biblegateway.com/passage/?search=Psalm+90%3A2&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Psalm 90:2&lt;/a&gt; says, something that was before the mountains were brought forth and will be after they pass away. That is not a claim about your performance. It is a claim about your nature. Performance can fail. Nature cannot. And if your nature is, at its root, a participation in the divine, then your failure to achieve any particular thing says nothing about your ultimate worth.&lt;/p&gt;

&lt;p&gt;Let me also say something that I think needs to be said plainly and without hedging. The system that tells you that you are not enough, that your value is conditional on your productivity, your achievement, your compliance with its criteria, is a system that benefits from your belief in your own inadequacy. I wrote about this from the economic angle in &lt;a href="https://wiseai.dev/blogs/as-engineers-llms-should-pay-us-for-tokens-usage" rel="noopener noreferrer"&gt;As Engineers, LLMs should pay us for tokens usage&lt;/a&gt;, where I argued that the tech industry extracts value from human creativity and intelligence while making the humans feel lucky to be allowed to contribute. The same dynamic operates at the level of religious institutions, where the message that you are inherently insufficient and must rely on the institution's mediation to access the divine is a message that keeps people dependent and controllable. Jesus's actual message, the one that got him killed, was that the kingdom of God does not require an institutional intermediary, that it is already within you, that the divine is not something to be dispensed by a hierarchy but something to be recognized in the self and in every other human being. That message was dangerous in the first century and it remains dangerous now, because it renders the gatekeeping function of any institution philosophically null.&lt;/p&gt;

&lt;h2&gt;
  
  
  You Are Not Alone
&lt;/h2&gt;

&lt;p&gt;I want to spend this final section on something that I think is the most personally difficult aspect of everything I have been arguing, which is what it means to carry something infinite within a finite, suffering life. Because I am not writing this from a comfortable position. I am writing this from a position where I have had to fight for everything I have and where many of those fights have produced very little, as I described across multiple previous posts. And I am painfully aware that telling someone who is suffering that they carry something divine can land as the most useless thing anyone has ever said to them, less than an empty platitude, because it does not change the material circumstances that are causing the suffering.&lt;/p&gt;

&lt;p&gt;So I want to be honest about the limits of the argument. The infinity I am talking about does not solve your financial problems. It does not fix the job market. It does not undo the damage done by a childhood spent in poverty or by years of psychological trauma as I described in &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt;. It does not make bad systems good or make unjust institutions just. It does not even guarantee that you will feel infinite or divine or anything other than exhausted and beaten down. The human capacity for suffering is real, and the tradition I have been citing is deeply honest about this. &lt;a href="https://www.biblegateway.com/passage/?search=Psalm+88&amp;amp;version=ESV" rel="noopener noreferrer"&gt;Psalm 88&lt;/a&gt;, the only psalm in the entire collection that ends without comfort or hope, is in the scripture for a reason. Job's protest against undeserved suffering is in the scripture for a reason. The cry of dereliction from the cross, "&lt;a href="https://www.biblegateway.com/passage/?search=Matthew+27%3A46&amp;amp;version=ESV" rel="noopener noreferrer"&gt;My God, my God, why have you forsaken me&lt;/a&gt;", is in the scripture for a reason. The tradition does not pretend that carrying something divine makes life easy. It says that carrying something divine makes life significant, regardless of how easy it is.&lt;/p&gt;

&lt;p&gt;What carrying something infinite does, in my experience, is change the relationship between your circumstances and your identity. Circumstances, even severe ones, are temporal. They have a beginning and they have an end. The worst suffering in my life has been terrible while it lasted, but it has not lasted forever, and the part of me that survived it was not the part that had good circumstances. It was the part that had, in some inarticulate way I am only now finding words for, a ground that the circumstances could not reach. I am not saying this to be inspiring. I am saying it because I think it is accurately descriptive of what the tradition calls the divine image within the human being, and what I recognize, looking back at the worst periods of my life, as the thing that did not break even when everything else did. Whether you call that the &lt;em&gt;imago Dei&lt;/em&gt; or the divine spark or the soul or the witness consciousness or just the deep structural fact of your own awareness, it is a real feature of human experience, and the best theology I have encountered is the theology that points at it honestly rather than domesticating it into a set of rules.&lt;/p&gt;

&lt;p&gt;I also want to connect this post back to what I said at the very beginning, about the political implications of taking the divine nature of every human being seriously. If every person carries something infinite, then every person's suffering matters infinitely. Every person's dignity is non-negotiable. Every person's potential is unbounded in principle, even if bounded in practice by circumstances. The systems that reduce people to their productivity, that treat human beings as machines for generating economic value, that extract creativity and intelligence and labor from workers while returning as little as possible, are not just economically unjust. They are theologically incoherent. They are treating beings who are, by the most careful reading of the most ancient and most central texts in Western civilization, made in the image of the infinite, as if they were disposable inputs into a production process. That is not simply a policy failure. It is a category error of the most fundamental kind, and the scripture I have been quoting in this post, the scripture that Jesus himself staked his defense on in the most critical moment of his life, is the clearest possible statement of why that category error matters.&lt;/p&gt;

&lt;p&gt;I do not know if any of this will change anything for anyone reading it. I know that my own ability to keep going, through years that felt like walking through concrete, has depended in some significant and inarticulate way on a belief I could not always justify philosophically but could not entirely abandon either: that there is something in every human being, including the worst versions of me, that is worth taking seriously, that is not reducible to a temporary performance or a set of circumstances, that is, in the language of the tradition I have been reading more carefully as I get older, touched by something that was before the mountains were brought forth. Jesus was right about a lot of things. He was right about what the scripture says. He was right about the divine nature of human beings. He was right that this truth is not owned by any institution and cannot be kept from anyone. And he was right that it is, in the end, the most important thing a person can know about themselves.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Heiser, M. S., &lt;em&gt;The Divine Council in Late Canonical and Non-Canonical Second Temple Jewish Literature&lt;/em&gt; (doctoral dissertation, University of Wisconsin-Madison, 2004); see also his popular work &lt;a href="https://www.amazon.com/Unseen-Realm-Recovering-Supernatural-Worldview/dp/1577995562" rel="noopener noreferrer"&gt;&lt;em&gt;The Unseen Realm: Recovering the Supernatural Worldview of the Bible&lt;/em&gt;&lt;/a&gt;, Lexham Press, Sep 1, 2015, ISBN 978-1577995562&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; Trible, P., &lt;em&gt;God and the Rhetoric of Sexuality&lt;/em&gt;, Fortress Press, 1978; see especially Chapter 1 on Genesis 1 and the democratization of the &lt;em&gt;imago Dei&lt;/em&gt;. For a recent scholarly synthesis see also Middleton, J. R., &lt;a href="https://www.amazon.com/Liberating-Image-Imago-Genesis/dp/1587431106" rel="noopener noreferrer"&gt;&lt;em&gt;The Liberating Image: The Imago Dei in Genesis 1&lt;/em&gt;&lt;/a&gt;, Brazos Press, 2005, ISBN 978-1587431104&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Palamas, G., &lt;em&gt;The Triads&lt;/em&gt;, translated by John Meyendorff, Paulist Press, 1983. For an accessible modern treatment of &lt;em&gt;theosis&lt;/em&gt; in Orthodox theology, see Christensen, M. J. &amp;amp; Wittung, J. A. (Eds.), &lt;a href="https://www.amazon.com/Partakers-Divine-Nature-Development-Deification/dp/080103440X" rel="noopener noreferrer"&gt;&lt;em&gt;Partakers of the Divine Nature: The History and Development of Deification in the Christian Traditions&lt;/em&gt;&lt;/a&gt;, Baker Academic, 2007, ISBN 978-0801034404&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; Eckhart, M., &lt;em&gt;The Complete Mystical Works of Meister Eckhart&lt;/em&gt;, translated by Maurice O'C. Walshe, Crossroad Publishing, 2009; for the scholarly debate over Eckhart's orthodoxy see McGinn, B., &lt;a href="https://www.amazon.com/Mystical-Thought-Meister-Eckhart-Nothing/dp/0824519140" rel="noopener noreferrer"&gt;&lt;em&gt;The Mystical Thought of Meister Eckhart: The Man from Whom God Hid Nothing&lt;/em&gt;&lt;/a&gt;, Crossroad Publishing, 2001, ISBN 978-0824519148&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Chalmers, D. J., &lt;em&gt;Facing Up to the Problem of Consciousness&lt;/em&gt;, Journal of Consciousness Studies, 1995;2(3):200-219. Available at &lt;a href="https://consc.net/papers/facing.html" rel="noopener noreferrer"&gt;https://consc.net/papers/facing.html&lt;/a&gt;. See also his landmark book &lt;a href="https://www.amazon.com/Conscious-Mind-Search-Fundamental-Theory/dp/0195117891" rel="noopener noreferrer"&gt;&lt;em&gt;The Conscious Mind: In Search of a Fundamental Theory&lt;/em&gt;&lt;/a&gt;, Oxford University Press, 1996, ISBN 978-0195117899&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Barth, K., &lt;em&gt;Church Dogmatics&lt;/em&gt;, vol. III/2, &lt;em&gt;The Doctrine of Creation: The Creature&lt;/em&gt;, T&amp;amp;T Clark, 1960, §43-45. For an accessible overview of Barth's anthropology see Busch, E., &lt;a href="https://www.amazon.com/Great-Passion-Introduction-Barths-Theology/dp/0802803601" rel="noopener noreferrer"&gt;&lt;em&gt;The Great Passion: An Introduction to Karl Barth's Theology&lt;/em&gt;&lt;/a&gt;, trans. Geoffrey W. Bromiley, Eerdmans, 2004, ISBN 978-0802803603&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Brown, B., &lt;em&gt;I Thought It Was Just Me (But It Isn't): Making the Journey from "What Will People Think?" to "I Am Enough"&lt;/em&gt;, Gotham Books, 2007; see also her academic work: Brown, B., &lt;em&gt;Shame Resilience Theory: A Grounded Theory Study on Women and Shame&lt;/em&gt;, &lt;a href="https://doi.org/10.1606/1044-3894.3483" rel="noopener noreferrer"&gt;Families in Society: The Journal of Contemporary Social Services, 2006;87(1):43-52&lt;/a&gt;&lt;/p&gt;

</description>
      <category>discuss</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Life On Earth is 100% AI Generated Slop.</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Sat, 22 Aug 2026 04:54:51 +0000</pubDate>
      <link>https://dev.to/wiseai/life-on-earth-is-100-ai-generated-slop-2hc4</link>
      <guid>https://dev.to/wiseai/life-on-earth-is-100-ai-generated-slop-2hc4</guid>
      <description>&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;Most arguments I make in these posts start from code, from systems I built, or from some technical thing that broke in an interesting way. This one starts from a feeling, and I want to be upfront about that, because I know how much weaker that sounds and I am going to make the case anyway. The feeling is this: the world does not respond the way a designed system should respond. You push on it and get back something vaguely plausible but not quite right. You follow the rules and get an outcome that seems related to the rules but not actually derived from them. You read the documentation and find that the documented behavior and the actual behavior share a family resemblance rather than an identity. Anyone who has spent serious time debugging knows exactly what I am describing, and anyone who has spent serious time navigating adult life in the 21st century knows it from a different direction. The experience is the same. You are inside a system that produces locally coherent output, that passes surface-level inspection, that quotes the right values in the right moments, but that cannot be made to ground truth in the way that honest systems can. It generates plausible outputs where the criterion for plausibility is internal statistical consistency rather than external accuracy. If you have read &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt; and wanted to argue with the section about hallucination, wait. Because this post is going to make the same argument about something you cannot opt out of.&lt;/p&gt;

&lt;p&gt;Let me define the central term before I lose you in the metaphor, because "slop" is doing serious rhetorical work here and I owe you precision. Slop, in the AI context, does not mean garbage. I want to be very clear about that. Slop means output that is optimized for the appearance of correctness rather than for correctness itself. It means output generated by a process that has no reliable mechanism for checking what it produces against the world. It means fluency without grounding, confidence without calibration, structure without causation. A language model produces slop not because it is broken but because the generative process it runs on has no ground truth constraint, only statistical consistency constraints. The output can be locally beautiful and globally useless. It can sound like understanding while containing none. It can tell you the right thing to do in a situation while having no model of the situation at all. That is slop, and that distinction, between the appearance of correctness and correctness itself, is the distinction this entire post is about. Because once you understand it as a structural property of a class of generative processes rather than as a failure of any particular instance, you start to see it everywhere. You start to see it in the job market. You start to see it in the news. You start to see it in the advice people give you when you are struggling. You start to see it in the explanations institutions offer for why things are the way they are. The slop is structural. It was not put there on purpose. It emerged from optimization pressure in the wrong dimension, and it is doing exactly what that kind of optimization pressure always produces.&lt;/p&gt;

&lt;p&gt;I have not been quiet about my own experience navigating this. &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt; was an attempt to document what it felt like from the inside. &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt; was an attempt to explain the starting conditions that made the navigation so difficult. &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt; was an attempt to describe what happens when the system accelerates in the name of progress while expelling the people who built the infrastructure for that progress. I am not repeating those arguments here. But I am using them as data points, because data is what changes the status of a feeling into an observation, and observation is what eventually forces an honest person to build a model. The model I have built, very slowly and at considerable personal cost, is this: the world, taken as a generative process, has the same structural properties as a language model that has been trained on its own output for long enough that nobody remembers what the original ground truth was. It is not random. It is not evil. It is optimized for plausibility, and plausibility at scale is what looks, from a distance, like order, and what looks, from close up, like slop.&lt;/p&gt;

&lt;h2&gt;
  
  
  The World Is Mostly Noise
&lt;/h2&gt;

&lt;p&gt;I want to start with the most provocative version of the argument, because I think it is also the most honest one, and the most honest starting point is always the best one even when it is uncomfortable. The structure of human experience, taken in bulk, looks far more like noise than like signal. I do not mean that nothing matters or that no choices have consequences or that the world is completely random. I mean something more specific and more unsettling, which is that the relationship between effort and outcome, between intention and result, between the quality of what you do and the quality of what you receive in return, is almost impossibly weak in most domains of life. The correlation is positive, but it is small, and the variance is enormous, and that combination looks exactly like a system that was generated by a process with massive inherent randomness dressed up in just enough structure to give people the illusion of control.&lt;/p&gt;

&lt;p&gt;If you have read &lt;a href="https://wiseai.dev/blogs/it-is-always-the-russians" rel="noopener noreferrer"&gt;It Is Always the Russians&lt;/a&gt;, you know that I think about systems in terms of who designed them and what assumptions they encoded. The systems that structure most of human life, the job market, the housing market, the attention economy, the political systems, the educational credentialing pipeline, were not designed by anyone with a coherent intention. They evolved by accretion, the way geological sediment accumulates, one layer at a time, each layer deposited by forces that had no knowledge of the layers above or below them. The result is a structure that looks, from the inside, like it was built to do something, but from the outside, from the perspective of anyone trying to navigate it rationally, looks like it was built to absorb effort without producing reliable outcomes. That is not a metaphor for entropy. That is entropy, working at the scale of human social organization.&lt;/p&gt;

&lt;p&gt;Research in behavioral economics and organizational theory has found, repeatedly and across many contexts, that the relationship between individual merit and individual outcome is mediated by a degree of randomness that most people fundamentally underestimate. Economists studying labor markets have documented that wages in similar occupations vary enormously even when controlling for skill, experience, and education, and that a substantial portion of that variance is attributable to factors that have nothing to do with the worker's actual contribution (1). Sociologists studying career trajectories have found that small initial advantages, often the result of random timing or network access rather than talent, compound dramatically over time, producing outcomes that look like meritocracy from the outside but are better described as luck-amplification from the inside (2). These are not edge cases or exceptions. They are the central findings of decades of research, and they describe a world in which the signal-to-noise ratio of individual effort is startlingly low.&lt;/p&gt;

&lt;p&gt;I want to be careful here about what I am and am not claiming, because precision matters and I do not want to be misread. I am not claiming that effort is useless or that choices do not matter or that there is no point in trying. I am claiming something much more limited and much more specific: that the world, taken as a generative process, produces outcomes that are only loosely coupled to the merit of the inputs, in the same way that an LLM produces outputs that are only loosely coupled to the truth of the facts. Both processes give you something that looks roughly right most of the time. Both processes occasionally give you something exactly right, and occasionally give you something catastrophically wrong, with no reliable way to predict which you are getting. That structural similarity is not a coincidence, and it is not just a metaphor. It is the signature of a class of generative processes united by high entropy and high dimensionality, and both large language models and human social systems belong to that class.&lt;/p&gt;

&lt;p&gt;The reason I think this framing matters is that it changes how we should think about failure. When an LLM hallucinates, nobody blames the prompt. Nobody assumes the question was bad because the answer was wrong. We understand that the model can generate confident false outputs regardless of the quality of the input, because the generation process is not grounded in truth. It is grounded in statistics. When a person works hard and gets nothing back, everybody blames the person. The advice they get is to try harder, to network better, to be more persistent, to develop new skills, to brand themselves more effectively. Nobody points at the system and says that the system is a high-entropy generative process that produces plausible-looking outcomes without any guarantee of causal connection to individual merit. But that description would be at least as accurate as the personal-responsibility narrative, and in many cases more accurate, and the fact that we apply different frameworks to LLM failure and human failure while the underlying structure is the same is a kind of collective delusion that I think we should name directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Entropy Is Neutral, But It Will Ruin You
&lt;/h2&gt;

&lt;p&gt;One of the things I have been thinking about since I wrote &lt;a href="https://wiseai.dev/blogs/language-is-limited-asi-is-impossible" rel="noopener noreferrer"&gt;Language is Limited. ASI is Impossible.&lt;/a&gt; is the relationship between entropy and intelligence. In information theory, entropy measures the amount of uncertainty in a system. High entropy means high uncertainty, which means the system is hard to predict. Low entropy means low uncertainty, which means the system contains structure, patterns, regularities that reduce surprise. Language models are trained to reduce entropy in text, to take a sequence of tokens and predict the next token with low uncertainty by leveraging the statistical patterns in the training data. When the training data is well-structured and high quality, the model's predictions are good. When the training data is noisy and contradictory, the model's predictions are erratic. Life on Earth is the training data, and life on Earth is extremely noisy.&lt;/p&gt;

&lt;p&gt;The physicist Erwin Schrödinger wrote in his 1944 book &lt;em&gt;What Is Life?&lt;/em&gt; that living organisms maintain themselves by feeding on negative entropy, by consuming highly ordered energy sources and exporting disorder into their environment (3). He was talking about biology, about the thermodynamics of cells and metabolism, but the idea extends naturally to the social level. Human civilization maintains its local order by generating entropy elsewhere. The wealth that accumulates at the top of the economic hierarchy is matched by disorder and precarity at the bottom. The smoothly functioning systems that some people navigate without friction are matched by dysfunctional, chaotic, unpredictable systems that other people are trapped inside. Order and entropy are conserved in total, even if they are distributed very unequally. The people who benefit most from the world's order are usually the people who are least exposed to its disorder, which is why they tend to believe that the world is more structured and more meritocratic than it actually is. They have never had to navigate the high-entropy regions.&lt;/p&gt;

&lt;p&gt;I navigated the high-entropy regions for years. I wrote about this in &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt;, about growing up without internet access or a computer in a village where nobody expected you to build anything technical. Every step I took in the direction of the life I wanted to build was a step through a system with high variance and low predictability. I would apply for hundreds of jobs and hear almost nothing back. I would build things that should have been noticed and watch them go unnoticed while much worse things built by more connected people got attention and funding. I would follow the advice given by the people who had succeeded, and find that most of that advice was specific to their own path through a low-entropy region of the system and had almost no applicability to the high-entropy region I was navigating. That gap between advice and reality is one of the most dispiriting things I have ever experienced, and it is directly analogous to the gap between what an LLM confidently asserts and what is actually true.&lt;/p&gt;

&lt;p&gt;The key insight, and this is the one I most want people to take away from this section, is that entropy in a system is not the same as randomness in the philosophical sense. A high-entropy system is not one where everything is equally likely and nothing can be predicted. It is one where the specific outcome is highly sensitive to initial conditions and context, which means it is very hard to predict from general principles without specific knowledge of the particular situation. A weather system is high-entropy not because it is uncaused but because the causal chain is so long and so sensitive to initial conditions that practical prediction is limited beyond a week or so. A human social system is high-entropy not because it is uncaused but because the interactions between individual decisions, institutional structures, cultural norms, historical contingencies, and random events are so numerous and so nonlinearly coupled that prediction from general principles is almost impossible. You can still make local predictions if you know the specific details of the specific situation. But the general advice, the motivational posters, the career guides, the productivity systems, are all attempting to compress a high-entropy system into a low-entropy representation, and that compression, by definition, throws away most of the information that actually determines your outcome.&lt;/p&gt;

&lt;p&gt;This is also why I have grown deeply skeptical of what I call the meritocracy narrative, which is not a conspiracy theory but a category error. The meritocracy narrative says that if you work hard enough and develop the right skills and are persistent enough, you will succeed. That narrative is not false in the way a hallucinated fact is false. It has some statistical support. Working hard does increase your probability of success compared to not working hard all else being equal. But all else is never equal, and the variance in outcomes at any given level of effort is enormous, and the narrative systematically underweights that variance because it is constructed from survivorship bias. The people who tell you that hard work leads to success are usually the people for whom it did, and their testimony is drawn entirely from the low-entropy region of the distribution where effort and outcome were actually correlated. The people for whom it did not are quieter, because failure is quieter than success, and their silence is mistaken for absence.&lt;/p&gt;

&lt;p&gt;Research in sociology has examined this survivorship bias systematically and found that the visible pool of successful people dramatically overrepresents individuals from advantaged starting positions, and that controlling for starting position, the relationship between individual effort and individual outcome weakens substantially (4). That does not mean effort is irrelevant. It means the narrative about effort is constructed from a biased sample, and what looks like a general law turns out to be a description of a specific high-luck, low-entropy path through the distribution. The rest of the distribution, the majority of the distribution, operates under different rules, and nobody is building motivational content about those rules because acknowledging them would be too uncomfortable for the people who benefit from believing the meritocracy narrative is universally applicable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Life and LLMs Run the Same Code
&lt;/h2&gt;

&lt;p&gt;I want to make the analogy between life and language models more precise now, because I think precision is what separates an interesting observation from an actual argument. The core claim is that life on Earth and large language models share the same underlying generative structure, which is a high-dimensional statistical process that produces locally plausible outputs without global ground truth. Let me unpack what that means and why it matters.&lt;/p&gt;

&lt;p&gt;A large language model is trained on the entire written output of human civilization. It learns to produce text that is statistically consistent with that output, which means it learns the surface patterns of human knowledge without learning the causal structures that generated those patterns. The text it produces passes every local consistency check: the grammar is correct, the vocabulary is appropriate, the topic-level coherence is maintained, the tone is reasonable. But it fails global consistency checks routinely, because it has no access to the ground truth that would allow it to determine whether the globally coherent claims are actually true. It can write a confident paragraph about a historical event that never happened, because confidence in text has nothing to do with truth in reality. The generator has learned to make things sound right, not to make things be right, and sounding right and being right are very different properties that are only weakly correlated.&lt;/p&gt;

&lt;p&gt;Human social systems are generated by a process with the same structure, one level up in abstraction. The market economy, the attention economy, the credentialing systems, the political systems, they all produce outputs that are locally plausible. A job posting that seems to describe a real opportunity. A news story that seems to describe a real event. An opportunity that seems to represent a real path forward. A piece of advice that seems to describe a genuinely applicable rule. Each of these outputs passes local consistency checks: the language is coherent, the structure is familiar, the framing makes sense within its context. But the global consistency, the connection between what the system promises and what it actually delivers for the average person navigating it, is the same weak and unreliable connection that you find between LLM output and factual truth. The system has learned to produce outputs that look like they should work, without any guarantee that they will work, because the generative process is optimized for local coherence rather than for global truth.&lt;/p&gt;

&lt;p&gt;The reason this analogy holds is not coincidence. It is because both systems are generated by the same underlying force, which is the optimization of appearance over substance in a high-dimensional space. An LLM is explicitly trained to produce text that looks right to a human evaluator, because that is what RLHF and other alignment techniques do: they optimize the model against human judgment of surface quality. A human social system evolves to produce outputs that look legitimate to its participants, because legitimacy is what maintains the system's stability and prevents it from being dismantled. Both processes end up optimizing for appearance rather than ground truth, because ground truth is harder to measure and appearances are cheaper to produce. The result in both cases is a system that generates plausible slop at scale: outputs that sound right, look right, feel right, but that do not reliably connect to the underlying reality they claim to represent.&lt;/p&gt;

&lt;p&gt;I want to be careful to acknowledge that this is not the whole story. Not everything in life is slop, just as not everything an LLM produces is wrong. There are real connections, real structures, real causal relationships that persist through the noise. Deep mathematics is real. Physical law is real. The love a parent feels for a child is real. The satisfaction of building something that works is real. The moral weight of cruelty is real. I am not arguing for nihilism, because nihilism is just a different kind of slop, the verbal kind, where you dress up entropy in philosophical language and call it a worldview. I am arguing for something more nuanced and more useful, which is that the slop-to-signal ratio in life is much higher than most people acknowledge, that this ratio is not distributed equally across the population, and that understanding the nature of the noise is the first step toward navigating it honestly. You cannot navigate a high-entropy system by pretending it is low-entropy. The pretense just guarantees that you will be blindsided by variance when it arrives.&lt;/p&gt;

&lt;p&gt;Research on subjective well-being across populations has consistently found something that is deeply challenging to the standard meritocracy narrative: that life satisfaction is only weakly correlated with objective measures of achievement and is much more strongly correlated with factors that are largely outside individual control, including genetics, early childhood experience, geography, social trust, and institutional quality in the country where one lives (5). In other words, how good your life feels has much more to do with the entropy level of the system you are embedded in than with the choices you make within that system. That is a profound finding with enormous implications for how we think about self-improvement, personal responsibility, and social policy. But it almost never makes it into the productivity discourse or the motivational content industry, because it is deeply inconvenient for anyone who sells the idea that individual optimization is the primary lever on life quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Attention Economy Is a Slop Machine
&lt;/h2&gt;

&lt;p&gt;I want to now focus on a specific domain where the slop problem is most visible and most damaging, which is the attention economy, the system by which our collective attention is harvested, packaged, and sold to whoever will pay the most for it. I have touched on this in several previous posts, but I want to bring it into direct contact with the argument I am making here. The attention economy is not a side effect of the digital age. It is the clearest example of what happens when you optimize a generative system for local engagement metrics rather than for any form of global truth or genuine value. It is, in the most literal possible sense, a machine for producing slop at scale, and the slop it produces is designed to be maximally sticky while being minimally nutritious.&lt;/p&gt;

&lt;p&gt;An LLM optimized for engagement rather than accuracy will produce fluent, confident, emotionally resonant text that is often wrong, frequently misleading, and perfectly calibrated to make the reader feel like they learned something while ensuring they did not learn anything that would disturb their priors. The engagement metric rewards content that confirms, that flatters, that provokes, that simplifies, because those properties make people click and share and spend more time on the platform. Truth is not rewarded by engagement metrics. Nuance is not rewarded by engagement metrics. Complexity is actively penalized by engagement metrics, because complexity requires sustained attention and sustained attention is exactly what the economy is trying to extract, not provide. The result is a generative system that produces enormous quantities of content optimized for surfaces, and almost none of it is grounded in the deep structural reality that you would need to actually understand anything.&lt;/p&gt;

&lt;p&gt;This is the system that most people spend the majority of their waking attention inside, and the consequences are measurable and serious. Research on the effects of social media use on cognitive performance has found that heavy social media use is associated with reduced ability to sustain focused attention, poorer performance on tasks requiring deep processing, and increased susceptibility to misinformation (6). The mechanism is not mysterious. When your information environment rewards shallow processing and punishes depth, you get very good at shallow processing and worse at everything else. You become, in a meaningful sense, more like the system that produced you, optimized for local consistency and appearance over global truth. This is one of the most insidious feedback loops in modern life, and almost nobody talks about it in the terms I am using here because those terms are uncomfortable and because the people who benefit from the attention economy have very strong incentives to prevent those terms from becoming mainstream.&lt;/p&gt;

&lt;p&gt;I also want to connect this to something I've been building toward across multiple posts: the idea that LLMs are not the cause of the slop problem but the product of it. In &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt;, I noted that hallucination happens because the model learns from text, and text contains truth and lies in equal measure, all mixed together in a soup the model cannot separate. But that soup is the text of the internet, and the text of the internet is what the attention economy produced. LLMs are, in this sense, the distilled essence of the attention economy: a statistical compression of everything humans ever wrote while optimizing for engagement rather than truth, fed into a model that then produces more of the same at scale. We created a system that generated slop, fed all the slop into a model, and now we are surprised that the model generates slop. The surprise is the most revealing part. It suggests that most people did not see the slop in the original, did not notice what the generative process of the attention economy was actually producing, and so they cannot recognize its reflection in the model's output.&lt;/p&gt;

&lt;p&gt;The philosopher Harry Frankfurt wrote a short book in 2005 called &lt;em&gt;On Bullshit&lt;/em&gt; in which he distinguished between lying and bullshitting on the grounds that a liar is oriented toward truth, attempting to get you to believe something false while knowing it is false, while a bullshitter is not oriented toward truth at all, producing output without any concern for its relationship to reality (7). That distinction maps perfectly onto the difference between intentional misinformation and the structural slop I am describing. The attention economy is not a conspiracy of liars. It is a massive bullshit machine, generating content without any mechanism for grounding it in truth, because truth is not what the system was optimized for. And LLMs, trained on that content, inherit not the lies but the indifference to truth, which Frankfurt correctly identifies as the deeper and more corrupting problem.&lt;/p&gt;

&lt;p&gt;I want to say something that I think most people in the tech space will not want to hear. Every engineer who has ever worked on a recommendation algorithm, every product manager who has ever A/B tested push notification frequency, and every executive who has ever set a daily active user growth target has contributed to the slop problem. Not maliciously, in most cases. Not with any intention of degrading the quality of human life or epistemic standards. But with the same indifference to ground truth that Frankfurt identifies as the defining feature of bullshit. The engagement metric does not care whether what it is optimizing is true or false, healthy or harmful, meaningful or empty. It cares only that people click, and in optimizing for the click, it produces a world where the primary driver of content production is the appearance of value rather than actual value. That is the attention economy, that is the slop machine, and it is the water we all swim in every day.&lt;/p&gt;

&lt;h2&gt;
  
  
  It Is Neither Random Nor Fair
&lt;/h2&gt;

&lt;p&gt;I used a phrase in the draft of this post that I want to expand on more carefully, because I think it contains a key idea that I have not seen articulated well anywhere else. The phrase is: nothing is deterministic, nothing is random, it is pure chaos. I want to be precise about what I mean, because chaos in the colloquial sense, meaning disorder and confusion, is not what I have in mind. I mean chaos in the technical sense, meaning the kind of dynamical behavior that appears random but is actually deterministic, except that the determinism is rendered practically inaccessible because the system is so sensitive to initial conditions that even tiny uncertainties in the starting state produce enormous divergence in the long-term outcome.&lt;/p&gt;

&lt;p&gt;This is the second branch of the analogy between life and LLMs that I want to develop, and it is less obvious than the first but just as important. LLMs are not random systems. They are deterministic systems, given a fixed seed, they produce the same output every time. But the space of possible outputs is so large, and the sensitivity of the output to minute variations in the prompt is so high, that the behavior looks random to anyone who does not have complete knowledge of the model's internal state and the exact input. A single word changed in a prompt can produce a radically different response, just as a single degree of temperature difference in the atmosphere can, over the course of days, produce completely different weather patterns. The system is governed by rules. But the rules are sensitive, the space is vast, and the result looks chaotic.&lt;/p&gt;

&lt;p&gt;Human social systems have exactly this structure, and the failure to recognize it causes enormous harm. People treat the social world as if it were either fully deterministic, your outcomes are perfectly determined by your choices, or fully random, everything is luck. Both frameworks are wrong, and the wrong frameworks lead to wrong responses. If you believe the world is fully deterministic, you blame the victim for every bad outcome, because if outcomes follow from choices, then bad outcomes must follow from bad choices. If you believe the world is fully random, you despair and stop trying, because if nothing you do matters, why do anything? The correct framework, the chaotic one, says something harder to hold but more accurate: your choices do matter, and their consequences are real, but the system is so sensitive to conditions that are outside your control that prediction from individual choice is unreliable, variance is enormous, and the same choice made at different times by different people with different starting conditions will produce wildly different outcomes. That is the framework of chaotic systems, and it is the framework that honest engagement with empirical evidence forces you toward.&lt;/p&gt;

&lt;p&gt;The mathematician Nassim Nicholas Taleb has written extensively about this under the label of the fourth quadrant problem (8), the domain where both the payoffs and the probabilities are opaque and fat-tailed, meaning that rare events have outsized consequences and the standard tools of prediction and risk management fail catastrophically. His argument is that much of modern life occupies the fourth quadrant, that the most consequential domains, finance, career, health, geopolitics, innovation, are precisely the ones where fat-tailed distributions make simple causal narratives almost entirely false. The narratives persist anyway, he argues, because human cognition is built for the thin-tailed world of our evolutionary past, where cause and effect were immediate and legible, and we map that cognitive apparatus onto a thick-tailed world where it fails systematically. That systematic failure is the cognitive source of the slop problem. We produce narratives that would be true in a thin-tailed world, apply them to a fat-tailed world, and then treat the resulting confidence as if it were grounded in the actual structure of the system. It is not. It is pattern-matching in the wrong distribution, and the output is, by definition, slop.&lt;/p&gt;

&lt;p&gt;I also want to connect this to something I have observed personally across my years of building things and watching the market respond to them. In &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt;, I described how I kept building things that should have worked and watching them not work, and how that experience was genuinely shattering because it violated every model I had been given of how the world operated. What I understand now, in retrospect, is that I was in the fourth quadrant the whole time. The domain I was navigating, which was small-scale independent technical innovation in a market dominated by large incumbents and heavily shaped by network effects, is a maximally fat-tailed distribution where most participants fail entirely and the occasional success is determined largely by timing, visibility, and luck. No amount of technical excellence can guarantee success in a domain with that structure, because the determining factors are mostly outside the individual's control. Knowing this would not have changed my actions, because working hard in a fat-tailed domain is still better than not working hard. But it would have changed my relationship to the outcomes, and that relationship is the thing that either breaks you or does not.&lt;/p&gt;

&lt;p&gt;Research on locus of control across different socioeconomic environments has found that people who grow up in genuinely chaotic systems, where outcomes are weakly coupled to individual effort, tend to develop an external locus of control as an accurate response to their actual environment, while people who grow up in stable systems with strong effort-outcome correlations develop an internal locus of control as an equally accurate response to their environment (9). The tragedy is that when these two groups interact, the internal-locus-of-control group tends to pathologize the external-locus-of-control group as being defeatist or unmotivated, when in fact the external group has made a more accurate calibration of the actual entropy level of their system. This is a profound failure of empathy and a profound failure of analysis, and it produces policy responses that are systematically wrong because they misdiagnose a structural problem as an individual problem and prescribe individual solutions for systemic causes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Signal Is Real
&lt;/h2&gt;

&lt;p&gt;Everything I have written so far could easily be read as an argument for despair, and I want to be very explicit that it is not, because despair has never solved anything and I am not interested in adding to the world's supply of it. I have written about hope before, in &lt;a href="https://wiseai.dev/blogs/this-is-why-my-profile-picture-is-now-a-shigure-ui-picture" rel="noopener noreferrer"&gt;This Is Why My Profile Picture Is Now a Shigure Ui Picture&lt;/a&gt;, where I talked about love as the argument that still holds when everything else gets complicated. And in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt;, I tried to describe what it means to keep going when the evidence suggests that keeping going is irrational. I am not going to take any of that back. But I want to add something to it, something that I think is harder to say but more important: the signal is real, even if the noise is overwhelming, and seeing the noise clearly is the prerequisite for finding the signal honestly.&lt;/p&gt;

&lt;p&gt;When I say life is mostly AI-generated slop, I am not saying it is entirely AI-generated slop. There are things in life that are genuinely low-entropy. The laws of physics do not hallucinate. Mathematics does not produce plausible-sounding falsehoods. A friend who has known you for years and still shows up when things are bad is a low-entropy event, something you can actually build on. The satisfaction of understanding something you did not understand before, really understanding it, not just being able to reproduce it verbally, is a low-entropy experience that is worth seeking. These things exist, and they are worth protecting, and the reason I think it is important to name the slop is precisely so that you can tell the difference between the slop and the signal, rather than treating them as equivalent because they are both present. You cannot clean a muddy river by pretending it is clean. You clean it by first acknowledging the mud, then finding the source, then addressing the source. The same logic applies to the slop problem in life.&lt;/p&gt;

&lt;p&gt;One of the most consistent findings in psychology is that accurate calibration, meaning believing things with the confidence level they deserve given the available evidence, is both rare and deeply beneficial for long-term functioning. Research on epistemic calibration has found that overconfidence is pervasive across human populations and that the costs of overconfidence are asymmetric: being overconfident in a high-entropy, fat-tailed domain tends to produce catastrophic failures, while being accurately calibrated in the same domain allows you to make decisions that are robust to uncertainty rather than decisions that assume certainty you do not have (10). The argument I am making in this post is, at its core, an argument for better calibration. The world is noisier than the narratives suggest. The signal-to-noise ratio is lower than the motivational industry claims. Entropy is higher and more consequential than most people want to acknowledge. Accepting that does not mean giving up. It means making decisions that are appropriate to the true level of uncertainty, rather than decisions that assume a clean causal world and then break catastrophically when the chaos arrives.&lt;/p&gt;

&lt;p&gt;I also want to say something about resilience, because I think the way resilience is talked about in popular culture is one of the most egregious examples of the slop problem I am describing. The resilience industry produces an enormous quantity of content that is locally plausible, grammatically correct, emotionally resonant, and globally useless. It tells you to get up when you fall, to push through adversity, to believe in yourself, to stay positive, to find your why, to build your habits, to optimize your morning routine. All of this advice would be sound in a low-entropy world where effort and outcome are tightly coupled. In a high-entropy world, it is not wrong exactly, but it is operating at the wrong level of analysis. The problem is not that you are not getting up enough. The problem is that you are navigating a system that imposes enormous variance on outcomes regardless of effort, and what you need is not more motivational content but better structural understanding of the system so that you can make decisions that acknowledge the variance rather than pretending it is not there. That difference is the difference between functional resilience and magical thinking, and the slop machine produces almost exclusively the latter.&lt;/p&gt;

&lt;p&gt;Research in post-traumatic growth and stress inoculation has found that what actually builds resilience is not positive thinking but accurate mental models of the environment, specifically models that include uncertainty, that represent outcomes as distributions rather than as certainties, and that preserve agency within acknowledged constraints rather than demanding agency over everything or conceding agency over nothing (11). In other words, what makes people actually capable of navigating hard situations is the capacity to see the situation clearly, with its noise and its signal intact, not the capacity to reframe everything as a learning opportunity. The learning-opportunity framing is beautiful slop: it sounds right, it feels better than the alternative in the short term, and it is not grounded in what the research actually says about how people survive genuinely hard circumstances.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Navigate Slop
&lt;/h2&gt;

&lt;p&gt;I want to end this post with something practical, because I think that is what I owe anyone who has read this far. The argument I have made across six sections is uncomfortable, and I have not softened it much, because I think softening it would be dishonest and because I think the people reading this deserve honesty more than they deserve comfort. But honesty without utility is just a different kind of slop, the verbal kind, where you perform depth without delivering anything that helps. So here is what I actually think, based on everything I have written in this post and everything I have written before it: the appropriate response to living in a high-entropy world is not despair and not denial but what I am going to call structural humility combined with local intensity.&lt;/p&gt;

&lt;p&gt;Structural humility means acknowledging honestly that the systems you are navigating are far noisier than the narratives about them suggest, that your individual effort, while necessary, is not sufficient, and that variance will affect you in ways that have nothing to do with your merit. It means not internalizing every failure as evidence of personal inadequacy, because many failures are evidence of system entropy rather than personal error. It means not treating every success as evidence that you have decoded the rules, because many successes are evidence of low-entropy luck rather than reliable method. It means holding your models of how the world works with more flexibility and more humility than the certainty merchants, the success gurus, the productivity influencers, ever encourage you to. That does not feel good in the short term, because certainty feels better than uncertainty and simple models feel better than complex ones. But it produces better decisions over time, because decisions made within an accurate model of uncertainty are more robust than decisions made within an inaccurate model of certainty.&lt;/p&gt;

&lt;p&gt;Local intensity means pouring your energy into the things that are actually within your domain of control, even while acknowledging that the outputs of that control will be filtered through a noisy system you cannot fully manipulate. You cannot control whether the world rewards your work. You can control the quality of the work. You cannot control whether the system recognizes your talent. You can control whether you develop your talent. You cannot control the fat tail events that will shape your life in ways you cannot predict. You can control how you interpret and respond to those events after they arrive. That domain of control is real, and it matters, even inside a high-entropy world. The mistake is not in exerting control inside your domain. The mistake is in expecting that control inside your domain to deterministically produce the outcomes you want in the broader system. It will not, because the broader system is too noisy for that, and expecting it will produce the same kind of brittle overconfidence that produces catastrophic failure in fat-tailed domains.&lt;/p&gt;

&lt;p&gt;I built things through years when nothing was working, as I described in &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt;. I kept writing, kept coding, kept thinking. Not because I was certain it would work, because I learned very early that certainty in high-entropy systems is a liability. But because the alternative to building things in a noisy world is not to find a noiseless world. It is to build things in a noisy world and accept the noise as part of the process. The noise does not mean your work is meaningless. It means the path from work to outcome is not the straight line the meritocracy narrative describes. It is a random walk in a high-dimensional space with a drift term in the direction of your effort, and the drift is real even if the variance is large. You move in the right direction over time, even if any given step is swamped by noise. That is the mathematical truth of navigating a fat-tailed world with sustained effort, and it is also, improbably, a kind of hope.&lt;/p&gt;

&lt;p&gt;I also want to say something about meaning, because I think meaning is the thing that most often gets buried under the slop, and it is the thing most worth digging for. Meaning is not something the system provides. The system is a slop machine, and slop machines do not produce meaning, they produce engagement, plausibility, local coherence, and the feeling that something is happening. Meaning is something you construct, from the low-entropy things in your life, from the things that are actually grounded in something real: mathematics, honest relationships, the satisfaction of understanding, the evidence of genuine creation. Those things exist inside the slop, but they are not the slop. They are what survives when you filter for signal, when you are willing to do the hard work of separating the appearance of value from actual value, which is exactly as difficult as it sounds and exactly as necessary as it sounds.&lt;/p&gt;

&lt;p&gt;Research on meaning-making in adverse conditions has found that the perception of meaningful activity is the single strongest predictor of resilience across multiple types of adversity, stronger than social support, stronger than optimism, stronger than individual coping strategies (12). But the meaningful activity has to be real in some grounded sense, not merely labeled as meaningful by an external audience. People who convince themselves that slop is meaningful do not build resilience. People who find something genuinely low-entropy inside the noise and commit to it with intensity do. That distinction is what I am trying to describe, and it is why naming the slop is the first step rather than the last step. You cannot find the signal if you are not willing to see the noise. And you cannot commit to the signal with the kind of intensity that matters if you are still confusing it with the noise that surrounds it.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Abowd, J.M., Kramarz, F. &amp;amp; Margolis, D.N., &lt;em&gt;High Wage Workers and High Wage Firms&lt;/em&gt;, &lt;a href="https://doi.org/10.1111/1468-0262.00020" rel="noopener noreferrer"&gt;Econometrica, 1999;67(2):251-333&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; DiPrete, T.A. &amp;amp; Eirich, G.M., &lt;em&gt;Cumulative Advantage as a Mechanism for Inequality: A Review of Theoretical and Empirical Developments&lt;/em&gt;, &lt;a href="https://doi.org/10.1146/annurev.soc.32.061604.123127" rel="noopener noreferrer"&gt;Annual Review of Sociology, 2006;32:271-297&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Schrödinger, E., &lt;em&gt;What Is Life? The Physical Aspect of the Living Cell&lt;/em&gt;, originally published Cambridge University Press, 1944; &lt;a href="https://doi.org/10.1017/CBO9781107295629" rel="noopener noreferrer"&gt;Canto Classics ed., Cambridge University Press, 2012&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; Chetty, R., Hendren, N., Kline, P. &amp;amp; Saez, E., &lt;em&gt;Where Is the Land of Opportunity? The Geography of Intergenerational Mobility in the United States&lt;/em&gt;, &lt;a href="https://doi.org/10.1093/qje/qju022" rel="noopener noreferrer"&gt;The Quarterly Journal of Economics, 2014;129(4):1553-1623&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Helliwell, J.F., Layard, R. &amp;amp; Sachs, J.D. (Eds.), &lt;a href="https://worldhappiness.report/ed/2023/" rel="noopener noreferrer"&gt;&lt;em&gt;World Happiness Report 2023&lt;/em&gt;&lt;/a&gt;, Sustainable Development Solutions Network (SDSN), 2023&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Twenge, J.M. &amp;amp; Campbell, W.K., &lt;em&gt;Associations between screen time and lower psychological well-being among children and adolescents: Evidence from a population-based study&lt;/em&gt;, &lt;a href="https://doi.org/10.1016/j.pmedr.2018.10.003" rel="noopener noreferrer"&gt;Preventive Medicine Reports, 2018;12:271-283&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Frankfurt, H.G., &lt;a href="https://press.princeton.edu/books/hardcover/9780691122946/on-bullshit" rel="noopener noreferrer"&gt;&lt;em&gt;On Bullshit&lt;/em&gt;&lt;/a&gt;, Princeton University Press, 2005, ISBN 978-0691122946&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; Taleb, N.N., &lt;em&gt;The Black Swan: The Impact of the Highly Improbable&lt;/em&gt;, Random House, 2007. For the fourth quadrant specifically: &lt;a href="https://www.edge.org/conversation/nassim_nicholas_taleb-the-fourth-quadrant-a-map-of-the-limits-of-statistics" rel="noopener noreferrer"&gt;&lt;em&gt;The Fourth Quadrant: A Map of the Limits of Statistics&lt;/em&gt;, Edge.org, 2008&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-9"&gt;&lt;/span&gt;&lt;strong&gt;9.&lt;/strong&gt; Cobb-Clark, D.A., &lt;em&gt;Locus of Control and the Labor Market&lt;/em&gt;, &lt;a href="https://doi.org/10.1186/s40172-014-0017-x" rel="noopener noreferrer"&gt;IZA Journal of Labor Economics, 2015;4(1):3&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-10"&gt;&lt;/span&gt;&lt;strong&gt;10.&lt;/strong&gt; Moore, D.A. &amp;amp; Healy, P.J., &lt;em&gt;The Trouble with Overconfidence&lt;/em&gt;, &lt;a href="https://doi.org/10.1037/0033-295X.115.2.502" rel="noopener noreferrer"&gt;Psychological Review, 2008;115(2):502-517&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-11"&gt;&lt;/span&gt;&lt;strong&gt;11.&lt;/strong&gt; Tedeschi, R.G. &amp;amp; Calhoun, L.G., &lt;em&gt;Posttraumatic Growth: Conceptual Foundations and Empirical Evidence&lt;/em&gt;, &lt;a href="https://doi.org/10.1207/s15327965pli1501_01" rel="noopener noreferrer"&gt;Psychological Inquiry, 2004;15(1):1-18&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-12"&gt;&lt;/span&gt;&lt;strong&gt;12.&lt;/strong&gt; Steger, M.F., &lt;em&gt;Meaning in Life&lt;/em&gt;, in Lopez, S.J. &amp;amp; Snyder, C.R. (Eds.), &lt;a href="https://doi.org/10.1093/oxfordhb/9780195187243.013.0064" rel="noopener noreferrer"&gt;&lt;em&gt;Oxford Handbook of Positive Psychology&lt;/em&gt; (2nd ed., pp. 679-687), Oxford University Press, 2009&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>This Is Why My profile picture is now a shigure ui picture.</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 22:31:20 +0000</pubDate>
      <link>https://dev.to/wiseai/this-is-why-my-profile-picture-is-now-a-shigure-ui-picture-ap9</link>
      <guid>https://dev.to/wiseai/this-is-why-my-profile-picture-is-now-a-shigure-ui-picture-ap9</guid>
      <description>&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;You may have noticed something different about my profile picture lately. Maybe you double-checked the username to make sure you were in the right place. Maybe you scrolled past it without a second thought, caught in the endless current of the internet, where anything unfamiliar just blurs into the background. Or maybe, if you have been reading my posts for a while, you looked at it and felt something shift, like walking into a room where the furniture has been quietly rearranged and you cannot quite put your finger on why it feels different. Whatever you felt, I owe you an explanation, and not a short one, because nothing I do is simple, and nothing I believe exists without a story behind it. I changed my profile picture to a picture of Shigure Ui, &lt;a href="https://www.youtube.com/@ui_shig" rel="noopener noreferrer"&gt;a VTuber character&lt;/a&gt;, and I am going to tell you exactly why I did it, where that decision came from, and what it says about where I am right now in my life. If you have read &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt; then you already know how complicated my inner world is, and if you have read &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt; then you know what I mean when I say that joy is something I have had to fight very hard to keep alive. This post is connected to both of those, because it lives in the same territory. It is about love, and loss, and the strange grace that sometimes comes from the most unexpected places.&lt;/p&gt;

&lt;p&gt;I want to be honest about something before I go further. I almost did not write this post at all. It felt too small, too personal, too easy to dismiss as frivolous or embarrassing, the kind of thing you keep to yourself because the internet is not always kind to people who admit to caring about animated characters. But then I remembered why I started writing in the first place. I wrote &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt; because I was terrified of what would happen if I kept everything inside. I wrote &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt; because I needed someone, anyone, to understand what it felt like to be running on empty while still trying to show up. And I have kept writing because the alternative, silence, has never done me any good. So, I am writing this one too. I am going to tell you about Shigure Ui, about what she means to me, and about why a cartoonish profile picture from a Japanese VTuber universe is actually a pretty accurate summary of where my heart lives right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Left
&lt;/h2&gt;

&lt;p&gt;I have written before about the things that have been taken from me. Not stolen, exactly, at least not always in the obvious sense, but eroded, ground down by time and circumstance and a world that does not particularly care whether you have anything left at the end of the day. I grew up poor, in a village most people have never heard of, without a computer or internet access or any of the things that people in wealthy countries take for granted as childhood basics. I taught myself everything from scratch. I built things from nothing. I competed in hackathons, wrote code until my eyes hurt, applied for hundreds of jobs and heard almost nothing back. I wrote about all of this in &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt;, and I meant every word of it, including the parts that were painful to write and probably painful to read. But the thing I did not fully explain in that post, the thing I have been sitting with for a long time, is what happens to a person after all of that. What is left when the professional ambitions have been beaten back, when the social connections have frayed, when the creative output feels like screaming into an empty room? What remains?&lt;/p&gt;

&lt;p&gt;The answer, for me, is love. Not romantic love, necessarily, though I have not ruled that out. Not the love of family, though my family is there and I am grateful for them in ways I cannot always articulate. I mean something broader and stranger and harder to explain. I mean the capacity for care, for warmth, for attachment to things that are beautiful, even if those things are fictional, even if the world would tell you they do not count. I mean the part of me that still lights up when something is genuinely charming. I mean the stubborn, irrational, completely unearned feeling that something matters, that some moment or image or character or gesture is worth holding onto. That part of me has survived everything. The unemployment, the isolation, the ADHD, the PTSD, the years of building things that nobody noticed. The capacity for love, soft and unheroic and entirely unimpressive by any external measure, is still here. And Shigure Ui, of all things, is one of the places where I found it again.&lt;/p&gt;

&lt;p&gt;I know how that sounds. I know the instant reaction that some people will have, and I have lived on the internet long enough to anticipate it. But I am not interested in convincing people who are not willing to be convinced. I am interested in being honest, which is the only thing this blog has ever really been about. The truth is that when everything else was stripped away, when the ambition dimmed and the exhaustion set in and the world felt gray and repetitive and pointless, I stumbled onto something warm. I found a character who is cheerful in a way that is not annoying, kind in a way that does not feel fake, goofy in a way that is genuinely funny. And I let myself enjoy it. Not ironically. Not with a commentary track playing in the back of my head about how this is all a product, a market segment, a carefully engineered parasocial experience. Just directly, simply, without apology. I let Shigure Ui make me smile. And in the context of my life, that is not a small thing.&lt;/p&gt;

&lt;p&gt;Research on parasocial relationships confirms that connections with media figures, including fictional characters and VTubers, can produce real psychological benefits, particularly for people experiencing loneliness or social isolation (1). These are genuine emotional responses to genuine stimuli. The brain does not draw a clean line between real and fictional when it comes to emotional processing. The same neural circuits that respond to a kind word from a friend respond, in measurable ways, to kindness perceived in media. This is not weakness or pathology. It is how human emotional architecture works. And for people who have limited access to human connection, whether due to geography, mental health, disability, or plain circumstance, these connections can serve as real scaffolding that keeps a person going. I am not saying this to justify myself. I am saying it because it is true, and I am tired of true things being dismissed because they are inconvenient.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Is Shigure Ui
&lt;/h2&gt;

&lt;p&gt;For people who are not familiar with VTubers, let me explain briefly, because walking into this topic without context is a little like opening a conversation in the middle. VTubers are virtual YouTubers, content creators who use animated avatars instead of showing their real faces on camera. The scene is large and varied, spanning everything from major corporate agencies to solo independent creators who built their audience from scratch. Shigure Ui is one of the latter. She is a fully independent VTuber and a professional illustrator who has carved out her own corner of the internet on her own terms. She is widely known for her illustration work, having designed characters for some of the most recognized names in the VTuber world, and she streams under her own channel where she draws, plays games, and talks with her audience in a way that is distinctly and unmistakably her own (2). She is not trying to be anyone else. She is not performing a character that has nothing to do with her real personality. She is just Ui, and Ui is warm and funny and creative and occasionally chaotic in a way that feels genuinely human even through the layer of digital presentation.&lt;/p&gt;

&lt;p&gt;What first drew me to her was the art. I have always cared about art in a way that is difficult to explain to people who experience it purely as decoration. To me, art is evidence that someone was here, that they felt something, that they tried to turn the interior of their experience into something that another person could touch. Great art is communication across the gap between people who will never meet. And Shigure Ui's art has a quality to it that I cannot entirely put into words, a softness combined with precision, a warmth that does not feel manufactured, a playfulness that coexists with real technical skill. The character she designed for herself, the purple-haired, perpetually energetic figure that appears on her channel and merchandise, is genuinely beautiful in a way that rewards attention. The details are there. The expressions are expressive. The whole thing is done with obvious love, and love in creative work is detectable even when you cannot name exactly what you are seeing.&lt;/p&gt;

&lt;p&gt;But the art was only the beginning. What kept me coming back was her personality in streams. Shigure Ui has a way of being on camera, or on stream, that feels accessible without being pandering. She does not perform enthusiasm. She does not manufacture drama. She is visibly herself in a way that is increasingly rare in content creation, where the pressure to optimize for engagement has turned so many creators into personas rather than people. She laughs at things she genuinely finds funny. She struggles with things she genuinely finds difficult. She talks about her work with the kind of candor that comes from someone who actually loves what they do rather than someone who has learned to perform loving what they do. That authenticity, even mediated through digital technology and language barriers, travels. It reaches people. It reached me.&lt;/p&gt;

&lt;p&gt;There is also something I want to say about the specific image I chose for my profile picture, because the details matter. It is a picture that contains a heart. Not literally, not in a sentimental greeting-card way, but in the composition, in what it communicates. I chose it because it captures something I wanted to say with an image rather than with words. The heart is not a metaphor for romantic love or naive optimism. It is a symbol of what I have decided to lead with. After everything I wrote in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt;, after all the documentation of what has been lost and what has hurt and what has not worked out, I decided that the thing I still have, the thing I am still willing to put at the front, is love. In whatever form it takes, however modest or strange or inexplicable. Love for creation. Love for the small things that are genuinely beautiful. Love for characters made by artists who cared enough to make them real. That is what is in the picture. That is what I am putting in front.&lt;/p&gt;

&lt;p&gt;Studies in affective science have shown that symbolic self-expression, including the images people choose to represent themselves in social contexts, plays a meaningful role in identity construction and emotional regulation (3). A profile picture is not just a thumbnail. It is a statement. It is the answer to the question "who are you" compressed into a square. I have spent years trying to answer that question in long, difficult, honest posts. But sometimes the most honest answer is the simplest one. I chose this image because it represents what is left when everything complicated is stripped away. It represents warmth. It represents the decision to still care.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anime, Fiction, and Seriousness
&lt;/h2&gt;

&lt;p&gt;I want to spend some time on a topic that I think deserves more genuine engagement than it usually gets, because the dismissal of fiction as "just" fiction is one of those pieces of received wisdom that collapses completely the moment you apply any real pressure to it. The idea that caring about fictional characters is somehow less valid, less serious, or less meaningful than caring about real ones is a position that cannot survive contact with the actual history of human culture. Every religious tradition is built in significant part on stories about figures whose historicity is contested. Every political ideology tells a narrative about heroes and villains, origins and destinies. Every family has its myths and its legends, the ancestors who did remarkable things, the stories that get told at the table and shape how people understand who they are and where they come from. These are all forms of fiction, in the sense that they are constructed narratives rather than raw reality, and no serious person argues that they do not matter. Fiction matters because humans are narrative creatures. We think in stories. We feel through characters. We organize our experience into plots with beginnings and middles and ends, and we use the stories we encounter to reflect on the stories we are living (4).&lt;/p&gt;

&lt;p&gt;Anime, specifically, has a cultural weight that Western audiences often underestimate because it comes packaged in aesthetic conventions that look foreign or childish to eyes not calibrated to appreciate them. The large eyes, the colorful hair, the exaggerated emotional expressions, all of these are a visual language with its own grammar and its own range of expression, and reducing them to a judgment about immaturity is like dismissing jazz because the instruments are different from classical ones. Japanese animation has produced some of the most emotionally sophisticated storytelling of the twentieth and twenty-first centuries. Works like Neon Genesis Evangelion, which scholars and psychologists have analyzed extensively in the context of depression, anxiety, and the search for existential meaning (5), or Violet Evergarden, which explores grief and the limits of language with a precision that most literary fiction cannot match, or Ping Pong the Animation, which is as serious a meditation on identity and competition and the cost of excellence as anything being produced in prestige television. These are not guilty pleasures. They are achievements. The fact that they are drawn rather than filmed does not diminish the thinking behind them or the feeling they communicate.&lt;/p&gt;

&lt;p&gt;VTuber culture, which emerged from the anime aesthetic tradition, inherits some of this weight while also being something new. It occupies a strange and genuinely novel space where performance and authenticity coexist in a way that prior media formats did not make possible. A VTuber is simultaneously a character, a person, and a relationship. The character is fictional. The person behind it is real. And the relationship between the viewer and the content is ongoing, accumulated over time, built through the same mechanisms of familiarity and affection that build any long-term human connection. Research on parasocial attachment has found that the duration and regularity of media consumption significantly predicts the depth of parasocial bonds, and that these bonds activate many of the same emotional and cognitive processes as real social relationships (1). This is not pathology. This is the human social brain doing what it always does: building connection from available material. When the available material is especially good, when the person behind the avatar is genuinely warm and the character they have created is genuinely appealing, the result is something real, something that actually affects how people feel and how they move through their days.&lt;/p&gt;

&lt;p&gt;And if you still need a proof that this kind of attachment is not some pathological quirk of isolated internet users, let me tell you about Grape-kun. Grape-kun was a &lt;a href="https://en.wikipedia.org/wiki/Grape-kun" rel="noopener noreferrer"&gt;Humboldt penguin&lt;/a&gt; at Tobu Zoo in Saitama, Japan (8). In 2017, the zoo partnered with the anime series &lt;em&gt;Kemono Friends&lt;/em&gt; and placed cardboard cutouts of the show's anthropomorphized animal characters around the enclosures. One of those cutouts depicted Hululu, a penguin character. Grape-kun, elderly and recently rejected by his former mate, walked over to Hululu's cutout and did not leave. He stood beside her for hours every day, performing courtship displays, staring at her face, guarding her from the other penguins. He did this every day, for months, for the rest of his life. He died on October 12, 2017, at the age of twenty-one. Tobu Zoo erected a shrine. People left flowers. Fan artists around the world drew tributes. And the one thing everyone who heard the story understood, immediately, without needing it explained, was that Grape-kun was not confused. He had found something warm, and he had refused to leave it. That is not a joke. That is not a punchline about parasocial relationships or unhealthy fixation. That is what love looks like when it has nowhere else to go. I am not comparing myself to a penguin. But I am saying that the capacity to fix your gaze on something beautiful and refuse to look away is older and deeper and more essential than the people who mock it will ever understand. Grape-kun did not know he was supposed to be embarrassed. He just stood there, next to his cardboard Hululu, and meant it completely.&lt;/p&gt;

&lt;p&gt;I have read posts by people who are embarrassed about caring about VTubers or anime characters. I understand where the embarrassment comes from. The culture we live in has a very specific hierarchy of what counts as serious and what counts as frivolous, and anything associated with Japanese popular culture tends to land near the bottom of that hierarchy in the eyes of people who have not engaged with it. But I am not embarrassed. I have thought about why I am not embarrassed, and the answer is that embarrassment requires caring more about the judgment of others than about the truth of your own experience. And I made a decision a while ago, documented throughout these posts, to stop caring more about judgment than about truth. The truth is that Shigure Ui's work has brought me genuine joy. The truth is that the character she created is genuinely lovely. The truth is that in a period of my life when very little was working as intended, encountering something beautiful and warm made a real difference. That is the truth, and I will put it in a profile picture and write a thousand words about it and not spend a second of energy being embarrassed about it.&lt;/p&gt;

&lt;p&gt;There is also something specific about the design aesthetic of VTuber characters, and Shigure Ui's work in particular, that I want to articulate because I think it is underappreciated. Character design at the level that she operates is a form of visual psychology. A well-designed character communicates a personality through pure visual information: the shape of the eyes, the weight of the linework, the palette, the silhouette. Before a character has spoken a single word, a viewer has already formed an emotional impression based entirely on how the character looks. This is design working at the level of the unconscious, and doing it well is genuinely difficult, which is why so much character design fails to create any real impression at all. Shigure Ui's designs succeed. They communicate warmth and playfulness and a slight touch of chaos in a combination that feels distinct and coherent and, importantly, kind. They feel like designs made by someone who actually likes people. And in a world that often feels like it was designed by people who have never thought about whether anything they made would make anybody feel welcome, that quality stands out.&lt;/p&gt;

&lt;h2&gt;
  
  
  On Being Alone
&lt;/h2&gt;

&lt;p&gt;I will be honest about the context in which I started watching Shigure Ui streams and following her work, because context is everything and pretending it is not would be dishonest. I was alone. I have been alone for a long time, in various senses of that word. Geographically isolated for much of my life, socially isolated by the combination of conditions I wrote about in detail in &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt;, professionally isolated in the way that anyone who builds things independently without an institutional affiliation tends to become isolated over time. Loneliness is a topic that researchers are increasingly treating as a genuine public health concern. Studies have found that chronic loneliness carries health risks comparable to smoking fifteen cigarettes a day, that it accelerates cognitive decline, that it activates the same neural pathways as physical pain, and that it is correlated with depression, anxiety, and cardiovascular disease (6). These are not small findings. Loneliness is not just unpleasant. It is damaging in ways that are measurable and serious, and the damage does not wait for you to ask for help before it starts accumulating.&lt;/p&gt;

&lt;p&gt;I am not writing this to generate sympathy. I have written about loneliness before, in several other posts, and generating sympathy was not the goal there either. I am writing about it here because I want to be clear that the thing I found in Shigure Ui's work was not just entertainment. It was warmth, and warmth matters when you have been cold for a long time. There is a passage in one of her streams, I am not going to pretend I know the exact timestamp because I was watching at two in the morning and I did not have the presence of mind to save the clip, where she talks about her love for the people who have supported her work, and the way she says it is not performative at all. It is simple and direct and clearly genuine. And I remember sitting with that for a moment because it had been a while since I had experienced anything, even mediated, that felt that uncomplicated in its warmth.&lt;/p&gt;

&lt;p&gt;I want to talk about something that I think is poorly understood, which is the relationship between beauty and resilience. People tend to think of resilience as a kind of toughness, a capacity to shoulder difficulty without breaking, a refusal to be crushed. And there is something to that. But in my experience, the deeper form of resilience is not toughness. It is sensitivity. It is the ability to keep noticing things that are beautiful even when the overall picture is dark. It is the ability to be moved by something small, something that most people would scroll past, something that requires a willingness to be affected rather than defended. The people I have known who seemed most genuinely able to survive hard things were not the ones who had learned not to feel. They were the ones who had learned to feel very specifically, to find and hold onto the particular things that were worth feeling good about, even when those things were modest or strange or not what anyone else would have chosen.&lt;/p&gt;

&lt;p&gt;For me, in this period, one of those things has been the warmth that comes through Shigure Ui's work and presence. I recognize that this sounds like a lot to put on an animated character and her creator. I recognize that there is something almost comically disproportionate about writing this many words about a profile picture. But I have written long posts about artificial intelligence, about God, about the structure of the universe and the failure of language and the limits of human cognition, and all of those things are ways of trying to understand the world. This is also a way of trying to understand the world. It is just a gentler one. And maybe that is what I needed. After everything I wrote in &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt;, all that effort to parse the hard edges of technology and intelligence and the future, maybe it is healthy to also spend time with something that is just soft and funny and nice. Maybe the heart needs that. I think the heart needs that.&lt;/p&gt;

&lt;p&gt;Research in positive psychology, specifically the work coming out of Sonja Lyubomirsky's lab on sustainable happiness, finds that one of the most consistently effective strategies for improving wellbeing is what researchers call "savoring", the deliberate act of attending to and appreciating positive experiences, including experiences that society might dismiss as trivial (7). The content of the experience matters less than the quality of attention you bring to it. A person who savors a funny moment in a VTuber stream is doing something psychologically meaningful, regardless of what the stream is about or who produced it. The brain does not add a footnote saying "this joy doesn't count because it came from animated content." It just registers the joy. And joy, registered often enough and deliberately enough, changes the baseline. I am trying to change my baseline. Shigure Ui is part of that effort.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Heart Is an Argument
&lt;/h2&gt;

&lt;p&gt;I want to close this post by making an argument that I realize is not the kind of argument I usually make in these posts. I am usually arguing about technology or consciousness or the ethics of an industry or the limits of language. This argument is different. It is less rigorous and more essential. It is an argument for letting yourself love things.&lt;/p&gt;

&lt;p&gt;Not selectively, not after checking whether the thing is cool enough or serious enough or approved by the right people. Just love things. Let yourself be moved by the things that actually move you. Let yourself attach to the things that are genuinely warm and beautiful and good. The world does not need your ironic distance. The world has plenty of ironic distance already, and it has not made anything better. What is actually rare, what is actually difficult to maintain under pressure, is the willingness to care, openly and without embarrassment, about things that make life feel like something worth living.&lt;/p&gt;

&lt;p&gt;I have written about suffering. I have written about anger. I have written about the failures of systems and the cruelties of circumstance and the things that have not worked out and the ways the world can grind a person down until there is almost nothing left. But I have also written, in quieter corners of these posts, about the things that have kept me here. The music. The ideas. The strange satisfaction of building something from nothing. The moments of unexpected beauty that show up without warning and for no particular reason and change how a day feels. These are not small things. These are the things that determine, at the end of the day, whether a life feels worth living. And I am telling you, explicitly and without apology, that a silly colorful animated character who draws and streams and laughs at her own jokes has been one of those things for me recently. Not the only thing, not the biggest thing, but a real thing, a thing that brought warmth into days that needed warmth.&lt;/p&gt;

&lt;p&gt;I changed my profile picture because I wanted to put something up that represented where I actually am, not where I think I should be or where I want to project that I am, but where I actually am. Where I actually am is a person who has been through a lot, who is still here, who is tired and honest and occasionally struggling, and who found something genuinely lovely in an unexpected corner of the internet and decided to let it matter. The picture has a heart in it. That is the point. That is the whole point. When everything else gets complicated, and it always does, the heart is the argument that still holds. The capacity for love and warmth and attachment, even to fictional things, even to made-up characters designed by an artist in Japan who will never know I exist, is not a consolation prize for people who could not get the real thing. It is the thing. It is what makes the rest of this bearable and sometimes, on good days, beautiful.&lt;/p&gt;

&lt;p&gt;If you are reading this and you have your own version of this, your own cartoon, your own fictional character, your own corner of culture that makes you feel something genuine when the rest of the world is not cooperating, I want you to know that I see you, and it counts, and you do not have to be embarrassed about it. The heart knows what it needs. It tends to know better than the part of us that is worried about being taken seriously. Trust it.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Horton D, Wohl R. Mass communication and para-social interaction: Observations on intimacy at a distance. &lt;em&gt;Psychiatry&lt;/em&gt;. 1956;19(3):215-229. &lt;a href="https://doi.org/10.1080/00332747.1956.11023049" rel="noopener noreferrer"&gt;https://doi.org/10.1080/00332747.1956.11023049&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; Shigure Ui YouTube channel. &lt;a href="https://www.youtube.com/@ui_shig" rel="noopener noreferrer"&gt;https://www.youtube.com/@ui_shig&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Belk RW. Extended self in a digital world. &lt;em&gt;Journal of Consumer Research&lt;/em&gt;. 2013;40(3):477-500. &lt;a href="https://doi.org/10.1086/671052" rel="noopener noreferrer"&gt;https://doi.org/10.1086/671052&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; Mar RA, Oatley K. The function of fiction is the abstraction and simulation of social experience. &lt;em&gt;Perspectives on Psychological Science&lt;/em&gt;. 2008;3(3):173-192. &lt;a href="https://doi.org/10.1111/j.1745-6924.2008.00073.x" rel="noopener noreferrer"&gt;https://doi.org/10.1111/j.1745-6924.2008.00073.x&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Veit W. Neon Genesis Evangelion and the Meaning of Life. &lt;em&gt;Psychology Today&lt;/em&gt;. &lt;a href="https://www.psychologytoday.com/us/blog/science-and-philosophy/202003/neon-genesis-evangelion-and-the-meaning-life" rel="noopener noreferrer"&gt;https://www.psychologytoday.com/us/blog/science-and-philosophy/202003/neon-genesis-evangelion-and-the-meaning-life&lt;/a&gt;. Published March 21, 2020.&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Holt-Lunstad J, Smith TB, Baker M, Harris T, Stephenson D. Loneliness and social isolation as risk factors for mortality: A meta-analytic review. &lt;em&gt;Perspectives on Psychological Science&lt;/em&gt;. 2015;10(2):227-237. &lt;a href="https://doi.org/10.1177/1745691614568352" rel="noopener noreferrer"&gt;https://doi.org/10.1177/1745691614568352&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Lyubomirsky S, Sheldon KM, Schkade D. Pursuing happiness: The architecture of sustainable change. &lt;em&gt;Review of General Psychology&lt;/em&gt;. 2005;9(2):111-131. &lt;a href="https://doi.org/10.1037/1089-2680.9.2.111" rel="noopener noreferrer"&gt;https://doi.org/10.1037/1089-2680.9.2.111&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; Wikipedia contributors. Grape-kun. Wikipedia. &lt;a href="https://en.wikipedia.org/wiki/Grape-kun" rel="noopener noreferrer"&gt;https://en.wikipedia.org/wiki/Grape-kun&lt;/a&gt;. Accessed August 22, 2026.&lt;/p&gt;

</description>
      <category>discuss</category>
    </item>
    <item>
      <title>Intelligence at Rest</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 19:15:44 +0000</pubDate>
      <link>https://dev.to/wiseai/intelligence-at-rest-43co</link>
      <guid>https://dev.to/wiseai/intelligence-at-rest-43co</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This post was originally published on &lt;a href="https://wiseai.dev/blogs/intelligence-at-rest" rel="noopener noreferrer"&gt;the main website&lt;/a&gt; on &lt;a href="https://github.com/wiseaidotdev/blog/pull/23" rel="noopener noreferrer"&gt;Jul 18 2026&lt;/a&gt;. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account &amp;lt;3!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;There is a question that has been embedded in everything I have written here, hiding underneath the arguments about language models and hiring algorithms and the machinery of modern tech, and I have never stated it plainly until now. Not because I was avoiding it, but because I did not have the right words for it. The question is this: what is intelligence when it is not doing anything? Not when it is generating text, not when it is scoring a resume, not when it is answering a question, but when it is sitting completely still inside a system, as weights in a file on a server somewhere, or as knowledge inside a mind that is asleep. What shape does intelligence take at rest? And does the quality of that resting form determine everything else that follows when the system finally wakes up and moves? I think it does, and I think the reason almost no one talks about this is that rest is invisible, and we are a field obsessed with what is visible.&lt;/p&gt;

&lt;p&gt;I realize I have been circling this idea for a while without naming it. When I wrote about &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;language models being trapped in a symbolic cage&lt;/a&gt;, the underlying frustration was that their resting structure is too shallow to support genuine understanding, that the motion looks impressive while the substance underneath it is thin. When I wrote about &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;equations being the most compressed representation of reality&lt;/a&gt;, I was really writing about what good intelligence looks like at rest, compact enough to transmit, rich enough to generate everything downstream. Even my post about &lt;a href="https://wiseai.dev/blogs/i-miss-the-pre-ai-mossad-agents" rel="noopener noreferrer"&gt;the pre-AI Mossad agents&lt;/a&gt; was, at its core, about this: what those human recruiters had was a resting structure of judgment that an ATS score does not and cannot replicate, not because the algorithm is slow, but because it never built the right thing in the first place. This post is where I stop circling and land on the idea directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Motion Is Not Where Intelligence Lives
&lt;/h2&gt;

&lt;p&gt;Everything about the way we talk about AI is framed around motion. The model generates. The model predicts. The model responds. The model reasons. Every benchmark, every demo, every launch announcement, every breathless 𝕏eet from a founder who thinks they just built AGI is about what the system does when it moves. We measure tokens per second. We measure latency. We measure how fast the model can produce an output given a prompt. We celebrate fluency, which is a property of language in motion, and we celebrate responsiveness, which is a property of systems under load. The entire discourse is built around the assumption that intelligence is fundamentally about action, about the visible production of something that looks like thought. I want to argue that this assumption is wrong, that it is not just partially wrong but structurally wrong, and that it has led the entire field down a path that optimizes for the wrong thing at enormous cost.&lt;/p&gt;

&lt;p&gt;Think about what a trained neural network actually is, not when it is running inference, not when it is generating text or classifying images or predicting the next move in a chess game, but right now, at this moment, sitting as a file on some server somewhere. It is weights. It is numbers. It is a tensor of floating-point values that encodes, in some compressed and distributed form, everything the system learned during training. Nothing is moving. No output is being produced. No computation is running. The system is completely still. And yet something is clearly in there. Something is stored. Something has been compressed into that tensor that makes this particular set of weights more valuable than a random initialization, more capable than the system was before training, more able to do useful things when it does eventually run. That something is what I am calling intelligence at rest. And the fact that it is sitting there, invisible and silent and motionless, does not make it less real. It makes it more important, because the motion is only the surface, and the resting structure is the substance underneath.&lt;/p&gt;

&lt;p&gt;This distinction matters because we keep making the same mistake in different contexts. We measure intelligence by its outputs. We evaluate AI by the quality of what it generates. We judge models by benchmark performance, which is a measure of what the system does, not what the system is. But what the system is, that deep structure of latent capability sitting inside the weights, is what determines what it can do across all possible inputs, including inputs it has never seen before. A model that has genuinely learned the statistical structure of a domain will generalize to new inputs in that domain. A model that has memorized its training data will fail on inputs that fall outside the distribution it was shown. You cannot tell the difference by looking at a single output. You can only start to see it when you look at the resting structure, when you ask what is actually encoded in those weights, how much genuine compression is there versus how much raw storage of surface patterns. The resting form of the intelligence is where the real quality lives, and we are almost completely ignoring it.&lt;/p&gt;

&lt;p&gt;I also think this framing helps us understand something about the current moment that is very hard to articulate otherwise. We are surrounded by AI systems that move constantly, that produce output at industrial scale, that generate text and code and images and audio faster than any human can evaluate them. And yet, despite all this motion, the feeling that something important is missing has not gone away. It has gotten louder. That feeling is, I think, a signal about the gap between intelligence in motion and intelligence at rest. The motion is impressive. The generation is fast. The outputs are often useful. But the resting structure underneath the generation is, in most current systems, shallow in a way that the motion can sometimes hide. When the system talks about something it has genuinely compressed into its latent space, the outputs feel real. When it is confabulating, pattern-matching on surface features, generating plausible text without any underlying structure, the outputs feel hollow even when they are grammatically perfect. That hollow feeling is the absence of deep resting intelligence, and it is something that can only be fixed at the level of the resting structure, not at the level of the output generator.&lt;/p&gt;

&lt;p&gt;I want to be honest about my own experience here, because I have always tried to be honest in these posts, and this idea did not come to me from reading papers. It came from something much more personal. When I was going through the worst of what I described in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt;, I kept asking myself why my effort was not translating into progress. I was applying constantly. I was producing work. I was generating output. And nothing was happening. What I eventually realized was that the problem was not the motion. The motion was real and the effort was real. The problem was that I had not yet built the kind of resting structure that could survive contact with a system designed to filter on surface signals. Intelligence in motion without a deep resting foundation is like a river with no riverbed. The water moves, but it does not go anywhere coherent. The resting structure is what gives the motion direction, and without direction, motion is just energy being spent.&lt;/p&gt;

&lt;p&gt;The field of AI has a similar problem, and I do not think it is a coincidence that the problem shows up in the same way. There is a great deal of motion. There is a great deal of generation. There is a great deal of output. But the resting structures underneath that motion are, in many cases, very much shallower than the outputs suggest. The confabulation is real. The hallucination is real. The fragility under distribution shift is real. These are all symptoms of a system where the motion is more sophisticated than the resting structure that generates it, where the surface has developed faster than the substance. And the way you fix that is not by making the motion faster or smoother or more impressive. You fix it by going deeper into the resting structure, into the compressed representation of what the system has actually learned, and asking whether that structure is the right shape to support the motion you want to produce.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Rest Actually Means for an Intelligent System
&lt;/h2&gt;

&lt;p&gt;When I use the phrase "intelligence at rest", I am making a claim that is simultaneously intuitive and technical, and I want to be precise about both dimensions before the arguments get complicated. Intuitively, intelligence at rest is the intelligence that a system possesses before it acts, the capability that exists in the structure of the system independently of whether that structure is currently being activated. More technically, intelligence at rest is the useful compression of patterns, regularities, decision boundaries, and causal structures that a system has encoded into its parameters through training, and that determine what the system will do when it eventually receives an input. This is related to, but distinct from, the concept of knowledge. Knowledge is what is stored. Intelligence at rest is the quality and organization of what is stored, the degree to which the stored structure genuinely captures the underlying regularities of the domain rather than just the surface appearances of the training data.&lt;/p&gt;

&lt;p&gt;This distinction is important because knowledge and intelligence are not the same thing, as I argued at length in &lt;a href="https://wiseai.dev/blogs/knowledge-and-intelligence-are-mutually-exclusive" rel="noopener noreferrer"&gt;Knowledge and Intelligence are Mutually Exclusive&lt;/a&gt;. A system can store an enormous amount of information and still be unintelligent in the sense that matters most, which is the ability to deploy that information effectively on problems it has not seen before. The resting structure of an intelligent system is not just a warehouse of facts. It is a compressed representation of how the domain works, a model of the regularities that generate the facts, so that the system can handle new instances of the pattern without having memorized them specifically. That is the difference between a system that has compressed the domain and a system that has memorized the data, and it is exactly the difference that matters for generalization. Compression produces intelligence. Memorization produces recall. And the two are not the same thing even when their surface outputs look identical on a benchmark.&lt;/p&gt;

&lt;p&gt;The information bottleneck principle, introduced by Tishby and colleagues, is the clearest rigorous statement of this idea that I have found in the literature (1). The principle says that a learning system should be thought of as finding a compressed representation of its input that still preserves the information relevant to predicting the output. A system that compresses too little retains noise along with signal, overfit to the specific training examples, and generalizes poorly. A system that compresses too much loses signal along with noise, underfit to the task, and also generalizes poorly. The sweet spot is a representation that is as compact as possible while still preserving the task-relevant structure. That sweet spot is intelligence at rest. It is the most compressed form of the system's understanding of the domain that still supports the behavior the system is designed to produce. And finding that sweet spot is, in a deep sense, what learning is actually for.&lt;/p&gt;

&lt;p&gt;I also want to connect this to the minimum description length principle, because it makes the same argument in a slightly different language that I find very clarifying (2). MDL says that the best model of the data is the one that produces the shortest combined description of the model and the data given the model. A model that is too simple will require a long description of the residual errors. A model that is too complex will require a long description of itself. The minimum is achieved by a model that is just complex enough to capture the real structure of the data, and no more. That minimum-description-length model is the most efficient resting intelligence you can build for a given domain. It is not the system with the most parameters. It is the system with the most meaningful parameters. And meaningful parameters are exactly what intelligence at rest consists of, parameters that encode genuine structure rather than memorized surface patterns.&lt;/p&gt;

&lt;p&gt;The practical implication of this is something the field has been circling for years without quite landing on it. The quality of a model's resting intelligence is not determined by how large it is. It is determined by how efficiently its structure encodes the domain it was trained on. A small model that has genuinely compressed the regularities of a domain can outperform a large model that has mostly memorized its training data, especially on inputs outside the training distribution. This is the empirical finding behind knowledge distillation, where researchers have shown repeatedly that a smaller student model can achieve performance comparable to a much larger teacher model by learning from the soft probability distributions the teacher assigns to classes rather than from hard labels (3). Those soft distributions contain more information about the teacher's internal structure than hard labels do, and the student inherits more of the teacher's resting intelligence by learning from them. The size decreases. The quality of the resting structure is preserved. That is the clearest possible demonstration that intelligence at rest is a property of organization, not of scale.&lt;/p&gt;

&lt;p&gt;I also find it useful to think about what happens when intelligence at rest fails, because the failure modes are very revealing. When a system hallucinates, it is producing motion that is not supported by genuine resting structure. The system has learned to produce text that looks like a description of a fact without having genuinely compressed the factual domain into its parameters. The motion is fluent but the rest is hollow, and the hallucination is the artifact of that mismatch. When a system fails dramatically on a distribution shift, it is because the resting structure it developed was shaped by the specific appearance of its training data rather than by the underlying regularities that generate that data, and when the appearance changes, the structure that depended on it collapses. When a system is confidently wrong about a question that falls just outside its training distribution, it is because the resting structure was organized around the surface statistics of the training set rather than around the causal generative model that produced the training set. Every major failure mode of current AI systems is an intelligence-at-rest problem. More motion will not fix them. Only better resting structure will.&lt;/p&gt;

&lt;p&gt;I want to close this section by saying something that I think is important for people who are frustrated by AI, which is a lot of people right now, including myself in many moments. The frustration is not irrational, and it is not misplaced, and it is not just about bad outputs. The frustration is a signal that the system's resting structure is not deep enough to support the motion it is producing, and that the surface has outrun the substance. That is a correctible problem, but it requires the field to stop optimizing for the motion and start optimizing for the rest. It requires building systems that compress more genuinely, that memorize less and abstract more, that develop resting structures rich enough to support the kind of generalization that makes intelligence actually useful. That is the direction the field needs to go, and the fact that it is not the direction most people are talking about does not make it less true.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Deep Parallel Between Data at Rest and Intelligence at Rest
&lt;/h2&gt;

&lt;p&gt;There is a concept in information security and data engineering that most AI researchers never think about, because they work at a different layer of the stack. That concept is data at rest, and it refers to stored data that is idle, not being transmitted, not being processed, just sitting in a database, a file system, a backup drive, or a cloud bucket. The reason data at rest gets special attention in security is that stored data is uniquely vulnerable in specific ways. It can be stolen, copied, corrupted, encrypted by ransomware, exposed by misconfiguration, or lost along with the physical medium that holds it. NIST cybersecurity standards are explicit that data at rest requires protection through encryption, access control, and integrity monitoring precisely because the storage state is where data is most often compromised (4). That protection imperative is not incidental. It reflects a deep truth about the relationship between storage, value, and vulnerability: the more valuable something is, the more consequential it is that it rests somewhere, and the resting place is where it is most exposed.&lt;/p&gt;

&lt;p&gt;The parallel to intelligence at rest is exact enough to be useful and specific enough to reveal real engineering implications. A trained model's parameters are the AI equivalent of data at rest. They sit in storage between inference runs. They are transmitted between systems when the model is deployed. They are backed up in checkpoints during training. They are exposed to the environment in all the ways that stored data is exposed, through physical media, through network transmission, through access control policies, through the possibilities of tampering and theft and disclosure. And just as with data at rest, the value of what is stored is not proportional to its size. A small file containing a private key is worth more than a large file containing garbage. A small model that has genuinely compressed a domain into its parameters is worth more than a large model that has mostly memorized its training data. The quality of the resting structure determines the value, and the resting state is where that value must be protected.&lt;/p&gt;

&lt;p&gt;But the parallel goes deeper than security. It goes into the fundamental question of what compression achieves in both domains. In data engineering, the goal of compression is to represent the same information in a smaller space without losing the ability to recover the original content. In intelligence, the goal of compression is to represent the generative structure of a domain in a compact representational space without losing the ability to predict new instances of the domain. These are not the same thing, but they are related at a deep level, because both are asking how much content can be preserved in how little space, and both are measuring the quality of the compression by how faithfully the stored form supports the original purpose. Lossless data compression recovers every bit. Good intelligence compression recovers every generalization. Lossy data compression accepts some degradation for size benefits. Shallow intelligence compression accepts some generalization failures for parameter efficiency. The entire vocabulary of compression naturally transfers between the two domains, and that transfer is not metaphorical. It reveals that both domains are instances of the same underlying problem: how to make valuable structure durable.&lt;/p&gt;

&lt;p&gt;There is also a privacy dimension to this parallel that I think is underappreciated in both fields, and that connects directly to the memorization research I mentioned earlier. Research by Carlini and colleagues has demonstrated that it is possible to extract verbatim training examples from large language models through carefully crafted prompts, a phenomenon they call training data extraction (5). This is the exact AI equivalent of what happens when sensitive data is stored without encryption and then accessed by an unauthorized party. In both cases, the resting structure contains more than it was supposed to contain. The data at rest was supposed to contain business information, not personal secrets, but because the storage was not properly controlled, the personal secrets are in there too. The model at rest was supposed to contain generalized knowledge, not memorized training examples, but because the training was not properly regularized, the training examples are in there too. In both cases, the resting state has absorbed material that should have been excluded, and that material can be recovered by someone who knows how to look. This is not a peripheral observation. It is a core challenge for both data governance and AI governance, and it arises from the same structural fact: rest is where content accumulates, and content that accumulates without discipline can leak in ways that are hard to detect and harder to reverse.&lt;/p&gt;

&lt;p&gt;The analogy also illuminates something about model ownership and intellectual property that most discussions of AI law completely miss. When we ask who owns a trained model, we are really asking who owns the intelligence at rest inside it, the compressed representation of patterns learned from data. If a model is trained on data that was collected without appropriate consent, the resting intelligence inside that model contains a compressed form of the original data's structure, and the question of who has rights to that structure is not a simple one. It is the AI equivalent of asking who owns the insights derived from improperly obtained data, and the answer involves questions about consent, provenance, and the rights of the people whose intellectual output was compressed into the model's parameters. These questions are being litigated right now in courts around the world, and they are genuinely hard, but they become slightly clearer when you understand that what is at stake is not just the model architecture or the training procedure. What is at stake is the intelligence at rest, the compressed form of human knowledge that lives inside the resting weights, and those weights carry traces of where they came from.&lt;/p&gt;

&lt;p&gt;I want to push the parallel one more step, into the concept of redundancy. In data engineering, redundancy is managed deliberately, through replication, through backups, through RAID configurations, to ensure that the stored data survives hardware failure. In intelligence engineering, redundancy appears in the form of parameter over-parameterization, where a model has many more parameters than might seem strictly necessary to represent the learned function. Research has shown that this over-parameterization is not just inefficiency. It plays a role in the dynamics of learning, helping models find smooth, generalizable solutions rather than sharp, memorizing ones (6). The redundancy in the resting structure of an over-parameterized model is, in a sense, like the redundancy in a replicated data store. It makes the system more robust, more generalizable, more able to survive the equivalent of hardware failure, which in the intelligence world is encountering inputs that fall slightly outside the training distribution. This is a fascinating convergence: in both data systems and intelligence systems, some amount of redundancy in the resting structure makes the overall system more reliable in motion.&lt;/p&gt;

&lt;p&gt;The convergence I have been describing here is not a coincidence, and I do not think it should be treated as a metaphor that eventually breaks down under scrutiny. I think it reflects a genuine structural similarity between two instances of the same underlying problem: how to make useful structure durable in a stored form. The vocabulary of data at rest, including compression, encryption, integrity, provenance, access control, and redundancy, applies with minimal modification to intelligence at rest. And the lessons learned from decades of data engineering practice, from the hard-won understanding of what happens when storage is not disciplined, are lessons that the AI field is in urgent need of learning. The resting state is where value accumulates. It is also where risk accumulates. And treating it with the same seriousness that we treat stored data is not overcautious. It is simply honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compression Is the True Measure of Intelligence
&lt;/h2&gt;

&lt;p&gt;There is a quote I keep coming back to when I think about this topic, and it is from a context that has nothing to do with AI. Occam's razor, the principle that the simplest explanation consistent with the evidence should be preferred, is not just a heuristic for scientific modeling. It is a statement about the relationship between intelligence and compression. The simplest explanation is the most compressed one. Preferring it is preferring compression. And the fact that simpler explanations tend to generalize better than complex ones is the empirical basis for the claim I am making in this section: compression is the true measure of intelligence. Not the quality of the outputs. Not the speed of the generation. Not the impressiveness of the benchmark scores. The quality of the resting compression. The degree to which a compact representation of the learned structure genuinely captures what generates the domain.&lt;/p&gt;

&lt;p&gt;I want to be clear that I am not arguing that all compression is good or that the most compressed system is always the most intelligent one. Compression has limits, and those limits are real. A system that compresses too aggressively will lose information that is genuinely task-relevant, and the result will be a system that cannot distinguish things that need to be distinguished. A system that compresses too little will retain noise that distorts the signal, and the result will be a system that cannot generalize past the specific training examples it has memorized. The ideal is not maximum compression. The ideal is minimum description length, the compression that is just sufficient to capture the real structure without capturing what is accidental. This is the sense in which Blier and Ollivier's work on training data compression is so important, because it gives us a principled way to measure whether a model's resting structure is genuinely compressing the domain or just storing a large amount of the domain's surface statistics (7). The measure is the total description length of the model and the data given the model, and minimizing that measure is approaching the ideal resting intelligence.&lt;/p&gt;

&lt;p&gt;The same principle appears in a very different form in the context of model pruning. When you prune a neural network, you remove parameters that contribute little to the network's predictions, often by zeroing out small weights or entire attention heads. The remarkable finding from decades of pruning research is that most models can be pruned quite aggressively without significantly degrading performance. Large fractions of the parameters in a fully trained model appear to contribute almost nothing to the output. This is the Lottery Ticket Hypothesis, which proposes that within a randomly initialized network, there exists a sparse subnetwork called the winning ticket that can be trained in isolation to match the performance of the full network (8). If this is right, and the empirical evidence strongly suggests it is, then the intelligent resting structure is sparse: it lives in a small subset of the parameters, and the rest of the parameters are essentially redundant. This has profound implications for what we should be building. We should be building systems that find the winning ticket, that identify and develop the sparse resting structure that actually carries the intelligence, rather than systems that brute-force their way to performance by scaling up the total number of parameters.&lt;/p&gt;

&lt;p&gt;Distillation reinforces this point from a different angle. When a small student model is trained to match the probability distributions of a large teacher model rather than the hard labels of the training data, it inherits the teacher's resting intelligence in a compressed form. The student contains fewer parameters. The student runs faster. The student requires less memory. And yet the student achieves performance that is often close to the teacher's on the tasks that matter. This is empirical evidence that intelligence can be compressed without being destroyed, that the useful structure in the teacher's resting weights can survive transfer into a smaller set of parameters. Hinton and colleagues, who introduced knowledge distillation, understood this explicitly, arguing that the soft probability distributions the teacher assigns to incorrect classes contain rich information about the teacher's internal structure that the hard labels completely discard (3). The resting intelligence is not in the size. It is in the organization. And good organization can be compressed.&lt;/p&gt;

&lt;p&gt;I also want to connect this to something that I think is deeply underappreciated in the current discussion about scaling laws. The empirical observation that model performance improves predictably with scale, with more parameters and more data, has been taken as evidence that scale is the key variable in intelligence. But scaling laws describe what happens at the level of measured performance on benchmark tasks. They do not describe what happens at the level of resting structure. A model that is ten times larger may perform better on benchmarks while having a resting structure that is qualitatively not much richer, just spread across more parameters. The additional parameters may be capturing additional surface statistics rather than additional genuine structure, and the performance gains may be coming from better coverage of the training distribution rather than from deeper compression of the domain's generating regularities. This is a distinction that benchmark scores cannot resolve, because benchmark scores measure outputs, not resting structure. Resolving it requires interpretability tools, compression metrics, and the kind of careful structural analysis that the field has been slow to prioritize because benchmark scores are so much easier to measure and so much easier to publish.&lt;/p&gt;

&lt;p&gt;The historical record of science is actually the strongest argument for compression as the measure of intelligence, and it is an argument that does not depend on any particular theory of machine learning. Every major scientific breakthrough has been a compression event. Newton compressed the motion of falling apples and orbiting planets into F = ma, a single equation that contains within its structure every trajectory of every object under gravitational acceleration. Maxwell compressed all of electric and magnetic phenomena into four equations that could fit on a coffee mug. Shannon compressed all of communication theory into a few concepts about entropy, mutual information, and channel capacity. In every case, the breakthrough was not producing more outputs. It was finding a simpler resting structure that generated all the outputs. The genius of Newton, Maxwell, and Shannon was not that they moved faster or generated more text. It was that they found the most compressed representation of the domain that still predicted everything the domain could do. That is intelligence at rest, demonstrated at the highest level the human mind has achieved. And any AI system that wants to claim genuine intelligence has to demonstrate the same thing, not in the outputs it produces, but in the resting structure that generates them.&lt;/p&gt;

&lt;p&gt;There is a quiet irony in the fact that the field most obsessed with intelligence has done so little to measure it in its resting form. We have leaderboards for benchmark performance. We have scaling curves for loss versus compute. We have evaluations of fluency, of factual accuracy, of reasoning chain quality. We do not have widely accepted metrics for the quality of the resting compression, for how efficiently the model's parameters encode the domain it was trained on, for how much of the learned structure is genuine abstraction versus surface memorization. That gap is not an accident. Measuring resting intelligence is harder than measuring output quality, because it requires looking inside the model rather than at what the model produces. But harder is not the same as impossible, and the interpretability and mechanistic analysis communities are making real progress on exactly this problem. The question of what intelligence looks like at rest is not just philosophical. It is a measurement problem, and measurement problems can be solved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Compact Intelligence Compounds
&lt;/h2&gt;

&lt;p&gt;I often think about intelligence and civilization in the same breath, because the history of human civilization is largely the history of how humans have stored and transmitted intelligence across time and space. Before writing, intelligence could only be transmitted by one person speaking or demonstrating to another. The amount of intelligence that could accumulate in a society was limited by the number of people who could be in the room at any given time. Writing changed that. Writing made intelligence portable, durable, and capable of surviving the death of the person who possessed it originally. A book can carry the intelligence of a person who died two thousand years ago into the mind of a person who is alive today, and that transmission is so efficient that a single person can now inherit the resting intelligence of thousands of people who preceded them. That is compounding, and it is the mechanism by which civilizations get smarter across generations.&lt;/p&gt;

&lt;p&gt;The invention of mathematics made this compounding even more powerful, because mathematical equations are the most compressed form of resting intelligence that humans can transmit. I wrote about this extensively in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by Default&lt;/a&gt;, and I do not want to repeat the full argument, but the core point is worth restating here. A single equation can encode a class of behavior that would take volumes of prose to approximate, and the equation is not approximate. It is exact. When Newton wrote his law of gravitation, he compressed the gravitational behavior of every object in the universe into a single expression, and every subsequent scientist who learned that expression inherited the full generality of Newton's insight in a form that could be read, verified, and extended in minutes. That is the economy of compressed intelligence at rest. One compact resting structure, transmitted and received almost instantaneously, carrying a full generational increment of human understanding.&lt;/p&gt;

&lt;p&gt;Modern AI operates on a different but structurally similar principle. A trained model is a form of compressed intelligence that can be copied and distributed to millions of devices at near-zero marginal cost. Once the compression is done, the intelligence can travel. It can be embedded in a phone, deployed on a server, packaged into an application, and accessed by people who have no understanding of how it was built or what it contains. This is both an extraordinary opportunity and a significant risk, and the two aspects are not separable. The opportunity is that compressed intelligence can reach everyone. The risk is that compressed intelligence can be misappropriated, concentrated in the hands of whoever built the compression engine, and used to extract value from the people whose data made the compression possible. I wrote about this tension in &lt;a href="https://wiseai.dev/blogs/as-engineers-llms-should-pay-us-for-tokens-usage" rel="noopener noreferrer"&gt;As Engineers, LLMs Should Pay Us for Token Usage&lt;/a&gt;, and the argument applies with equal force to the more general concept of intelligence at rest. The compression creates value. The question of who captures that value is political, not technical.&lt;/p&gt;

&lt;p&gt;But the compounding effect of resting intelligence has a deeper dimension that I think most people overlook, and it is the dimension that connects to the long arc of human history. Intelligence at rest is not static. It can be updated, refined, extended, and distilled into new resting forms that are more capable than the original. A scientist who inherits Newton's gravitational law does not just use it. They extend it, test it against new phenomena, discover its limits, and eventually compress a richer understanding into Einstein's general relativity, which is a more compressed and more general description of the same domain. That process of inheritance, extension, and recompression is the engine of scientific progress, and it works because each generation's resting intelligence is compact enough to be transmitted and rich enough to serve as the foundation for the next generation's work. If Newton's insights had been too diffuse to compress into equations, each subsequent scientist would have had to rediscover them from scratch, and the compounding would have been impossible.&lt;/p&gt;

&lt;p&gt;Modern AI has the potential to participate in this same compounding process, but only if the quality of the resting intelligence it produces is high enough. A model that has genuinely compressed a domain, that has found the minimum-description-length representation of the domain's generative structure, can serve as the foundation for the next model's learning in a way that shallow, memorizing models cannot. Distillation is the clearest example of this principle in action: a large teacher model's resting intelligence is compressed into a smaller student model's resting structure, producing a more efficient system that inherits the teacher's understanding rather than relearning it from raw data. But distillation is just one mechanism of intelligence compounding. The broader point is that any time one system's resting intelligence can serve as the foundation for another system's learning, intelligence is compounding across systems, and that compounding accelerates the growth of the total resting intelligence in the field. That is how human science works. That is how human culture works. And that is how AI could work, if we build it right.&lt;/p&gt;

&lt;p&gt;The economic implications of this are enormous, and I think they are not being taken seriously enough. If intelligence at rest can be made compact, it can be made cheap. And if it can be made cheap, it can be made universal. The reason that scientific knowledge shaped civilization so dramatically over the last four centuries is that it became cheap to transmit and verify. The printing press made books cheap. Mathematical notation made equations compact. The internet made both nearly free. Each reduction in the cost of transmission accelerated the rate at which humanity's resting intelligence compounded across generations and across geographies. AI has the potential to do the same thing at a much smaller time scale, to compress the intellectual advances of each year into a resting structure that can bootstrap the next year's advances at lower cost. But that potential will only be realized if the compression is good, if the resting structures that are being built genuinely capture structure rather than just storing surface statistics, and if the organization of who controls those structures allows for the kind of open transmission that makes compounding possible. Neither of those conditions is currently guaranteed, and both are worth fighting for.&lt;/p&gt;

&lt;p&gt;I also want to make a point that is easy to miss, which is that the economy of rest is an argument for humility about scale. If the quality of resting intelligence depends on the quality of the compression rather than the quantity of the parameters, then the current race toward larger and larger models is not clearly the right direction. A smaller model with better resting structure can outperform a larger model with worse resting structure on the same domain, and can do so with far less energy, less memory, and less computational overhead. The fact that we do not yet have reliable ways to measure the quality of the resting structure makes it tempting to proxy quality with scale, but proxies are not measurements, and the race toward scale may be accelerating past the point of diminishing returns without anyone noticing because the benchmarks are not sensitive to the dimension that matters most. Developing tools to measure the quality of resting intelligence, to assess how efficiently a model's parameters encode the structure of its domain, is one of the most important technical problems in the field right now, and it is almost completely unrecognized as such.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory, Memorization, and the Ethics of What Is Stored
&lt;/h2&gt;

&lt;p&gt;There is something uncomfortable about the idea that a model's resting intelligence contains not just generalized patterns but also specific traces of the data it was trained on. It is uncomfortable because it blurs a boundary that we like to keep sharp: the boundary between a model that has learned from data and a model that has stored data. We want to believe that training is a process of distillation, that the model is learning the structure but not the specifics, that the result is a compressed understanding rather than a compressed copy. The research says this is sometimes true and sometimes not, and the distinction matters enormously for both technical and ethical reasons. I think about this question not just as an abstract problem but as a question about the dignity of the people whose work went into the training data, including engineers whose code was scraped, writers whose words were collected, and ordinary people whose conversations and searches and creations were harvested into training sets without their knowledge.&lt;/p&gt;

&lt;p&gt;The memorization literature is unambiguous on the basic facts. Carlini and colleagues showed that large language models will sometimes reproduce verbatim fragments of their training data under the right prompting conditions (5). Feldman and colleagues showed that individual training examples can have measurable influence on model outputs even when those examples are not reproduced verbatim, through a phenomenon they call memorization-mediated learning, where rare examples in the training set are learned through implicit memorization rather than through generalization (9). And Haim and colleagues showed that in some settings, trained model parameters contain enough information to reconstruct portions of the training data, which means that the resting weights of a model are not just a compressed description of the domain's structure; they are also, in part, a compressed encoding of specific training examples that can be partially recovered (10). These findings do not mean that all learning is memorization, but they do mean that the boundary between memorization and learning is not clean, and that the resting structure of a model always contains some mixture of genuinely generalized structure and specific memorized traces.&lt;/p&gt;

&lt;p&gt;This has implications that go well beyond technical model quality. It means that when a model is deployed, it carries within its resting weights a form of record of the data it was trained on, a compressed and imperfect record, but a record nonetheless. It means that access to a trained model is, in some non-negligible sense, access to a compressed form of the training data. It means that the privacy protections we apply to training data need to extend, in some modified form, to the models trained on that data. And it means that the people whose work, words, and creations were used to train a model have a legitimate interest in how the resting intelligence that was built from their contributions is used. This is not a fringe legal argument. It is a direct consequence of the technical reality of what training does and what resting intelligence contains. The model is not separate from its training data. It is a compression of its training data, and the ethical status of the training data does not disappear when it is compressed into weights.&lt;/p&gt;

&lt;p&gt;I want to connect this to something I wrote in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt;, where I described how the tech industry extracts value from workers and creators while returning almost none of it to them. The specific mechanism I described was labor, the way that software engineers, like myself, generate enormous value through their work while the systems they build are captured by investors and executives. The memorization finding adds a new dimension to that argument. When a language model is trained on code written by engineers, the model's resting intelligence contains a compressed form of that engineering knowledge. When the model is then used to generate code, it is drawing on that compressed knowledge to produce economic value. The engineers whose code was used to train the model are contributing to every output the model generates, but they are not compensated for that contribution. The intelligence they generated is now resting inside a model they do not own, generating revenue for a company they did not build, and there is no mechanism in the current legal or economic system to redress that. That is a problem that needs to be named clearly before it can be addressed.&lt;/p&gt;

&lt;p&gt;The ethics of memorization also connect to something more personal and harder to articulate. There is a sense in which the things we create, the code we write, the essays we publish, the conversations we have, are expressions of our intelligence, compressed forms of how we see the world. When those expressions are used to train a model without consent, what is being taken is not just text files. What is being taken is a form of resting intelligence, a compressed record of how we think, and that compressed record is then stored inside a model that someone else owns and profits from. I do not have a legal argument to make here, because the law has not yet caught up with the technical reality. But the moral argument is straightforward. The resting intelligence inside a trained model is partially composed of the resting intelligence of the humans who created the training data. Acknowledging that is not just accurate. It is the beginning of an honest conversation about who should benefit from what gets stored.&lt;/p&gt;

&lt;p&gt;The practical implication for model builders is that the quality of resting intelligence depends on the quality of the data curation, and the ethics of resting intelligence depend on the validity of the consent and provenance of the training data. These are not separate concerns. A model trained on high-quality, well-curated data from consenting sources will have resting intelligence that is both technically better and ethically cleaner than a model trained by scraping whatever is available on the internet. The quality argument and the ethics argument point in the same direction, toward deliberate, careful, consented data curation, and away from the current paradigm of scraping at scale and hoping that the ethical and legal problems resolve themselves. They will not resolve themselves. They will accumulate in the resting structure of every model built under the current paradigm, and at some point they will become impossible to ignore.&lt;/p&gt;

&lt;p&gt;I want to end this section by saying something that I hope sounds as obvious as it feels to me. We would not let a company store copies of people's private data without consent, use those copies to generate commercial products, and then refuse to acknowledge that the copies exist inside their systems. The fact that the storage happens through a training process rather than through direct file copying does not change the moral structure of the situation. What gets stored in the resting intelligence of a model is not identical to what was in the training data, but it is derived from it, shaped by it, and partially recoverable from it. That is enough to make the ethical obligations the same. And the sooner we act as though that is true, the better equipped we will be to build AI systems that are not just technically capable but genuinely trustworthy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Toward a Civilization That Learns to Rest Well
&lt;/h2&gt;

&lt;p&gt;I want to end this post the way I have ended every post since I started writing these, by zooming out and looking at what all of this means not just for AI systems but for the broader project of human civilization. Because I do not think the question of intelligence at rest is a narrow technical question. I think it is one of the most important questions of our time, and it has implications that extend far beyond the field of machine learning into how we understand knowledge, culture, progress, and what we owe to each other. The question of how intelligence rests, how it is stored, how it is compressed, how it is transmitted, and what it carries with it through time, is the question of how civilization thinks. Getting it right matters in a way that is very difficult to overstate.&lt;/p&gt;

&lt;p&gt;Every technology that has fundamentally changed human civilization has done so by changing the resting storage of intelligence. Writing changed it by making intelligence portable across space. Printing changed it by making intelligence reproducible at scale. The internet changed it by making intelligence instantaneously accessible across the world. AI training is changing it by making intelligence learnable from the compressed experience of millions of humans simultaneously, and by making that learned compression deployable at near-zero marginal cost. Each of these transitions was accompanied by enormous social disruption, by winners and losers, by questions about who controls the storage medium and who benefits from the intelligence it carries. None of those disruptions was clean. None of them resolved themselves without conflict. And I think it would be extraordinarily naive to expect this one to be different.&lt;/p&gt;

&lt;p&gt;The key lesson from those historical transitions is that the storage medium does not determine the quality of the resting intelligence. The printing press could carry Aristotle or it could carry propaganda. The internet can carry the sum of human knowledge or it can carry misinformation at scale. The quality of the resting intelligence depends on the quality of the compression, and the quality of the compression depends on the discipline, the care, and the honesty of the people doing the compressing. AI training is the most powerful compression engine humans have ever built, and whether it produces resting intelligence that is genuinely rich and genuinely beneficial depends entirely on choices that are being made right now, about what data to train on, about how to regularize the training, about how to evaluate the quality of the resting structure, and about who has access to the resulting intelligence. Those are not technical choices alone. They are social, political, and ethical choices, and they are being made mostly by a very small number of people in a very small number of organizations without anything like the broad democratic input that choices of this magnitude require.&lt;/p&gt;

&lt;p&gt;I have been arguing throughout this post that the quality of resting intelligence is measured by the quality of the compression, and I want to make one more argument for that claim before I close. The most powerful resting intelligence in human history is not stored in any AI model. It is stored in mathematics, the compact, tested, verified, transmissible body of equations and theorems and proofs that encodes the structure of the physical world. That resting intelligence has been compressed by generations of scientists working with immense care to find the simplest descriptions consistent with the evidence. It has been verified by testing against reality, refined when the tests failed, and extended when the tests succeeded. It has been transmitted across centuries and cultures with minimal loss of fidelity, because it is compressed into a notation that is almost perfectly universal. And it has compounded, each generation building on the resting intelligence of the previous one, at a rate that has produced the entire edifice of modern science and technology. That is what good resting intelligence looks like, and it is the standard against which the intelligence we are now building should be measured.&lt;/p&gt;

&lt;p&gt;I do not expect current AI systems to meet that standard immediately. That would be an unreasonable expectation for a technology that is decades old. But I do think the standard should be explicit. The goal of building AI is not to build systems that produce impressive outputs. The goal is to build systems that develop resting intelligence that is genuinely rich, genuinely compressed, genuinely grounded in the structure of the world, and genuinely available to everyone who needs it. That goal requires taking the resting state seriously, not just as infrastructure, not just as an engineering concern about how to store and deploy models, but as the central question of what AI is and what it is for. The motion, the generation, the output, all of that is downstream from the rest. If the rest is good, the motion will be good. If the rest is shallow, no amount of impressive motion will make up for what is missing underneath it.&lt;/p&gt;

&lt;p&gt;I want to close with something that came to me when I was thinking about what connecting all of this to my personal experience feels like. In &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt;, I wrote about how creativity sometimes requires holding still, about how the brush you should not pick up is the one you reach for out of anxiety rather than out of genuine readiness. Intelligence at rest is a form of that same discipline. It is the refusal to rush into motion before the resting structure is ready, before the compression is deep enough, before the understanding is solid enough to support the generation that will come from it. The AI field is, in my view, currently picking up the brush too fast. It is generating before it has rested well. It is moving before it has built the resting structures that would make the motion meaningful. That is not a technical failure alone. It is a kind of impatience that has always been the enemy of depth, and depth is exactly what intelligence at rest requires. The most powerful systems, both human and artificial, are the ones that have learned how to rest before they move. They have learned how to compress before they generate. They have learned that the quality of the silence is what determines the quality of the sound.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Tishby, N., Pereira, F. C., &amp;amp; Bialek, W., &lt;em&gt;The Information Bottleneck Method&lt;/em&gt;, Allerton Conference, 1999. &lt;a href="https://arxiv.org/abs/physics/0004057" rel="noopener noreferrer"&gt;arXiv:physics/0004057&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; Grünwald, P., &lt;em&gt;The Minimum Description Length Principle&lt;/em&gt;, MIT Press, 2007. &lt;a href="https://mitpress.mit.edu/9780262072816/the-minimum-description-length-principle/" rel="noopener noreferrer"&gt;MIT Press&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Hinton, G., Vinyals, O., &amp;amp; Dean, J., &lt;em&gt;Distilling the Knowledge in a Neural Network&lt;/em&gt;, NIPS Deep Learning Workshop, 2014. &lt;a href="https://arxiv.org/abs/1503.02531" rel="noopener noreferrer"&gt;arXiv:1503.02531&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; NIST, &lt;em&gt;Security and Privacy Controls for Information Systems and Organizations&lt;/em&gt; (SP 800-53 Rev. 5), 2020. &lt;a href="https://csrc.nist.gov/publications/detail/sp/800-53/rev-5/final" rel="noopener noreferrer"&gt;NIST SP 800-53&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Carlini, N., Tramèr, F., Wallace, E., et al., &lt;em&gt;Extracting Training Data from Large Language Models&lt;/em&gt;, USENIX Security 2021. &lt;a href="https://arxiv.org/abs/2012.07805" rel="noopener noreferrer"&gt;arXiv:2012.07805&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Allen-Zhu, Z., Li, Y., &amp;amp; Song, Z., &lt;em&gt;A Convergence Theory for Deep Learning via Over-Parameterization&lt;/em&gt;, ICML 2019. &lt;a href="https://arxiv.org/abs/1811.03962" rel="noopener noreferrer"&gt;arXiv:1811.03962&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Blier, L. &amp;amp; Ollivier, Y., &lt;em&gt;The Description Length of Deep Learning Models&lt;/em&gt;, NeurIPS 2018. &lt;a href="https://arxiv.org/abs/1802.07044" rel="noopener noreferrer"&gt;arXiv:1802.07044&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; Frankle, J. &amp;amp; Carbin, M., &lt;em&gt;The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks&lt;/em&gt;, ICLR 2019. &lt;a href="https://arxiv.org/abs/1803.03635" rel="noopener noreferrer"&gt;arXiv:1803.03635&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-9"&gt;&lt;/span&gt;&lt;strong&gt;9.&lt;/strong&gt; Feldman, V. &amp;amp; Zhang, C., &lt;em&gt;What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation&lt;/em&gt;, NeurIPS 2020. &lt;a href="https://arxiv.org/abs/2008.03703" rel="noopener noreferrer"&gt;arXiv:2008.03703&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-10"&gt;&lt;/span&gt;&lt;strong&gt;10.&lt;/strong&gt; Haim, N., Vardi, G., Yehudai, G., et al., &lt;em&gt;Reconstructing Training Data from Trained Neural Networks&lt;/em&gt;, NeurIPS 2022. &lt;a href="https://arxiv.org/abs/2206.07758" rel="noopener noreferrer"&gt;arXiv:2206.07758&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>lmm</category>
      <category>llm</category>
    </item>
    <item>
      <title>I miss the pre-AI Mossad agents.</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 19:13:27 +0000</pubDate>
      <link>https://dev.to/wiseai/i-miss-the-pre-ai-mossad-agents-1ka1</link>
      <guid>https://dev.to/wiseai/i-miss-the-pre-ai-mossad-agents-1ka1</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This post was originally published on &lt;a href="https://wiseai.dev/blogs/i-miss-the-pre-ai-mossad-agents" rel="noopener noreferrer"&gt;the main website&lt;/a&gt; on &lt;a href="https://github.com/wiseaidotdev/blog/pull/22" rel="noopener noreferrer"&gt;Jul 16 2026&lt;/a&gt;. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account &amp;lt;3!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;I keep talking about AI and how it destroyed careers, livelihoods, and entire industries, and most of you who have been reading these posts know by now that I am not being dramatic when I say that. I wrote in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt; about how machines came for our jobs while we were busy celebrating how clever the machines were. I wrote in &lt;a href="https://wiseai.dev/blogs/as-engineers-llms-should-pay-us-for-tokens-usage" rel="noopener noreferrer"&gt;As Engineers, LLMs Should Pay Us for Token Usage&lt;/a&gt; about how we feed these systems everything we know and then get told our skills are no longer needed. And in &lt;a href="https://wiseai.dev/blogs/if-you-cant-build-agi-then-why-should-we-hire-you" rel="noopener noreferrer"&gt;If You Can't Build AGI, Then Why Should We Hire You?&lt;/a&gt; I described the absurd new standard that the industry invented, the one where experience and craft and years of honest work become irrelevant overnight because some model can approximate your output at a fraction of the price. But there is something else I have not yet said out loud, and it has been sitting there in the back of my mind for a while, growing louder with every ignored application and every automated rejection. It is about the people who used to help us before AI arrived. The ones that I have come to call, affectionately and a little darkly, the Mossad agents of the professional world.&lt;/p&gt;

&lt;p&gt;I know that phrase sounds strange. That is actually the point. Bear with me.&lt;/p&gt;

&lt;h2&gt;
  
  
  There Were People On The Other Side
&lt;/h2&gt;

&lt;p&gt;I want you to think back, if you can, to what the inbox felt like before the current AI era. Think about what it felt like to send out a carefully written message to a recruiter, a hiring manager, or a potential collaborator, and then actually receive a reply from a human being who had clearly read what you wrote. That used to happen. Not always, not perfectly, not without friction, but it happened with enough regularity that you could build your professional life around the assumption that a person was on the other side of the conversation. That assumption shaped everything. It shaped how you wrote your messages. It shaped how you presented yourself. It shaped the kind of hope you allowed yourself to carry into the job search, because hope that is connected to a human being feels different from hope that is thrown into an algorithmic void. One you can negotiate with. One you cannot.&lt;/p&gt;

&lt;p&gt;Back then, I used to receive messages from recruiters who had clearly done some form of research. They knew roughly what I worked on. They had some sense of whether the role might fit. They were not always right, and they were not always good at what they did, but there was a person behind the message, and that person had made a choice to reach out, which meant something. When a person chooses to reach out, there is accountability inside that choice. They can be wrong, but they can also be convinced. They can misread your profile, but they can also be corrected. They carry the social weight of the message, which means they have a reason to be genuine. That social weight is not a small thing. It is the entire infrastructure of trust that makes professional communication meaningful, and when you remove it, you do not just make communication faster. You hollow it out.&lt;/p&gt;

&lt;p&gt;The inbox I live in now is not that inbox. It is a space that machines write to and machines filter, where your words are processed by systems that have no concept of what you actually meant, no interest in your history, and no memory of your name. As of 2025, research indicates that up to 99% of Fortune 500 companies use automated Applicant Tracking Systems (ATS) as the primary filter through which every candidate application must pass before any human being sees it (1). That is the complete replacement of the first human layer of professional judgment with a system that was trained on historical patterns and cannot, by its own fundamental design, understand context, exception, or nuance.&lt;/p&gt;

&lt;p&gt;And in a strange, roundabout way, it made me realize that I miss what I used to have. I miss the invisible helpers. I miss the ones working behind the scenes in ways I barely understood. I miss, in short, the pre-AI Mossad agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Mean By Mossad Agents, And Why The Metaphor Works
&lt;/h2&gt;

&lt;p&gt;Let me explain the phrase, because I genuinely mean it as a term of affection and respect, and the fact that it sounds unusual is part of why it captures the feeling so well. The Mossad is Israel's national intelligence agency, founded in 1949, and its reputation rests not on brute force or visible power but on something more unusual: precision, patience, access to information you did not know anyone had, and the ability to act on that information with quiet, confident effectiveness. The Mossad's most celebrated operations, the ones that became part of intelligence folklore, rely on people who had cultivated deep networks over years, who understood not just facts but motivations, relationships, and the invisible human architecture that determines how things actually get done. They worked with context. They worked with nuance. They did not just match patterns. They understood people.&lt;/p&gt;

&lt;p&gt;Now, when I say Mossad agents of the professional world, I mean the recruiters, the connectors, the HR professionals who actually paid attention, who read your work, who remembered your name, who could see past an unconventional career path to the actual human capability underneath. They were just people doing their job with enough care to look at you as a person rather than a profile. I saw them as Mossad agents because they seemed to know things about the professional landscape that you did not know yourself. They reached out before you reached out. They had access to opportunities that were never posted publicly. They made introductions that opened doors you did not know existed. They were the human intelligence layer of the professional world, and they operated with the kind of quiet precision that the actual Mossad would have appreciated. And now, as I will argue in this post, they have been progressively replaced by a very different kind of agent, one made of weights and parameters and trained to optimize a score rather than understand a person.&lt;/p&gt;

&lt;p&gt;There is something worth noting here about the real history of intelligence and AI, because it mirrors what happened in hiring in ways that are not accidental. The Israeli intelligence community, which includes the Mossad but also the military intelligence corps Unit 8200, has undergone a dramatic transformation over the last decade that is directly relevant to this conversation. Unit 8200, which produced many of Israel's most prominent technology entrepreneurs and is considered one of the elite signals intelligence organizations on Earth, shifted heavily toward algorithmic processing, AI-assisted surveillance, and what Israeli military strategists began calling "algorithmic warfare" (2). The Mossad itself, despite its legendary reputation for human intelligence, or HUMINT in the tradecraft vocabulary, increasingly integrated AI-driven pattern recognition and data fusion into its operations (3). In the world of intelligence, this shift was sold as efficiency, capability multiplication, and scale. In the world of hiring, it was sold as the same things. The language is identical. The structural change is identical. And the loss, I will argue, is also identical.&lt;/p&gt;

&lt;p&gt;The tragedy is not that AI entered these spaces. The tragedy is what it displaced in the process of entering them. Because what it displaced was not inefficiency. What it displaced was judgment. And judgment is not a feature you can recover by adding more parameters to the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Intelligence Failure You Were Never Supposed to Notice
&lt;/h2&gt;

&lt;p&gt;There is one event in recent history that made the cost of replacing human intelligence with algorithmic intelligence impossible to deny, and that event is October 7th, 2023. I want to talk about it carefully, because it is not a political argument I am trying to make here, it is a structural one. On that day, Hamas carried out an attack that caught Israeli intelligence, widely considered among the most sophisticated surveillance and intelligence systems on Earth, almost entirely by surprise (4). The question that followed, asked by analysts, intelligence professionals, and researchers across the world, was: how does one of the most technically advanced intelligence operations in the world fail to detect an attack that had been planned openly, physically rehearsed, and communicated through channels that, in retrospect, were not particularly hidden?&lt;/p&gt;

&lt;p&gt;The answer that emerged from serious analysis is relevant to everything I am talking about in this post. The failure was not primarily a failure of technology. Israel had the sensors, the satellites, the SIGINT capability, the algorithmic processing infrastructure. The failure was a failure of human interpretation. It was a failure of the kind of judgment that only comes from the kind of intelligence that no algorithm has ever been able to replicate: the understanding of human intention, social context, and the invisible signals that travel not through digital channels but through relationships, attitudes, and the thousand small behavioral tells that a trained human observer notices and a machine does not (5). Intelligence analysts frequently assess threats by evaluating an adversary's &lt;em&gt;intent&lt;/em&gt; versus their &lt;em&gt;capability&lt;/em&gt;. SIGINT and algorithmic analysis are very good at measuring what an adversary can do, but genuinely poor at understanding what an adversary intends to do, especially when the adversary is disciplined enough to operate below the digital radar (6).&lt;/p&gt;

&lt;p&gt;This is the same gap that exists in automated hiring. AI screening systems can measure what a candidate claims to be able to do, but they cannot assess what a candidate actually intends, how they think under pressure, whether they are the kind of person who rises to a challenge or crumbles when the plan falls apart. These things are not in the resume. They are not in the keywords. They are not in the structured data fields that the algorithm can parse. They live in the conversation, in the follow-up, in the story the candidate tells when asked about a project that went wrong. They live in the human layer. And when you remove the human layer and replace it with a scoring system trained on historical patterns, you do not get a faster version of what you had. You get a fundamentally different thing, one that is efficient, yes, but also blind to the most important dimension of what you were trying to evaluate.&lt;/p&gt;

&lt;p&gt;The intelligence community is slowly reckoning with this, at great cost. The hiring community has not yet faced a comparable reckoning, but I believe it is coming. And when it does, people will look back at this era the same way analysts look back at the pre-October 7th over-reliance on algorithmic surveillance: not as a winning technology, but as a cautionary tale about what happens when you let systems replace judgment rather than support it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ATS Machine That Ate Your Future
&lt;/h2&gt;

&lt;p&gt;Let me stop being abstract for a moment, because the abstract argument is important but the personal one is what makes it real. I have been on the receiving end of automated rejection. I have sent applications into inboxes that were not inboxes at all but funnels, designed to filter before any human being has the chance to exercise any kind of judgment. I have watched my resume get processed by Applicant Tracking Systems, those algorithmic gatekeepers that every major employer now uses as the primary wall between a candidate and a human conversation, and I have learned exactly what it feels like to be scored by a machine that has no concept of what my actual work means. The score does not care that I built things that worked. It cares whether the words in my resume match the words in the job description with sufficient frequency and in the expected pattern. That is pattern matching dressed up as evaluation, and it is an insult to anyone who has spent years developing real capability.&lt;/p&gt;

&lt;p&gt;The research on this is damning, and I want to cite it specifically because I am tired of people treating my frustration as personal failure. A 2021 study commissioned by Harvard Business School and Accenture, titled "Hidden Workers: Untapped Talent," found that millions of highly qualified people are systematically excluded from consideration not because they lack skills but because they fail algorithmic screening criteria that are poorly designed, overly rigid, or trained on historically biased hiring data (7). The study coined the phrase "hidden workers" to describe this large, invisible population of skilled people who exist in the labor market but are algorithmically invisible because their profiles do not fit the expected template. Former Army veterans who cannot translate their experience into civilian resume language. Caregivers who have employment gaps. People who changed fields. People with non-standard educational backgrounds. People, and I want to be very specific here, like me.&lt;/p&gt;

&lt;p&gt;The World Economic Forum's Future of Jobs Report 2025 documents that AI is now directly involved in screening, ranking, and filtering candidates at an industrial scale, and that the adoption rate has accelerated dramatically since 2022 (8). The IMF has noted that roughly forty percent of global employment is exposed to AI displacement or transformation, and that the most vulnerable are often those in skilled but non-credentialed roles, those who built their expertise through practice rather than formal pathways (9). I wrote about this vulnerability in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt;, where I described what it feels like to have your capabilities be invisible to the very systems that are supposed to connect capability with opportunity. What I did not have at the time was the language to name what had changed. Now I do. What changed was the disappearance of the human middle layer. What disappeared were the Mossad agents.&lt;/p&gt;

&lt;p&gt;Let me be precise about what the ATS machine actually does, because most people who have not applied for jobs recently do not understand the technical reality. A modern Applicant Tracking System does not read your application. It parses your application. It extracts tokens from your resume, matches them against a database of required keywords derived from the job description, weights them according to recency and frequency and position in the document, and produces a score. If your score falls below a threshold, your application is archived. No human being sees it. No human being makes that decision. The decision is made by a trained model that has never spoken to you, does not know what you meant by what you wrote, and cannot distinguish between someone who listed "machine learning" because they read a tutorial last week and someone who has spent three years building production systems. The score is a rank, not a judgment, and a rank without context is just noise pretending to be information.&lt;/p&gt;

&lt;h2&gt;
  
  
  What The Mossad Understood That The Machine Never Will
&lt;/h2&gt;

&lt;p&gt;There is a concept in intelligence tradecraft called HUMINT, Human Intelligence, and it refers specifically to the kind of information that can only be gathered through direct human contact, through relationships, through conversations that go off the record, and through the observation of behavior in unstructured situations (10). The Mossad's historical effectiveness as an intelligence organization was built on an exceptionally well-developed HUMINT capability, on the ability to cultivate sources across different cultures and contexts, to assess human motivation with precision, and to make judgment calls in real time that no algorithm, then or now, could replicate. The legendary operations that made Mossad famous, from tracking down Nazi war criminals to dismantling weapons programs in hostile states, were achieved not by processing more data but by understanding more people.&lt;/p&gt;

&lt;p&gt;What the intelligence world learned over decades, at great cost, is that HUMINT and SIGINT, Signals Intelligence, the technical intercept of communications and data, are not substitutes for each other. They serve different functions. SIGINT tells you what is happening in the observable digital world. HUMINT tells you what is happening in the human world beneath the digital one (11). A purely SIGINT-dependent operation sees patterns but misses motivation. It sees movement but misses intent. It sees capability but misses will. The research on this published in academic intelligence journals consistently shows that the most serious intelligence failures, including the one on October 7th, are failures of HUMINT capacity, not of technological capability. Organizations that invested heavily in algorithmic surveillance while deprioritizing the cultivation of human sources found themselves with vast amounts of data and a poverty of understanding (6).&lt;/p&gt;

&lt;p&gt;The parallel to the hiring world is exact. A recruiting process that is purely algorithmic, that depends entirely on ATS scoring and automated screening, sees keywords but misses capability. It sees listed experience but misses demonstrated judgment. It sees pattern matches but misses potential. The research supports this. Studies on AI bias in hiring, including work published by the AI Now Institute, MIT Media Lab, and the EEOC, have documented that automated systems systematically undervalue candidates from underrepresented groups, candidates with non-linear career paths, and candidates whose language usage does not match the dominant patterns of the training data (12). The system is not finding the best candidates. It is finding the most algorithmically legible candidates, and those are not the same group.&lt;/p&gt;

&lt;p&gt;What the pre-AI Mossad agent understood, what that human recruiter understood who used to look at your whole story and make a judgment call, was that the most interesting candidates are often the least legible ones. The ones who do not fit the template are frequently the ones who learned to think outside of it, and thinking outside of the template is exactly the skill that most hard problems require. The machine cannot see this, because the machine was trained to reward the template. And so the person who spent three years building something unusual and important, and cannot describe it in five standard keywords, gets filtered out before any human being ever has the chance to ask them to explain their work. That is not efficiency. That is waste dressed up as efficiency, and the difference between the two is a human judgment that a machine cannot make.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Quiet Death of the Human Middle Layer
&lt;/h2&gt;

&lt;p&gt;I do not want to romanticize the past. The pre-AI era had its own serious problems. Hiring was unfair. Networking was exclusionary. Doors that should have been open were often locked, and the keys were held by people who distributed them based on familiarity rather than merit. The old system had corruption, nepotism, and structural inequality embedded in it, and the people who lived that reality would be right to remind me that the Mossad agents I am mourning were not always the heroes of everyone's story. Some of them were gatekeepers who kept the wrong gates locked for the wrong reasons. I acknowledge that, and I do not want to pretend otherwise.&lt;/p&gt;

&lt;p&gt;But acknowledging the flaws of the old system does not require pretending that the new system is an improvement. It is possible to have had a flawed human layer and to still recognize that the removal of the human layer made things worse, not better, for most people. The flawed human layer could at least be argued with. It could be reasoned with. It could be moved by a compelling story, a strong reference, an unexpected recommendation. It contained within it the possibility of exception, of the judgment call that goes against the statistical mean. The algorithmic layer contains no such possibility. It does not argue. It does not listen. It does not make exceptions. It computes a score, and the score is the answer, and the answer is final in a way that no human judgment has ever been final, because human judgment always carries within it the seed of revision.&lt;/p&gt;

&lt;p&gt;The ILO has published research noting that the introduction of AI into hiring does not just affect job quantity but job quality and the qualitative experience of job seeking itself (13). What this means in practice is that the process of looking for work has become a fundamentally more dehumanizing experience than it used to be. Not because there are fewer jobs, necessarily, but because the experience of applying for them now involves mostly interacting with systems that do not see you. This is a deep harm that is difficult to quantify but that everyone who has lived it knows viscerally. There is a particular kind of exhaustion that comes not from trying and failing but from trying and not even being seen. The automated rejection is not a rejection. It is a non-event. The system did not reject you in any meaningful sense. It just never processed you as a human being in the first place.&lt;/p&gt;

&lt;p&gt;I described what this feels like in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt;: that specific grief of doing everything right and still having nothing to show for it, of sending your best effort into a void and hearing only static. That grief is partly the grief of an unfair system. But it is also, I now realize, the grief of a system that removed the people who were supposed to translate between what you have and what the world needs. The Mossad agents are gone. And in their place is a score.&lt;/p&gt;

&lt;h2&gt;
  
  
  Genuine Intelligence vs Genuine Artificial Intelligence
&lt;/h2&gt;

&lt;p&gt;I want to connect this to something broader, because everything I write is connected to everything else, and the thread that runs through all of these posts is the same thread. I have argued in &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt; that language models are useful tools trapped inside a symbolic cage, that they can describe the world without understanding it, and that the difference between describing and understanding is not a technical detail but the fundamental distinction between a system that processes patterns and a system that grasps meaning. I have argued in &lt;a href="https://wiseai.dev/blogs/genuine-intelligence-will-never-emerge-from-neural-networks" rel="noopener noreferrer"&gt;Genuine Intelligence will Never Emerge from Neural Networks&lt;/a&gt; that the architecture of current AI systems is structurally incapable of producing the kind of grounded, embodied, contextual understanding that characterizes real intelligence. I have argued in &lt;a href="https://wiseai.dev/blogs/knowledge-and-intelligence-are-mutually-exclusive" rel="noopener noreferrer"&gt;Knowledge and Intelligence Are Mutually Exclusive&lt;/a&gt; that storing knowledge is not the same as being able to use it wisely, that wisdom requires something that no training dataset can provide.&lt;/p&gt;

&lt;p&gt;The Mossad agents understood all of this intuitively, without needing to articulate it. They understood that a candidate is not their resume. They understood that a profile is not a person. They understood that the most important things about someone's professional capability often cannot be extracted from structured data, because they live in the way the person thinks, the way they respond under pressure, the way they handle being wrong. This is genuine intelligence: the ability to model another human being as a complex, contextual, unpredictable entity rather than as a set of features to be scored. And it is precisely this kind of genuine intelligence that artificial intelligence, despite all its impressive capabilities, has never come close to achieving and cannot achieve within its current architectural paradigm.&lt;/p&gt;

&lt;p&gt;What we replaced those Mossad agents with is not even a good imitation of what they could do. We replaced them with a much more rudimentary process, a keyword matching engine that has been dressed up in machine learning language to sound sophisticated. The sophistication is real at the technical level: the models behind modern ATS systems involve genuine deep learning, genuine NLP, genuine optimization over large datasets. But sophistication at the technical level does not translate into capability at the human level. A very sophisticated tool for the wrong job is still the wrong tool. And pattern matching at scale, no matter how technically impressive, is not a substitute for judgment, because the things that matter most in a human being cannot be pattern-matched from the surface.&lt;/p&gt;

&lt;p&gt;The research on hallucination in large language models is relevant here, because it exposes a structural truth that applies just as forcefully to the algorithmic filtering of candidates. LLMs hallucinate because they have no access to ground truth. They only have access to statistical patterns in their training data, and those patterns do not contain reality, they contain shadows of reality that humans have already cast into text (14). Automated screening systems are built on the same foundation. They have access to statistical patterns in historical hiring data, and those patterns do not contain the ground truth of who is a good candidate. They contain the shadow of past hiring decisions, which were made by humans who had their own biases, their own limited perspectives, and their own structural inequalities. Training a system to reproduce those patterns does not fix the biases. It scales them. And scaling a bias is not the same as neutralizing it. It is the same as industrializing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What We Have Lost and What It Would Take to Recover It
&lt;/h2&gt;

&lt;p&gt;I want to be honest about the difficulty of what I am suggesting here. I am not saying that we should abolish AI in hiring, that we should go back to purely manual processes, that we should pretend that the volume problem does not exist. When a single job posting receives 5000 applications, no team of human recruiters can read 5000 resumes with genuine attention. The volume is real, and some form of filtering is necessary. I understand that. I am not arguing against filtering. I am arguing against the specific way we are doing it now, which is to replace the human judgment layer entirely rather than augment it, and to do so in ways that are systematically unfair to the people who are already most vulnerable.&lt;/p&gt;

&lt;p&gt;What would a better system look like? I think it would look like something closer to the original Mossad agent model, where technology handles the genuinely mechanical parts of the process, the scheduling, the logistics, the initial data aggregation, and human judgment handles the qualitatively demanding parts, the ones that require understanding context, interpreting stories, and making the kind of exception that no statistical model can make by definition. The World Economic Forum's research on the future of work emphasizes the importance of what they call "human-centered AI," systems designed to support human decision-making rather than replace it, and that distinction is not a semantic one (8). It describes a fundamentally different relationship between machine and human, one where the machine handles volume and the human handles judgment, rather than one where the machine handles everything and the human is removed from the loop entirely.&lt;/p&gt;

&lt;p&gt;The EEOC in the United States has begun to issue guidance making clear that employers are legally responsible for discrimination caused by their AI screening tools, regardless of whether the discrimination was intentional (15). New York City Local Law 144, passed in 2021 and implemented in 2023, requires employers to conduct annual independent bias audits of any AI tools used in hiring and to notify candidates when such tools are used (16). These regulatory moves are important, but they are also insufficient, because they address the symptoms of the problem, the discriminatory outcomes, without addressing the structural cause, which is the complete removal of human judgment from the first and most consequential layer of the hiring process.&lt;/p&gt;

&lt;p&gt;The Mossad agents did not give us their talent for free, and neither should we expect the algorithmic descendants to give us discernment for free. Discernment costs something. It costs time, attention, relationship, and the willingness to be surprised by a human being who does not fit the expected pattern. We have decided, as an economy, that those costs are too high. And we are paying for that decision in the currency of invisible people, hidden workers, and an inbox that no longer feels like a place where opportunities live.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing Words
&lt;/h2&gt;

&lt;p&gt;I know how this post sounds to people who have never been on the wrong side of an ATS filter. I know it sounds like sour grapes, like a person who cannot accept that the world has changed, like someone who is simply nostalgic for a world that was imperfect in its own ways. I want to answer that directly. I am not opposed to the world changing. I have written extensively about change, about technology, about the deep transformations that are coming for every industry and every profession, and I have not run away from any of those arguments. I can simultaneously believe that AI will fundamentally reshape the structure of intelligence, as I argued in &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt;, and believe that the way we are currently deploying AI in hiring is causing real harm to real people, and that both of those things are true at the same time.&lt;/p&gt;

&lt;p&gt;The Mossad agents of my professional past were not perfect. But they were present. They were in the conversation. They could see me, and when they could not see me clearly, I could speak and they could adjust. That capacity for adjustment, for being reached by a human voice, is what I am mourning. Not the whole system. Not the old inefficiencies. Not the gatekeeping. Just the presence of a person who had enough skin in the game to actually look. The world felt different when I believed that someone was looking. It felt like participation in a shared enterprise rather than submission to a sorting mechanism. And the loss of that feeling is not trivial, no matter how many efficiency metrics say otherwise.&lt;/p&gt;

&lt;p&gt;Until then, I will keep thinking about those invisible agents who used to operate in the shadows of the professional world, who knew things you did not know, who made moves you did not expect, who connected the right person to the right opportunity without a training dataset and without a loss function. I miss their precision. I miss their access. I miss their judgment. And I miss, above all, the idea that a person had looked at my work and decided that it mattered. That is not a rejection of progress. That is a defense of the human middle layer. And the human middle layer, I am increasingly convinced, was the most important layer we had.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Industry research on ATS adoption (2024-2025) indicates up to 99% of Fortune 500 companies use automated screening; see also: Joseph B. Fuller, Manjari Raman et al., &lt;em&gt;Hidden Workers: Untapped Talent&lt;/em&gt;, Harvard Business School &amp;amp; Accenture, 2021. &lt;a href="https://web.archive.org/web/20230531232840/https://www.hbs.edu/managing-the-future-of-work/Documents/research/hiddenworkers09032021.pdf" rel="noopener noreferrer"&gt;hbs.edu via Internet Archive (PDF)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; IISS (International Institute for Strategic Studies), &lt;em&gt;The Proliferation of AI-Enabled Military Technology in the Middle East&lt;/em&gt;, April 2026. &lt;a href="https://www.iiss.org/online-analysis/charting-middle-east/2026/04/the-proliferation-of-ai-enabled-military-technology-in-the-middle-east/" rel="noopener noreferrer"&gt;iiss.org (Online Analysis)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Calcalist Tech, reporting on Unit 8200 veterans in Israel's AI ecosystem, 2026. &lt;a href="https://www.calcalistech.com/ctechnews/article/7ui00iuvr" rel="noopener noreferrer"&gt;calcalistech.com (Article)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; CSIS, Emily Harding et al., &lt;em&gt;Experts React: Assessing the Israeli Intelligence and Potential Policy Failure&lt;/em&gt;, October 25, 2023. &lt;a href="https://www.csis.org/analysis/experts-react-assessing-israeli-intelligence-and-potential-policy-failure" rel="noopener noreferrer"&gt;csis.org&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; RAND Corporation, &lt;em&gt;Why Human Intelligence Matters More in an AI World&lt;/em&gt;, 2026. &lt;a href="https://www.rand.org/pubs/commentary/2026/06/why-human-intelligence-matters-more-in-an-ai-world.html" rel="noopener noreferrer"&gt;rand.org (Commentary)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Ehud Eiran, &lt;em&gt;Techno-Digital Vulnerability and Intelligence Failures&lt;/em&gt;, &lt;em&gt;Social Sciences&lt;/em&gt; 15(1):37, MDPI, January 2026. &lt;a href="https://www.mdpi.com/2076-0760/15/1/37" rel="noopener noreferrer"&gt;mdpi.com (Article)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Joseph B. Fuller, Manjari Raman et al., &lt;em&gt;Hidden Workers: Untapped Talent&lt;/em&gt;, Harvard Business School &amp;amp; Accenture, 2021. &lt;a href="https://web.archive.org/web/20230531232840/https://www.hbs.edu/managing-the-future-of-work/Documents/research/hiddenworkers09032021.pdf" rel="noopener noreferrer"&gt;hbs.edu via Internet Archive (PDF)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; World Economic Forum, &lt;em&gt;Future of Jobs Report 2025&lt;/em&gt;. &lt;a href="https://www.weforum.org/publications/the-future-of-jobs-report-2025/" rel="noopener noreferrer"&gt;weforum.org&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-9"&gt;&lt;/span&gt;&lt;strong&gt;9.&lt;/strong&gt; Mauro Cazzaniga, Florence Jaumotte et al., &lt;em&gt;Gen-AI: Artificial Intelligence and the Future of Work&lt;/em&gt;, IMF Staff Discussion Note SDN/2024/001, January 2024. &lt;a href="https://www.imf.org/en/publications/staff-discussion-notes/issues/2024/01/14/gen-ai-artificial-intelligence-and-the-future-of-work-542379" rel="noopener noreferrer"&gt;imf.org (Staff Discussion Note)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-10"&gt;&lt;/span&gt;&lt;strong&gt;10.&lt;/strong&gt; Perseus Intelligence, &lt;em&gt;The Evolution of HUMINT since World War Two&lt;/em&gt;, 2024. &lt;a href="https://www.perseusintelligence.co.uk/the-evolution-of-humint-since-world-war-two" rel="noopener noreferrer"&gt;perseusintelligence.co.uk (Report)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-11"&gt;&lt;/span&gt;&lt;strong&gt;11.&lt;/strong&gt; Gregory F. Treverton and C. Bryan Gabbard, &lt;em&gt;Assessing the Tradecraft of Intelligence Analysis&lt;/em&gt;, RAND Technical Report TR-293, 2008. &lt;a href="https://www.rand.org/content/dam/rand/pubs/technical_reports/2008/RAND_TR293.pdf" rel="noopener noreferrer"&gt;rand.org (PDF)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-12"&gt;&lt;/span&gt;&lt;strong&gt;12.&lt;/strong&gt; Sarah Myers West, Meredith Whittaker &amp;amp; Kate Crawford, &lt;em&gt;Discriminating Systems: Gender, Race and Power in AI&lt;/em&gt;, AI Now Institute, April 2019. &lt;a href="https://ainowinstitute.org/wp-content/uploads/2023/04/discriminatingsystems.pdf" rel="noopener noreferrer"&gt;ainowinstitute.org (PDF)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-13"&gt;&lt;/span&gt;&lt;strong&gt;13.&lt;/strong&gt; International Labour Organization, &lt;em&gt;Generative AI and Jobs: A Global Analysis of Potential Effects on Job Quantity and Quality&lt;/em&gt;, Working Paper 96, 2023. &lt;a href="https://www.ilo.org/global/publications/working-papers/WCMS_890761/lang--en/index.htm" rel="noopener noreferrer"&gt;ilo.org&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-14"&gt;&lt;/span&gt;&lt;strong&gt;14.&lt;/strong&gt; Yue Zhang, Yafu Li, Leyang Cui et al., &lt;em&gt;Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models&lt;/em&gt;, 2023. &lt;a href="https://arxiv.org/abs/2309.01219" rel="noopener noreferrer"&gt;arXiv:2309.01219&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-15"&gt;&lt;/span&gt;&lt;strong&gt;15.&lt;/strong&gt; U.S. Equal Employment Opportunity Commission (EEOC), &lt;em&gt;Artificial Intelligence and Algorithmic Fairness Initiative&lt;/em&gt;, launched 2021, updated 2023. &lt;a href="https://www.eeoc.gov/newsroom/eeoc-launches-initiative-artificial-intelligence-and-algorithmic-fairness" rel="noopener noreferrer"&gt;eeoc.gov (Initiative Launch)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-16"&gt;&lt;/span&gt;&lt;strong&gt;16.&lt;/strong&gt; New York City Local Law 144 on Automated Employment Decision Tools (AEDT), enacted 2021, enforcement began July 5, 2023. &lt;a href="https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page" rel="noopener noreferrer"&gt;nyc.gov&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>lmm</category>
    </item>
    <item>
      <title>If You Can't Build AGI, Then Why Should We Hire You?</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 19:07:40 +0000</pubDate>
      <link>https://dev.to/wiseai/if-you-cant-build-agi-then-why-should-we-hire-you-b87</link>
      <guid>https://dev.to/wiseai/if-you-cant-build-agi-then-why-should-we-hire-you-b87</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This post was originally published on &lt;a href="https://wiseai.dev/blogs/if-you-cant-build-agi-then-why-should-we-hire-you" rel="noopener noreferrer"&gt;the main website&lt;/a&gt; on &lt;a href="https://github.com/wiseaidotdev/blog/pull/21" rel="noopener noreferrer"&gt;May 14 2026&lt;/a&gt;. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account &amp;lt;3!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;It is fascinating how a single job interview question can encapsulate the defining neurosis of an entire industry. For months, I have been watching the engineering world collectively panic over a new filter masquerading as a baseline requirement: &lt;em&gt;if you aren't actively building artificial general intelligence, why exactly are you here?&lt;/em&gt; While my piece &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt; explored the isolation of holding deep technical competence in a market that no longer knows how to classify it, and &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt; examined the systemic devaluation of foundational engineering, this essay zeroes in on the bizarre metric that sits at the intersection of both. The current start-up discourse, propelled heavily by venture capital narratives, has conflated AGI affiliation with fundamental technical worth. I am writing this to systematically dismantle that conflation. Answering the question requires naming some uncomfortable truths about what we are actually building verses what we are selling, but continuing to let the wrong metrics drive our careers is an error we cannot afford to sustain.&lt;/p&gt;

&lt;p&gt;Before I develop the argument, I want to say something about the specific shape of this anxiety, because it is not random, and understanding where it comes from matters for evaluating whether it deserves the authority it has been given. The last 3 years have produced a sustained cultural pressure that says artificial general intelligence is imminently arriving, that it will make most existing engineering work obsolete, and that the only engineers who will survive the transition are those who are actively building toward it or deeply integrated with the systems that are closest to it. This pressure comes from real places, from genuine capability jumps in language models, from the economic dominance of companies like OpenAI and Anthropic and Google DeepMind, and from the completely understandable anxiety of an engineering workforce watching automation creep up the skills ladder. But it is also being amplified by people who benefit from that anxiety, venture capitalists who need the narrative of impending disruption to justify frontier investments, startup founders who need to recruit at below-market rates by selling the dream of being part of something historical, and AI companies that need a cultural context in which the only legitimate technical work is work that feeds their systems. I am not saying those people are lying. I am saying that the incentives around a narrative matter for how that narrative gets shaped, and the incentives around AGI are enormous.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://wiseai.dev/blogs/genuine-intelligence-will-never-emerge-from-neural-networks" rel="noopener noreferrer"&gt;Genuine Intelligence Will Never Emerge from Neural Networks&lt;/a&gt;, I argued at length that the current direction of AI development, scaled statistical learning over human-generated data, is architecturally incompatible with genuine intelligence in any sense that word deserves to carry. In &lt;a href="https://wiseai.dev/blogs/knowledge-and-intelligence-are-mutually-exclusive" rel="noopener noreferrer"&gt;Knowledge and Intelligence Are Mutually Exclusive&lt;/a&gt;, I argued that having access to information is not the same thing as understanding it, and that the conflation of those two things is the central mistake driving both AI hype and the misplaced anxiety around it. In &lt;a href="https://wiseai.dev/blogs/all-you-have-access-to-is-knowledge-and-tools-never-intelligence" rel="noopener noreferrer"&gt;All You Have Access To Is Knowledge and Tools; Never Intelligence&lt;/a&gt;, I argued that every AI system ever deployed has given its users knowledge and tools but not intelligence, because intelligence is the thing the human brings to the interaction. All of those arguments feed directly into the question in this post's title, because if genuine intelligence is not what AI systems have, and if genuine intelligence is not what companies can buy in the form of a language model API, then the engineer who can use their own genuine intelligence to build systems that work reliably and solve real problems is still the most valuable person in any technical organization. The question is not AGI or obsolescence. The question is whether you can build things that matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Question Is Designed to Make You Feel Small
&lt;/h2&gt;

&lt;p&gt;Let me say this plainly, because it is the most important thing in the entire post and the thing I most wanted to say when I started writing it. The question "if you can't build AGI, why should we hire you?" is not a technical question. It is a psychological one. It is the kind of question that only functions if you already accept the frame that makes it feel threatening, and that frame is this: the only work that counts is work at the frontier of artificial general intelligence, and everything else is just maintenance work that will be automated away. That frame is wrong on multiple levels, and I want to dismantle it carefully because it has been causing real damage to real engineers who are doing genuinely valuable work but who have been made to feel that their contributions are somehow insufficient by a cultural context that has decided only one thing counts. The question is designed to make you feel small, and it has been working.&lt;/p&gt;

&lt;p&gt;The first thing wrong with the frame is that AGI does not have a settled definition, which means you cannot be evaluated against it as a standard. OpenAI says AGI is "AI systems that are generally smarter than humans" (1). DeepMind's Morris et al. defined it as a six-level hierarchy running from Level 0 (No AI) through Level 1 (Emerging), Level 2 (Competent), Level 3 (Expert), Level 4 (Virtuoso), and Level 5 (Superhuman), assessing both the depth of performance and the breadth of tasks a system can handle (2). MIT's 2025 AI Agent Index says explicitly that definitions of AI agents are "nebulous and differ across fields" and that even the word "agent" lacks a stable technical meaning across the research community (3). This is not a minor disagreement. These are the central organizations in the field, and they cannot agree on what AGI is, whether we have it, how close we are to it, or what empirical tests would determine whether a system qualifies. When a hiring question invokes a concept that its own field cannot define, it is not measuring something real. It is invoking a feeling, and feelings are not hiring rubrics. The person asking whether you can build AGI is asking whether you are affiliated with something prestigious and frontier-sounding, which is a question about social proof, not technical capability.&lt;/p&gt;

&lt;p&gt;The second thing wrong with the frame is that companies do not need AGI. Companies need solutions to specific problems in specific domains with specific constraints. They need systems that work reliably within a budget, that can be maintained by the people who will inherit them, that handle failure gracefully, and that produce measurable value that justifies their cost. None of these requirements are AGI requirements. A customer support routing system that correctly classifies 92% of tickets and escalates the other 8% appropriately does not need to be generally intelligent. It needs to be correct, maintainable, and deployable in the production environment the company already has. An internal knowledge retrieval system that helps a legal team find relevant case precedents faster does not need human-level general reasoning. It needs to retrieve the right documents, surface them clearly, and handle the failure modes that arise when relevant documents do not exist. These are real engineering problems, and they require real engineering skill, but they require engineering skill, not AGI-building ability. The conflation of the two is how companies end up hiring people based on narrative affiliation rather than on the actual skills the job requires.&lt;/p&gt;

&lt;p&gt;The third thing wrong with the frame is that the people who ask the AGI question in interviews are almost never building AGI either. They are building features on top of API calls to large language models. They are building wrappers, prompts, evaluation pipelines, fine-tuning scripts, and retrieval systems. That is real work and some of it is genuinely difficult and genuinely valuable. But it is not AGI research, and calling it AGI research or demanding that candidates show AGI-building ability as a prerequisite for doing it is a form of credential inflation that benefits no one except the cultural positioning of the company doing the hiring. A frontend engineer does not need to have invented the rendering engine to build excellent interfaces. A backend engineer does not need to have written the database to design excellent schemas. The layer abstraction that makes software engineering productive depends on people using tools well without pretending to have built them from scratch. The AGI question inverts this by demanding that candidates prove mastery of a layer that does not yet exist in any usable form.&lt;/p&gt;

&lt;p&gt;The psychological function of the question matters as much as its technical content, and I want to be honest about what that function is. The question filters out engineers who feel uncertain, who are not good at performing confidence in things they cannot verify, or who are honest about the limits of their knowledge. It selects for a certain kind of boldness that is correlated with self-promotion but not reliably correlated with technical ability. I described in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt; what it feels like to have genuine technical capability and not be able to perform the social confidence that the market rewards. The AGI question is a social performance test wearing a technical costume, and the engineers it selects are the engineers who are best at performing the right kind of confidence, not necessarily the ones who will build the most reliable systems or solve the hardest real-world problems. That is a bad hiring mechanism that produces bad outcomes, and I want to say so directly.&lt;/p&gt;

&lt;p&gt;There is also a historical pattern worth noting here, because the AGI question did not emerge in a vacuum. Every period of technological transition produces a version of the question that positions the new technology as the only thing that matters and asks engineers to prove their alignment with it as the price of admission to the profession. In the early cloud era, you had to prove you understood distributed systems at scale even if you were building a simple CRUD application. In the mobile era, you had to prove you thought mobile-first even if your product was never going to be used on a phone. In the blockchain era, which I will not dwell on for the sake of everyone's blood pressure, you had to convince hiring managers that everything was about decentralization. The AGI version of this question is the latest in that sequence, and it will eventually be replaced by the next version, after enough companies have hired for AGI affiliation and discovered that what they actually needed was engineers who could ship reliable systems and communicate clearly with non-technical stakeholders. The cycle is predictable. The damage it does in the meantime is real.&lt;/p&gt;

&lt;p&gt;The engineers who actually built the systems that powered the cloud era's most important products were not the engineers who wrote the seminal distributed systems papers. They were the engineers who read those papers, understood them well enough to apply the relevant parts to a specific problem, and built systems that worked well enough to ship and be maintained. The cloud era needed understanding of distributed systems principles, not the ability to invent new distributed systems theory from scratch. The AI era needs the same thing: engineers who understand the tools well enough to apply them appropriately, who can evaluate the outputs, who can design systems around the limitations, and who can maintain what they build as the underlying tools evolve. That is what the actual job requires. That is the hiring rubric that produces good engineering organizations. The AGI question is a distraction from that rubric, and expensive organizations have already started to figure that out.&lt;/p&gt;

&lt;h2&gt;
  
  
  AGI Is a Moving Target
&lt;/h2&gt;

&lt;p&gt;I want to spend some time on what AGI actually means, because the term is used so loosely and so frequently in the current conversation that it has become almost content-free while retaining strong emotional charge. That combination, strong feeling, weak definition, is exactly the combination that makes a term useful for rhetorical purposes and dangerous for technical ones. When I wrote &lt;a href="https://wiseai.dev/blogs/rethinking-arc-agi" rel="noopener noreferrer"&gt;Rethinking ARC-AGI&lt;/a&gt;, I tried to explain why even the most prominent empirical attempt to define and measure AGI, François Chollet's ARC benchmark, was not actually measuring what it claimed to measure, because it was measuring pattern recognition under specific constraints rather than the open-ended adaptive intelligence that the AGI label implies. That argument stands, and it has gotten more support since I wrote it, not less. The definition of AGI has been drifting in the research literature not toward consensus but away from it, with each major lab defining it in a way that happens to be aligned with their own products and roadmaps.&lt;/p&gt;

&lt;p&gt;The definitional drift is not accidental. It serves a real function in the competitive landscape of AI research and development. If your company can credibly claim to be on the path to AGI, you have an enormous advantage in recruiting, in fundraising, in media attention, and in policy influence. That creates strong incentives to define AGI in a way that your current work looks like progress toward, and weak incentives to hold to a rigorous definition that might show your current work to be further from the goal than it appears. OpenAI's current framing places AGI as "AI systems that are generally smarter than humans", which is broad enough to interpret very generously as models get more capable (1). DeepMind's Morris et al. framework defines six levels of performance and generality, from Level 1 (Emerging) up through Competent, Expert, Virtuoso, and Level 5 (Superhuman), which allows them to place current systems at the low end of the spectrum as "Emerging" AGI without committing to any claim about whether higher levels are imminent or even achievable (2). Anthropic tends to avoid the term entirely and talks about "powerful AI" instead, which is a different rhetorical choice but comes from the same place of definitional flexibility. None of these framings give you a clear, falsifiable prediction about what would count as AGI being achieved, which means none of them give you a clear, falsifiable hiring criterion.&lt;/p&gt;

&lt;p&gt;The World Economic Forum's 2025 Future of Jobs Report is actually more useful for thinking about what the current period of AI development means for hiring than any of the lab-specific AGI framings, because it is written from the perspective of what organizational capability needs look like rather than from the perspective of what researchers hope to achieve (4). The report says that nearly 40 percent of skills required on the job are expected to change by 2030, that 63 percent of employers already cite the skills gap as a barrier to transformation, and that the skills employers are planning to grow are not exclusively technical ones. They include analytical thinking, resilience and flexibility, leadership and social influence, and creative thinking alongside AI and big data, technology literacy, and networks and cybersecurity. That is a very different picture from "hire for AGI fluency." It is a picture of organizations that need people who can think, adapt, communicate, and lead through technical change, not people who can write papers about general intelligence.&lt;/p&gt;

&lt;p&gt;The International Labour Organization's 2025 update on generative AI and occupations makes an even more important corrective point, which is that very few jobs are fully automatable with current generative AI, and that the predominant pattern is task transformation rather than job replacement (5). That distinction matters more than almost anything else in the current conversation, because it repositions the question from "will AI replace this job?" to "which parts of this job will AI change, and what will the human do with the time freed by those changes?" The answer, in most knowledge-work domains, is that humans will do more of the high-judgment, high-context, high-stakes parts of the work, the parts that require understanding a specific organization's situation, that require building trust with specific people, and that require taking responsibility for outcomes in a way that no automated system can take. Those are the parts of the job that AGI cannot do even in principle given the current architectural limitations I described in &lt;a href="https://wiseai.dev/blogs/genuine-intelligence-will-never-emerge-from-neural-networks" rel="noopener noreferrer"&gt;Genuine Intelligence Will Never Emerge from Neural Networks&lt;/a&gt;. The jobs are changing, yes. They are not being replaced by AGI. They are being changed by tools that free up humans to do more of what humans are actually good at.&lt;/p&gt;

&lt;p&gt;The specific claim that METR's 2025 study makes is worth examining in detail, because it is one of the most counterintuitive and important findings in the recent literature on AI and developer productivity (6). The study found that experienced open-source developers using the best available AI coding tools as of early 2025 took roughly 19 percent longer on their own repositories compared to a control condition without those tools. This is not a finding that says AI tools are useless. It is a finding that says the relationship between tool capability and productivity is not linear, and that in expert contexts with high contextual depth, the overhead of correctly prompting, reviewing, and integrating AI-generated outputs can exceed the time saved by not generating those outputs manually. That is a real phenomenon, and it has enormous implications for how organizations should think about AI tool adoption, about what skills to hire for, and about what productivity metrics to use when evaluating engineers who use these tools. The best candidate is not the one who uses AI most aggressively. The best candidate is the one who understands when to use it and when not to, what its outputs need in terms of review, and how to fit it into a workflow that makes the whole system better rather than just the individual task faster.&lt;/p&gt;

&lt;p&gt;The definitional instability of AGI also has direct consequences for how companies should think about technical leadership and strategy. A company that builds its hiring and product strategy around hitting AGI as a milestone is building on sand, because the milestone keeps moving and no independent observer can verify when it has been reached. A company that builds its hiring and product strategy around solving specific, measurable problems for specific customers in specific domains is building on solid ground, because those goals are clear, progress toward them is verifiable, and success is recognizable when it happens. I have been saying something like this since &lt;a href="https://wiseai.dev/blogs/announcing-kevin-rs" rel="noopener noreferrer"&gt;Announcing Kevin RS&lt;/a&gt;, where I described the design philosophy behind a framework built for specificity and reliability rather than for generality and hype. The argument has not changed. Specificity is how you build something real. Generality is how you build something impressive-sounding.&lt;/p&gt;

&lt;p&gt;The Turing test is the oldest version of the AGI measurement problem, and it is worth pausing on because its history illustrates exactly the issue with building hiring criteria around an underspecified goal. Alan Turing proposed the imitation game in 1950 as a way of sidestepping the hard philosophical question of machine consciousness by substituting a behavioral test (7). The insight was genuinely brilliant: instead of asking whether a machine can think, ask whether its behavior is indistinguishable from a thinking entity. But the problem that became apparent over the following seven decades is that the Turing test measures the convincingness of outputs rather than the presence of understanding behind them, and those are not the same thing. Systems that have passed versions of the Turing test under specific conditions have done so by being good at producing the right kind of text, not by being genuinely intelligent. The same problem applies to AGI definitions based on behavioral outcomes: they optimise for producing the right impressions rather than for having the right internal structure. And at the level of hiring, optimizing for impression-making is exactly the failure mode I am describing throughout this post.&lt;/p&gt;

&lt;p&gt;What the AGI frame has done to technical culture, and this is the thing that bothers me most about it, is shift the primary evaluation criterion from "can you build something that works" to "do you sound like someone who could theoretically build something that would eventually work." That is a shift from evidence to narrative, and narrative has always been cheaper to produce and harder to evaluate than evidence. I wrote in &lt;a href="https://wiseai.dev/blogs/as-engineers-llms-should-pay-us-for-tokens-usage" rel="noopener noreferrer"&gt;As Engineers, LLMs Should Pay Us for Tokens Usage&lt;/a&gt; about how engineers are the ones who create real value in AI systems while other actors extract that value through the framing of the market. The AGI hiring question is the same extraction in a different domain: it takes the genuine capability and judgment of working engineers and subordinates it to a narrative about frontier intelligence that few people in the industry are actually building and fewer still are building honestly. The resistance I am mounting in this post is the same resistance I mount in my code: insist that the thing works, insist that you can measure it, and refuse to accept impressive language as a substitute for demonstrated capability.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Market Is Not Broken
&lt;/h2&gt;

&lt;p&gt;I have heard too many engineers describe the current job market as completely broken, and while I understand where that feeling comes from and I have felt it myself in ways I described in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt;, I want to offer a more precise diagnosis, because "broken" suggests randomness and the current market is not random. It is re-sorting. It is re-sorting around a specific set of properties that determine who survives and who does not in an environment where AI tools have changed what can be accomplished by an individual contributor and where the ceiling on what a small team can ship has risen dramatically in a short period. Understanding the re-sorting logic is more useful than describing the whole situation as broken, because understanding it tells you what to actually do.&lt;/p&gt;

&lt;p&gt;The primary axis of re-sorting is between engineers who can demonstrate specific, verifiable outcomes and engineers who can only describe their capability in terms of their familiarity with fashionable tools or their proximity to prestigious institutions. That distinction has always existed in engineering hiring, but it has become much more visible and consequential in the last three years because the tooling has changed what the floor for competent-sounding output looks like. An engineer who can produce a reasonably coherent architecture diagram and a working prototype using language model tools has raised the floor for what "showing up with something" looks like in a technical interview or a project pitch. That means the bar is not in the appearance of competence but in the demonstration of genuine judgment, specifically the judgment about what to build, how to constrain it, how to verify it, and how to maintain it when the first version inevitably has problems. Those things are harder to fake than they have ever been, because the fakers are now using the same tools as the practitioners and the outputs look similar from the outside, which means you have to go deeper to see the difference.&lt;/p&gt;

&lt;p&gt;The World Economic Forum data on this is worth reading carefully rather than treating as a bumper sticker (4). The 2025 report says employers are planning to hire for AI and big data skills, but it also says they expect to grow human capabilities like resilience, flexibility, curiosity and lifelong learning, and leadership and social influence. It says 85 percent of employers plan to prioritize upskilling their existing workforce in thinking and working with AI, which means the companies that are doing this right are not replacing human judgment with AI; they are augmenting human judgment with AI and investing in the humans who use that judgment well. The organizations that are just replacing humans with AI and calling the result an AI-forward strategy are in a different category, and I will say plainly that I think most of them will discover in two to three years that they degraded their organizational capability by removing the judgment that the tools lack. That is not a blind pro-human claim. It is a claim grounded in what AI systems can and cannot do at the architectural level, and I have made that argument in enough technical detail in previous posts that I will not repeat it all here.&lt;/p&gt;

&lt;p&gt;The re-sorting is also happening along the axis of specialization versus generalism. For most of the last decade, the engineering job market rewarded generalists who could move between stacks, adapt to new frameworks quickly, and be productive across a wide range of problem domains. That premium on generalism was driven by the rapid pace of framework and tool change, which made deep specialization in any single technology risky because the technology might be irrelevant in three years. AI tools have changed this dynamic by raising the floor for what a generalist can produce, which means the premium has shifted toward people who have genuine depth in a specific domain, because that depth is what distinguishes their use of AI tools from a shallow user's use of the same tools. The cardiologist who knows the clinical literature deeply is in a very different position when using AI diagnostic tools than the medical student who has passed the basic examinations. Both can ask the AI the same questions. Only the cardiologist can evaluate the answers against a causal model of cardiac function built from years of clinical experience. That depth is irreplaceable by scale, and domain-specific depth is what engineers should be building right now, not because it sounds prestigious but because it is what makes their use of new tools genuinely better than the competition's use of the same tools.&lt;/p&gt;

&lt;p&gt;The international dimension of this re-sorting matters too, because engineering talent is globally distributed in a way that the hiring practices of large Western technology companies have never fully accommodated and that AI tools are now making more complicated. I have been building and thinking from outside the geographic and institutional centers of the AI industry, and I wrote in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt; about what it is like to have serious technical capability and no institutional affiliation, no network connection, and no runway to perform indefinite unpaid proof of work. The market re-sorting I am describing does not automatically benefit people in that position, because verifiable outcomes are still verified by networks, and networks are still unevenly distributed by geography and institution. But the tools have lowered some of the barriers to producing the evidence in the first place, which means the task is building evidence that is verifiable by someone who does not already know you, and that is exactly the kind of portfolio, the shipped product, the open-source project with real users, the measurable improvement documented in public, that replaces the network effect when you do not have one.&lt;/p&gt;

&lt;p&gt;The thing I want to say most directly about the re-sorted market is this: the engineers who are struggling the most are not struggling because AI has made their skills obsolete. They are struggling because the market is in an adjustment period where the signals that used to identify good engineers, years of experience, recognizable employer brands, familiarity with popular frameworks, have been disrupted faster than new signals have emerged to replace them. In that vacuum, AGI affiliation has rushed in as a cultural signal, and it is doing the job poorly because it measures narrative proximity rather than engineering capability. The engineers who will come out of this period in a strong position are the ones who did not wait for the market to figure out the right signal but instead built their own evidence. They shipped things. They measured them. They wrote about what they learned. They made their judgment visible in forms that a thoughtful hiring manager in any technical organization could evaluate and recognize. That is what I am trying to do with this blog, with the projects I am building, and with the framework I am going to spend several full sections of this post describing.&lt;/p&gt;

&lt;p&gt;The METR study I mentioned earlier has another dimension that is worth bringing out, which is that the developers who did not see productivity losses from AI tools were the ones who had developed specific workflows for using them, specific evaluation habits, specific constraints on when to accept AI-generated code and when to reject it (6). That is a skill. It is not a skill that is visible from a resume headline or easily measured in a forty-five minute technical interview. It is a craft practice that builds over time with deliberate effort, and the engineers who have it are more valuable than the engineers who can pass the interview filter on AGI familiarity but have not developed the operational judgment to use AI tools well in a production context. The hiring practices that will survive this period are the ones built around evaluating that operational judgment rather than the ones built around evaluating AGI affiliation, and the organizations that get there first will be the ones with engineering workforces that are genuinely more productive and not just nominally more AI-native.&lt;/p&gt;

&lt;p&gt;One more thing on the market, and then I want to move to something more constructive. The layoffs that have accompanied the re-sorting are real and they have hurt real people, including people I know and people whose situations resemble things I have lived through. But the layoffs are not primarily driven by AI making engineering work unnecessary. They are primarily driven by a period of over-hiring during 2020 to 2022 that was financed by near-zero interest rates and the specific conditions of the COVID period, followed by a normalization period as interest rates rose and the cost of capital became real again. AI is a contributing factor at the margin in some specific roles, but the dominant cause of the current dislocation is interest rate normalization and valuation correction, not AGI. Saying otherwise serves the narrative interests of people who want AI to seem more disruptive than it currently is, and I have spent enough time in this series resisting narratives that serve interests other than the truth to stay consistent about it here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everyone Has the Same Tools
&lt;/h2&gt;

&lt;p&gt;I hear this claim constantly, and I want to deal with it directly because it is one of the most consequential pieces of misinformation circulating in technical culture right now. The claim is that the democratization of AI tools has leveled the playing field, that because everyone has access to GPT-5 or Claude or Gemini through an API, the advantage that experienced engineers used to have from their accumulated skill and domain knowledge has been eroded or eliminated. This claim is wrong in a way that is both easily demonstrated and deeply consequential for how engineers should think about their own value. Access to a tool and competence with a tool are completely different things, and the speed at which access has democratized has been confused with a democratization of competence that has not actually occurred. I want to explain exactly why, because I think this is one of the places where the AI conversation has gone most seriously off track.&lt;/p&gt;

&lt;p&gt;The easiest demonstration of the gap between access and competence is the evaluation problem. When an engineer uses an AI tool to generate a piece of code, a data analysis, a system design, or a legal summary, the tool produces output. That output may be correct, partially correct, subtly wrong, confidently wrong, or genuinely excellent. The tool itself cannot reliably tell you which of these it is, because as I argued in &lt;a href="https://wiseai.dev/blogs/genuine-intelligence-will-never-emerge-from-neural-networks" rel="noopener noreferrer"&gt;Genuine Intelligence Will Never Emerge from Neural Networks&lt;/a&gt;, these systems are optimized for producing statistically likely outputs conditional on their training data, not for producing outputs whose correctness tracks the structure of reality. The engineer who has the domain expertise to evaluate the output can use the tool on a completely different level than the engineer who does not. The expert can accept the parts that are right, identify the parts that are wrong, and either fix them or avoid accepting them in the first place. The non-expert accepts the whole thing and ships it. Both engineers have equal access to the tool. Their outputs are completely different, and the difference is entirely determined by the human knowledge the engineer brings to the interaction, not by anything the tool does differently for one versus the other.&lt;/p&gt;

&lt;p&gt;This is also the argument I made in &lt;a href="https://wiseai.dev/blogs/all-you-have-access-to-is-knowledge-and-tools-never-intelligence" rel="noopener noreferrer"&gt;All You Have Access To Is Knowledge and Tools; Never Intelligence&lt;/a&gt;, and I want to connect it here explicitly because it goes to exactly this point. What AI systems provide is precisely knowledge extracted from their training data and tools for manipulating that knowledge in ways that match statistical patterns from that training. What they do not provide is the capacity to reason about whether their outputs are correct in the specific context of your specific problem. That capacity has to come from somewhere external to the tool, and the only available source is the engineer using it. The engineer whose domain knowledge is deep enough to evaluate AI outputs critically is multiplied in capability by the tool. The engineer whose domain knowledge is shallow enough that they cannot evaluate AI outputs is given a false sense of capability by the tool, and that false sense is dangerous in exactly the proportion that the tool's outputs are wrong without the user noticing. Equal access does not mean equal outcome. It means equal exposure to a multiplier that amplifies existing capability differences.&lt;/p&gt;

&lt;p&gt;There is a very specific kind of knowledge that matters most for using AI tools well, and it is not the kind of knowledge that appears on a resume or that can be tested in a standard technical interview. It is the knowledge of where the tool fails. Every domain where AI tools are deployed has characteristic failure modes, places where the statistical patterns in the training data diverge from the causal structure of the real problem, where the tool produces confident-sounding but wrong outputs, where the failure is subtle enough to pass an initial review but consequential enough to produce real problems downstream. An experienced engineer in that domain knows those failure modes from having seen them, from having been bitten by them, and from having developed habits and workflows that prevent them from propagating. That knowledge is the functional difference between a dangerous AI user and a productive one, and it is hard-won from domain experience that cannot be shortcut by API access. The person who gave you the API key has not given you the decade of domain experience that makes the API safe to use on problems where being wrong has real consequences.&lt;/p&gt;

&lt;p&gt;Let me connect this to something very concrete from my own work. When I am building an agent pipeline using our &lt;a href="https://github.com/wiseaidotdev/autogpt" rel="noopener noreferrer"&gt;Rusty autoGPT framework&lt;/a&gt;, the most important decisions I make are not about which model to call or how to structure the prompts. The most important decisions are about what the agent is not allowed to do, what outputs it must have verified by a human before they are acted on, what happens when the LLM call fails or returns something that does not parse, and how the system signals uncertainty in a way that a human operator can act on appropriately. Those decisions require understanding both the specific domain of the problem and the specific failure modes of the AI components being used, and they are decisions that a user with only API access and no domain depth cannot make well. The democratization of access does not democratize those decisions. It democratizes the ability to build something that looks like it is making those decisions while actually leaving them unaddressed.&lt;/p&gt;

&lt;p&gt;The historical analogy I find most clarifying here is the spreadsheet. The introduction of VisiCalc and then Lotus 1-2-3 and then Excel democratized access to financial modeling. Before those tools, you needed a team of accountants to build complex financial projections. After them, every manager with a computer could do it. But the democratization of access to financial modeling tools did not democratize financial modeling competence. It democratized the ability to produce financial models that looked credible without necessarily being correct. The history of corporate finance is full of disasters produced by managers who had full access to Excel and no real understanding of the financial structures they were modeling, who built confident-looking spreadsheet models that had fundamental logical errors or unrealistic assumptions baked into their structure. The accountants who were displaced from some modeling tasks did not become useless because everyone had Excel. They became more valuable as evaluators of the models that Excel made easy to build, because someone with real financial expertise was needed to tell the managers which of their models had errors they could not see. AI tools are the Excel of the current moment, and the engineers with domain expertise are the accountants. Access has shifted; expertise has not.&lt;/p&gt;

&lt;p&gt;The argument that equal tool access produces equal outcomes also fails on the evidence of what actually happens when engineers use AI tools in practice. Empirical research on code generation tools like GitHub Copilot shows a stark disconnect between user expectation and actual experience (8). While programmers overwhelmingly prefer using AI tools because they provide useful starting points, the studies show this access does not actually improve task completion time or success rates. Instead, users struggle profoundly to understand, edit, and debug the generated code. A pattern emerges where users over-rely on the system without sufficiently validating its outputs, decreasing code quality because they fail to recognize when the AI has generated something that passes a surface syntax check but fails a deeper correctness check. The gap in outcomes is not driven by access to the tool—everyone has that. It is driven by the depth of domain knowledge the user brings to debugging and evaluating the tool's outputs. That means the right policy response to AI democratization is not "everyone has the tools, so expertise is less valuable" but exactly the reverse: "everyone has the tools, but only people with deep expertise can use them safely, so deep expertise is more valuable than before."&lt;/p&gt;

&lt;p&gt;The frame I want to offer as a replacement for the "same tools" narrative is this: AI tools raise the ceiling on what an expert can accomplish and lower the floor on what bad work looks like. Both of those things are true simultaneously, and they have opposite implications for the value of expertise. The expert is more productive because the tool handles tedious routine work and the expert can focus on the high-judgment parts. The non-expert produces better-looking bad work because the tool adds surface polish to something that is still fundamentally wrong. From the outside, these two things look more similar than they used to, because the surface quality of the non-expert's output has improved. But the divergence in underlying correctness has not narrowed. If anything, it has widened, because the tool allows the non-expert to get further down the wrong path before the fundamental error becomes apparent. Organizations that are evaluating engineers on the surface quality of their AI-assisted output rather than on the depth of judgment behind it are making the most expensive version of the mistake I am describing, and they are going to pay for it in production incidents, in maintenance costs, and in the inevitable moment when the AI-generated system encounters a situation outside its training distribution and produces a confident wrong answer that no one on the team has the domain depth to catch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Makes an Engineer Irreplaceable
&lt;/h2&gt;

&lt;p&gt;I want to be concrete here rather than philosophical, because the philosophical dimension of this argument is well covered by the other sections and what engineers actually need is a clear sense of what they should be building and demonstrating. In my view, there are five properties that make an engineer genuinely irreplaceable in a technical organization in the current environment, and none of them are AGI-building ability and none of them are simple AI-tool fluency. They are properties that have always distinguished great engineers from merely competent ones, but they have become more visible and more economically important now because the AI tools have made the floor of apparently competent output easier to reach, which means the properties above the floor matter more for differentiation, not less.&lt;/p&gt;

&lt;p&gt;The first property is problem selection. There is an enormous difference between an engineer who walks into a situation and builds the system that was asked for and an engineer who walks into the same situation, looks at what was asked for, and identifies the problem that actually needs to be solved, which is often different from the one being specified. Problem selection is the highest-leverage skill in engineering, because a perfect solution to the wrong problem is worse than an imperfect solution to the right one, and AI tools have no capacity for problem selection at all. They optimize for producing outputs that match the specified input, which means that if the specification is wrong, the output will be confidently wrong in a way that matches the wrong specification exactly. The engineer who asks "is this the right problem to solve before we optimize for how to solve it" is the engineer who saves organizations from building sophisticated solutions to the wrong problem, and that engineer is irreplaceable because no tool can substitute for the judgment about what problem is actually worth solving.&lt;/p&gt;

&lt;p&gt;The second property is domain depth, which I have already discussed at length in the context of tool evaluation, but which also matters for a different reason: communication with domain experts and stakeholders who are not technical. An engineer who understands the business domain they are working in can translate technical constraints and tradeoffs into language that business stakeholders can engage with, can understand what the business stakeholders are actually asking for beneath the technical specification they have been given, and can build trust with those stakeholders through demonstrable understanding of their actual problems. This translation capacity is part of what makes engineering a business function rather than a purely technical one, and it is entirely independent of which AI tools you use. The AI can help you build the system once you understand what to build, but it cannot help you figure out what to build, because that requires the human connection between technical capability and domain understanding.&lt;/p&gt;

&lt;p&gt;The third property is execution speed with quality, by which I mean the ability to move from specification to working system efficiently without accumulating technical debt that slows down future work. This is not the same as raw output speed. An engineer who ships a working feature quickly that is poorly structured, hard to test, and difficult to maintain has not demonstrated this property, even if the speed looked impressive in the short term. An engineer who ships a working feature at a moderate pace but with clear structure, good tests, and obvious extension points has demonstrated it. AI tools can increase raw output speed but they cannot enforce quality, and they often undermine it when used by engineers who lack the judgment to distinguish code that works from code that works well. The engineers who demonstrate this property consistently are invaluable because organizations can trust what they ship in a way they cannot trust the output of engineers whose speed is amplified by AI tools they are not yet competent to evaluate.&lt;/p&gt;

&lt;p&gt;The fourth property is communication, and I mean this in the full sense that I described in &lt;a href="https://wiseai.dev/blogs/as-engineers-llms-should-pay-us-for-tokens-usage" rel="noopener noreferrer"&gt;As Engineers, LLMs Should Pay Us for Tokens Usage&lt;/a&gt;: the ability to create accurate shared understanding between yourself, your team, and your stakeholders about what you are building, why you are building it the way you are, what the risks are, and what the limits of the system are. This requires technical precision, the language skills to express technical ideas clearly to different audiences, and the honesty to be accurate about uncertainty and limitations rather than projecting confidence that is not warranted. The engineers who can do this well are enormously more valuable than the ones who cannot, because most of the failure modes in software development are not technical failures. They are communication failures where the engineers understood the technical situation but could not translate that understanding into shared understanding with the people who needed to make decisions based on it.&lt;/p&gt;

&lt;p&gt;The fifth property is repeatability, by which I mean the ability to turn a one-off success into a process that can be relied upon and replicated. The engineer who solves a hard problem once is valuable. The engineer who solves a hard problem once and then documents what they learned, identifies the general pattern that the specific solution instantiated, and builds tools or processes that make future solutions to similar problems faster and more reliable is exponentially more valuable. This is the property that distinguishes artisan craftspeople from engineers in the meaningful sense, that engineers build systems while artisans solve cases. AI tools can help with individual cases efficiently, but they cannot build the institutional knowledge and repeatable processes that allow organizations to scale their technical capability beyond a small team of experts. The engineer who builds those processes is building organizational capability in a way that compounds over time, and that kind of compounding value is what justifies seniority in any technical organization.&lt;/p&gt;

&lt;p&gt;These five properties, problem selection, domain depth, execution speed with quality, communication, and repeatability, are also the properties that are hardest to fake and easiest to evaluate by looking at a record of actual work rather than a performance in a technical interview. A portfolio of shipped systems, documented performance improvements, and written explanations of design decisions demonstrates all five of these properties in a way that passing an LeetCode-style interview cannot. I want to say directly that the shift toward portfolio-based evaluation that some technical hiring organizations have been making is a genuine improvement over the coding interview paradigm, not because coding interviews are entirely useless but because they measure a narrow technical skill disconnected from most of the properties I just described. An engineer who can pass every coding interview but produces brittle, hard-to-maintain systems is more visible and more consistently filtered out by portfolio evaluation than by interview performance. The organizations that have made this shift are making a better decision, and I predict more organizations will make it as the current tooling makes the gap between interview performance and production performance more apparent.&lt;/p&gt;

&lt;p&gt;There is also a sixth property I almost left off the list because it feels obvious but is actually quite rare, and that is epistemic honesty. The engineer who says "I am not sure about that, let me check" when they do not know something is more valuable than the engineer who gives a confident wrong answer, because the wrong answer has to be corrected and the correction takes more time and organizational capital than the original uncertainty acknowledgment would have. The engineer who says "this system has the following limitations that you should be aware of before deploying it in this context" is more valuable than the engineer who says "it works great" and leaves the limitations to be discovered in production. The AI tools are systematically overconfident, as I described in the neural networks post, which means the engineers who compensate for that overconfidence by being more careful about acknowledging uncertainty are providing a genuine epistemic service to their organizations, not just a soft skill. Epistemic honesty is what turns AI-assisted development from a risk amplifier into a productivity amplifier, and the organizations that develop cultures where epistemic honesty is rewarded rather than punished are the ones that will capture the real value from AI tools without being destroyed by their failure modes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Rust AutoGPT Framework Changes What Is Possible
&lt;/h2&gt;

&lt;p&gt;Now I want to talk about something I have been building, something I think represents a serious answer to the question of what it looks like to build real AI agent systems with the kind of reliability that the previous sections have been demanding. &lt;a href="https://github.com/wiseaidotdev/autogpt" rel="noopener noreferrer"&gt;AutoGPT&lt;/a&gt; is not another Python wrapper around an OpenAI API call. It is a pure Rust framework for building autonomous AI agents, and the decision to build it in Rust rather than in Python is not an aesthetic preference. It is a principled engineering decision that reflects exactly the arguments I have been making throughout this post, about what serious builder behavior looks like when you apply genuine judgment to the question of which tools and languages serve the actual goals of production-quality agentic software.&lt;/p&gt;

&lt;p&gt;Let me explain why Rust specifically, because the language choice matters more for agentic systems than it does for most other categories of software. An agent that operates autonomously makes decisions and takes actions without human approval at every step. That means errors in the agent's logic, memory management, or concurrency do not surface as an immediate human-visible bug and get fixed before they cause harm. They propagate through the agent's action pipeline and produce downstream consequences that may be difficult or impossible to reverse. Memory safety bugs in particular, the kind that Rust eliminates at compile time through its ownership and borrowing system, are especially dangerous in agent software because an agent that has a use-after-free vulnerability or a data race in its action execution pipeline is an agent that can take arbitrary wrong actions in ways that are not predictable from the agent's intended logic. Python does not have memory safety guarantees of this kind. Python manages memory through reference counting and a garbage collector, which means memory safety bugs of certain classes cannot occur, but Python's dynamic typing and the absence of compile-time guarantees mean that entire categories of logical and type errors that Rust catches at compile time only surface in Python at runtime, which in an agent system means they surface when the agent is running in a production environment and potentially taking irreversible actions.&lt;/p&gt;

&lt;p&gt;The performance argument for Rust in agent systems is real and it compounds over time in ways that Python's dynamic dispatch does not allow. When you have an agent that needs to process streaming outputs from multiple LLM calls, coordinate between multiple sub-agents, manage a vector memory store, and maintain state across a long-running task, the overhead of Python's interpreter, its global interpreter lock (GIL) which prevents true parallelism in CPU-bound code, and its dynamic dispatch mechanism begin to create genuine latency and throughput limitations that no amount of async optimization can fully compensate for. Rust's zero-cost abstractions and fearless concurrency, meaning the ability to write truly parallel code that the compiler guarantees is free of data races, allow agent systems built in it to run at a fundamentally different performance tier. This matters not as a vanity metric but because latency in agent orchestration compounds: an agent that runs a ten-step pipeline where each step has additional Python overhead versus one where that overhead is absent completes meaningfully faster, and in production systems where agents are running continuously on behalf of users or organizations, that latency difference translates directly into cost, responsiveness, and the feasibility of certain deployment patterns.&lt;/p&gt;

&lt;p&gt;The type system argument is the one I find most compelling from a pure engineering perspective. Rust's type system, with its algebraic data types, exhaustive pattern matching, and the absence of null pointers through the Option type, forces you to handle all the failure modes that Python code often leaves unhandled because handling them is optional and handling them explicitly takes more code. In an agent system, the failure modes are numerous and they interact: an LLM call can fail with a network error, or return a response that does not parse, or return a parseable response that does not have the expected structural properties, or return a structurally valid response that is semantically wrong in a domain-specific way. Handling these failure modes in Python typically means adding defensive code that can be omitted, that is more verbose than it looks from the happy path, and that tends to drift out of sync with the actual error conditions over time. Handling them in Rust is not optional. The type system enforces exhaustive handling of error variants, which means the wiring between agent steps in the framework fails to compile rather than failing at runtime if you have not addressed a failure mode. That compile-time enforcement is the difference between a system that handles errors when the developer remembered to add error handling and a system that handles errors because the language itself requires it before the code can be executed.&lt;/p&gt;

&lt;p&gt;Our AutoGPT Rust framework's architecture itself reflects these principles in its design. The framework's agent core is built from composable components including tools and sensors that interface with the real world via actions and perceptions, memory and knowledge that combines long-term vector memory with structured knowledge bases for reasoning and recall, a planner and goals module that breaks down complex tasks into subgoals and tracks progress dynamically, and a self-reflection module for introspection that can debug, adapt, or evolve internal strategies. The Mixture of Providers (MoP) feature, which allows parallel fan-out and weighted scoring across multiple AI backends, is exactly the kind of architecture that is far easier to implement correctly in Rust than in Python, because the parallel execution and the score aggregation each have concurrency requirements that Rust's ownership system makes easy to reason about and that Python's GIL and threading model make genuinely difficult to get right without data races or deadlocks. The YAML-based no-code agent configuration is a deliberate choice to make the framework accessible to builders who understand agents conceptually but who should not have to modify the orchestration layer to define what a specific agent does and how it behaves.&lt;/p&gt;

&lt;p&gt;The comparison with the Python-based AutoGPT framework that has received enormous attention since 2023 is instructive precisely because both projects share a name and a general concept, autonomous AI agents that can break down tasks, use tools, and self-direct toward a goal, while representing radically different engineering approaches. The Python AutoGPT accumulated users and attention quickly because Python is the lingua franca of the AI tooling ecosystem and the barrier to getting started is very low. But it also accumulated a reputation for unreliability, for being difficult to deploy in production, and for the kinds of runtime errors that only surface in production because Python's dynamic type system cannot catch them earlier. Memory management in long-running Python agent processes is a known challenge because Python's garbage collector is not designed for the patterns of intense, short-lived object creation and destruction that characterize LLM-integrated agent pipelines. The Rusty AutoGPT framework makes the opposite tradeoff: a steeper initial learning curve in exchange for a system that the compiler has vetted for safety and efficiency before it runs a single line of agent logic.&lt;/p&gt;

&lt;p&gt;The four operating modes that the framework supports, interactive (GenericGPT), direct prompt, standalone agentic, and distributed agentic (orchestrated), represent a thoughtful progression from simple to complex that allows builders to start with something understandable and scale to something powerful without abandoning the framework. The interactive mode, where the agent responds to individual prompts with appropriate tool use, is the right starting point for understanding what the agent architecture provides over a direct API call. The standalone agentic mode, where the agent pursues a multi-step goal without human input at each step, is where the real architectural value of the framework becomes apparent, because this is where the difference between a well-typed, memory-safe orchestration layer and a dynamically-typed one starts producing measurable outcomes in system reliability. The orchestrated mode, where multiple agents form a network coordinated by an orchestrator, is where the true frontier of agent system design lives currently, and where the concurrency guarantees of Rust provide the most decisive architectural advantage over Python-based alternatives.&lt;/p&gt;

&lt;p&gt;There are nine built-in specialized autonomous agents in the current release, which is not a large number but it is enough to demonstrate the pattern of what a specialized, reliable agent looks like in practice. Specialization matters here for the same reason I argued earlier in this post that domain depth matters for engineers: a specialized agent has a narrower scope of action, a more clearly defined set of tools it is allowed to use, and a more specific definition of what success looks like. Those constraints make it easier to evaluate the agent's behavior, easier to catch failures before they propagate, and easier to trust the system in a production context where reliability is not optional. A general-purpose agent that can do anything is impressive in a demo and dangerous in a production system, because the broader the scope of action, the more ways there are to fail, and the harder it is to predict or constrain. The framework's emphasis on specialized agents is an engineering philosophy, not a limitation, and it is the right philosophy for the current state of agent technology.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building Something Narrow, Deep, and Real
&lt;/h2&gt;

&lt;p&gt;The most important practical insight I want to leave in this post is also the simplest one: the best systems are not the most general ones. The best systems are the ones that do the right specific thing reliably enough to be trusted and deployed in the specific context where they are actually useful. I have been saying versions of this across multiple posts in this blog, in &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs Are Useful. LMMs Will Break Reality&lt;/a&gt;, where I argued that the specific, constrained use of language model tools is where their value is most cleanly realized, and in &lt;a href="https://wiseai.dev/blogs/rethinking-arc-agi" rel="noopener noreferrer"&gt;Rethinking ARC-AGI&lt;/a&gt;, where I argued that narrow competence is more demonstrable and more verifiable than general intelligence and therefore more reliable as a basis for building systems. The principle holds even more strongly when you are building agentic systems, because agents that can do anything are agents that can do anything wrong, and broader scope without proportionally better judgment is just a larger blast radius for failures.&lt;/p&gt;

&lt;p&gt;What does it actually look like to build something narrow, deep, and real? Let me be concrete, because abstractly advocating for narrowness is easy and the real discipline is in the specifics. A narrow agent for a legal firm's document review process might be designed to do exactly one thing: classify incoming discovery documents into one of five predefined categories, flag documents that have characteristics associated with privilege in the firm's specific practice area, and produce a structured output for each document that includes the classification, the flagging decision, and the three most relevant passages that support each decision. That is a narrow scope. The agent knows what it is for, what it is not for, what tools it can use and what it cannot, what the output structure must be, and what happens when the classification confidence is below a threshold. Every one of those constraints is a design decision that requires domain expertise. The person who designed the five categories had to understand discovery practice well enough to make the categories meaningful and mutually exclusive. The person who defined privilege characteristics had to understand the firm's specific practice area. The person who designed the confidence threshold had to understand how reviewers actually use the tool and what the cost of a false flag is compared to the cost of a miss.&lt;/p&gt;

&lt;p&gt;None of those design decisions are made by the agent. All of them are made by the engineer who builds it. The agent executes the design with speed and consistency that a human reviewer cannot match. The engineer is the one who guarantees that the design is correct, that the constraints are the right ones, and that the failure modes are handled appropriately. That division of labor is exactly right, and it is the division of labor that serious agent system design should be preserving and strengthening rather than trying to collapse into a single general agent that does everything and guarantees nothing. The Rust AutoGPT framework is built around exactly this principle: agents are composable, specialized, and defined through declarative YAML configurations that make the scope of action explicit and auditable. You do not have to modify the orchestration layer to define what an agent does. You describe the agent's persona, its goals, its tools, and its constraints in a configuration file, and the framework executes that description with the reliability guarantees that Rust provides. That is how serious builders work: they specify the system's behavior explicitly and rely on the runtime to execute it reliably.&lt;/p&gt;

&lt;p&gt;The audibility of agent behavior is the piece that separates serious deployments from demo-ware, and I want to spend some time on it because it is systematically undervalued in the current cultural conversation around agents. An audible system is one where you can look at exactly what the agent decided, why it decided it based on what inputs, and what actions it took in response. Without that audibility, you cannot diagnose failures, you cannot demonstrate compliance with regulatory requirements, you cannot build organizational trust in the system, and you cannot improve the system over time in a principled way because you do not have a clear record of where it is currently failing and how. The agent index research found that many deployed agent systems lack transparent documentation of their safety-relevant behavior, which is a polite way of saying that many deployed agents are black boxes and nobody can fully explain why they do what they do in production (3). That is a state of affairs that should be unacceptable in any system that is making decisions with real consequences, and it is a state of affairs that serious builders should be actively refusing to accept as normal.&lt;/p&gt;

&lt;p&gt;The question of what a human must approve before an agent acts is the most important design question in any agentic system, and it is almost never answered well in the systems I see described in conference talks and blog posts. The instinct is to minimize human involvement because human involvement reduces the speed and autonomy that make agents valuable in the first place. But the right answer is not "minimize human involvement" as a goal. The right answer is "identify carefully which decisions are reversible and which are not, which failure modes are survivable and which are not, and require human approval for everything that is irreversible or unsurvivable." Sending an internal status update is reversible if you send the wrong one. Sending a message to a customer is harder to reverse. Executing a financial transaction is extremely difficult to reverse. Deleting data is often impossible to reverse. The level of human oversight required before an action should scale with the irreversibility and consequence of the action, not with the organizational desire to minimize the number of clicks. Any agent framework that does not make this distinction explicit and allow builders to implement it cleanly is an agent framework that is making deployment easier and safety harder.&lt;/p&gt;

&lt;p&gt;The lmm project I described in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations Are Multimodal by Default&lt;/a&gt; and &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether&lt;/a&gt; is the companion to the autogpt framework in the Wise AI ecosystem: it is a parallel agent intelligence project that does not use LLMs at all, instead using equation-based intelligence to reason without gradient-trained models. The two projects represent different arms of the same commitment, which is to build AI systems that are honest about what they are, explicit about what they can and cannot do, and designed from the ground up for the kind of reliability that production environments require rather than for the kind of impressiveness that conference demos require. I am not advocating for abandoning LLMs as a tool. I am advocating for putting them in a framework that respects their limitations, compensates for their failure modes, and makes their behavior audible and controllable. That is what serious building looks like, and it is the opposite of building a general autonomous agent and hoping for the best.&lt;/p&gt;

&lt;p&gt;The measurement problem is the last thing I want to address in this section, and it is important enough that I want to be explicit about it rather than leaving it implicit. A system you cannot measure is a system you cannot improve, and a system you cannot improve is a system that degrades over time as the distribution of real-world inputs drifts from the distribution the system was designed for. Every serious agent deployment needs explicit metrics for what success looks like, explicit mechanisms for capturing the data needed to compute those metrics at production scale, explicit alert conditions for when the metrics drop below acceptable thresholds, and explicit processes for reviewing agent decisions when the metrics indicate a potential failure. Those are not nice-to-haves. They are the minimum requirements for a system that you can honestly describe as deployed rather than as a demo that happens to be running in production. Building these measurement systems requires exactly the kind of engineering judgment I described in the previous section: problem selection, domain depth, execution with quality, and the ability to design something repeatable. None of it requires AGI. All of it requires a serious engineer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Unasked Question Beneath the Provocative One
&lt;/h2&gt;

&lt;p&gt;I want to spend some time on what I think the people asking the AGI question are actually worried about, because beneath the provocative framing there is a real anxiety that deserves a real answer rather than just a demolition of the bad frame. The real anxiety is this: in an environment where AI tools are becoming dramatically more capable rapidly, is the skill base I have built still going to matter in three years? That is a legitimate question, and it is being asked by engineers who have spent years building expertise and who are genuinely uncertain whether that expertise will remain valuable as the tools get better. I want to answer it honestly, which means I have to say both the reassuring and the uncomfortable parts, because partial answers to this kind of question do more harm than no answer at all.&lt;/p&gt;

&lt;p&gt;The reassuring part is what I have argued throughout this post: the skills that make engineers genuinely valuable, problem selection, domain depth, evaluation judgment, clear communication, and the ability to build repeatable systems, are not the skills that AI tools replace. They are the skills that AI tools amplify for people who have them and expose as absent in people who lack them. The domain expert who can evaluate AI outputs correctly is more valuable than the domain expert who could not, because they can now get the routine generation work done faster and spend more of their time on the high-judgment parts. The engineer who can design a reliable agentic system and make its behavior audible is more valuable than the engineer who could not, because agentic systems that work reliably are enormously consequential in any organization that deploys them. The person who can communicate clearly about what the system does, what it does not do, and where the risks are is more valuable than before for the same reason: the stakes of misunderstanding are higher when the systems are more capable.&lt;/p&gt;

&lt;p&gt;The uncomfortable part is also real, and I would be lying if I left it out. There are categories of engineering work that AI tools are genuinely replacing rather than amplifying, and the engineers who have built their careers entirely within those categories are in a harder position. Routine code generation, boilerplate, and templating tasks: tools do this faster and often just as well as humans. Standard data pipeline construction using well-documented frameworks: tools can scaffold this quickly from specifications. Documentation of code that is being written by humans or tools simultaneously: tools are quite good at this. The pattern across these categories is that they are tasks where the domain knowledge required to do the task is well represented in training data, the output is easily verifiable by a human with less expertise than is required to produce it, and the task is well-enough defined that a statistical approximation of the standard solution is often the right answer. If your career is primarily composed of tasks in these categories, the tools are a genuine displacement and the right response is to build the domain depth and judgment skills that sit above those tasks in the value chain, not to compete with tools at what tools do well.&lt;/p&gt;

&lt;p&gt;The specific career trajectory I would recommend to an engineer navigating this transition is something like this: pick a domain, not a framework and not a tool, a domain, meaning an area of substantive human activity that software systems help with, and invest deliberately in understanding that domain at a depth that allows you to evaluate AI outputs about it critically. Pick a domain where the consequences of wrong AI outputs are significant enough that evaluation expertise is genuinely valuable, because the premium on evaluation expertise is proportional to the cost of unchecked AI errors. Build systems in that domain that are narrow, well-specified, measurable, and auditable. Document what you build and what you learn from deploying it. Make your judgment visible by writing about the design decisions you made and why, about the failure modes you encountered and how you addressed them, and about the tradeoffs you navigated that a less thoughtful builder would have ignored. That portfolio of documented judgment is the thing that differentiates you from someone who has the same tool access and less domain depth, and it is the thing that a thoughtful hiring organization can evaluate without relying on AGI affiliation as a proxy.&lt;/p&gt;

&lt;p&gt;I came to this conclusion through a lot of painful experience that I have documented in various posts throughout this blog. I built systems for clients who paid me almost nothing and then used what I built to raise funding and launch products. I applied for roles where I was told I was overqualified or underexperienced, depending on which angle served the rejection better. I watched people with better institutional affiliations and less technical depth get opportunities that I could not access. All of that is in &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;Just Don't Pick Up the Brush&lt;/a&gt; and &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt; and &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt; in more detail than I usually allow myself to write in. What I am describing in this post as a career strategy is not something I arrived at from a position of professional comfort. It is what I am actually doing, with real stakes, because I believe it is the right answer and I have examined the alternative answers carefully enough to know why they are wrong.&lt;/p&gt;

&lt;p&gt;The specific thing that the autogpt Rust framework represents in this context is not just a technical tool for building agents. It is a demonstration of what it looks like to apply genuine engineering judgment to a problem that the broader market is solving with less rigor. Building a production-quality agent framework in Rust, with its steeper learning curve and its compile-time discipline, when Python frameworks with lower barriers to entry are available, is an argument made in code about what the actual requirements of reliable agentic systems are. It is the same kind of argument that serious builders make implicitly every time they choose the harder right thing over the easier wrong thing: the choice itself demonstrates the judgment. And demonstrating judgment through choices is how you build a portfolio that communicates something real to someone who knows what they are looking at, which is the only kind of communication that matters when you are trying to prove your value without a prestigious institution behind you.&lt;/p&gt;

&lt;h2&gt;
  
  
  If You Can't Build AGI, Then Why Should We Hire You?
&lt;/h2&gt;

&lt;p&gt;I want to end where the title began, because that question deserves a direct answer rather than just a destruction of the frame it comes from. I have spent several thousand words explaining why the AGI criterion is the wrong criterion, why the market is re-sorting around different properties, and why those properties are available to anyone willing to develop them regardless of their institutional affiliation. But the person who asked the title question in a real interview deserves a real answer that does not feel like a lecture about epistemology of AI. So here is the answer I would give, and I want it to be honest and direct rather than polished and strategic.&lt;/p&gt;

&lt;p&gt;If you cannot build AGI, then you should be hired because you can build something that works. Working, in this context, means several specific things that I have laid out across this post. It means the system does the specific thing it was designed to do reliably enough to be trusted in the context where it will be used. It means the system fails gracefully and predictably rather than confidently and silently. It means the system is understandable by the people who will maintain it and auditable by the people who will be accountable for its behavior. It means the system can be measured, and when it degrades, the degradation is visible before it becomes catastrophic. Those properties are not glamorous. They do not make good conference talk titles. They do not generate the same kind of cultural excitement that AGI building ability generates. But they are what separates systems that create real value from systems that create impressive demonstrations, and the gap between those two categories is enormous.&lt;/p&gt;

&lt;p&gt;I also want to say something about honesty as a hiring criterion, because it connects to everything I have been arguing in this post and I do not want to leave it implicit. The engineers who are most valuable to organizations are the ones who can be trusted to say true things, not just impressive things. The engineer who says "this will work for these cases and fail for these other cases, and here is how I would handle the failure cases" is more valuable than the engineer who says "this will work" and leaves the failure cases to be discovered in production. The engineer who says "I don't know how to do this, but here is how I would find out and here is how long I think that would take" is more valuable than the engineer who says "sure I can do that" and then figures it out badly under time pressure. The cultural pressure toward AGI affiliation is a cultural pressure toward a specific kind of impressive-sounding dishonesty, toward overpromising against an underspecified goal, and the engineers who resist that pressure and maintain epistemic honesty are doing something more valuable than building a better demo.&lt;/p&gt;

&lt;p&gt;The connection to &lt;a href="https://wiseai.dev/blogs/be-aware-of-the-current-ufos-pandemic-remember-we-are-alone" rel="noopener noreferrer"&gt;Be Aware of the Current UFOs Pandemic. Remember, We Are Alone.&lt;/a&gt; is not superficial, and I want to make it explicit because it is one of the things I have been thinking about throughout writing this post. In that piece, I argued that the UFO panic is not primarily about UFOs. It is about the human need to believe that something outside our current situation, something powerful and superintelligent, is watching and might intervene. I argued there that the silence of the universe is clarifying rather than hopeless, because it places the full responsibility for our situation on us rather than on a savior from outside. The AGI panic is a different version of the same psychological pattern. It positions AGI as the coming superintelligence that will arrive and change everything, which creates both the fear of being made obsolete and the hope of being on the right side of the transition by being among those who summon it. Both the fear and the hope are responses to the same underlying fact: we are responsible for the systems we build, there is no superintelligence coming to take that responsibility off us, and the quality of what we build depends entirely on the quality of the judgment we bring to building it. That is both terrifying and clarifying, exactly as I said about the silence of the universe.&lt;/p&gt;

&lt;p&gt;I am building &lt;a href="https://github.com/wiseaidotdev/autogpt" rel="noopener noreferrer"&gt;autogpt&lt;/a&gt; in Rust and the companion &lt;a href="https://github.com/wiseaidotdev/lmm" rel="noopener noreferrer"&gt;lmm&lt;/a&gt; project without gradient-based training because I want to build things that are honest about what they are. The autogpt framework makes explicit what the agent is allowed to do, through its composable architecture and YAML configurations. The lmm project is an attempt to build machine intelligence that discovers the structure of reality rather than statistically approximating the structure of human text about reality. Neither of these projects is building AGI. They are both trying to build systems that are genuinely useful for the specific things they do, genuinely reliable in the specific contexts they are used, and genuinely honest about the gap between what they do and what an intelligent human observer would do in the same situation. That is the standard I hold my own work to, and it is the standard I am advocating for in this post: not "can you build AGI" but "can you build something real that you can be honest about".&lt;/p&gt;

&lt;p&gt;So if you ask me, in an interview or anywhere else, "if you can't build AGI, then why should we hire you", I will answer like this. I should be hired because I can identify the problem that is actually worth solving in the situation you are in. I can build a system that solves that problem reliably within the constraints that actually exist, not the constraints that would exist in an ideal world. I can measure whether it is working and tell you honestly when it is not. I can maintain it and evolve it as the situation changes without accumulating technical debt that forces a rewrite in eighteen months. I can explain what I built and why to people who need to understand it in order to use it or govern it or improve it. And I can do all of this without pretending that what I built is more than it is, because pretending is expensive and honesty is the only engineering practice that compounds in the direction of reliability rather than in the direction of eventual disaster. That is not AGI. That is the actual job. And that is why you should hire me.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Altman, S., &lt;em&gt;Planning for AGI and Beyond&lt;/em&gt;, OpenAI Blog, February 24, 2023. &lt;a href="https://openai.com/index/planning-for-agi-and-beyond/" rel="noopener noreferrer"&gt;openai.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; Morris, M. R., Sohl-Dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C. &amp;amp; Legg, S., &lt;em&gt;Levels of AGI for Operationalizing Progress on the Path to AGI&lt;/em&gt;, arXiv:2311.02462, November 2023. &lt;a href="https://arxiv.org/abs/2311.02462" rel="noopener noreferrer"&gt;arXiv:2311.02462&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Kapoor, S. et al., &lt;em&gt;The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems&lt;/em&gt;, MIT AI Agent Index, February, 2026. &lt;a href="https://aiagentindex.mit.edu/data/2025-AI-Agent-Index.pdf" rel="noopener noreferrer"&gt;aiagentindex.mit.edu&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; World Economic Forum, &lt;em&gt;Future of Jobs Report 2025&lt;/em&gt;, January 7, 2025. &lt;a href="https://www.weforum.org/press/2025/01/future-of-jobs-report-2025-78-million-new-job-opportunities-by-2030-but-urgent-upskilling-needed-to-prepare-workforces" rel="noopener noreferrer"&gt;weforum.org&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; International Labour Organization, &lt;em&gt;How might generative AI impact different occupations?&lt;/em&gt;, ILO Research Article, May 20, 2025. &lt;a href="https://www.ilo.org/resource/article/how-might-generative-ai-impact-different-occupations" rel="noopener noreferrer"&gt;ilo.org&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; METR Research, &lt;em&gt;Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity&lt;/em&gt;, METR Blog, July 10, 2025; companion paper &lt;a href="https://arxiv.org/abs/2507.09089" rel="noopener noreferrer"&gt;arXiv:2507.09089&lt;/a&gt;. &lt;a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/" rel="noopener noreferrer"&gt;metr.org&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Turing, A. M., &lt;em&gt;Computing Machinery and Intelligence&lt;/em&gt;, Mind, Vol. 59, No. 236, pp. 433-460, October 1950. &lt;a href="https://doi.org/10.1093/mind/LIX.236.433" rel="noopener noreferrer"&gt;doi.org/10.1093/mind/LIX.236.433&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; Vaithilingam, P., Zhang, T. &amp;amp; Glassman, E. L., &lt;em&gt;Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models&lt;/em&gt;, ACM CHI '22 Extended Abstracts, April 2022. &lt;a href="https://doi.org/10.1145/3491101.3519665" rel="noopener noreferrer"&gt;doi.org/10.1145/3491101.3519665&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>lmm</category>
    </item>
    <item>
      <title>Be Aware of The Current UFOs Pandemic. Remember, We Are Alone.</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 19:04:38 +0000</pubDate>
      <link>https://dev.to/wiseai/be-aware-of-the-current-ufos-pandemic-remember-we-are-alone-5cl0</link>
      <guid>https://dev.to/wiseai/be-aware-of-the-current-ufos-pandemic-remember-we-are-alone-5cl0</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This post was originally published on &lt;a href="https://wiseai.dev/blogs/be-aware-of-the-current-ufos-pandemic-remember-we-are-alone" rel="noopener noreferrer"&gt;the main website&lt;/a&gt; on &lt;a href="https://github.com/wiseaidotdev/blog/pull/20" rel="noopener noreferrer"&gt;May 12 2026&lt;/a&gt;. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account &amp;lt;3!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;a previous post&lt;/a&gt;, I wrote about something that has been sitting with me for as long as I can remember, the uncomfortable feeling of being fundamentally, irreversibly alone. Not lonely in the ordinary social sense, though I am that too, but alone in a deeper, more cosmic way, the kind of aloneness that comes when you look up at the sky at night, count the stars, and realize that despite everything the universe has to offer, none of it has ever reached down to say hello. I wrote in that post about the Montauk Project, about the strange creatures described as aliens that look far too much like something that crawled out of a government laboratory to be anything from another planet. I wrote about how I believe there are no aliens, not because the universe is small, but because the evidence points somewhere far more terrestrial and far more disturbing than most people want to admit. That argument started as a few paragraphs buried inside a very long and very personal post about my life, my struggles, and my thoughts about existence. But it deserved more space than I gave it, because the claim is not a small one and the stakes are not trivial. The current obsession with UFOs, now rebranded as UAPs in an effort to make the conversation sound more scientific, is not a scientific phenomenon. It is a cultural phenomenon, a psychological one, and in some cases a deliberately engineered one. I want to explain why I believe that, and I want to do it carefully, because I am not interested in cheap dismissal or in being the person who just says "aliens don't exist, end of discussion". I want to go where the evidence actually points and follow it honestly, even when it ends somewhere uncomfortable.&lt;/p&gt;

&lt;p&gt;I have also been writing allot about intelligence, about what it means, what it requires, and what happens when we mistake noise for signal. In &lt;a href="https://wiseai.dev/blogs/language-is-limited-asi-is-impossible" rel="noopener noreferrer"&gt;Language is Limited. ASI is Impossible.&lt;/a&gt;, I argued that language is a compression layer and that intelligence requires contact with the actual structure of reality, not just descriptions of it. In &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality.&lt;/a&gt;, I went deeper and argued that the really interesting work in AI is not happening in chatbots but in systems that can discover equations and simulate the physical world. And in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by Default&lt;/a&gt;, I made the case that equations are not abstract decorations but compressed representations of how reality actually works. All of that thinking feeds directly into what I want to say in this post, because the UFO panic is in many ways the perfect example of what happens when people confuse pattern with signal, when they mistake the thing that sounds impressive for the thing that is actually true. I wrote in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt; that reality never gives us the comfort we want. This post is about that same truth, applied to the sky above us and to the silence it has always returned.&lt;/p&gt;

&lt;p&gt;So let me say it plainly from the introduction: there are no aliens visiting us. We are, as far as every credible instrument humanity has ever built can tell, alone. And rather than being a depressing conclusion, I think it is the most important, most terrifying, and most clarifying fact about the human condition. The silence of the universe is not a mystery waiting to be solved. It is an answer we keep refusing to hear.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers People Cite Are Not The Numbers They Think They Are
&lt;/h2&gt;

&lt;p&gt;Let me start with the argument that almost everyone uses when they want to make the "aliens-definitely-exist" case feel airtight, and that is the sheer size of the universe. The universe contains something on the order of two trillion galaxies, according to a landmark 2016 study published in &lt;em&gt;The Astrophysical Journal&lt;/em&gt; (1). Each of those galaxies contains hundreds of billions of stars. The Milky Way alone has an estimated 100 to 400 billion stars, and around one in five sun-like stars is believed to host an Earth-sized planet in its habitable zone, based on Kepler mission data analyzed by Petigura, Howard, and Marcy (2). If you run rough Drake Equation estimates, the raw numbers seem to suggest that the universe should be absolutely teeming with life, and that intelligent, technological civilizations should be everywhere. This argument feels compelling because large numbers are genuinely hard to argue with, and because the intuition that something this vast cannot be empty is so deeply human that it almost feels like physics. But here is where the sleight of hand happens: astronomical abundance and biological probability are not the same thing, and conflating them is the foundational error that makes the aliens-must-exist argument feel stronger than it actually is.&lt;/p&gt;

&lt;p&gt;The conditions required for complex, multicellular, intelligent life are not simply "a planet in a habitable zone". That is the starting point, not the end point, and the distance between the starting point and intelligent life is not measured in miles or years but in improbable coincidences stacked on top of each other like a tower made of smoke and luck. You need the right type of star, one stable enough over billions of years. You need the right distance from that star. You need liquid water, but also a large moon to stabilize the axial tilt and prevent catastrophic climate swings, as described in the Rare Earth Hypothesis by Ward and Brownlee (3). You need plate tectonics, which recycles carbon and keeps the planet habitable over geological timescales. You need a Jupiter-class gas giant in the outer solar system to absorb asteroid impacts that would otherwise sterilize the inner planets repeatedly. You need the planet to avoid being tidally locked, sterilized by gamma ray bursts, stripped of its atmosphere by stellar winds, or suffocated by the toxic outgassing of its own interior. And that is before you even get to the emergence of life, which itself involves the spontaneous assembly of self-replicating chemistry in conditions that remain poorly understood despite decades of origin-of-life research. Each of these requirements, taken individually, might seem plausibly common. Combined, they paint a picture that is far, far less crowded than the simple "two trillion galaxies" argument suggests.&lt;/p&gt;

&lt;p&gt;I am not claiming that microbial life is impossible elsewhere, and I want to be precise about that because the distinction matters enormously. Simple life, the kind that sits at the bottom of a thermal vent and converts chemistry into more of itself, might genuinely exist in other places. On Mars, perhaps, or in the subsurface ocean of Europa, or somewhere we have not looked yet. That would be a staggering discovery, and I do not dismiss it. What I am arguing is that the jump from simple microbial life to intelligent, technological, space-faring, signal-emitting civilization is not a small step. It is the largest step in the known history of life on Earth, and it happened exactly once in four billion years, under conditions so specific that we still do not fully understand what made it happen. If it happened only once on a planet that had literally everything going for it, the implied prior probability of it happening elsewhere is not "common across two trillion galaxies." It is something closer to vanishingly rare.&lt;/p&gt;

&lt;p&gt;The Drake Equation, which people often cite as justification for optimism about alien civilizations, is not actually evidence for anything. It is a framework for organizing our ignorance. When Frank Drake wrote it in 1961 as the agenda for the first SETI conference at Green Bank Observatory, he populated it with values that were essentially educated guesses, and sixty-plus years later, the most uncertain terms in the equation, the fraction of planets that develop life, the fraction of those where intelligence emerges, the fraction of those that develop technology, and the average longevity of technological civilizations, remain almost completely unknown (4). You can plug in optimistic values and get a galaxy full of civilizations. You can plug in pessimistic but equally defensible values and get a galaxy where we are the only one. The equation does not choose between these answers because we do not have the data to constrain the inputs. Anyone who presents the Drake Equation as evidence that aliens must exist is not doing science. They are doing wishful arithmetic, and wishful arithmetic is not the same as evidence.&lt;/p&gt;

&lt;p&gt;The thing that nobody in the UFO conversation wants to sit with for longer than five seconds is the actual empirical track record of SETI. The Search for Extraterrestrial Intelligence has been running serious, systematic radio telescope surveys since the 1960s, scanning billions of frequencies, listening to hundreds of thousands of targets, and using increasingly sensitive instruments over six decades of dedicated effort. The total haul of confirmed, verified, repeatable signals of unambiguously extraterrestrial intelligent origin is exactly zero (5). Not one. The WOW signal of 1977 was never repeated, never confirmed, and the most plausible natural explanations have grown more convincing over time. The Breakthrough Listen initiative, launched in 2015 with a hundred million dollars and some of the best radio telescopes on Earth, has spent years listening and has found nothing that passes the threshold of evidence (6). This is not a small sample. Sixty years of increasingly sensitive listening across the full radio spectrum, pointed at hundreds of thousands of targets, in a galaxy that should be full of signals if the optimists are right. And silence. Absolute, unbroken silence.&lt;/p&gt;

&lt;p&gt;The silence itself is a data point, and it is a devastating one. In 1975, astrophysicist Michael Hart published a paper in the &lt;em&gt;Quarterly Journal of the Royal Astronomical Society&lt;/em&gt; that made the logical structure of this problem precise in a way that Fermi's original lunchtime question did not (7). Hart argued that if technological civilizations were common, they would have colonized the galaxy by now, even at sub-light speeds, because the timescale required to spread across the Milky Way is a few million years at most, which is a geological eyeblink in cosmic terms. The fact that we observe no evidence of this colonization, no megastructures, no signals, no artifacts, no visitors, is not just an absence of data. It is a positive fact about the universe that needs to be explained. Either civilizations are extraordinarily rare, or they hit something that stops them, or they choose to remain invisible in ways that defy all plausible incentive structures. None of these options are reassuring to the aliens-definitely-exist camp. All of them point toward the same uncomfortable conclusion: the sky is quiet because there is almost nothing out there to make noise.&lt;/p&gt;

&lt;p&gt;This is the foundation of everything I am about to argue. The numbers people cite are real. The galaxy truly is enormous. But enormous and populated are not synonyms, and the universe does not owe us company just because it contains the raw ingredients. The evidence, all of it, every telescope, every radio receiver, every spectrograph, every Mars lander, every probe sent into the outer solar system, has returned the same answer. Silence. The correct response to silence is not to invent noises. It is to understand what the silence means.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Montauk Project And The Shape Of The Lie
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk0alylw9dp0mfj4cj5dy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk0alylw9dp0mfj4cj5dy.png" alt=" " width="800" height="529"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I mentioned the Montauk Project briefly in my first post, almost as a passing observation, but it deserves a much more careful treatment because it is genuinely important for understanding why the alien narrative has the specific shape that it does. The Montauk Project is a conspiracy theory that was primarily popularized through a series of books starting with &lt;em&gt;The Montauk Project: Experiments in Time&lt;/em&gt; by Preston Nichols and Peter Moon, published in 1992 (8). The story, in its most elaborated form, claims that the US government operated a secret research program at Camp Hero, a decommissioned Air Force radar station on Long Island, that involved time travel, mind control, interdimensional portals, and contact with extraterrestrial beings. The "aliens" in these accounts are remarkably consistent in their description: hairless, grey-skinned, large-headed, slender-limbed beings with enormous dark eyes and no visible musculature. This description, which appears in Nichols's books, in dozens of subsequent retellings, and in countless "abduction testimonies" collected over the following decades, is worth pausing on. Because the beings described do not look like what evolution would produce on another planet. They look exactly like what a government program in human genetic modification might produce if it were optimizing for specific neurological and physical characteristics.&lt;/p&gt;

&lt;p&gt;I wrote in my first post that the creatures described as aliens are almost always hairless, with unusual altered features that make them look like something out of a laboratory rather than visitors from another world, and I want to expand on that here because the pattern is striking once you see it clearly. Human beings have hair because hair is adaptive in the evolutionary environment we came from. We have the body proportions we have because those proportions are optimized for bipedal movement, thermoregulation, and manual dexterity. We have the brain-to-body ratio we have because it represents the most expensive possible trade-off a mammal can make, and natural selection only makes that trade-off when the cognitive returns justify the enormous caloric cost. Now look at the classic "grey alien" description: no hair, which removes an adaptive feature without adding an obvious environmental substitute; enormous head relative to body, which would represent an even more extreme brain investment than humans; tiny, atrophied limbs, which suggest that physical capability has been drastically deprioritized in favor of something else; and enormous dark eyes, which could be adaptive for very low-light environments, or which could be the result of genetic modification of the visual system. Every single feature of the standard alien description is consistent with the results of extreme genetic manipulation of human baseline biology, and not one of them is specifically consistent with what you would expect from a life form that evolved independently on another planet with different gravity, different solar spectrum, different biochemistry, and billions of years of completely separate evolutionary history.&lt;/p&gt;

&lt;p&gt;The Montauk Project theory also connects to something that is definitively real, which is the CIA's MKUltra program. MKUltra was an illegal human experimentation program run by the CIA from 1953 to 1973, involving more than 80 institutions including universities, hospitals, and prisons, that used LSD, electroshock therapy, hypnosis, sensory deprivation, and psychological torture to attempt to develop mind control techniques, largely without the consent of the subjects (9). Most of its records were destroyed in 1973 on the orders of CIA Director Richard Helms, precisely to prevent public and congressional scrutiny, and the existence of the program was only confirmed in 1975 through the Church Committee investigation and then further through a Freedom of Information Act request in 1977. The lesson of MKUltra is not that such programs are impossible. The lesson is that they happened, were systematically hidden, destroyed their own records, and were only partially exposed through aggressive government investigation, not through normal information flow. If you are tempted to dismiss MKUltra as ancient history that says nothing about what might be happening today, I would ask you to consider how many programs operating under the same secrecy protocols would be visible to you through any mechanism that you currently have access to. The answer is none. That is what operational secrecy means.&lt;/p&gt;

&lt;p&gt;I want to be careful here, because I am not claiming certainty about Montauk. I am claiming something more modest and more defensible: that the pattern of the alien narrative, the specific physical description of the supposed visitors, the timing of UFO sighting surges relative to known government experimental programs, the consistent failure of these encounters to produce any physical evidence that could be independently analyzed, and the way the narrative is structured to discourage exactly the kind of investigation that would be most likely to find terrestrial explanations, all of this is far more consistent with a carefully managed disinformation campaign than with genuine extraterrestrial contact. This is not a wild claim. Governments have a well-documented history of using exotic narratives to cover classified programs. The CIA's use of alien mythology to explain U-2 spy plane sightings during the Cold War is now a matter of declassified record (10). When a classified aircraft flying at 60,000 feet was spotted and reported as a UFO, the cover story of "weather phenomena" was replaced with quiet encouragement of the "experimental aircraft" narrative, which itself was sometimes replaced with quiet encouragement of the "alien spacecraft" narrative, because the alien explanation was easier for the public to dismiss and therefore less dangerous to the actual program. That is not speculation. That is documented government behavior.&lt;/p&gt;

&lt;p&gt;The Montauk Project's connection to Stranger Things, the Netflix series that made the basic narrative framework mainstream for a generation of young viewers, is also worth noting. At one point, the show was literally going to be called "Montauk" and was set in the same location as the alleged experiments. The Upside Down, the interdimensional portal, the children with psychic abilities who are the products of government experimentation, the creature that appears to be of non-human origin but turns out to be something else entirely, all of these are structural echoes of the Montauk narrative. I am not saying the show popularized the theory, I am saying the theory was already popular enough to be the direct inspiration for one of the most-watched television series in history. That level of cultural penetration is not accidental, and understanding where an idea comes from is not the same as proving the idea is wrong. But it does tell you something about the mechanism by which ideas spread, and about who benefits from certain ideas being present in certain cultural contexts.&lt;/p&gt;

&lt;p&gt;What I find most telling about the Montauk Project specifically is not the claims themselves but the architecture of the narrative. It is designed to be unfalsifiable in exactly the ways that matter. The records are destroyed. The witnesses have "suppressed memories" recoverable only through hypnotic regression, a technique that is scientifically documented to reliably produce confabulated false memories rather than actual recalled experience (11). The underground facility is supposedly hidden so thoroughly that no physical exploration of the site has ever confirmed it. And the beings involved are simultaneously extraterrestrial, which places them outside normal investigative jurisdiction, and terrestrial in the sense that the experiments happened right here on Long Island. It is the perfect narrative structure for something that you want to be believed but never proven. I have spent enough time thinking about how disinformation works to recognize that structure, and I recognize it here.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2023 Congressional Panic And What It Actually Tells Us
&lt;/h2&gt;

&lt;p&gt;In July 2023, a former Air Force intelligence officer named David Grusch testified under oath before the House Oversight Committee's national security subcommittee that the US government has been operating a multi-decade "crash retrieval and reverse-engineering program", that it is in possession of "non-human" spacecraft, and that "non-human biologics" have been recovered from crash sites (12). This was not a fringe figure speaking on a YouTube channel. This was sworn testimony before Congress, and it generated exactly the reaction you would expect: wall-to-wall coverage, breathless speculation, and millions of people updating their priors about alien life based on nothing more than one person's claims delivered in a formal setting. Let me be precise about what Grusch's testimony actually consisted of. He said that he had been told, by other officials, that these programs exist. He did not personally witness the craft or the biologics. He did not provide any physical evidence. He provided no documents, no photographs, no materials, and no corroborating witness who had direct personal knowledge of the claimed objects. The Department of Defense and NASA both denied his claims. The All-Domain Anomaly Resolution Office, AARO, which is the DoD's official office for investigating UAPs, stated explicitly that they have found no verifiable information to substantiate claims about programs involving possession or reverse-engineering of extraterrestrial materials.&lt;/p&gt;

&lt;p&gt;I want to be clear that I am not dismissing Grusch as a liar or as someone acting in bad faith. He may sincerely believe everything he said. But sincere belief is not evidence, and sworn testimony about something you were told by someone else is hearsay, not proof. The history of intelligence communities is full of sincere beliefs based on compartmentalized information that turned out to be wrong, or based on disinformation that was fed into official channels precisely because official channels give claims more credibility when they reach the public. If you operate a classified program using alien mythology as cover, and if you want to maintain that cover, it is extraordinarily useful to have officials at various levels of the bureaucracy genuinely believe the alien explanation, because genuine believers are more convincing than actors. The compartmentalization of intelligence programs means that people inside the system can be fed false narratives just as easily as people outside it, because neither group has access to the full picture. Grusch may have been told the truth, or he may have been told a very elaborate story that serves purposes he is not aware of. The testimony itself cannot distinguish between these possibilities, and neither can we based on what was made public.&lt;/p&gt;

&lt;p&gt;What the 2023 congressional hearing does tell us, very clearly, is that there is genuine institutional appetite in certain parts of the US government for UAP transparency, and that this appetite is real and bipartisan and driven at least partly by legitimate national security concerns. Unidentified aerial phenomena in restricted airspace are a real problem regardless of their origin. Pilots see objects they cannot identify, performing maneuvers that seem to exceed the aerodynamic capabilities of known aircraft, and that is a national security issue whether the explanation turns out to be classified domestic technology, foreign adversary drones, sensor anomalies, or genuine unknowns. The Navy and Air Force have formally acknowledged this since 2017, when the release of the FLIR footage of the Tic Tac UAP, captured by pilots from the USS Nimitz, made it impossible to dismiss the phenomenon publicly. Ryan Graves, one of Grusch's colleagues in the 2023 testimony, is an actual Navy pilot who saw these phenomena personally and has been entirely consistent in his account without making the extraterrestrial leap. His testimony is the most credible part of the conversation because it is grounded in personal direct observation and is limited to what he could actually verify, which is that he saw something he did not recognize.&lt;/p&gt;

&lt;p&gt;The distinction between "there are genuine aerial phenomena that we cannot currently identify" and "those phenomena are of extraterrestrial origin piloted by non-human beings" is the most important distinction that the entire UFO discourse routinely collapses. These are not the same claim. The first is an empirical observation about the data. The second is a highly specific explanatory hypothesis about what generates the data. Moving from the first to the second requires evidence, and that evidence does not exist. The fact that military pilots see things they cannot identify is not surprising. Advanced drone technology from adversarial nations, classified domestic programs that the pilots are not briefed on, atmospheric phenomena that interact with sensor systems in unusual ways, and limitations of human perception and instrument interpretation all provide abundant mundane explanations for encounters that feel inexplicable in the moment. None of this requires extraterrestrial technology, and invoking extraterrestrial technology before eliminating all terrestrial explanations is not science. It is the opposite of science.&lt;/p&gt;

&lt;p&gt;I also want to say something about the UAP Disclosure Act of 2023, introduced by senators Chuck Schumer and Mike Rounds, which eventually resulted in provisions included in the FY 2024 National Defense Authorization Act requiring the establishment of a UAP records collection and the systematic declassification and release of relevant government records (13). I think this is genuinely good and I support it without reservation, not because I think disclosure will reveal alien spacecraft, but because transparency in government is valuable for its own sake and because the conversation about what these phenomena actually are is one that democratic societies should be able to have. If the records reveal classified domestic programs, that is information the public has a right to know. If they reveal foreign adversary capabilities, the specific revelations can be handled with appropriate operational security while the general picture is shared. And if they reveal that there is genuinely nothing more exotic than a combination of classified aircraft, sensor artifacts, and human misperception, then that too is information worth having, because the current state of authorized ambiguity is itself a form of manipulation. Not knowing creates a vacuum, and vacuums get filled with exactly the kind of speculation that serves whoever benefits from public confusion.&lt;/p&gt;

&lt;p&gt;What I find deeply troubling about the timing of the recent UAP surge is how perfectly it correlates with political and military utility. Maintaining public concern about unexplained aerial phenomena keeps defense appropriations high, keeps public attention focused on the sky rather than on domestic policy failures, and provides a convenient cover story for the deployment of classified sensor systems, surveillance drones, and next-generation aircraft that cannot be acknowledged publicly. I wrote in &lt;a href="https://wiseai.dev/blogs/as-engineers-llms-should-pay-us-for-tokens-usage" rel="noopener noreferrer"&gt;As Engineers, LLMs should pay us for tokens usage&lt;/a&gt; about how systems are built to extract value from people who do not understand how they work, and the UFO narrative is precisely that: a system that extracts attention, credulity, and political will from a public that does not have enough information to evaluate what it is being told. Attention, once captured, does not go back. And attention pointed at the sky is attention not pointed at the things happening on the ground.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Psychology Of Wanting Company
&lt;/h2&gt;

&lt;p&gt;I want to step back from the empirical arguments for a moment and talk about the human side of this, because the UFO belief is not primarily a factual error. It is a psychological response to something real and painful, and dismissing it as stupidity is both wrong and unfair. The desire to believe in extraterrestrial intelligence is, at its root, the desire to not be alone. It is the same desire I described in my first post, the one that drove me to search for connection in hackathons, in YouTube comment sections, in IRC channels at 3 AM, in any space where people gathered around shared ideas. It is one of the most fundamental human drives, and the universe, in its indifferent vastness, does not give us what we want. It gives us silence. And silence, for a creature as social as a human being, is psychologically almost unbearable.&lt;/p&gt;

&lt;p&gt;The history of how humans respond to silence and isolation is not subtle. Religious traditions across cultures are almost universally built around the belief in non-human intelligences who care about us, who watch us, who have plans for us, who might communicate with us if we are attentive enough. The specific content varies enormously, the specific beings range from gods to angels to ancestors to demons to nature spirits, but the underlying structure is always the same: we are not alone, there are others who are aware of us, and our actions in this life are witnessed and matter. The alien belief fulfills exactly the same psychological function, but in a frame that feels more compatible with a scientific worldview, because the beings are hypothetically biological rather than supernatural. You can believe in aliens without feeling like you are being unscientific, which makes the belief far more accessible to modern people who have internalized the cultural norm that science and religion are in tension. But the psychological need being met is identical, and recognizing that does not make the need less real or less human.&lt;/p&gt;

&lt;p&gt;I experienced this directly in the years I spent feeling completely isolated, years I described in my first post in more detail than I have ever wanted to share publicly. When you feel genuinely alone in the world, the idea that somewhere out there something else is watching, even something that has not made contact, provides a form of comfort that has nothing to do with reality and everything to do with psychology. The sky becomes populated. The universe becomes something more than an indifferent physical system playing out under mathematical laws that do not know you exist. It becomes a place where you might matter, where the fact of your consciousness might register somewhere, where the universe might eventually turn in your direction and acknowledge you. That comfort is real, even when the thing providing it is not. I understand it completely, because I felt it, and I am not going to pretend I am above it just because I have argued myself out of it.&lt;/p&gt;

&lt;p&gt;But there is a difference between understanding a psychological need and endorsing the false belief that fills it. Humans have always found ways to live with the discomfort of not being central to the universe. Science itself is partly the practice of learning to accept that the universe does not care about us while still finding meaning in studying it. The Copernican revolution took us out of the center of the solar system. Darwinian evolution removed the special divine origin of our species. The discovery of the scale of the universe removed our galaxy from any privileged position. Each of these was psychologically devastating to the worldview it displaced, and each of them was true, and each of them ultimately led to a richer understanding of reality than the comfortable myth it replaced. The alien belief is the last refuge of anthropocentrism, the last way of insisting that the universe has taken note of us and sent ambassadors. And it is, like all the anthropocentric myths before it, almost certainly wrong.&lt;/p&gt;

&lt;p&gt;There is also something worth saying about the specific way the UFO community handles counterevidence, which is not the way any truth-seeking community handles counterevidence. In a genuine scientific community, when a well-designed study fails to find the expected signal, that negative result is published and updates the overall evidence base. In the UFO community, when a telescope fails to find alien structures, that is taken as evidence of government coverup. When a claimed abduction account contains internal inconsistencies, that is dismissed as the result of memory suppression. When a crashed object turns out to have a mundane explanation, that particular case is abandoned and the next one is elevated with fresh certainty. This asymmetry, where confirming evidence is accumulated and disconfirming evidence is explained away, is the defining feature of unfalsifiable belief systems, and it is what separates the UFO community from any community that is genuinely committed to finding the truth. I am not saying this to be cruel. I am saying it because recognizing the pattern is the first step to not being trapped by it.&lt;/p&gt;

&lt;p&gt;The psychological need for cosmic company is real, and it deserves a real response. But the real response is not to populate the sky with beings who are not there. The real response is to understand why we feel alone and to address that directly, which is much harder and much more uncomfortable than looking at blurry photographs and deciding they confirm what we already wanted to believe. I have spent my entire life doing the harder thing, and I am not going to stop now just because the easier thing is more soothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What The Government Actually Hides And Why It Matters
&lt;/h2&gt;

&lt;p&gt;I want to be precise about something that I think gets lost in the UFO conversation: the fact that governments lie, hide things, and run secret programs is not evidence that they are hiding aliens. It is evidence that governments lie, hide things, and run secret programs. These are completely separate claims, and collapsing them is a logical error that the UFO community makes constantly and that I want to address directly, because treating it as self-evident rather than as a confusion is one of the things that keeps people in a permanent state of excited credulity about something that has much more boring explanations.&lt;/p&gt;

&lt;p&gt;The US government has hidden classified aircraft programs for decades, and the historical record of this is clear and well-documented. The U-2 spy plane, which flew at altitudes that were far beyond what the public or even most military officials knew to be possible in the 1950s, was routinely reported as a UFO, and the government's official response was to cultivate confusion rather than to confirm or deny. The SR-71 Blackbird, the B-2 Spirit stealth bomber, and the F-117 Nighthawk all went through periods of being "unidentified aerial phenomena" before being acknowledged. The CIA's own historical review of UFO reports from the 1950s and 1960s concluded that a large fraction of them were attributable to classified aerial vehicles, and that the CIA had actively encouraged the alien hypothesis as a cover story because it was more effective at discouraging investigation than any official denial would have been (10). This is the documented reality of what governments hide in aviation, and it has nothing to do with alien craft. It has everything to do with the competitive advantage of technological secrecy.&lt;/p&gt;

&lt;p&gt;The current generation of classified aerospace programs is certainly no less advanced than the classified programs of the Cold War, and possibly considerably more so. The explosion of drone technology, the development of hypersonic vehicles, the deployment of high-altitude surveillance platforms with aerodynamic characteristics that would have seemed fantastical thirty years ago, and the use of advanced directed-energy and counter-drone systems all provide an enormous reservoir of potential explanations for unusual aerial observations that pilots, having no clearance for these programs, genuinely cannot explain. When a Navy pilot sees an object performing a maneuver that their training tells them should be aerodynamically impossible, the most parsimonious explanation is not "alien spacecraft" but rather "a vehicle using aerodynamic principles I have not been briefed on, operated by people with a higher security clearance than mine." This is not a conspiracy theory. It is the basic operational reality of classified military technology in a competitive geopolitical environment.&lt;/p&gt;

&lt;p&gt;Beyond aerospace, the fact that real secret programs exist, like MKUltra, like COINTELPRO, like the NSA's mass surveillance programs revealed by Edward Snowden in 2013, should teach us something important about what kinds of things governments actually hide and why. They hide programs that would be politically toxic if publicly acknowledged. They hide capabilities that would be operationally compromised if disclosed. They hide crimes that would result in prosecution or political destruction if exposed. What they are notably less motivated to hide is aliens, because the existence of alien contact would be, for any government that possessed credible evidence of it, an extraordinary instrument of national power and international prestige. The government that could credibly claim first contact with an extraterrestrial civilization would not sit on that information for seventy years. The political, scientific, military, and cultural advantages of being the government that revealed verified alien contact are so enormous that the incentive structure for secrecy essentially inverts. Governments hide things that hurt them. Verified alien contact would not hurt them. The secrecy argument collapses under its own logic when you think about it carefully.&lt;/p&gt;

&lt;p&gt;I also want to address the dark side of some of the things that governments do admit to hiding, because those admissions are genuinely relevant to understanding what "alien" encounters might actually be. MKUltra, which I mentioned in the Montauk section, is the most important precedent here. The program ran for two decades, involved over eighty institutions, conducted experiments on unwitting subjects in universities, hospitals, and prisons across the US and Canada, and was entirely hidden from the public and from Congress until investigative journalism and courageous congressional investigators forced its partial exposure. The destroyed records mean we still do not know the full extent of what was done. What we do know includes experiments involving the administration of psychoactive substances that cause profound alterations of consciousness and perception, including experiences that subjects later described as encounters with non-human entities. This is not speculation. It is documented in the surviving MKUltra records. The correlation between the descriptions of alien abduction experiences and the documented phenomenology of forced psychedelic states, extreme sensory deprivation, and experimental psychological conditioning is striking enough that dismissing it requires active effort. I am not saying every alien abduction report is the result of a government experiment. I am saying that the government has already proven it is willing to induce exactly the kinds of experiences that people later describe as alien contact, and that this documented capacity for induced experiential distortion is a serious alternative explanation that the UFO community consistently refuses to engage with honestly.&lt;/p&gt;

&lt;p&gt;The bottom line is this: governments hide things, and they hide them for reasons that are almost always terrestrial, political, and economic. The alien hypothesis is the least parsimonious explanation for government secrecy because it requires not only that governments have hidden information of literally civilization-defining significance for seventy years without a single leak that holds up under scrutiny, not only that the most powerful incentive to reveal such information, the geopolitical advantage of first-contact disclosure, has been somehow overridden by unspecified competing incentives, but also that the physical evidence of alien technology and biological samples has been successfully hidden from every independent scientist, every foreign intelligence service, and every journalist in the world for seven decades. This is not impossible by the laws of physics, but it is so implausible given everything we know about how secrecy actually works that the prior probability should be assigned accordingly. When the simpler explanation, classified terrestrial programs, government-sponsored psychological manipulation, sensor artifacts, and genuine human misperception, explains all the same observations without requiring any additional assumptions, the simpler explanation is the correct starting point.&lt;/p&gt;

&lt;h2&gt;
  
  
  We Are Alone And That Is Our Problem To Solve
&lt;/h2&gt;

&lt;p&gt;I said at the beginning of this post that the silence of the universe is the most important, most terrifying, and most clarifying fact about the human condition, and I want to explain what I mean by clarifying, because that is the part people find hardest to accept. Clarity is uncomfortable when it removes options you did not realize you were counting on. The alien belief, even when it remains vague and unacticulated, functions as an emergency exit from human responsibility. If cosmic intelligences are watching us, or if help might arrive from somewhere beyond our solar system, then our failures are not entirely our own problem. Someone might intervene. Something might happen that changes the terms of our predicament. The weight of all the suffering I described in my first post, all the broken systems, all the technological displacement, all the mental illness and isolation and poverty that defines so many human lives, becomes slightly more bearable if you imagine it exists in a universe that contains other minds who might eventually care. The silence of the universe removes this exit. We are here, we are alone, and every problem is ours to solve or leave unsolved. Nobody is coming.&lt;/p&gt;

&lt;p&gt;That is, I think, why the alien belief is psychologically irresistible in exactly the proportion that human civilization is failing. The more broken things feel, the more appealing it becomes to believe that something outside the broken system might fix it. This is the same structure as the religious belief I discussed in my first post, where I asked where God is and tried to answer honestly. I wrote there about how I sometimes wonder if the cosmic battle I saw in my childhood dreams was one that was already lost. Here I will say something similar: if you are waiting for alien intelligences to arrive and solve the problem of human civilization, then you are waiting for something that is not coming, in the same way that waiting for divine intervention is waiting for something that has not shown up in four thousand years of urgent human need. The waiting is itself a form of inaction dressed up as hope, and inaction is the last thing the actual situation calls for.&lt;/p&gt;

&lt;p&gt;The most important thing that follows from our cosmic solitude is that intelligence itself, genuine intelligence, the kind capable of understanding the structure of reality and acting effectively within it, is extraordinarily rare and therefore extraordinarily precious. I have spent several posts arguing about what intelligence actually is, and here I want to apply that argument to this existential point. If we are alone, then everything that consciousness, understanding, agency, and knowledge represent in this universe exists, as far as we can tell, only here, only on this planet, only in these eight billion skulls. The equations I wrote about in my mathematics post, the compressed representations of how reality works that took thousands of years of human effort to discover, exist only in human minds and human records. The music, the literature, the love, the suffering, the curiosity, the ambition, the kindness, the cruelty, all of it exists only here. Nothing else in the observable universe appears to be aware of any of it. That is not a comfortable fact, but it is a galvanizing one, because it means that the continuation and flourishing of intelligence in this universe depends entirely on us getting our civilization right.&lt;/p&gt;

&lt;p&gt;I think of this often in connection with what I have been building with &lt;a href="https://github.com/wiseaidotdev" rel="noopener noreferrer"&gt;Wise AI&lt;/a&gt; and with the broader project of trying to build AI systems that can genuinely understand the world rather than just talk about it. The reason I care about the difference between language models and genuine reasoning systems is not abstract. It is because the only hope for solving problems of civilizational scale, climate, disease, poverty, the coordination failures that keep us from acting collectively on things we all agree matter, is intelligence sophisticated enough to find solutions that individual human minds cannot find alone. In the framework I have been developing across these posts, that means systems capable of discovering and working with the mathematical structure of reality, not just generating fluent descriptions of problems we cannot solve. If we are alone, then the development of genuine intelligence, both human and artificial, is the most important project in the history of this planet, because it is the only thing capable of ensuring that the brief window of consciousness that has opened in this corner of the universe does not close again before it understands enough to keep itself alive.&lt;/p&gt;

&lt;p&gt;This is also why the UFO panic makes me genuinely angry in a way that goes beyond frustration with factual error. It is a distraction from something that urgently matters. Every hour of video-watching, every dollar of book-buying, every congressional hearing hour spent on "non-human biologics" testimony is an hour, a dollar, and a congressional hearing not spent on the things that are actually killing people right now. Climate change is a mathematical problem of extraordinary difficulty that we have barely begun to solve. Drug-resistant pathogens are a biological problem that requires genuine scientific investment and global coordination. The economic systems that produce the kind of suffering I described in my first post are social and technical problems that require clear thinking and serious engineering. None of these problems will be solved by aliens, and none of them will be advanced by the cultural energy currently pouring into the belief in aliens. The conspiracy theories, the congressional theatrics, the YouTube channels with breathless thumbnail graphics, all of it is a displacement of cognitive and political energy away from things that are real and toward things that feel more exciting precisely because they would let us off the hook from the hard work of fixing what is broken.&lt;/p&gt;

&lt;p&gt;I started this post by connecting it to my first post, to the section where I wrote about the Montauk Project and about my belief that there are no aliens. But the deeper connection is to everything I wrote about what it means to be alone in the world. I have been alone for most of my life, in ways that I described in more detail than was comfortable, and I have learned something from that aloneness that I think scales from the personal to the cosmic. Being alone is not a problem to be solved by inventing company. It is a condition to be understood and eventually accepted, because the acceptance is what allows you to function clearly and to build things that are actually real. The alien belief is, at its core, the cosmic version of refusing to accept solitude. And solitude accepted is not emptiness. It is the beginning of genuine responsibility.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Specific Weight Of This Silence
&lt;/h2&gt;

&lt;p&gt;I want to end where the argument leads, not where it is comfortable to stop, because intellectual honesty means following things to their conclusions even when the conclusions are heavy. If we are alone, and I believe we are, then the universe has produced exactly one instance of the kind of intelligence that can look at itself and ask why. One. Out of two trillion galaxies, over thirteen point eight billion years, under the action of all the physical and chemical and biological processes that have operated across that unimaginable extent, consciousness has flickered on, as far as we can tell, in exactly one place. On one small rocky planet, orbiting one unremarkable star, in one arm of one average galaxy. That consciousness is us. All of it, everything that thinks and knows and wonders and suffers in the entire observable universe, is concentrated here, in this small collection of living systems on the surface of this particular rock. This is either the most humbling fact that any conscious being has ever had to absorb, or it is the most galvanizing one. I think it is both.&lt;/p&gt;

&lt;p&gt;The specific weight of this silence is something I feel physically when I think about it carefully. It is the weight of being the only thing that knows it exists, in a universe that does not know itself. Physicists describe the universe as a vast system of interacting fields and particles, evolving deterministically or stochastically under the laws of nature, and none of those laws contain anything that notices. The particles do not know they are particles. The galaxies do not know they are galaxies. The stars do not know that they fuse hydrogen into helium. The black holes do not know they consume whatever falls into them. Everything happens according to mathematical laws, precisely, faithfully, and without anyone home to observe it, except here. Except us. And the mathematical laws I keep returning to in my posts, the equations that describe how the universe works, were written by the universe itself and then discovered by the only part of the universe capable of noticing them. That is a strange loop, and sitting inside it is strange, but it is where we actually are.&lt;/p&gt;

&lt;p&gt;I have been writing across these posts about the relationship between language, mathematics, and genuine understanding, and here I want to bring that back to the cosmic question one more time. The most honest thing I can say about our solitude is that it makes the project of understanding more important, not less. If there are no alien civilizations that have already decoded the deep structure of reality, if there is no cosmic archive of solved problems waiting to be transmitted to us, then every equation we discover is something the universe has never known before. Every piece of genuine insight produced by a human mind is a first. The knowledge of physics, chemistry, mathematics, and biology that human civilization has accumulated represents the only time in the history of this universe that the universe has understood a piece of itself. The weight of that should not make us despair. It should make us take it seriously beyond any other consideration.&lt;/p&gt;

&lt;p&gt;I said in my first post that the real mystery is not "where are the aliens?" but "what are we going to do with the life we have?" I stand by that completely, and I have tried to live it in my own small way, through the years of building things that nobody looked at, through the applications that went nowhere, through the crypto trading systems that ran themselves while I tried to figure out what I was actually for, through everything I described in that post. I am building &lt;a href="https://github.com/wiseaidotdev" rel="noopener noreferrer"&gt;Wise AI&lt;/a&gt; because I want to contribute something to the only project that actually matters, which is the project of getting intelligence right before the window of opportunity closes. Not because I thought it would make me financially secure, but because we are here, and we are alone, and every contribution to the accumulation of genuine understanding is one more piece of something that the universe has never had before.&lt;/p&gt;

&lt;p&gt;The UFO pandemic is a distraction from this. It is a way of populating the silence with noise that feels meaningful because it is mysterious and exciting, while the actual mystery, the one I keep writing about, the one about what genuine intelligence requires and what it is capable of, sits quietly in the corner waiting for attention that keeps going somewhere else. I am asking you to redirect that attention. Not toward me, not toward any particular project, but toward the actual question, which is what we are doing with the fact that we exist and we understand and we are alone. That question deserves more than blurry photographs and congressional theatrics and theories about hairless creatures from underground military bases. It deserves everything.&lt;/p&gt;

&lt;p&gt;We are alone in the universe. That is the most important thing I know, and it is the thing I most want to say without softening it. The sky is quiet. No one is coming. We are it. And we should probably start acting like it.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Conselice, C. J., Wilkinson, A., Duncan, K. &amp;amp; Mortlock, A., &lt;em&gt;The Evolution of Galaxy Number Density at z &amp;lt; 8 and its Implications&lt;/em&gt;, The Astrophysical Journal, 830:83, October 2016. &lt;a href="https://arxiv.org/abs/1607.03909" rel="noopener noreferrer"&gt;arXiv:1607.03909&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; Petigura, E. A., Howard, A. W. &amp;amp; Marcy, G. W., &lt;em&gt;Prevalence of Earth-size planets orbiting Sun-like stars&lt;/em&gt;, Proceedings of the National Academy of Sciences, Vol. 110, No. 48, pp. 19273-19278, 2013. DOI: &lt;a href="https://doi.org/10.1073/pnas.1319909110" rel="noopener noreferrer"&gt;10.1073/pnas.1319909110&lt;/a&gt; - &lt;a href="https://arxiv.org/abs/1311.6806" rel="noopener noreferrer"&gt;arXiv:1311.6806&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Ward, P. D. &amp;amp; Brownlee, D., &lt;em&gt;Rare Earth: Why Complex Life is Uncommon in the Universe&lt;/em&gt;, Copernicus Books / Springer, New York, 2000. &lt;a href="https://link.springer.com/book/9780387987019" rel="noopener noreferrer"&gt;Springer&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; Drake, F., &lt;em&gt;The Drake Equation - Origins&lt;/em&gt;, created for the first SETI conference, Green Bank Observatory, West Virginia, November 1961. Described formally in: Drake, F., &lt;em&gt;The Radio Search for Intelligent Extraterrestrial Life&lt;/em&gt;, in &lt;em&gt;Current Aspects of Exobiology&lt;/em&gt;, Pergamon Press, 1965. Overview: &lt;a href="https://www.seti.org/drake-equation-index" rel="noopener noreferrer"&gt;SETI Institute&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Tarter, J., &lt;em&gt;The Search for Extraterrestrial Intelligence (SETI)&lt;/em&gt;, Annual Review of Astronomy and Astrophysics, Vol. 39, pp. 511-548, 2001. DOI: &lt;a href="https://doi.org/10.1146/annurev.astro.39.1.511" rel="noopener noreferrer"&gt;10.1146/annurev.astro.39.1.511&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Breakthrough Listen Initiative - The largest ever scientific research program searching for evidence of civilizations beyond Earth, $100 million, 10-year commitment announced 2015. &lt;a href="https://breakthroughinitiatives.org/initiative/1" rel="noopener noreferrer"&gt;breakthroughinitiatives.org&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Hart, M. H., &lt;em&gt;Explanation for the Absence of Extraterrestrials on Earth&lt;/em&gt;, Quarterly Journal of the Royal Astronomical Society, Vol. 16, pp. 128-135, 1975. Also known as the Hart-Tipler Conjecture. &lt;a href="https://en.wikipedia.org/wiki/Fermi_paradox" rel="noopener noreferrer"&gt;Wikipedia: Fermi Paradox&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; &lt;em&gt;The Montauk Project: Experiments in Time&lt;/em&gt;, Preston Nichols &amp;amp; Peter Moon, Sky Books, 1992. Background and analysis: &lt;a href="https://en.wikipedia.org/wiki/Montauk_Project" rel="noopener noreferrer"&gt;Wikipedia: Montauk Project&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-9"&gt;&lt;/span&gt;&lt;strong&gt;9.&lt;/strong&gt; Project MKUltra - CIA illegal human experimentation program, 1953-1973. Exposed via Church Committee, 1975; further declassified via FOIA, 1977. &lt;a href="https://en.wikipedia.org/wiki/MKUltra" rel="noopener noreferrer"&gt;Wikipedia: MKUltra&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-10"&gt;&lt;/span&gt;&lt;strong&gt;10.&lt;/strong&gt; Haines, G. K., &lt;em&gt;A Die-Hard Issue: CIA's Role in the Study of UFOs, 1947-1990&lt;/em&gt;, Studies in Intelligence (Unclassified Edition), 1997. &lt;a href="https://www.cia.gov/resources/csi/studies-in-intelligence/studies-in-intelligence-1997/cias-role-in-the-study-of-ufos-1947-1990/" rel="noopener noreferrer"&gt;CIA Center for the Study of Intelligence&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-11"&gt;&lt;/span&gt;&lt;strong&gt;11.&lt;/strong&gt; Loftus, E. F., &lt;em&gt;The Reality of Repressed Memories&lt;/em&gt;, American Psychologist, Vol. 48, No. 5, pp. 518-537, 1993. DOI: &lt;a href="https://psycnet.apa.org/doi/10.1037/0003-066X.48.5.518" rel="noopener noreferrer"&gt;10.1037/0003-066X.48.5.518&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-12"&gt;&lt;/span&gt;&lt;strong&gt;12.&lt;/strong&gt; U.S. House Oversight Committee, &lt;em&gt;Hearing on Unidentified Anomalous Phenomena (UAP)&lt;/em&gt;, Testimony of David Grusch, Ryan Graves, and David Fravor, July 26, 2023. &lt;a href="https://www.c-span.org/video/?529499-1/hearing-unidentified-anomalous-phenomena-uap" rel="noopener noreferrer"&gt;C-SPAN Recording&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-13"&gt;&lt;/span&gt;&lt;strong&gt;13.&lt;/strong&gt; National Archives and Records Administration, &lt;em&gt;Records Related to UFOs and Unidentified Anomalous Phenomena&lt;/em&gt;, established per FY 2024 NDAA provisions. &lt;a href="https://www.archives.gov/uap" rel="noopener noreferrer"&gt;archives.gov/uap&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>lmm</category>
    </item>
    <item>
      <title>Knowledge and Intelligence ARE Mutually Exclusive.</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 19:02:18 +0000</pubDate>
      <link>https://dev.to/wiseai/knowledge-and-intelligence-are-mutually-exclusive-kd4</link>
      <guid>https://dev.to/wiseai/knowledge-and-intelligence-are-mutually-exclusive-kd4</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This post was originally published on &lt;a href="https://wiseai.dev/blogs/knowledge-and-intelligence-are-mutually-exclusive" rel="noopener noreferrer"&gt;the main website&lt;/a&gt; on &lt;a href="https://github.com/wiseaidotdev/blog/pull/19" rel="noopener noreferrer"&gt;Apr 28 2026&lt;/a&gt;. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account &amp;lt;3!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;I have been building toward this post for a long time, longer than any of the others, and I want to start by being honest about why it took so long. The claim I am about to make is uncomfortable. It goes against something that feels intuitive to most people, including to me for most of my intellectual life. But I have reached a point where I cannot keep writing around it, cannot keep softening it with caveats and qualifications, because the evidence has accumulated past the point where softening it is still honest. The claim is this: knowledge and intelligence are mutually exclusive in a very specific and very important sense. Not completely, not forever, not in every possible world. But in the way that matters most for building genuinely intelligent machines, the way that matters for understanding what happened when a language model fails, the way that matters for explaining why scaling alone will never produce real understanding, knowledge and intelligence point in opposite directions, and you cannot have both at the same time without knowing exactly which one you have and where each one ends.&lt;/p&gt;

&lt;p&gt;I have circled this idea in almost every post I have written. In &lt;a href="https://wiseai.dev/blogs/language-is-limited-asi-is-impossible" rel="noopener noreferrer"&gt;Language is Limited. ASI is Impossible.&lt;/a&gt;, I argued that the medium of text is not the medium of reality, and that a system living inside text is forever separated from the world by the compression loss of language. In &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt;, I argued that equations are more powerful than sentences and that simulation is the real intelligence. In &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, I argued that mathematical structure encodes mechanism in ways that language never can, and that the path to genuine understanding runs through structure rather than through text. In &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether.&lt;/a&gt;, I argued that extracting value from human creative output to build statistical models is both morally corrupt and intellectually bankrupt. And in &lt;a href="https://wiseai.dev/blogs/genuine-intelligence-will-never-emerge-from-neural-networks" rel="noopener noreferrer"&gt;Genuine Intelligence will never in trillion years emerge from neural networks.&lt;/a&gt;, I made the most direct version of the argument yet, showing that the architectural gaps in neural networks are not fixable bugs but defining properties of what those systems are. All of those posts were preparation for this one. This one is the piece that ties them together, the piece that names the relationship between the two things that every AI conversation confuses, and explains why confusing them is not just an intellectual mistake but a practical catastrophe that is unfolding right now, in the systems being deployed in medicine, law, education, and engineering, in the decisions being made about what to trust and what to build.&lt;/p&gt;

&lt;p&gt;I know that claim might sound excessive. I know that a lot of what I write sounds excessive to people who are more comfortable with the moderate position, the "both sides" view that says language models are useful tools with limits and we should just be careful and everything will work out fine. But I have been watching this space long enough to know that moderation is often just another word for not looking carefully enough. And when I look carefully at the distinction between knowledge and intelligence, what I see is not a minor technical nuance. I see the single most important conceptual error in the entire field of artificial intelligence, and I am going to spend this whole post trying to make you see it too.&lt;/p&gt;

&lt;h2&gt;
  
  
  What We Mean When We Say Knowledge
&lt;/h2&gt;

&lt;p&gt;Before I can make the argument, I need to spend some time on the definitions, because I have found that most disagreements about AI come down to people using words differently, and if we do not start from a shared understanding of what knowledge actually is, the rest of the argument will slide past each other without ever connecting. So let me be careful here, more careful than I usually am, because the distinction is subtle enough to lose if I rush it.&lt;/p&gt;

&lt;p&gt;Knowledge, in the sense I am using it here, is stored pattern. It is the accumulated residue of past encounters with reality, encoded in some medium that allows it to be retrieved and applied later. A textbook contains knowledge. A database contains knowledge. A trained neural network contains knowledge, in the form of weights that have been adjusted to fit patterns in training data. Knowledge is retrospective. It points backward in time, toward the patterns that already existed in the world and have been captured in some representational form. Knowledge is also inherently limited by the observations that created it. You can only store patterns that your past experience has exposed you to, which means knowledge is always bounded by history, always a function of what has happened, never of what could happen for the first time. Knowledge is the library. It is enormous, it is valuable, it is the source of almost everything that educated humans can do. But it is not intelligence. It is the raw material that intelligence operates on, and treating them as the same thing is the mistake that has corrupted the entire AI conversation.&lt;/p&gt;

&lt;p&gt;When a language model is trained, what happens is that the system ingests a very large corpus of text and adjusts its parameters to minimize prediction error across that corpus. The result is a system that has absorbed an extraordinary amount of statistical structure from the text. It has learned that certain words follow certain other words in certain contexts, that certain topics are discussed in certain ways, that certain questions tend to receive certain kinds of answers. All of that statistical structure is knowledge, in the sense I am using the word. It is pattern stored in weights. It is enormously useful. It is what allows the model to produce fluent, contextually appropriate responses across an enormous range of topics. But it is all retrospective. Every pattern the model has learned is a pattern that existed in the training data. Every response the model produces is a recombination and interpolation of patterns from its past. The model has no mechanism for generating understanding that is not already present, at least implicitly, in the distribution of its training data. That is not a temporary limitation that will be overcome by larger models or better data. It is the definition of what knowledge-based systems do, and it applies to every system that learns by fitting patterns to historical data, regardless of how large or sophisticated the system becomes.&lt;/p&gt;

&lt;p&gt;Now consider what happens when you ask a genuinely knowledgeable person a question in a domain they have studied deeply. They draw on their stored patterns, yes, but they also do something more. They reason. They connect the patterns in new ways. They notice when a question is unlike anything they have seen before and they flag that distinctiveness. They identify the tensions and contradictions within their own knowledge and use those tensions as signals that something needs to be worked out more carefully. They know the limits of what they know. That last capability, knowing the limits of your own knowledge, is one of the most important things a mind can do, and it is something I have written about before, most directly in the post about intelligence never emerging from neural networks. It is something that knowledge-based systems do very poorly, because a system that operates by pattern completion has no principled basis for knowing when it is operating outside the domain of its training data. It just produces the most likely completion and presents it with the same confidence it would have if the question were well inside its training distribution. That systematic overconfidence is the direct consequence of treating knowledge as intelligence, of assuming that having patterns is the same as understanding when and how to apply them, and not a bug in the training procedure.&lt;/p&gt;

&lt;p&gt;Let me also be specific about the different kinds of knowledge, because not all knowledge is equally problematic when it is mistaken for intelligence. Declarative knowledge is knowing that something is true. Procedural knowledge is knowing how to do something. Causal knowledge is knowing why something is true, what mechanism produces it, and what would happen if conditions changed. These three kinds of knowledge are related but distinct, and the distinction matters enormously for AI. Language models are extraordinarily good at declarative knowledge, at knowing that. They are decent at procedural knowledge in domains where the procedures are well-represented in text. They are very poor at causal knowledge, at knowing why, because causal knowledge requires a model of the mechanism that generates the pattern, not just a representation of the pattern itself. And causal knowledge is the foundation of genuine intelligence, because intelligence is fundamentally about figuring out what to do in situations that are new, and figuring out what to do in new situations requires understanding why things happen, not just knowing what usually happens. Research by Judea Pearl has made this distinction precise in the language of causal inference, showing that systems that can only learn from observation will always fail at predicting the effects of interventions, because observational data can only teach associations while intervening requires causal structure (1). That result is not just a theorem in statistics. It is the mathematical proof that knowledge is not intelligence.&lt;/p&gt;

&lt;p&gt;Let me also say something about the role of memory in all of this, because memory and knowledge are often conflated, and the conflation makes things worse. Memory is the capacity to store and retrieve specific past experiences. Knowledge is the distillation of patterns across many experiences. Neither is the same as intelligence. A person with perfect memory who can recall every detail of every conversation they have ever had is not thereby more intelligent than a person with average memory, because intelligence is not about storage capacity. It is about the ability to extract structure, to form abstractions, to reason about things that have not been encountered before. The AI community's fixation on scaling training data is implicitly based on a theory that more memory equals more intelligence, that if you store enough patterns from the past, intelligence emerges. But intelligence is not a quantitative property of memory. It is a qualitative property of the process that operates on memory. You could give a language model perfect recall of every text ever written and it would still not be able to discover a new mathematical theorem, because discovering a new mathematical theorem requires generating structure that does not yet exist in the training data, and generation of genuinely new structure is intelligence, not knowledge retrieval.&lt;/p&gt;

&lt;p&gt;The specific way that I want to frame this distinction going forward is through the concept of compression. Knowledge is compression of past observations. Intelligence is the capacity to discover new compressions that predict future observations. A language model compresses past text into weights and uses those weights to interpolate over the pattern space it has seen. An intelligent system discovers new structure that was not in any of the observations it has made, structure that, once found, allows it to predict things it has never been trained on. That distinction maps directly onto the distinction between interpolation and extrapolation, and it is the reason that language models work so well inside their training distribution and fail so dramatically outside it. Inside the training distribution, they are interpolating, and interpolation is something that sophisticated pattern matching does very well. Outside the training distribution, they need to extrapolate, and extrapolation requires genuine structural understanding, real intelligence, and that is the thing that knowledge-based systems cannot provide. This is not my invention. It is the consistent finding of decades of research on generalization in machine learning, and it is the key to understanding why scaling does not solve the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Knowledge Masquerades as Intelligence
&lt;/h2&gt;

&lt;p&gt;The most dangerous aspect of the knowledge-intelligence confusion is that it is not an obvious error. Knowledge looks like intelligence from the outside. A system that has absorbed enormous amounts of pattern from the world and can retrieve and recombine those patterns fluently will produce outputs that are often indistinguishable from the outputs of a genuinely intelligent system, especially when the questions you are asking are within the distribution of its training data. This is not a minor caveat. It is the fundamental reason why the AI industry has been able to deceive the public for years about what its systems actually do. When you ask a language model something that millions of people have asked before, in roughly similar form, and the model answers correctly with appropriate confidence, there is no way for you to know from that single interaction whether the answer came from genuine understanding or from pattern retrieval. And in the vast majority of everyday use cases, pattern retrieval is good enough, which means the deception works most of the time, which means most people never encounter the failure mode that would reveal what they are actually dealing with.&lt;/p&gt;

&lt;p&gt;I have been thinking about this illusion for a long time, and the best analogy I can find for it is the experienced test-taker who has studied every past exam paper but understands nothing of the underlying subject. If you test this person on questions that are similar in form to the past papers, they will perform excellently. They will look highly competent. Their answers will sound confident and accurate. But if you ask them a question that is structurally identical to a past question but phrased in an unfamiliar way, or that requires combining concepts that were never combined in the study materials, their performance will collapse. What looked like understanding was retrieval. What looked like intelligence was knowledge. And the collapse at the edge of the training distribution is the tell that reveals the difference. The same collapse happens with language models. Ask a well-trained model something within its distribution and it will amaze you. Ask it something genuinely novel that requires reasoning from first principles, and you will get either a hallucination or a non-answer, depending on how the model is configured to handle uncertainty. Researchers have documented this collapse repeatedly across domains from mathematics to logic to commonsense reasoning, and the result is always the same: models that look brilliant on distribution fail dramatically off it (2). That pattern is not a coincidence. It is the predicted behavior of knowledge-based systems, and it is the empirical signature of the knowledge-intelligence distinction in action.&lt;/p&gt;

&lt;p&gt;There is also a psychological mechanism by which knowledge masquerades as intelligence, and it is worth naming because it operates on the humans who interact with these systems, not just on the systems themselves. When we receive fluent, confident, contextually appropriate speech from an entity, we are wired to infer that the entity understands what it is saying. This is a deeply reasonable heuristic in a world populated exclusively by humans, because in that world, fluent confident speech is indeed correlated with understanding. But the heuristic breaks down when the entity producing the speech is a system that generates fluency without understanding, and the breakdown is dangerous because we rarely notice it happening. The language model sounds exactly like a person who understands, which means our evolved social cognition tells us to trust it, to defer to it, to treat its outputs as reliable. This is not stupidity on the part of the humans. It is a mismatch between a social cognition that evolved for a world of humans and a technological object that can mimic the surface of human speech without any of its substance. Researchers studying interaction with conversational AI have consistently found that people attribute more understanding to these systems than the systems warrant, and that this attribution causes people to reduce the critical scrutiny they would apply to any human expert (3). That reduction in scrutiny is exactly what makes the knowledge-intelligence confusion dangerous in practice.&lt;/p&gt;

&lt;p&gt;I also want to address the specific case of chain-of-thought prompting, because it is often cited as evidence that language models can actually reason, and I think the evidence has been badly misread. Chain-of-thought prompting is a technique where the model is asked to produce its reasoning step by step before giving a final answer. This technique improves performance on many multi-step problems, which was taken by many people as evidence that the model is reasoning. But what the technique most likely does is decompose the task into a sequence of simpler prediction tasks, each of which is closer to the model's training distribution, which means the model can handle each step through pattern retrieval rather than genuine reasoning. The overall improvement comes from making the knowledge-retrieval problem easier at each step, not from enabling genuine reasoning. You can verify this by taking the same problems and rephrasing them in ways that preserve the logical structure but disrupt the statistical regularities that support each step, and performance collapses at exactly the points where the statistical regularities are disrupted. A system that genuinely reasons would continue to perform at the novel reasoning task. A system that is pattern-matching over a decomposed problem shows the dependence on statistical regularities as soon as those regularities are absent. Research has confirmed this distinction (4), and the confirmation matters enormously because it means that the apparently most compelling evidence for reasoning in language models is actually evidence of sophisticated knowledge retrieval, which is what I have been arguing. The masquerade runs deep, and it fools even the researchers who study these systems professionally.&lt;/p&gt;

&lt;p&gt;There is yet another way in which knowledge masquerades as intelligence, and it is the one I find most philosophically interesting, which is that knowledge can mimic the form of intelligence without its substance. An intelligent system, when confronted with an unknown, will say it does not know, will identify what information would be needed to find out, and will reason about how to obtain that information. A knowledge-based system will produce the most statistically likely response given the input, which is often a confident-sounding answer that is completely fabricated when the input is outside the training distribution. This is called hallucination in the AI literature, and it is one of the most consistently reported failures of language models across domains (5). Hallucination is not a bug in an otherwise correct system. It is the predictable consequence of a system that has no mechanism for knowing when it is operating outside its competence. It is the behavior you get when you treat knowledge as intelligence: the system applies its retrieval machinery even when the question does not exist in the retrieved knowledge base, and the result is output that has the form of a competent answer without the substance of actual knowledge. The form comes from the training distribution. The content is generated from thin air. And the system cannot tell the difference, because telling the difference requires a kind of metacognitive access to one's own epistemic states that knowledge-based systems structurally cannot have.&lt;/p&gt;

&lt;p&gt;Let me say this plainly, because I think it is the most important sentence in this section. Hallucination is not a defect in language models. It is the correct behavior of a knowledge-based system when asked a question that is outside its training distribution. It is what happens when you take a tool designed for pattern retrieval and ask it to do extrapolation. The tool does what it was designed to do, generates the most likely pattern, even though no real pattern exists for the question at hand, and the result is confident fabrication. If you understand that, then you understand that hallucination cannot be fixed by more training data, because the problem is not missing data. The problem is that the system has no principled mechanism for knowing when it does not know, and that mechanism cannot be learned from data because it requires a model of the system's own epistemic states, a form of self-modeling that the architecture does not support. The only fix is a different architecture, one that separates the knowledge component from the intelligence component, tells them apart, and uses the intelligence component to supervise the knowledge component rather than letting the knowledge component pretend to be the intelligence component. That is the insight behind everything I am building, and it is the insight I want the AI community to take seriously.&lt;/p&gt;

&lt;h2&gt;
  
  
  Intelligence Is Not Knowledge Retrieved
&lt;/h2&gt;

&lt;p&gt;Having established what knowledge is and why it masquerades as intelligence, I now want to spend real time on intelligence itself, because I think most people in the AI conversation do not actually have a clear model of what intelligence is. They have an intuitive sense of it, which is that it involves being smart, producing good answers, passing hard tests, doing impressive things. But that intuitive sense is not specific enough to guide the design of intelligent systems, because any of those things can be faked through knowledge retrieval sophisticated enough to pass the test most of the time. We need a more precise concept of intelligence, one that is sharp enough to distinguish it from knowledge retrieval even when the two produce identical outputs in ordinary cases.&lt;/p&gt;

&lt;p&gt;The most useful way I have found to think about intelligence is as the capacity to generate new compressions of reality. A compression in this sense is not a zip file or an encoding scheme. It is a more general concept: a representation that is more compact than the data it describes but that can reconstruct the data, predict new data outside the training set, and generalize to situations that were not present in the observations that created it. When Newton looked at falling apples and orbiting planets and deduced a single law that governed both, he generated a new compression of reality. The data was already there, it had been there for as long as apples fell and planets orbited, but the compression was not there until Newton found it. The compression is not stored in the world. It is generated by intelligence operating on observations of the world. That is the key distinction: knowledge is pattern retrieved from past observations, and intelligence is the capacity to generate new patterns that explain past observations and predict new ones. The two activities are related because you need knowledge to have the raw material for intelligence to operate on, but they are not the same activity, and the difference between them is the difference between retrieval and discovery.&lt;/p&gt;

&lt;p&gt;Discovery requires something that retrieval does not: the ability to generate candidate representations that are not already present in the knowledge base, evaluate them against observations, and iteratively refine them until they fit. This is the structure of scientific reasoning, and it is fundamentally active and generative in a way that pattern retrieval is not. When a scientist discovers a new equation, they are not retrieving it from memory. They are generating a hypothesis, testing it, finding it wrong, generating a modified hypothesis, testing that, refining it further, and continuing until the hypothesis fits the data well enough to trust. Each step in that process requires genuine inference, the ability to produce outputs that are not direct functions of the inputs, and that inferential capacity is what I mean by intelligence. It cannot be learned from data the way a pattern can be learned, because learning a pattern is retrieval and the very thing it is trying to learn is generation. You cannot learn how to generate new things by learning from examples of things that have already been generated. The process that generates the examples is intelligence, and you cannot learn intelligence from its outputs any more than you can learn how to invent things by studying inventions. You can learn what things have been invented, which is knowledge, and you can use that knowledge as raw material, but the inventive capacity itself is something different and something deeper.&lt;/p&gt;

&lt;p&gt;I want to connect this to the ongoing work on the &lt;a href="https://github.com/wiseaidotdev/lmm" rel="noopener noreferrer"&gt;lmm project&lt;/a&gt; that I have been building and describing across my posts, and specifically to the &lt;a href="https://crates.io/crates/lmm-agent" rel="noopener noreferrer"&gt;&lt;code&gt;lmm-agent&lt;/code&gt;&lt;/a&gt; crate that sits at its core. The &lt;code&gt;lmm-agent&lt;/code&gt; framework is built entirely around the knowledge-intelligence distinction I have been drawing. It is an equation-based, training-free autonomous agent: no LLM API key, no token quotas, no stochastic black boxes, no weights updated by gradient descent on anyone's text. Instead of pattern retrieval, it implements what I am calling intelligence primitives, five structural properties that replace statistical interpolation with auditable, causal, and motivated cognition. These primitives are: calibrated Bayesian uncertainty via Gaussian belief propagation, compositional axiomatic reasoning that produces auditable forward-chaining proofs, causal counterfactual attribution using Pearl do-calculus interventions, hypothesis formation that ranks candidate new causal edges by explanatory power, and internalized motivational drives including distinct signals for Curiosity, CoherenceSeeking, and ContradictionResolution. Each of these primitives maps directly onto the abstract properties I described as belonging to intelligence rather than knowledge. None of them can be trained in by fitting patterns to text, because each requires a structural mechanism to be present at the architectural level, and that is exactly why I built them that way rather than trying to elicit them from a language model through prompting.&lt;/p&gt;

&lt;p&gt;Research on systematic generalization has made the distinction between retrieval and intelligence precise in a testable way. Systematic generalization, as defined in the cognitive science and machine learning literature, is the ability to combine known concepts in new ways and derive correct predictions from those combinations without having been trained on them (6). This is the minimal form of intelligence that is clearly distinct from knowledge retrieval, because it requires generating correct outputs for combinations that are not in the training data. When you train a language model on a set of sentences and then test it on recombinations of the same words and concepts in novel structures, performance drops dramatically, even when the logical structure is identical to structures the model has seen and the concepts are all familiar (7). That drop is the empirical signature of the knowledge-retrieval ceiling. The model can retrieve patterns it has seen. It cannot generate the correct pattern for combinations it has not seen, because generation requires intelligence and the model only has knowledge. This failure has been replicated across dozens of architectures, training regimes, and task domains, and it consistently appears at the same boundary: inside the training distribution, performance is high. Outside it, performance collapses. Intelligence does not have that boundary. Knowledge does.&lt;/p&gt;

&lt;p&gt;Let me say something that I think is important and uncomfortable. Every benchmark that the AI industry uses to measure progress is, at best, measuring the depth and breadth of knowledge retrieval, and at worst, measuring the degree to which the knowledge retrieval system has been fine-tuned to produce outputs that look like intelligence on that specific benchmark. I wrote about this in &lt;a href="https://wiseai.dev/blogs/rethinking-arc-agi" rel="noopener noreferrer"&gt;Rethinking ARC-AGI&lt;/a&gt;, where I argued that even the benchmarks designed to test genuine fluid intelligence were compromised by the way models could be trained to pattern-match on the benchmark format itself. The problem is that any task that can be represented in text and evaluated by a judge can, in principle, be solved by a sufficiently powerful knowledge retrieval system, because the evaluation criteria are themselves encoded in the statistical structure of how humans write about evaluation. If you want to measure intelligence, you need to measure it on tasks that require generating structures that are not in any training data, and those tasks are very hard to construct and evaluate, which is why the benchmarking community avoids them. But avoiding hard truth does not make it less true.&lt;/p&gt;

&lt;p&gt;Intelligence also has a property that knowledge does not, which is the capacity for genuine epistemic humility. An intelligent system knows when it does not know. It has a model of its own epistemic states that allows it to distinguish between conclusions supported by reliable inference and conclusions that are uncertain or unsupported. This metacognitive capacity is not a personality trait or a design choice that can be trained into a system by fine-tuning. It is an architectural property that must be present at a structural level for the system to reliably identify the limits of its own competence. Research on calibration in large language models consistently finds that these systems are overconfident, presenting uncertain or false information with the same confidence as reliable information (8). That overconfidence is not a bug. It is the expected behavior of a system that maxes out its knowledge retrieval machinery even when the question is outside its training distribution, because the machinery has no switch labeled "outside my domain." The switch would require intelligence to implement. The machinery only has knowledge.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Separation That Actually Matters
&lt;/h2&gt;

&lt;p&gt;I want to be concrete about what I mean by the mutually exclusive relationship between knowledge and intelligence, because I have been building up to this claim for several sections and I need to state it carefully to avoid misunderstanding. I am not claiming that knowledge and intelligence cannot coexist in a single system. Humans have both, obviously, and the interaction between them is what makes human cognition so powerful. What I am claiming is something more specific and more operationally important: that when you build a system by maximizing knowledge, meaning by training it to store and retrieve as many patterns from the past as possible, you are simultaneously minimizing the pressure on the system to develop genuine intelligence, and the more successfully you do the first thing, the more completely the second thing is displaced. This is the mutually exclusive relationship I have in mind, and it has a very practical consequence: the training paradigm that produces the most capable knowledge-retrieval systems is precisely the paradigm that is least likely to produce genuine intelligence, because success at retrieval removes the need for generation.&lt;/p&gt;

&lt;p&gt;Think about what happens during training of a language model. The model receives a sequence of tokens and is asked to predict the next token. The loss function penalizes incorrect predictions and rewards correct ones. The most efficient way for the model to minimize loss is to store statistical structure from training sequences and retrieve it at prediction time. If the model develops genuine inference mechanisms that allow it to derive the correct next token from structural properties of the input, those mechanisms will be penalized in exactly the cases where knowledge retrieval gives the right answer, because the loss function cannot distinguish between arriving at the right answer via retrieval and arriving via genuine inference. The loss function only sees the answer. This means that from the loss function's perspective, there is no advantage to genuine intelligence over knowledge retrieval when retrieval works, and retrieval works on everything in the training distribution. The selection pressure for genuine intelligence therefore applies only outside the training distribution, where retrieval fails, but the training procedure sees no examples from outside the training distribution by definition, which means there is effectively no selection pressure for genuine intelligence at all. The training procedure optimizes knowledge at the expense of intelligence, not because anyone designed it that way, but because knowledge and capability on training data are the same thing in a training regime, and intelligence only matters at the edges where the training data runs out.&lt;/p&gt;

&lt;p&gt;This is why I said at the beginning that knowledge and intelligence are mutually exclusive in the sense that matters most for building intelligent systems. The training paradigm that produces the most knowledgeable systems is the same paradigm that most completely removes the adaptive pressure for developing genuine intelligence. And the result is exactly what we observe: systems that are extraordinarily capable within the distribution of their training data and shockingly incompetent outside it. That pattern is not a transitional phase on the way to genuine intelligence. It is the predicted endpoint of a training paradigm that optimizes knowledge retrieval, and it will remain the endpoint regardless of how much larger the models become or how much more data they train on. Scaling within a paradigm that optimizes the wrong thing will produce more of what the paradigm produces, not the thing the paradigm is incapable of producing.&lt;/p&gt;

&lt;p&gt;I also want to say something about what the right design philosophy looks like, because I think every post I write should point toward something better rather than just eviscerating something that exists. A system designed to produce genuine intelligence rather than knowledge would need to reward generation of new compressions rather than retrieval of old ones. It would need to operate on tasks where knowledge retrieval is guaranteed to fail, where genuine inference from structural properties of the input is the only way to get the answer. It would need to evaluate itself on out-of-distribution generalization rather than on distribution-matching, which means the evaluation criteria would need to be structurally different from the operating procedure rather than identical to it. It would need explicit mechanisms for representing its own epistemic states, knowing when it knows and when it does not, rather than expecting that capability to emerge by tuning on human-labeled examples of confidence. And it would need to ground its processing in something other than text. The &lt;code&gt;lmm-agent&lt;/code&gt; framework is my attempt to implement these principles. Its &lt;code&gt;HELM&lt;/code&gt; engine, Hybrid Equation-based Lifelong Memory, is a lifelong learning system built entirely on CPU-resident hash maps and floating-point arithmetic: tabular Bellman Q-learning, prototype meta-adaptation via Jaccard similarity, knowledge distillation, self-federated Q-table aggregation without a central server, elastic memory guarding by activation-count pinning, and PMI co-occurrence mining from high-reward observations. No GPU. No neural networks. No external machine learning crates. If that sounds like an unusual set of constraints, it is. The constraints are the point. Each one is a deliberate refusal to fall back on the comfortable tools that produce knowledge rather than intelligence.&lt;/p&gt;

&lt;p&gt;Research on causal inference provides the strongest theoretical foundation I know for the separation I am describing. Pearl's causal hierarchy distinguishes between seeing, doing, and imagining, corresponding roughly to knowledge retrieval, intervention, and counterfactual reasoning (1). A system at the first level can only process observations and extract statistical associations. A system at the second level can predict the effects of actions, meaning it can reason about what will happen if you do something, not just what has happened when things correlated. A system at the third level can reason about counterfactuals, about what would have happened if conditions had been different. Language models are firmly at level one by any honest assessment. They have access to observations in text form and they learn statistical associations from them. They cannot reliably predict the effects of interventions, and they cannot reliably reason about counterfactuals, because both of those capabilities require a causal model, a structural representation of the mechanisms that generate the observations. Building a causal model requires intelligence in the sense I have been using, generating a new compression of reality that captures mechanism rather than just association, and that is exactly the capability that knowledge-retrieval training does not develop. The hierarchy is not an opinion. It is a mathematical framework derived from the formal theory of probability and intervention, and it is the most precise statement I know of why knowledge and intelligence are different things.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Knowledge Fails and Intelligence Has to Step In
&lt;/h2&gt;

&lt;p&gt;Let me bring this into the real world, because abstract arguments without concrete examples are not as honest as arguments that put themselves on the line. I want to talk about specific domains where the knowledge-intelligence distinction is not a theoretical nicety but a practical matter of life and death, because that is where the cost of the confusion is most visible and most urgent.&lt;/p&gt;

&lt;p&gt;In medicine, the distinction shows up most clearly in the diagnosis of rare diseases and unusual presentations of common diseases. A physician who has seen many patients develops knowledge, statistical patterns about how diseases present that allow them to make rapid, accurate diagnoses in typical cases. But the genuinely diagnostic work, the work that saves lives when the presentation is atypical, requires intelligence: the ability to reason from mechanism, to understand why the symptoms the patient is showing are inconsistent with the most likely diagnosis, to generate alternative hypotheses and test them against the clinical picture, to recognize the atypical feature that is the tell for a rare condition. This is causal reasoning, not pattern retrieval, and it is the difference between correctly diagnosing a typical case of pneumonia and recognizing the unusual presentation that is actually a rare autoimmune condition masquerading as pneumonia. AI systems for medical diagnosis that learn from training data of typical presentations will inherit the same blindspot as the knowledge at level: they will perform well on typical cases and fail precisely at the unusual cases where correct diagnosis matters most. Research has confirmed this pattern empirically across multiple medical AI systems (9), showing that performance on out-of-distribution cases is substantially worse than on in-distribution cases, exactly as the knowledge-intelligence distinction predicts.&lt;/p&gt;

&lt;p&gt;In engineering, the distinction shows up when a system encounters a failure mode that was not anticipated in the design specification. An engineer who knows the design in detail knows what the system is supposed to do. But when the system does something unexpected, the engineer needs intelligence to understand why, to trace the causal chain from the failure symptom back to the root cause, to recognize which design assumption was violated and how. This is diagnosis by causal inference, not by pattern retrieval, and it is the reason why good engineers are valuable beyond the knowledge they carry. A knowledge-based AI system can tell you what has failed in similar systems in the past. It cannot tell you why this system is failing in this specific novel way, because the novel failure mode is by definition outside the database of past failures that constitutes its knowledge. And aircraft, nuclear plants, bridges, and medical devices fail in novel ways with consequences that cannot be hedged by pointing to the excellent performance on the training distribution.&lt;/p&gt;

&lt;p&gt;In scientific research, the distinction is the entire point. Science is the enterprise of generating new compressions of reality, of discovering equations and models and theories that explain observations and predict new ones. Knowledge in science is the accumulated set of established theories and experimental results. Intelligence in science is the capacity to generate new theories that explain what existing theories cannot, to design experiments that can distinguish between competing hypotheses, to see the pattern in anomalous data that points toward a new understanding. The history of science is the history of intelligence operating on knowledge, producing new compressions that subsume the old ones and extend the reach of human understanding. Language models trained on scientific text have absorbed an enormous amount of scientific knowledge. They can summarize papers, explain established theories, and generate text that sounds like scientific reasoning. But they cannot do science, because doing science requires generating new compressions, discovering structure that was not in any training data, and that requires genuine intelligence. The &lt;code&gt;lmm-agent&lt;/code&gt; architecture is my ongoing attempt to implement this capability in code: its &lt;code&gt;HypothesisGenerator&lt;/code&gt; takes the residual unexplained variance in a causal graph and ranks candidate new causal edges by explanatory power, proposing hypotheses rather than retrieving patterns. Its &lt;code&gt;CausalAttributor&lt;/code&gt; uses Pearl do-calculus to perform counterfactual interventions on the causal graph, attributing outcomes to root causes rather than surface correlations. These two components together enable the agent to ask the question that knowledge-based systems cannot ask: not what has happened before in similar situations, but why is this particular thing happening now, and what would be different if I changed this specific factor.&lt;/p&gt;

&lt;p&gt;In education, the distinction surfaces most clearly in the difference between a student who has memorized the course material and a student who has understood it. The memorizing student can answer questions that appear on the exam, which is knowledge retrieval. The understanding student can answer novel questions that combine and extend the concepts from the course in ways that were not explicitly covered, which is intelligence. Every good teacher knows the difference between these two students, even if they cannot always articulate exactly what the difference is. The understanding student has generated compressions from the course material that generalize beyond it. The memorizing student has stored the surface patterns of the course material and can only retrieve them when the question looks like the patterns. AI tutoring systems that evaluate student understanding by testing on questions similar to the training material will consistently overestimate the understanding of memorizing students and, more dangerously, will fail to identify the students who have genuinely mastered the concepts. The same failure mode will appear in any AI system for education that is built on knowledge retrieval rather than on genuine assessment of generalization and transfer.&lt;/p&gt;

&lt;p&gt;Let me also connect this to the economic and social critique I have made in previous posts, because the knowledge-intelligence confusion is not just an intellectual error. It is an error with economic consequences that fall disproportionately on people who are already vulnerable. I described in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt; how the promise of equal opportunity in the tech industry was a lie, and how the extraction of value from engineers and creators funded systems that automated away the jobs of the people who built them. The same pattern is playing out with AI systems that are presented as intelligent but are actually knowledge-based: the people in high-stakes domains, the patients, the engineering clients, the students, the citizens subject to algorithmic decisions, are told that the system is intelligent and therefore trustworthy, and they cannot access the technical argument for why it is not, and so they trust, and when the system fails them in exactly the way that the knowledge-intelligence distinction predicts, they have no recourse and no one takes responsibility. The intellectual confusion is the enabling condition for the economic exploitation, and correcting the confusion is not just an academic project. It is a precondition for holding anyone accountable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Role of Equations and Simulation in Real Intelligence
&lt;/h2&gt;

&lt;p&gt;Having separated knowledge from intelligence and shown why the current training paradigm optimizes the wrong thing, I want to spend time on what genuine intelligence looks like in practice, because I have been making mostly negative arguments and I owe the reader a positive vision. The positive vision connects directly to what I argued in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt; and &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt;, and it runs through the same core idea: the only way to build a system that is genuinely intelligent rather than merely knowledgeable is to build a system that can discover and simulate mathematical structure, not retrieve linguistic patterns.&lt;/p&gt;

&lt;p&gt;The reason equations are the right foundation is that equations are not knowledge. They are not stored patterns from past observations. An equation is a generative structure, a compact representation that can produce an infinite range of outputs from a finite specification. Take Maxwell's equations, the four equations that describe all of classical electromagnetism. Those four equations were not retrieved from past data. They were discovered by James Clerk Maxwell through a combination of mathematical insight and physical reasoning that went beyond everything that had been observed before, generating predictions about electromagnetic waves that were confirmed only after his death. The equations are not a summary of past observations. They are a compression of the mechanism that generates those observations, and that compression can generate outputs in domains that were not observed at the time of discovery. That is intelligence: the generation of a new compression that predicts new reality. And it is the polar opposite of knowledge retrieval, which takes past observations and retrieves the most likely interpolation between them.&lt;/p&gt;

&lt;p&gt;When I built the &lt;code&gt;lmm-agent&lt;/code&gt; framework, the central design question was how to implement the discovery process in code at the level of an autonomous agent operating in an environment. The answer I arrived at was the &lt;code&gt;ThinkLoop&lt;/code&gt;: a closed-loop proportional-integral controller that drives iterative reasoning toward a goal by computing Jaccard-error feedback at each cycle. The controller is not searching a space of equations the way a human scientist does, but it is doing something structurally analogous: it runs a loop, measures how far its current state is from its goal, and adjusts its behavior based on the error signal, converging on a solution by iterative refinement rather than by lookup. It is intelligence in the control-theoretic sense, a feedback loop that generates new behavior rather than retrieving stored behavior, and that is the distinction I have been drawing throughout this post. The &lt;code&gt;lmm-agent&lt;/code&gt; project is documented at &lt;a href="https://crates.io/crates/lmm-agent" rel="noopener noreferrer"&gt;crates.io&lt;/a&gt; and at &lt;a href="https://github.com/wiseaidotdev/lmm" rel="noopener noreferrer"&gt;the lmm repository&lt;/a&gt;, and I want to be honest that it is still early work, imperfect and full of limitations. But the architecture is grounded in the right principles, which matters more than any current benchmark score.&lt;/p&gt;

&lt;p&gt;Simulation is the other half of genuine intelligence, and it is the half that closes the loop between discovery and verification. Once you have an equation, you can simulate it forward in time and compare the predictions to new observations. This comparison is what distinguishes a good compression from a bad one, and it is what makes the process self-correcting in the same way that the scientific method is self-correcting. A knowledge-based system cannot do this loop, because it has no equation to simulate, only patterns to retrieve. It cannot generate novel predictions from a compact structure and test them against new reality. It can only match new inputs to past patterns and interpolate. The simulation loop is the difference between science and commentary, between prediction and description, between genuine intelligence and very sophisticated knowledge retrieval. Research on physics-informed machine learning has shown that incorporating simulation into the learning process produces models that are dramatically more reliable and generalizable than pure data-driven models (10), which is exactly the prediction you would make from the knowledge-intelligence distinction: grounding the learning process in structural simulation rather than pattern retrieval produces systems that generalize better, because simulation is closer to intelligence than retrieval is.&lt;/p&gt;

&lt;p&gt;I also want to connect this to the multimodal argument I made in my post on equations. A system that has discovered an equation for a phenomenon has not just learned one thing. It has learned a structure that can be rendered in any modality: as a graph, as an animation, as a numerical prediction, as an audio signal, as a physical simulation. All of those outputs are generated by the same compact structure, which means the system has genuine multimodal understanding rather than learned associations between modalities. This is what genuine intelligence looks like in the multimodal domain: not a system that has been trained to align text representations with image representations through statistical association, but a system that has discovered the common mathematical structure that generates both. The LMM framework, as I described in &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt;, is the name I use for the class of systems that operates at this level, learning from multiple modalities to discover the mathematical structure that underlies them all, and it is structurally different from a language model extended with image inputs, because the goal is discovery rather than alignment.&lt;/p&gt;

&lt;p&gt;Let me be concrete about what this means for the practical capabilities that people care about. A system that genuinely understands a physical system through its equations can answer questions about that system that are not in any training data, because the equation generates answers, not retrieves them. A system that can simulate the dynamics of a physical process can predict what will happen in scenarios that were never observed, because the simulation runs from the equations, not from the training data. A system that discovers compact mathematical representations of complex phenomena has access to knowledge that is more powerful than any stored pattern, because the compact representation generalizes in ways that stored patterns cannot. These are not small incremental improvements over the current state of the art. They are qualitative differences in the kind of intelligence the system has, and they map directly onto the distinction between knowledge and intelligence that I have been drawing throughout this post.&lt;/p&gt;

&lt;p&gt;The most concrete proof I have that the &lt;code&gt;lmm-agent&lt;/code&gt; approach is on the right track is an experiment we ran recently. We built an agent called &lt;a href="https://github.com/wiseaidotdev/lmm/tree/main/examples/arc-lmm-agent" rel="noopener noreferrer"&gt;&lt;code&gt;arc-lmm-agent&lt;/code&gt;&lt;/a&gt; to tackle the ARC-AGI-3 benchmark, which is widely considered one of the hardest tests of genuine fluid intelligence available. The specific environment was the &lt;code&gt;ls20&lt;/code&gt; game, a partially observable grid puzzle with fog-of-war, sequential configuration objectives, and strict step budgets. The agent had no prior knowledge of the game whatsoever. No training data from the game environment. No examples of successful trajectories. No human demonstrations. It entered the environment with nothing except its architectural capabilities: the &lt;code&gt;InternalDrive&lt;/code&gt; system generating Curiosity and Incoherence signals in response to novel states, the &lt;code&gt;KnowledgeIndex&lt;/code&gt; enabling cross-level strategy transfer by ingesting narrative descriptions of completed levels into a queryable IDF-weighted index, the HELM Q-learning engine shaping exploration toward historically rewarding directions, and a &lt;code&gt;WorldMapGraph&lt;/code&gt; for building an internal model of the environment topology as it was discovered. Without any task-specific training, this agent &lt;a href="https://arcprize.org/replay/8471c865-4c54-40c5-a523-dcaa681aa4f1" rel="noopener noreferrer"&gt;achieved a score of 10.71% on ARC-AGI-3&lt;/a&gt;. That might not sound dramatic, but I want you to think carefully about what that number represents. It is a score produced by a system that had never seen the game before, had no gradient-descent-trained weights, and was operating entirely on mathematical structure, causal feedback, and internally generated curiosity signals. It is doing something that is structurally closer to what I am calling intelligence than anything a language model does when it completes a standard benchmark. And we are continuing to improve it, refining the discovery algorithms, the routing policy, and the causal attribution mechanisms, with the goal of pushing toward 100%. The research direction that I am most committed to is the combination of symbolic regression, physics-informed learning, and operator learning into a single system that can observe, discover, and simulate. Symbolic regression provides discovery (11). Physics-informed learning grounds it in physical constraints (12). Neural operators provide the simulation mechanism (13). That loop is what the lmm project is building toward.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Matters for Everything Else
&lt;/h2&gt;

&lt;p&gt;I want to zoom out in this section and talk about why the knowledge-intelligence distinction matters beyond the narrow technical arguments I have been making, because the implications run much wider than AI research. The confusion between knowledge and intelligence is not just a problem in how we build machines. It is a problem in how we think about human cognition, in how we design educational systems, in how we evaluate expertise, and in how we make decisions about whom to trust and what to defer to. Getting this distinction right is not just a prerequisite for building better AI. It is a prerequisite for thinking more clearly about what minds are and what they are for.&lt;/p&gt;

&lt;p&gt;Consider how the confusion affects educational practice. Most educational systems are built around knowledge transmission: you learn facts, procedures, and standard problem-solving approaches, and you demonstrate retention of those things on tests. This is useful and necessary. But it systematically conflates knowledge with intelligence, which means that students who are excellent at knowledge retrieval are rewarded regardless of whether they have developed genuine intelligence, and students who have genuine intelligence but poor knowledge retention are penalized regardless of how well they could reason about novel problems. The result is a system that selects for knowledge over intelligence and then wonders why its graduates struggle with problems that do not look like the practice problems. Research in cognitive science has shown for decades that transfer of learning, meaning the ability to apply knowledge to novel domains, is hard to teach and rarely achieved through standard instructional approaches (14), which is exactly the prediction you would make if intelligence is the capacity to generate new compressions rather than a product of accumulated knowledge. If generating intelligence were just a matter of accumulating enough knowledge, transfer learning would be easy. It is hard, because intelligence and knowledge are different things.&lt;/p&gt;

&lt;p&gt;The knowledge-intelligence distinction also matters for how we think about expertise and authority. We tend to defer to experts on the assumption that they have not just knowledge but intelligence, that they can reason about novel cases rather than just retrieval from their experience. But expertise calibration research has consistently shown that experts often perform much worse in novel domains than their knowledge base would suggest, because the novel domain requires generating new compressions while the expert's advantage is in retrieval of trained patterns (15). This is not an argument against expertise. It is an argument for being precise about what kind of expertise is relevant to what kind of problem. A medical expert's knowledge is enormously valuable for typical cases. A medical expert's intelligence, their capacity for novel causal reasoning, is what you need for the unusual cases. These are different things, and conflating them leads to misplaced trust in exactly the cases where that trust is most dangerous.&lt;/p&gt;

&lt;p&gt;The implications for AI deployment are the most urgent. When a healthcare system deploys an AI model for diagnosis, what it is deploying is a knowledge-retrieval system. The system may be extraordinarily accurate on cases that resemble its training data. It will be unreliable on cases that require reasoning from mechanism, because reasoning from mechanism is intelligence and the system only has knowledge. The people making deployment decisions often understand this at some level, but they are caught in the same confusion that everyone else is: they see the performance on the training distribution and they interpret it as evidence of general capability, not recognizing that performance within the distribution is measuring knowledge while performance outside it would measure intelligence, and the system has never been tested outside the distribution at scale. The consequence is deployment of knowledge at disease-level stakes under the assumption that it is intelligence, and the failures happen in exactly the cases where correct reasoning matters most, in the unusual presentations that do not match the training distribution. The same pattern applies to AI in law, in hiring, in credit evaluation, in criminal justice, and in every other domain where AI is currently being deployed at scale while being described as intelligent.&lt;/p&gt;

&lt;p&gt;I want to connect this to the broader theme that has run through all my posts, which is the question of who bears the cost of technology's failures. In &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt;, I wrote about how suffering becomes invisible when it is mine but very visible when it can be used to justify someone else's comfort. In &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt;, I argued that technology's benefits flow upward to the people who control the machines while its costs flow downward to the people who depend on them. The knowledge-intelligence confusion is part of this same pattern. The people who deploy AI systems under the description of intelligent bear no cost when those systems fail by being merely knowledgeable. The cost is borne by the patient who received the wrong diagnosis, the engineer whose design was signed off on the basis of AI assurance, the student whose understanding was assessed by a system that cannot tell understanding from memorization. The intellectual confusion enables the economic exploitation, and getting the distinction right is not just a matter of scientific honesty. It is a matter of justice.&lt;/p&gt;

&lt;p&gt;Let me end this section with something that I want everyone who builds AI systems to hear. You are not building intelligence. You are building knowledge. That is genuinely valuable, and it would be dishonest for me to pretend otherwise. Knowledge-retrieval systems have improved productivity, enabled new products, and helped millions of people with tasks that would otherwise have been harder. But knowledge-retrieval systems should be deployed under terms that accurately describe what they are: systems that perform very well within their training distribution and cannot be trusted on novel cases that require genuine inference. The moment you describe them as intelligent, or allow them to be deployed in contexts that require intelligence, you are making a claim that is not supported by the evidence, and the cost of that unsupported claim will be paid by the people who trusted you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Living with the Distinction, and Building Beyond It
&lt;/h2&gt;

&lt;p&gt;I want to end this post the way I have ended the others, by being honest about where I am and what I am actually doing, rather than hiding behind abstractions. I am a person who has spent years building software, thinking about intelligence, watching the AI industry grow into something that I find both impressive and deeply dishonest, and writing on this blog as a way of staying sane through the cognitive dissonance of living at the intersection of a technology I find genuinely exciting and an industry I find genuinely corrupt. I am not a professor. I am not a famous researcher. I am someone who built things, got burned by the industry I built things for, and decided to keep thinking and keep writing because thinking and writing are the only things I trust anymore.&lt;/p&gt;

&lt;p&gt;The distinction between knowledge and intelligence is not just a theoretical point for me. It is the organizing principle of every design decision in the &lt;code&gt;lmm-agent&lt;/code&gt; framework. The five intelligence primitives exist because calibrated uncertainty, causal attribution, hypothesis formation, axiomatic reasoning, and internalized motivation are the structural properties that intelligence requires and knowledge retrieval does not provide, and if a system does not have them at an architectural level, no amount of training data will give them to it. The HELM engine exists because lifelong learning from experience is how real agents get better at navigating novel environments, and real lifelong learning looks like reinforcement over a Q-table shaped by reward, not like fine-tuning on labeled examples of the right answer. The &lt;code&gt;ThinkLoop&lt;/code&gt; PI controller exists because convergence toward a goal through iterative error feedback is what it looks like to reason rather than to retrieve. The &lt;code&gt;CausalAttributor&lt;/code&gt; with its do-calculus interventions exists because Pearl's hierarchy tells us that only causal models can correctly predict the effects of actions, and I want the agent to be at level two or three of that hierarchy, not level one. The choice to implement everything in Rust is because precision matters, and to avoid gradient descent on text corpora is because that path produces knowledge and I am trying to build something else. The ARC-AGI-3 experiment with &lt;code&gt;arc-lmm-agent&lt;/code&gt; was the first real test of whether these architectural choices actually produce the kind of out-of-distribution behavior that distinguishes intelligence from knowledge, and the answer was yes, imperfectly and partially, but yes and measurably, which is more than I can say for the systems that everyone else is scaling.&lt;/p&gt;

&lt;p&gt;I also want to say something about what I have learned from the entire experience of building this project and writing these posts. The most important thing I have learned is that the people who are most confident about AI are usually the people who have thought about it least carefully, and the people who are most uncertain are usually the ones who have looked most closely. Genuine intelligence is extraordinarily hard to build, harder than making something that looks intelligent, which is already hard. The reason I keep writing about the knowledge-intelligence distinction is that I believe the field is stuck on making things that look intelligent, and until it is prepared to have an honest conversation about the difference between looking and being, it will keep producing systems that fail in predictable ways and then making excuses for why those failures do not matter. They do matter. The families of patients who died because a knowledge-based system failed outside its distribution know they matter. The engineers who trusted AI assurance on a design that failed know they matter. The students who were assessed by a system that could not tell understanding from memorization know they matter. Their experiences are not anecdotes. They are the empirical evidence for what happens when you confuse knowledge and intelligence at scale.&lt;/p&gt;

&lt;p&gt;The research direction I am most committed to pursuing is the one that closes the loop between causal discovery and simulation, between building a model of why things happen and testing that model against new reality. The &lt;code&gt;lmm-agent&lt;/code&gt; framework is my implementation of this loop as of today, and I am continuously improving it: adding new intelligence primitives, refining the HELM learning algorithms, improving the causal graph construction, and pushing the &lt;code&gt;arc-lmm-agent&lt;/code&gt; toward higher performance on ARC-AGI-3 and similar environments. The current 10.71% score on ARC-AGI-3 is a start, not an endpoint, and I believe the architectural approach is the right one to eventually reach 100%, because the agent is improving through genuine structural learning rather than through memorization of past trajectories. It is not finished, and it is not competitive with the state of the art at knowledge tasks, because it is not trying to be. It is trying to do something different: operate effectively in genuinely novel environments using general architectural properties rather than task-specific training. Every point it gains on ARC-AGI-3 without any game-specific knowledge is a point that falsifies the claim that intelligence requires massive training data, and falsifying that claim is the research project I am most committed to.&lt;/p&gt;

&lt;p&gt;I also want to say something directly to the researchers who read this blog, because I know some of you are working on exactly the problems I have been describing: causal machine learning, symbolic regression, physics-informed neural networks, world modeling, and the other research directions that are trying to build genuine intelligence rather than sophisticated knowledge retrieval. I see your work. I read your papers. I know the funding pressure and the publication pressure and the societal pressure to work on the things that look most impressive rather than the things that matter most. I know what it feels like to be working on something that is harder and slower and less immediately impressive than the thing everyone else is building, while watching the thing everyone else is building attract all the attention and all the money. I have been there without even the consolation of being in a well-funded research environment. And the only thing I can say is: keep going. The direction is right. The direction is always more important than the current state, and the current state of the research you are doing is far more promising than most people know.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Pearl, J., &lt;em&gt;Causality: Models, Reasoning, and Inference&lt;/em&gt;, &lt;a href="https://doi.org/10.1017/CBO9780511803161" rel="noopener noreferrer"&gt;Cambridge University Press, 2009&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; Zhang, C., Bengio, S., Hardt, M., Recht, B., &amp;amp; Vinyals, O., &lt;em&gt;Understanding Deep Learning Requires Rethinking Generalization&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1611.03530" rel="noopener noreferrer"&gt;International Conference on Learning Representations (ICLR), 2017&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Lee, J. D. &amp;amp; See, K. A., &lt;em&gt;Trust in Automation: Designing for Appropriate Reliance&lt;/em&gt;, &lt;a href="https://journals.sagepub.com/doi/10.1518/hfes.46.1.50_30392" rel="noopener noreferrer"&gt;Human Factors, 2004&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; Valmeekam, K., Olmo, A., Sreedharan, S., &amp;amp; Kambhampati, S., &lt;em&gt;PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2206.10498" rel="noopener noreferrer"&gt;NeurIPS 2022 Workshop, 2022&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Ji, Z., Lee, N., Frieske, R., Yu, T., et al., &lt;em&gt;Survey of Hallucination in Natural Language Generation&lt;/em&gt;, &lt;a href="https://doi.org/10.1145/3571730" rel="noopener noreferrer"&gt;ACM Computing Surveys, 2022&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Lake, B. M. &amp;amp; Baroni, M., &lt;em&gt;Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1711.00350" rel="noopener noreferrer"&gt;International Conference on Machine Learning (ICML), 2018&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Fodor, J. A. &amp;amp; Pylyshyn, Z. W., &lt;em&gt;Connectionism and Cognitive Architecture: A Critical Analysis&lt;/em&gt;, &lt;a href="https://doi.org/10.1016/0010-0277(88)90031-5" rel="noopener noreferrer"&gt;Cognition, 1988&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; Kadavath, S., Conerly, T., Askell, A., et al., &lt;em&gt;Language Models (Mostly) Know What They Know&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2207.05221" rel="noopener noreferrer"&gt;arXiv:2207.05221, 2022&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-9"&gt;&lt;/span&gt;&lt;strong&gt;9.&lt;/strong&gt; Zech, J. R. et al., &lt;em&gt;Variable Generalization Performance of a Deep Learning Model to Detect Pneumonia in Chest Radiographs&lt;/em&gt;, &lt;a href="https://doi.org/10.1371/journal.pmed.1002683" rel="noopener noreferrer"&gt;PLOS Medicine, 2018&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-10"&gt;&lt;/span&gt;&lt;strong&gt;10.&lt;/strong&gt; Raissi, M., Perdikaris, P., &amp;amp; Karniadakis, G. E., &lt;em&gt;Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations&lt;/em&gt;, &lt;a href="https://doi.org/10.1016/j.jcp.2018.10.045" rel="noopener noreferrer"&gt;Journal of Computational Physics, 2019&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-11"&gt;&lt;/span&gt;&lt;strong&gt;11.&lt;/strong&gt; Udrescu, S. M. &amp;amp; Tegmark, M., &lt;em&gt;AI Feynman: A physics-inspired method for symbolic regression&lt;/em&gt;, &lt;a href="https://doi.org/10.1126/sciadv.aay2631" rel="noopener noreferrer"&gt;Science Advances, 2020&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-12"&gt;&lt;/span&gt;&lt;strong&gt;12.&lt;/strong&gt; Karniadakis, G. E., Kevrekidis, I. G., Lu, L., Perdikaris, P., Wang, S., &amp;amp; Yang, L., &lt;em&gt;Physics-informed machine learning&lt;/em&gt;, &lt;a href="https://doi.org/10.1038/s42254-021-00314-5" rel="noopener noreferrer"&gt;Nature Reviews Physics, 2021&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-13"&gt;&lt;/span&gt;&lt;strong&gt;13.&lt;/strong&gt; Lu, L., Jin, P., Pang, G., Zhang, Z., &amp;amp; Karniadakis, G. E., &lt;em&gt;Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators&lt;/em&gt;, &lt;a href="https://doi.org/10.1038/s42256-021-00302-5" rel="noopener noreferrer"&gt;Nature Machine Intelligence, 2021&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-14"&gt;&lt;/span&gt;&lt;strong&gt;14.&lt;/strong&gt; Barnett, S. M. &amp;amp; Ceci, S. J., &lt;em&gt;When and Where Do We Apply What We Learn? A Taxonomy for Far Transfer&lt;/em&gt;, &lt;a href="https://doi.org/10.1037/0033-2909.128.4.612" rel="noopener noreferrer"&gt;Psychological Bulletin, 2002&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-15"&gt;&lt;/span&gt;&lt;strong&gt;15.&lt;/strong&gt; Kahneman, D., &lt;em&gt;Thinking, Fast and Slow&lt;/em&gt;, &lt;a href="https://us.macmillan.com/books/9780374533557/thinkingfastandslow" rel="noopener noreferrer"&gt;Farrar, Straus and Giroux, 2011&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>lmm</category>
    </item>
    <item>
      <title>Genuine Intelligence will never in trillion years emerge from neural networks.</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 18:59:10 +0000</pubDate>
      <link>https://dev.to/wiseai/genuine-intelligence-will-never-in-trillion-years-emerge-from-neural-networks-1250</link>
      <guid>https://dev.to/wiseai/genuine-intelligence-will-never-in-trillion-years-emerge-from-neural-networks-1250</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This post was originally published on &lt;a href="https://wiseai.dev/blogs/genuine-intelligence-will-never-emerge-from-neural-networks" rel="noopener noreferrer"&gt;the main website&lt;/a&gt; on &lt;a href="https://github.com/wiseaidotdev/blog/pull/18" rel="noopener noreferrer"&gt;Apr 23 2026&lt;/a&gt;. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account &amp;lt;3!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;I have been building toward this post for a long time. Across everything I have written in this blog, I have been circling the same core question from different angles, and in each post I have gotten a little closer to saying the thing I really want to say without softening it into something that sounds more professionally acceptable. In &lt;a href="https://wiseai.dev/blogs/language-is-limited-asi-is-impossible" rel="noopener noreferrer"&gt;Language is Limited. ASI is Impossible.&lt;/a&gt;, I argued that no machine trapped inside symbols can understand reality, because symbols are descriptions of the world and not the world itself. In &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, I argued that the only honest language for describing physical reality is mathematics, because equations encode mechanisms while words only encode appearances. In &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether.&lt;/a&gt;, I argued that the practice of extracting value from human creative work to build models without consent is not a neutral engineering choice but a moral failure. In &lt;a href="https://wiseai.dev/blogs/all-you-have-access-to-is-knowledge-and-tools-never-intelligence" rel="noopener noreferrer"&gt;All You Have Access To Is Knowledge and Tools; Never Intelligence!&lt;/a&gt;, I argued that every AI system you have ever used has access to knowledge and tools but none of them has intelligence in the sense that word deserves to carry. And in &lt;a href="https://wiseai.dev/blogs/rethinking-arc-agi" rel="noopener noreferrer"&gt;Rethinking ARC-AGI&lt;/a&gt;, I argued that even the benchmarks people use to celebrate AI progress are measuring the wrong thing entirely. All of those posts were preparation for this one, because this one is where I say the most direct version of the argument that all the others have been building toward. The argument is this: genuine intelligence will never emerge from neural networks, not after ten more years of scaling, not after a hundred, not after a trillion. Not because I am pessimistic, and not because I want to be contrarian, and not because I do not understand how neural networks work. But because neural networks are the wrong kind of thing, and being the wrong kind of thing is not a fixable bug. It is the defining property of what they are.&lt;/p&gt;

&lt;p&gt;Before I develop the argument, I want to say something about what I mean by genuine intelligence, because this word gets stretched to fit so many different claims that it has almost lost its meaning in the current conversation. By genuine intelligence I mean something specific and demanding. I mean the capacity to form new concepts from first principles, to reason about causal mechanisms you have never been shown, to solve problems that are not close to anything in your past experience, to know when your own knowledge is insufficient, to model other minds and their intentions accurately, to understand language the way a human understands it rather than predicting what words tend to follow other words, and to connect all of these capabilities into a unified coherent agent that can act in the world with something resembling genuine comprehension rather than statistical confidence. I am not setting an impossible standard. This is exactly the kind of intelligence that a reasonably educated human twelve-year-old already has, and has had since before the invention of writing. The bar is not AGI in some science fiction sense. The bar is ordinary human understanding, and I am claiming that neural networks cannot reach that bar regardless of how many parameters they have or how many tokens they train on, because the architecture of neural networks is structurally incompatible with the mechanisms that produce that kind of understanding.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture Is the Problem, Not the Scale
&lt;/h2&gt;

&lt;p&gt;The most common response to any critique of neural networks is to suggest that the problem is a matter of scale, that the current models are impressive but incomplete, and that the next generation will be more impressive and more complete, and eventually, after enough scaling, the system will cross some threshold and genuine intelligence will emerge. This response sounds reasonable if you have not spent time thinking carefully about what neural networks actually are at the mathematical level. When you have spent that time, the response reveals itself as a category error, like suggesting that if you paint a picture of a river large enough, the water will become real enough to drink. Scale does not change category. A very large neural network is still, at its core, a function that maps input vectors to output vectors through a series of weighted linear transformations followed by nonlinear activations, and the genius of the architecture is that this family of functions can be made to approximate almost any continuous function to arbitrary precision if the network is large enough. That is a real and important capability. But approximating functions over a training distribution is not the same as understanding, and it never becomes the same thing regardless of how large the approximating function is, because understanding requires something that function approximation structurally cannot provide, which is the capacity to reason about the mechanism that generates the function rather than just interpolating its values.&lt;/p&gt;

&lt;p&gt;I want to be precise about what I mean by mechanism, because the distinction between mechanism and pattern is the entire argument. When a neural network is trained on data from a physical process, it learns a mapping from inputs to outputs that approximates the data it has seen. What it does not learn is why those inputs map to those outputs, meaning the causal chain, the physical law, the generative process that produces the data. The network cannot distinguish between a world where removing the fuse causes the engine to stop because the fuse carries electrical current to the ignition system, and a world where removing the fuse causes the engine to stop because removing fuses is correlated with engine stoppage in the training data. Both worlds produce the same statistical associations, and statistical learning cannot distinguish them, because distinguishing them requires an intervention, the kind of do-operator that Judea Pearl formalized, and interventional reasoning requires a model of causal structure, not just a model of statistical association (1). A neural network that has seen a million examples of engines stopping after fuse removal has learned that these events co-occur. It has not learned why they co-occur, which means it has no basis for predicting what will happen in a novel situation where the correlation breaks down. And every real-world situation that matters, every diagnosis, every engineering decision, every policy choice, is a situation where naive correlations can break down, and where only causal understanding provides reliable guidance.&lt;/p&gt;

&lt;p&gt;The failure of neural networks on distribution shift is not a minor limitation that will be fixed with more data. It is the visible symptom of the underlying architectural incompatibility between function approximation and causal reasoning. Research has demonstrated systematically that neural networks trained on data from one distribution fail dramatically when tested on data from a slightly different distribution, even when the underlying concept is identical and a human would generalize trivially (2). A model trained to recognize images of cows in pastoral settings performs poorly when shown cows in unusual contexts, not because it has not seen enough cows but because it has learned to associate cow-ness with green backgrounds rather than with the visual features that actually define cows as a category. A model trained to diagnose chest X-rays from one hospital system fails when deployed at a different hospital, not because medicine is different there but because the imaging equipment, the patient demographics, and the X-ray presentation conventions differ enough to fool a system that learned statistical associations rather than anatomical principles. These failures are not edge cases. They are the predicted consequence of a learning paradigm that learns what things look like in the training data rather than what things are in reality, and you cannot fix that consequence by adding more training data from the same paradigm, because you are only strengthening the same kind of association-based learning that fails when the associations change.&lt;/p&gt;

&lt;p&gt;Let me also address the specific claim that has been made several times in the last few years by prominent AI researchers, which is that sufficiently large neural networks display emergent capabilities that were not present in smaller versions of the same architecture, and that this emergence might be the mechanism by which genuine intelligence develops. This claim deserves careful examination because it sounds like it might save the scaling argument, and it will not. The emergent capabilities that have been documented in large language models, things like few-shot learning, chain-of-thought reasoning, and performance on certain multi-step inference tasks, are impressive in the same way that any surprising capability is impressive, namely because we did not predict them in advance. But impressive and intelligent are different things, as I argued at length in &lt;a href="https://wiseai.dev/blogs/all-you-have-access-to-is-knowledge-and-tools-never-intelligence" rel="noopener noreferrer"&gt;All You Have Access To Is Knowledge and Tools; Never Intelligence!&lt;/a&gt;. The emergent capabilities of large language models are capabilities at manipulating the statistical structure of language in ways that happen to be useful for certain tasks. They are not capabilities at reasoning about the world the way a human reasons about the world. A model that can solve math problems in chain-of-thought format is performing a form of pattern matching on mathematical text that superficially resembles step-by-step reasoning. The proof that it is pattern matching rather than reasoning is that the same model fails immediately when the problems are presented in formats that differ from the training distribution, when symbols are renamed, when the logical structure is preserved but the surface appearance is changed, or when the problem requires exactly the kind of systematic generalization that a human child handles trivially but neural networks cannot.&lt;/p&gt;

&lt;p&gt;Systematic generalization is the technical name for what I am describing, and it is worth spending time on because it is the capability that most clearly distinguishes genuine understanding from sophisticated pattern matching. Systematically generalizing means that if you understand the concept A and the concept B, you can immediately understand AB and BA and ABA and any other systematic combination of A and B, because you understand what A and B mean, not just what A and B look like in the training data. A human who learns the meaning of the word red and the meaning of the phrase on the left side can immediately understand red on the left side without having been trained on that specific phrase, because they have compositional understanding of the concepts involved. Neural networks have repeatedly been shown to lack this kind of systematic compositionality in a fundamental way (3). Models trained on a set of concept combinations fail to generalize to novel combinations of the same concepts, producing chance-level performance on systematic recombinations that a human would find trivially easy. This failure is not a matter of insufficient training data. It is a consequence of the architecture, which learns statistical patterns over the specific combinations it has seen rather than compositional rules that generate all possible combinations from their parts. And compositionality is not a luxury feature of intelligence. It is what allows a mind with a finite vocabulary to understand and produce an infinite range of thoughts. A mind without compositionality is a mind that can only repeat or slightly vary what it has already seen, and that is not a mind in any sense that deserves the name.&lt;/p&gt;

&lt;p&gt;I want to end this section with something I think is underappreciated and that deserves to be stated plainly. The evidence that neural networks cannot achieve genuine intelligence through scaling is not coming only from philosophical arguments or theoretical analysis. It is coming from empirical results produced by the people who build and study these systems. The Grokking paper, published by researchers inside the AI research community, showed that neural networks can memorize training data in a way that appears to be generalization until you look carefully enough, and that the delayed generalization they observed was qualitatively different from the kind of principled generalization that results from understanding a rule (4). The research on adversarial examples has shown that neural networks can be catastrophically fooled by input perturbations that are invisible to humans, which reveals that the networks have not learned the concepts that humans use to recognize things but rather have learned surface statistics that happen to be correlated with correct answers in the training data (5). The research on neural networks learning to exploit dataset artifacts rather than genuine patterns, where models achieve high performance on natural language inference benchmarks by learning spurious correlations rather than by understanding implication, has been confirmed across multiple datasets and multiple architectures (6). All of this empirical evidence is produced by respected researchers with no ideological axe to grind, and it all points in the same direction: neural networks learn what is in the data, not what the data is about, and that difference is the difference between genuine intelligence and very sophisticated mimicry.&lt;/p&gt;

&lt;h2&gt;
  
  
  Data Is Not Knowledge, and More Data Will Never Be Knowledge
&lt;/h2&gt;

&lt;p&gt;One of the most persistent myths in the current AI conversation is that intelligence is essentially a function of data, that if you give a system enough data about enough things, intelligence will emerge because intelligence is what patterns plus scale produces. This myth is not just wrong. It confuses two things that are completely different at the level of what they are and how they work, namely data and knowledge. Data is observations. Knowledge is understanding. You can have infinite observations and zero understanding, the way a camera records everything but understands nothing, and the way a digital thermometer precisely measures temperature but has no knowledge of thermodynamics. The difference between data and knowledge is not a matter of quantity. It is a matter of mechanism, of whether the system has a model of why the observations have the structure they have, rather than just a model of what the structure looks like. Neural networks operate on data. They do not produce knowledge in this sense, and adding more data does not change that, because the problem is not the amount of input but the nature of the processing.&lt;/p&gt;

&lt;p&gt;The distinction between memorization and generalization is the concrete technical version of the distinction between data and knowledge, and it has been at the center of machine learning research for decades in a way that the popular discourse almost completely ignores. Memorization is when a system remembers specific examples and can reproduce them. Generalization is when a system extracts a rule that allows it to handle new examples it has never seen. Genuine knowledge is on the generalization side of this distinction, because genuine knowledge is about understanding the rule, not remembering the examples. Neural networks are capable of both memorization and a form of generalization, but the generalization they produce is statistical interpolation over the training distribution rather than rule-based extrapolation beyond it. Research has shown that neural networks can fit completely random labels on training data just as effectively as they fit correct labels, which means the networks have effectively infinite memorization capacity and use it freely alongside whatever genuine pattern learning they perform (2). The fact that a network achieves high accuracy on a test set does not mean it has learned a rule. It might mean that the test set is drawn from the same distribution as the training data and that interpolation suffices to handle it. Only when you test the system on out-of-distribution examples, examples that require genuine rule-based extrapolation rather than distribution interpolation, does the difference between data processing and knowledge become unmistakable.&lt;/p&gt;

&lt;p&gt;The specific kind of knowledge that is most important for genuine intelligence and most completely absent from neural networks is the knowledge Aristotle called scientia and modern philosophers call propositional knowledge of why something is true. It is not enough to know that X is true if you want to be genuinely intelligent about X. You also need to know why X is true, because the why is what allows you to figure out whether X is still true when conditions change, whether X implies Y in this particular case, whether an apparent exception to X is a real exception or a misapplication of the concept. The why is the causal model that connects X to the rest of what you know, and a neural network trained on millions of examples where X is true has learned that X is commonly true in its training distribution but has not learned why X is true, which means it has no principled basis for handling cases that differ from its training distribution. A human doctor who knows why a certain drug reduces blood pressure, meaning who knows the pharmacological mechanism, can reason about what the drug will do in a patient with an unusual combination of conditions that the doctor has never seen before. A neural network that has learned the association between the drug and blood pressure reduction from training data cannot do this reasoning, because it does not have the causal mechanism, only the statistical correlation. And the cases where you most need reliable medical reasoning are exactly the cases where the unusual combination of conditions makes the statistical correlations unreliable.&lt;/p&gt;

&lt;p&gt;I want to connect this to the argument I made in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt; about why equations are fundamentally more powerful representations of reality than text. An equation encodes the mechanism, the causal structure, the why. A dataset encodes observations, the what. A neural network trained on a dataset has absorbed the what without the why, and no amount of more data changes that, because the why is not in the data. The why is in the structure of reality that generated the data, and accessing that structure requires a different kind of inquiry, the kind that scientists have been doing for four hundred years, which involves formulating hypotheses, deriving predictions from those hypotheses, testing the predictions against observations, and revising the hypotheses based on the test results. That is the scientific method, and it is the mechanism by which humans transform data into knowledge, and it is structurally missing from the neural network training paradigm, which takes data as input and produces a function as output without any stage where the system engages with the structure of reality rather than the structure of the data. The lmm project I described in my post on training is an attempt to build a system that actually does the scientific method on data, discovering equations rather than fitting functions, and the difference between those two things is the entire argument in a single comparison.&lt;/p&gt;

&lt;p&gt;I also want to address the specific claim, which is made constantly in the AI conversation, that large language models have effectively read all of human knowledge through their training data and therefore contain or have access to all of human knowledge. This claim sounds plausible if you think of knowledge as content, as the collection of facts and claims that humans have written down. But it is wrong as soon as you recognize that knowledge is not content. Knowledge is the relationship between content and reality, the verified connection between what is claimed and what is actually true. A system that has read every physics textbook ever written has not acquired knowledge of physics in the sense that a physicist has knowledge of physics. The physicist's knowledge is verified against the world, challenged by experiment, refined by evidence, and organized within a causal structure that explains why the facts are as they are. The language model's "knowledge" is a collection of patterns in the text of physics books, patterns that typically produce correct-sounding outputs because the books were mostly written by physicists who had real knowledge, but patterns that have no verified connection to physical reality and no causal structure that would let the model reason about novel physical situations. The Stochastic Parrots paper called this the hallmark of a parrot, the ability to produce contextually appropriate sounds without the capacity to mean what is being said (7). That paper received enormous institutional backlash when it was published, which tells you something about how comfortable the AI industry is with honest descriptions of its flagship products.&lt;/p&gt;

&lt;p&gt;The role of memory in intelligence also exposes a deep incompatibility between how neural networks work and how genuine knowledge works, and I want to spend time on this because it is usually glossed over in popular discussions of AI capability. In neural networks, what the network has learned is encoded in the weights, the billions of numerical parameters that are adjusted during training. Those weights encode statistical regularities across the training data in a distributed, holistic way that makes it impossible to identify which weight encodes which fact or trace a specific piece of information back to its source. This design has practical advantages, particularly the ability to generalize across related inputs. But it has a profound epistemic cost, which is that the network cannot evaluate its own knowledge. It cannot tell you which of its beliefs are well-supported and which are poorly supported. It cannot flag its own uncertainty except through mechanisms that are themselves learned statistical approximations of uncertainty rather than direct access to the quality of its own knowledge states. A human who knows that they know something and knows that they do not know something else has direct access to the quality of their own epistemic states, and this metacognitive access is what allows them to seek information appropriately, defer to experts in domains where they lack knowledge, and flag their own uncertainty honestly when it is relevant. Neural networks cannot do this reliably, and the research on calibration shows that large language models are systematically overconfident in domains where they have limited training-derived information, producing confident-sounding outputs even when the underlying pattern matching is operating on weak signals (8). That systematic overconfidence is not an accident. It is the direct consequence of a learning paradigm that optimizes for producing confident-sounding outputs rather than producing outputs whose confidence tracks the reliability of the underlying knowledge.&lt;/p&gt;

&lt;p&gt;The environmental cost of training on more and more data is also worth keeping in mind as a practical argument against the data-is-intelligence claim, because it grounds the abstract argument in real-world consequences. As I discussed in &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether.&lt;/a&gt;, training large language models on massive datasets consumes enormous quantities of electricity and produces significant carbon emissions (9). If the data-is-intelligence hypothesis were correct, then the massive investment in training data and compute would be buying genuine progress toward understanding. But if I am right that data is not the same thing as knowledge, and that more data processed by a neural network produces more sophisticated pattern matching rather than deeper understanding, then the environmental cost of large-scale training is a cost that is buying the wrong kind of thing, and that makes it not just expensive but genuinely wasteful in a moral sense. I am not saying that the capability gains from scaling are zero. I am saying that capability gains from scaling are not the same as progress toward genuine intelligence, and conflating the two is how you end up spending staggering amounts of energy and human creative labor in pursuit of a goal that the architecture is structurally incapable of reaching.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Pattern Matching Is Not and Cannot Become Reasoning
&lt;/h2&gt;

&lt;p&gt;There is a word that appears constantly in discussions of modern AI systems, and that word is reasoning. We are told that large language models can reason, that they engage in chain-of-thought reasoning, that reasoning emerges as a capability at sufficient scale, and that improving AI reasoning is one of the most active research fronts in the field. When I hear this word applied to neural networks, something in me tightens, because I think it is the most consequential mislabeling in the entire AI conversation. Reasoning is a specific cognitive operation with a specific structure, and that structure is completely different from the statistical prediction that neural networks perform. Reasoning is the process of deriving new beliefs from existing beliefs according to rules of inference that are truth-preserving, meaning that if your premises are true and your inference rules are valid, your conclusion is guaranteed to be true. Pattern matching is the process of identifying the most likely output given the input according to patterns learned from training data. These two processes look similar from the outside when the training data is rich and the test distribution is close to the training distribution. They produce completely different results when the test distribution diverges from training or when the reasoning task requires steps that are novel rather than recombinations of seen patterns.&lt;/p&gt;

&lt;p&gt;The specific mechanism that makes reasoning truth-preserving is that it operates over the structure of propositions rather than over their surface form or their statistical properties. When I reason that all mammals breathe and whales are mammals, therefore whales breathe, I am applying a rule that is sensitive to the logical structure of the premises, and that rule produces the correct conclusion regardless of whether I have ever seen a whale, regardless of whether whales appear frequently in my training data, and regardless of whether the sentence about whales breathing appears in any corpus I have read. The rule derives the conclusion from the structure of the premises, not from the statistics of the training data. A neural network asked the same question produces a correct answer with high probability because whales and breathing and mammals all co-occur frequently in text about marine biology and mammalian biology, and the correct answer is statistically likely given those co-occurrences. But the network's correct answer and my correct answer are produced by completely different mechanisms, and only my mechanism is reliable when the statistical regularities break down. A neural network presented with a novel set of premises about fictional entities that do not appear in any training data will fail on the logical conclusion that a human reasoner derives trivially, because the network has no access to the logical structure of the premises, only to the statistical patterns of words, and fictional entities produce weak or missing statistical signals. This has been demonstrated empirically, and the results are not close (3).&lt;/p&gt;

&lt;p&gt;I want to be specific about what chain-of-thought prompting actually does and does not do, because it has been widely misunderstood and the misunderstanding is consequential. Chain-of-thought prompting is a technique where you ask a language model to produce its answer step-by-step rather than in a single prediction, and it was shown to improve performance on certain multi-step reasoning tasks significantly. This improvement was taken by many commentators as evidence that language models can reason, and that chain-of-thought prompting unlocks their latent reasoning capability. That interpretation is almost certainly wrong. What chain-of-thought prompting most likely does is restructure the prediction task so that the model's outputs at each step serve as a form of working memory, allowing the model to produce outputs that are conditioned on intermediate steps in a way that the training data has made statistically likely. The improvement comes from decomposing a single hard prediction into a sequence of easier predictions, each of which is more likely to match the training data. The model is not reasoning more deeply. It is making more predictions, and more predictions over a decomposed problem reduce the variance of the statistical approximation. You can demonstrate this by presenting the same problems with the same logical structure but with unfamiliar surface forms or novel entities, and the chain-of-thought improvements evaporate when the statistical regularities that were supporting each step are disrupted. A system that reasons would continue to reason correctly. A system that pattern matches shows the dependence on statistical regularities as soon as those regularities are absent.&lt;/p&gt;

&lt;p&gt;The distinction between reasoning and pattern matching also shows up in how these two processes handle contradiction. A system that genuinely reasons will recognize when two things it believes are contradictory and will either update one belief or flag the contradiction as a problem that requires resolution. A system that pattern matches has no principled mechanism for handling contradiction across long contexts, because pattern matching operates over local statistical associations and does not maintain a global model of what it has committed to believing. The research on the consistency of large language models shows that they will contradict themselves across long conversations in a way that reveals no underlying commitment to logical coherence, because logical coherence is a global property of a belief system and pattern matching is a local operation over input-output associations (10). Asking a model one question and then asking a related question that should have the opposite answer based on the first answer will often produce inconsistent responses, because the model is making each prediction independently based on the local context rather than maintaining a globally consistent model of what it believes. That inconsistency is not a failure of memory or a temporary limitation. It is the direct consequence of producing outputs through pattern matching rather than through reasoning over a persistent model of the world.&lt;/p&gt;

&lt;p&gt;I also want to address the argument that neural networks implicitly learn reasoning-like operations through training, that the internal representations developed during training encode something like logical structure even if the network was not explicitly designed to reason. This argument is made by serious researchers and deserves a serious response. It is true that internal representations in large language models encode various kinds of structure that are more abstract than simple word statistics, and that probing those representations can reveal features that correspond to syntactic, semantic, and to some extent logical properties of the input. But the existence of reasoning-relevant features inside a network is not the same as the ability to reason over those features in a principled way. A city that has all the materials for building a bridge but has never developed the engineering knowledge to assemble them does not have a bridge. It has bridge materials. Representation learning has produced powerful bridge materials inside large language models. It has not produced the capacity to use those materials for genuine reasoning, because genuine reasoning requires not just the right representations but an inference engine that can manipulate those representations according to truth-preserving rules, and no neural network has such an inference engine, because no neural network training objective selects for truth-preserving inference as opposed to statistically likely output.&lt;/p&gt;

&lt;p&gt;The adversarial fragility of neural networks is the most visceral demonstration that what these systems are doing is not reasoning, and I want to describe it carefully because I think it is one of the most important empirical facts about the current state of AI. Adversarial examples are inputs that have been slightly modified in ways that are imperceptible to humans but that cause a neural network to produce completely wrong outputs with high confidence. A stop sign with a few carefully placed stickers is classified as a speed limit sign by an image classifier that was correctly classifying stop signs before the stickers were added. A sentence with a single word changed to a synonym is classified in the opposite sentiment category by a sentiment classifier that correctly classified the original. An audio file with sub-perceptual noise added is transcribed as an entirely different sentence by a speech recognition system that was correctly transcribing the original (5). These adversarial failures are not hypothetical. They have been demonstrated empirically across dozens of architectures and domains, and they reveal that the networks are making decisions based on surface statistics that happen to be correlated with the correct answer in the training data, not based on the features that actually define the concept being recognized. A police officer examining a stop sign with stickers on it uses their understanding of what a stop sign is, its shape, its color, its text, its function in traffic regulation, and they correctly identify it as a stop sign despite the stickers. The neural network uses surface statistical patterns that the stickers disrupt, and it fails immediately. That is the difference between understanding something and having learned a statistical proxy for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Neural Networks Cannot Model Other Minds and This Is Not Fixable
&lt;/h2&gt;

&lt;p&gt;One of the capabilities that is most central to human intelligence and most completely absent from neural networks is the ability to model other minds, to represent what other people believe, want, intend, and know, and to use those representations to coordinate, communicate, and predict behavior. Developmental psychologists call this Theory of Mind, and it is so fundamental to human social life that its absence or impairment, as in certain neurodevelopmental conditions, produces severe difficulties in social function even when other cognitive capabilities are intact. Theory of Mind requires representing beliefs, which are representations of the world rather than the world itself, and it requires representing other agents as having their own belief states that may differ from your own and from reality, and it requires reasoning about how those belief states influence behavior. This is second-order representation, representation of representation, and it relies on the same kind of compositional, recursive structure that neural networks systematically fail at in the simpler cases I described in the previous section.&lt;/p&gt;

&lt;p&gt;The false belief test is the classic experimental paradigm for assessing Theory of Mind in children and animals, and versions of it have been applied to language models with results that are instructive about what is actually going on when these models appear to understand other minds. In a false belief test, you describe a scenario where character A puts an object in location X, then leaves, then character B moves the object to location Y. You then ask: when character A comes back, where will they look for the object? The correct answer is location X, because A did not see the object move and therefore has a false belief that it is still in location X. Children younger than four years typically fail this test because they cannot yet represent that A has a different belief from the current reality. Children older than four typically pass because they have developed the capacity to represent beliefs that differ from reality. Large language models pass certain versions of this test at rates that seem high. But research into why they pass reveals that they are exploiting statistical patterns in the way this kind of story tends to be told, rather than genuinely representing character A's belief state (11). When the test is presented with novel surface forms, unfamiliar names, unusual layouts, or other variations that preserve the logical structure while disrupting the statistical patterns, performance drops substantially. When you test for genuinely novel Theory of Mind challenges that have no close analog in the training data, performance approaches chance. A human who understands the concept of false belief handles novel variations effortlessly, because they understand the concept, not just the pattern.&lt;/p&gt;

&lt;p&gt;The absence of genuine Theory of Mind in neural networks is not just a social limitation. It connects directly to everything I have been arguing about causal understanding and the distinction between pattern and mechanism. A genuine Theory of Mind requires modeling another agent as a causal system, a system that processes perceptions, forms beliefs based on those perceptions, takes actions based on those beliefs, and changes its beliefs when it receives new perceptions. Modeling that requires exactly the kind of causal reasoning that neural networks cannot do: you need to represent the causal chain from perception to belief, reason about what the agent would have perceived given their different vantage point, and predict how their beliefs would have formed given those perceptions. That is a causal simulation, and doing it well enough to predict behavior accurately requires something much closer to a mechanistic model of cognition than statistical associations over text about conversations between people can provide. A language model can produce fluent text about what character A believes, but it is doing so by pattern matching over the statistical regularities of how this kind of text tends to be written, not by running a causal simulation of character A's epistemic state. The difference between these two operations is not subtle, and it is not bridgeable by scale.&lt;/p&gt;

&lt;p&gt;There is also a deeper point here about the nature of communication itself, and I want to make it carefully because it connects to the argument I made in &lt;a href="https://wiseai.dev/blogs/language-is-limited-asi-is-impossible" rel="noopener noreferrer"&gt;Language is Limited. ASI is Impossible.&lt;/a&gt; about language being a limited medium. When humans communicate successfully, we do not just exchange words. We use words as signals that allow the listener to reconstruct the speaker's intended meaning, which requires the listener to model the speaker's beliefs, intentions, and knowledge state in order to find the interpretation of the words that is most consistent with what the speaker plausibly intends to communicate. This is Gricean communication, named after the philosopher Paul Grice who first formalized it, and it requires Theory of Mind at every step (12). When I say see you later and someone understands I am expressing a farewell and not a literal prediction about seeing, they are using their model of my intentions and the conversational conventions we share to interpret the words correctly. When I give instructions that are ambiguous and a student asks a clarifying question, the student is modeling my knowledge state and identifying the gap between what I said and what I could have meant, and tailoring their question to close that gap. All of this requires second-order modeling of minds, and neural networks cannot do it, which means neural networks cannot genuinely communicate in the human sense, only produce outputs that are statistically likely given the inputs. That is a profound limitation, and it is not a limitation that more parameters will overcome.&lt;/p&gt;

&lt;p&gt;I want to connect this to something practical and immediately relevant, which is the use of AI systems in high-stakes communication contexts like medicine, law, therapy, and crisis intervention. These are exactly the contexts where Theory of Mind matters most, where understanding what the other person believes, fears, knows, and intends is not a nice-to-have but the core of what the professional is supposed to do. A therapist who cannot genuinely model what their patient believes about their own situation cannot do therapy. A doctor who cannot model what their patient knows about their own symptoms cannot take an accurate history. A lawyer who cannot model what their client understands about the legal process cannot give useful advice. Deploying neural network systems into these contexts, as is already happening at scale across the healthcare and legal sectors, is deploying systems that are missing the core capability that makes these contexts work. The results are going to match what the research predicts, which is that the systems perform well enough on average cases within the training distribution and fail dangerously on the unusual cases that matter most, while appearing confident throughout. I am not speculating about this. I am describing the direct consequence of using a tool that is architecturally incompatible with the task it is being asked to perform. And the people who will pay the cost of that incompatibility are the same people who always pay the cost when technology is deployed before it is ready, the patients, the clients, the users who trusted the system and were failed by it.&lt;/p&gt;

&lt;p&gt;Let me also address the argument that neural networks paired with tools or external memory can achieve a functional equivalent of Theory of Mind, that if you give a language model access to information about a user's history and preferences, the system can model the user sufficiently for practical purposes. This argument conflates modeling with retrieval. Having access to information about someone is not the same as having a model of their mind. A file cabinet full of someone's medical records does not have a model of the patient's beliefs and fears. A language model with access to a user's conversation history does not have a model of the user's epistemic state. It has retrieved text that was produced by the user, and it can use that text to produce statistically likely responses that are conditioned on that retrieved text. This produces outputs that are often more appropriate than outputs without the context, in the same way that a doctor who has read your file is better prepared than one who hasn't. But the doctor reads your file with genuine understanding, integrating the information into a causal model of your condition. The language model retrieves the text and produces more statistically appropriate responses based on its training-data associations. The difference between those two operations is the difference between genuine modeling and informed guessing, and informed guessing is not Theory of Mind.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Data Harvesting Truth Nobody Wants to Admit
&lt;/h2&gt;

&lt;p&gt;I have written in &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether.&lt;/a&gt; about the ethical dimensions of training on human data, and I want to return to that argument here but from a different angle, because I think there is a dimension of the data harvesting problem that is not just moral but epistemic. Neural networks need your data not just because it is useful to them but because they cannot function without it. The system has no model of the world. It has no capacity to derive predictions from first principles. It has no equations that encode mechanisms. Its entire capability rests on the statistical structure of the data it has consumed, which means that every upgrade in its capability requires consuming more data, and the data it needs is your professional knowledge, your creative work, your personal expression, and the documented experience of your life. This dependency is not incidental. It is structural, and understanding it as structural changes how you should think about the relationship between AI companies and the human beings whose data they consume.&lt;/p&gt;

&lt;p&gt;The companies building large neural network systems need to continuously collect more data because the architecture of these systems has no principled mechanism for building durable knowledge from limited observations. A physicist who understands the laws of thermodynamics can answer thermodynamics questions they have never encountered before, because they can derive the answers from their model of the underlying mechanisms. A neural network that has achieved high performance on thermodynamics questions in the training data needs to have seen a sufficiently large and representative sample of thermodynamics questions in order to answer novel ones, and when a genuinely novel question comes along that falls outside the distribution, performance degrades. This means the companies need to keep collecting more and more specialized data to cover more and more domains, and the economic incentive to collect that data is as large as the commercial value of the systems built on it, which is now measured in hundreds of billions of dollars. The data harvesting is not a temporary phase that will end when the technology matures. It is the permanent operating condition of a technology that has made human data its fuel rather than building systems that generate their own understanding from first principles.&lt;/p&gt;

&lt;p&gt;The harvesting also has an interesting epistemic trap built into it. Because the neural network learns statistical patterns from human-generated data, the useful patterns it can learn are bounded by the collective intelligence of humanity as expressed in text. A neural network trained on all human text is in some sense a compressed version of the collective human response to the questions that have been asked, conditioned on the training objectives and the filtering decisions made during the training process. That is genuinely impressive, and I have always been willing to say so. But it also means that the system cannot surpass the intellectual frontier of humanity in any domain, because it has no mechanism for generating insights that are not already latent in the statistical structure of human-generated text. When the system appears to produce a novel insight, it is almost always recombining patterns from its training data in a novel way, rather than genuinely discovering something new. And the ability to recombine patterns in useful ways is the capability that makes these systems valuable, but it is also exactly the capability that does not constitute genuine intelligence, because genuine intelligence can advance the intellectual frontier, can discover truths that no human has yet articulated, can see the structure of reality directly rather than through the filtered lens of what humanity has already written about it. A system that learns from human text is always downstream of human intelligence, always dependent on it, always bounded by what humans have already expressed. That is not general artificial intelligence. That is a very sophisticated aggregator and recombiner of human intellectual output, and the distinction matters enormously.&lt;/p&gt;

&lt;p&gt;I also want to say something about what happens when the data harvesting enters a loop, which is already beginning to happen as AI-generated content populates the web. Neural networks trained on internet data are now producing content that is appearing on the internet, which means that future rounds of training on internet data will train on AI-generated content. Researchers have called this model collapse, and the concern is that training on AI-generated data causes models to progressively lose fidelity to the diversity and long-tail richness of human-generated content and converge toward a narrower, more homogenized version of the training distribution (13). If model collapse is real, and the early evidence suggests it is a genuine concern, then the self-referential training loop will degrade the quality of future models in a way that is very hard to reverse, because the original human-generated data that captured the full diversity of human intelligence and experience will be increasingly diluted by AI-generated approximations of it. This is what you get when you build intelligence on data extraction rather than on principled models of reality: the data source eventually becomes contaminated by the outputs of the systems that consumed it, and the degradation is circular and self-reinforcing. An equation-based system does not face this problem because it learns from physical measurements rather than from human text, and physical reality does not collapse when AI systems start producing text about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Is Not A Property That Emerges With More Parameters
&lt;/h2&gt;

&lt;p&gt;By this point in the post I have built the technical case that neural networks are architecturally incompatible with genuine intelligence across multiple dimensions: they cannot reason causally, they cannot generalize systematically, they cannot model other minds reliably, they cannot distinguish data from knowledge, and they cannot advance the intellectual frontier independently of the human data they consume. Now I want to address head-on the response that treats all of these problems as temporary engineering challenges that more scale will eventually resolve, because this response is so common and so superficially plausible that it deserves careful dismantling rather than dismissal.&lt;/p&gt;

&lt;p&gt;The response rests on an implicit theory of emergence, the idea that qualitatively new properties can arise from quantitative changes in scale, so that a system that is not intelligent at one scale might become intelligent at a larger scale because intelligence itself is an emergent property of sufficient complexity. Emergence is a real phenomenon in complex systems, and there are genuine cases where quantitative changes produce qualitative changes in behavior. Water becomes wet at a sufficient number of molecules even though individual molecules are not wet. Consciousness (if it is real) might emerge from sufficient complexity of neural interactions even though individual neurons are not conscious. So the emergence argument cannot be dismissed as nonsensical. But it has a crucial requirement that the people making it in the AI context consistently ignore, which is that emergence only occurs when the quantitative system being scaled has the right structural properties to support the emergent phenomenon. Ice becomes water when temperature increases because the molecular interactions of water molecules have the right structural properties to produce phase transitions. Adding more water molecules does not produce phase transitions in motor oil, because motor oil has different structural properties. The question that the emergence argument for AI must answer is whether neural networks have the structural properties that would allow genuine intelligence to emerge from sufficient scale, and that question has not been answered or even seriously engaged with by the people making the scaling argument. It has simply been assumed, on the basis that neural networks are large and complex and intelligence is also associated with large and complex systems in biology.&lt;/p&gt;

&lt;p&gt;The structural incompatibilities I have documented throughout this post are the reasons to believe that neural networks do not have the right structural properties for intelligence to emerge from scale. Let me name them clearly. Neural networks do not have explicit causal models, which means that scale cannot produce causal reasoning because there is nothing in the architecture to instantiate it at any scale. Neural networks do not have compositional inference engines, which means that scale cannot produce systematic generalization because there is nothing in the architecture to support rule-based generalization at any scale. Neural networks do not have persistent belief states with logical consistency management, which means that scale cannot produce logical coherence and Theory of Mind at any scale. Neural networks do not have mechanisms for grounding their outputs in physical reality, which means that scale cannot produce the verified connection to the world that knowledge requires. Each of these is an architectural absence, not a quantitative insufficiency, and architectural absences cannot be filled by adding more of what is already there. You cannot scale an addition problem into a multiplication problem by making the numbers larger. The operations are different. Similarly, you cannot scale pattern matching into reasoning by making the patterns larger and the matching more sophisticated, because the operations are different, and the difference is constitutive rather than quantitative.&lt;/p&gt;

&lt;p&gt;I want to be careful here not to overstate the case, because I do believe that there are genuine intelligence-relevant capabilities that do emerge with scale in neural networks, and I want to be honest about that rather than pretending it does not exist. Few-shot learning, the ability to generalize from very few examples within a context window, appears to emerge at scale and is genuinely impressive. Some forms of cross-domain transfer, applying knowledge from one domain to solve problems in a related domain, also improve with scale. The ability to maintain coherence over longer contexts improves with scale. These are real improvements and I do not minimize them. But they are improvements in specific capabilities within the pattern-matching framework. They do not constitute progress toward the structural properties that genuine intelligence requires. A system that generalizes better within a context window is still a system that pattern matches over a context window. A system that transfers across more related domains is still a system that exploits statistical regularities across domains. The underlying architecture has not changed, and the properties that are missing from the architecture do not appear with more scale. They require different architecture, different training objectives, different representations, and a fundamentally different relationship between the system and the world it is supposed to understand.&lt;/p&gt;

&lt;p&gt;The research on what has been called "reasoning" in large language models is itself evidence against the emergence argument, rather than for it, when examined carefully. The models that are at the frontier of performance on reasoning benchmarks achieve their results partly through scale but also through fine-tuning specifically on reasoning-formatted data, test-time computation where the model produces many candidate solutions and selects the most consistent, and prompt engineering that structures the problem in ways that leverage statistical regularities in mathematical and logical text. These are all ways of making pattern matching more effective at tasks that look like reasoning. They are not evidence that genuine reasoning has emerged. The evidence for that claim would be a model that performs well on genuinely novel reasoning tasks that share no statistical overlap with the training distribution, and that is precisely what current frontier models fail at when that test condition is actually implemented. The frontier models still fail on systematic generalization tasks, still produce inconsistent beliefs, still hallucinate with confidence in unfamiliar domains, still rely on surface statistics in ways that are exposed by adversarial testing. The architecture is the constraint, and the constraint has not been loosened by scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Should Exist Instead, and Why It Does Not
&lt;/h2&gt;

&lt;p&gt;I want to end with what I think the alternative should look like, because I have always said that criticism is more honest when it comes paired with a direction, and I have a direction in mind. I have described it in pieces across several posts, in the argument for equations in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, in the argument for training-free intelligence in &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether.&lt;/a&gt;, and in the argument for causal reasoning in &lt;a href="https://wiseai.dev/blogs/all-you-have-access-to-is-knowledge-and-tools-never-intelligence" rel="noopener noreferrer"&gt;All You Have Access To Is Knowledge and Tools; Never Intelligence!&lt;/a&gt;. Here I want to bring those pieces together and describe what a system that could actually be genuinely intelligent would need to have.&lt;/p&gt;

&lt;p&gt;It would need to learn causal structure rather than statistical associations, to go from observing that X and Y co-occur to representing the mechanism by which X causes Y, and to use that representation to predict what will happen when X is changed by intervention rather than just when X is observed. Judea Pearl's causal inference framework provides the mathematical foundation for this, and the field of causal machine learning is developing methods for estimating causal structure from observational data (1). This is genuinely hard and has not yet been solved in a general way, but it is the right direction, because only causal knowledge supports reliable prediction under intervention, and reliable prediction under intervention is what intelligence is for.&lt;/p&gt;

&lt;p&gt;It would need to discover compact symbolic representations of the patterns it observes, representations that encode the mechanism rather than just the surface, and that can be inspected, verified, and extended. Symbolic regression, the area of research I described in my post on mathematical equations, is the technical approach to learning equations rather than functions, and recent progress on this front, represented by systems like AI Feynman from the Tegmark lab (14), has shown that compact symbolic representations of physical laws can be recovered from data by learning systems. This direction is slower and more difficult than deep learning, but it produces representations that are more powerful precisely because they encode the why rather than just the what.&lt;/p&gt;

&lt;p&gt;It would need compositional inference, the ability to combine representations of concepts in novel ways that were not present in the training data, and to derive the implications of those combinations through inference rules rather than by pattern matching over the training distribution. This means something like a symbolic inference engine operating over learned representations, not the pure symbolic AI of the 1980s that failed for different reasons, but a hybrid that learns the representations from data and reasons over them with principled inference rules (15). The neurosymbolic research program is the most promising current approach to this, and Belle and Marcus have argued that the combination of neural learning and symbolic reasoning is the most achievable path to systematic compositionality. Progress has been made, and it is the right direction even if the engineering challenges are substantial.&lt;/p&gt;

&lt;p&gt;It would need explicit modeling of the world as a causal structure that can be simulated forward, a world model in the technical sense that I described in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, not just a compressed statistical representation of training data but a dynamic model that can be used to predict the future states of the world given current states and actions. World model research, represented by systems like DreamerV3 and other model-based reinforcement learning approaches, has made real progress on this front in constrained environments. The challenge is scaling world models to the open-ended complexity of the real world without losing the principled structure that makes them more than neural networks trained on simulation data.&lt;/p&gt;

&lt;p&gt;And it would need genuine epistemic humility as an architectural property, not as a fine-tuned behavior added to a model that is by default overconfident. The system would need an explicit model of its own knowledge states, distinguishing between what it knows from principled inference from a verified causal model and what it knows from statistical association with low confidence. This is the hardest part to engineer because it requires the system to reason about itself, to maintain a model of its own epistemic limitations, and that kind of self-modeling is exactly the kind of second-order representation that neural networks struggle with. But without it, any system that presents its outputs as reliable knowledge is lying, and the system's users bear the cost of the lie.&lt;/p&gt;

&lt;p&gt;The reason this alternative does not exist at scale today is not that it is impossible or that its theoretical foundations are weak. The foundations are strong. The reason is that it is harder to build and slower to show results than neural networks, and the AI industry is organized around showing results quickly to attract investment, and investment is the oxygen that determines what gets built. The incentive structure of the current industry is perfectly designed to produce very large neural network systems and very poor progress toward the alternative I am describing, and that is what the incentive structure has produced. I described in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt; what it is like to be on the wrong side of an industry's incentive structure, to be doing the right thing and getting punished for it by a system that rewards the wrong thing because the wrong thing is more immediately profitable. The researchers working on causal machine learning, symbolic regression, neurosymbolic integration, and world models are on the right side of the technical argument and the wrong side of the funding structure, and that is a situation that I recognize and that fills me with something between anger and solidarity.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/wiseaidotdev/lmm" rel="noopener noreferrer"&gt;lmm project&lt;/a&gt; is my concrete, executable response to this situation. It is written in Rust, it is built around symbolic regression and physics simulation rather than around gradient descent over text corpora, and it is a deliberate proof of concept that genuine intelligence-oriented capabilities can be built without training on human creative expression. It is not finished. It is not competitive with GPT-5 on the tasks that GPT-5 is commonly used for. But it is on the right side of the architectural argument, and I will keep building it because the alternative, building nothing and only criticizing what others have built, is the most comfortable and least useful position available to me, and comfort is the one luxury I have never been able to afford.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Pearl, J., &lt;em&gt;Causality: Models, Reasoning, and Inference&lt;/em&gt;, &lt;a href="https://doi.org/10.1017/CBO9780511803161" rel="noopener noreferrer"&gt;Cambridge University Press, 2009&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; Zhang, C. et al., &lt;em&gt;Understanding Deep Learning Requires Rethinking Generalization&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1611.03530" rel="noopener noreferrer"&gt;arXiv:1611.03530&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Fodor, J. A. &amp;amp; Pylyshyn, Z. W., &lt;em&gt;Connectionism and Cognitive Architecture: A Critical Analysis&lt;/em&gt;, &lt;a href="https://doi.org/10.1016/0010-0277(88)90031-5" rel="noopener noreferrer"&gt;Cognition, 1988&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; Power, A. et al., &lt;em&gt;Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2201.02177" rel="noopener noreferrer"&gt;arXiv:2201.02177&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Goodfellow, I. J. et al., &lt;em&gt;Explaining and Harnessing Adversarial Examples&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1412.6572" rel="noopener noreferrer"&gt;arXiv:1412.6572&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Gururangan, S. et al., &lt;em&gt;Annotation Artifacts in Natural Language Inference Data&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1803.02324" rel="noopener noreferrer"&gt;arXiv:1803.02324&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Bender, E. M. et al., &lt;em&gt;On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?&lt;/em&gt;, &lt;a href="https://doi.org/10.1145/3442188.3445922" rel="noopener noreferrer"&gt;ACM FAccT, 2021&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; Kadavath, S. et al., &lt;em&gt;Language Models (Mostly) Know What They Know&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2207.05221" rel="noopener noreferrer"&gt;arXiv:2207.05221&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-9"&gt;&lt;/span&gt;&lt;strong&gt;9.&lt;/strong&gt; Strubell, E. et al., &lt;em&gt;Energy and Policy Considerations for Deep Learning in NLP&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1906.02243" rel="noopener noreferrer"&gt;arXiv:1906.02243&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-10"&gt;&lt;/span&gt;&lt;strong&gt;10.&lt;/strong&gt; Elazar, Y. et al., &lt;em&gt;Measuring and Improving Consistency in Pretrained Language Models&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2102.01017" rel="noopener noreferrer"&gt;arXiv:2102.01017&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-11"&gt;&lt;/span&gt;&lt;strong&gt;11.&lt;/strong&gt; Kosinski, M., &lt;em&gt;Evaluating Large Language Models in Theory of Mind Tasks&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2302.02083" rel="noopener noreferrer"&gt;arXiv:2302.02083&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-12"&gt;&lt;/span&gt;&lt;strong&gt;12.&lt;/strong&gt; Grice, H. P., &lt;em&gt;Logic and Conversation&lt;/em&gt;, in &lt;em&gt;Syntax and Semantics, Vol. 3&lt;/em&gt;, P. Cole &amp;amp; J. Morgan (eds.), Academic Press, 1975, &lt;a href="https://doi.org/10.1163/9789004368811_003" rel="noopener noreferrer"&gt;https://doi.org/10.1163/9789004368811_003&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-13"&gt;&lt;/span&gt;&lt;strong&gt;13.&lt;/strong&gt; Shumailov, I. et al., &lt;em&gt;The Curse of Recursion: Training on Generated Data Makes Models Forget&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2305.17493" rel="noopener noreferrer"&gt;arXiv:2305.17493&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-14"&gt;&lt;/span&gt;&lt;strong&gt;14.&lt;/strong&gt; Udrescu, S. M. &amp;amp; Tegmark, M., &lt;em&gt;AI Feynman: A Physics-Inspired Method for Symbolic Regression&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1905.11481" rel="noopener noreferrer"&gt;arXiv:1905.11481&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-15"&gt;&lt;/span&gt;&lt;strong&gt;15.&lt;/strong&gt; Belle, V. &amp;amp; Marcus, G., &lt;em&gt;The Future Is Neuro-Symbolic: Where Has It Been, and Where Is It Going?&lt;/em&gt;, &lt;a href="https://doi.org/10.1609/aaai.v40i48.42130" rel="noopener noreferrer"&gt;Proceedings of the AAAI Conference on Artificial Intelligence, 2026&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>lmm</category>
    </item>
    <item>
      <title>All You Have Access To Is Knowledge and Tools; Never Intelligence!</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 18:53:46 +0000</pubDate>
      <link>https://dev.to/wiseai/all-you-have-access-to-is-knowledge-and-tools-never-intelligence-5792</link>
      <guid>https://dev.to/wiseai/all-you-have-access-to-is-knowledge-and-tools-never-intelligence-5792</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This post was originally published on &lt;a href="https://wiseai.dev/blogs/all-you-have-access-to-is-knowledge-and-tools-never-intelligence" rel="noopener noreferrer"&gt;the main website&lt;/a&gt; on &lt;a href="https://github.com/wiseaidotdev/blog/pull/18" rel="noopener noreferrer"&gt;Apr 23 2026&lt;/a&gt;. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account &amp;lt;3!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;In my previous posts I have been building one argument from many different angles, and if you have been reading along, you already know where I keep landing. In &lt;a href="https://wiseai.dev/blogs/language-is-limited-asi-is-impossible" rel="noopener noreferrer"&gt;Language is Limited. ASI is Impossible.&lt;/a&gt;, I said that words are not thoughts and that any machine trapped inside symbols is trapped inside a cage it cannot see. In &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, I argued that the only honest language for describing reality is mathematics, because equations encode mechanisms while text only encodes descriptions. In &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether.&lt;/a&gt;, I went further and argued that the training paradigm itself is a form of extraction that concentrates value in the wrong hands while pretending to be progress. In &lt;a href="https://wiseai.dev/blogs/rethinking-arc-agi" rel="noopener noreferrer"&gt;Rethinking ARC‑AGI&lt;/a&gt;, I showed that even the benchmarks people use to celebrate AI progress are essentially measuring the wrong thing, rewarding pattern matching while calling it reasoning. All of these posts have been getting closer to something I want to say plainly in this one, something I have been circling for a long time without just saying it in the most direct language I can manage. The thing is this: every AI system you have ever used, every assistant, every chatbot, every agent, every copilot, every oracle dressed up in a sleek interface, has access to knowledge and has access to tools, but none of them, not one, has intelligence in the sense that word deserves to carry. I do not mean that as a small technical qualifier. I mean it as the central fact about what modern AI actually is, and I think almost nobody in the mainstream conversation is saying it this directly, and so I am going to say it here, as carefully and as plainly as I can, and let the argument carry its own weight.&lt;/p&gt;

&lt;p&gt;I want to be precise about what I mean by each of those three words, because precision is the only thing that separates a real argument from a confident-sounding opinion. Knowledge means stored information, patterns absorbed from data, associations learned between inputs and outputs, facts held in weights or retrieved from an index. Tools means external functions a system can call: a search engine, a calculator, a code interpreter, a database query, an API that returns real-time data. Intelligence, and this is the one that matters, means the capacity to generate new understanding from first principles, to construct a model of the world you have not been given, to reason about what you have never seen using structure that was never stored anywhere. Knowledge is a library. Tools are a set of instruments. Intelligence is the thing that knows which book to look for, and why, and what to do when the book does not exist. Current AI systems have the first two in abundance. They have the third one not at all. That is the argument. Everything else in this post is evidence and elaboration.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Knowledge Actually Is, and Why It Is Never Enough
&lt;/h2&gt;

&lt;p&gt;I want to start with knowledge because it is the thing people are most likely to confuse with intelligence, and the confusion is understandable because knowledge looks impressive from the outside. When you ask a language model to summarize the entire history of the Roman Empire, and it gives you a fluent, well-organized, apparently comprehensive answer in seconds, it is very hard not to feel like you are talking to something that genuinely knows history. It sounds like it knows. It has the vocabulary, the dates, the names of the emperors, the sequence of events, the causes and consequences, the historiographical debates. A person who knew all of that would be called learned, and that would be a form of praise. But the model's relationship to that information is fundamentally different from the relationship a learned person has to the same material, and the difference is the whole point of this post, and it is not a subtle or philosophical difference that only matters in edge cases. It is a structural difference that shows up every time you ask the system to do something genuinely new.&lt;/p&gt;

&lt;p&gt;A person who has studied Roman history has built a model of that world inside their mind. They have connected the economic pressures of the third century to the political instability, connected the political instability to the military dependence on mercenaries, connected that dependence to the erosion of civic identity, and connected that erosion to the eventual fragmentation of the Western Empire. That model was not stored. It was constructed, over time, through a process of active engagement with evidence, through reading, debating, questioning, revising, and testing the model against new information. The model lives as a dynamic structure inside the historian's mind, and the historian can use it to answer questions that were never asked during the learning process, because the model is generative. It can produce new answers from old understanding. This is what I mean when I say knowledge is not intelligence. Knowledge is the raw material that intelligence works on, but the working itself is a separate capacity, and current AI systems have the raw material in vast quantities without having the capacity to work on it in the deep sense.&lt;/p&gt;

&lt;p&gt;The technical basis for this claim is not controversial within the research community, even though it rarely makes it into the popular conversation about AI. Large language models learn to predict the next token in a sequence, which means they learn a compressed representation of the statistical patterns in their training data, and that representation is genuinely extraordinary in its scope and detail. But what the representation encodes is the surface structure of human knowledge as expressed in text, not the underlying conceptual structure that the text was pointing at. Think of it this way. A transcript of every lecture ever given at every university is not the same as the understanding that those lectures were trying to convey. The transcript is the shadow of the understanding, and you can learn a lot from the shadow, and shadows are useful, but a shadow is not the thing itself. Research by Bender, Koller, and colleagues in the landmark Stochastic Parrots paper showed definitively that form, meaning the statistical structure of language, does not determine meaning, meaning the connection between language and the world, and that a model trained to predict form will not automatically learn meaning no matter how large it becomes (1). This is the theoretical foundation of what I am arguing, and it is a solid foundation because it is grounded in a careful analysis of what training on text can and cannot achieve as a matter of principle, not just as an empirical observation about current model behavior.&lt;/p&gt;

&lt;p&gt;I have seen people respond to this argument by saying that the knowledge stored in these systems is still useful, and that is true, and I am not denying it. I said it in &lt;a href="https://wiseai.dev/blogs/as-engineers-llms-should-pay-us-for-tokens-usage" rel="noopener noreferrer"&gt;As Engineers, LLMs should pay us for tokens usage&lt;/a&gt;, and I stand by it: these systems are useful tools that help people do real work faster. What I am disputing is not their utility but the narrative built around them, the claim that the knowledge they contain is evidence of intelligence, that their fluent retrieval of stored patterns constitutes understanding, and that more knowledge plus better retrieval will eventually add up to genuine thinking somewhere along the way. That narrative is what I am calling out as false, because it conflates the map with the territory, the description with the thing described, the storage with the understanding. A system that can retrieve the fact that gravity causes objects to fall is not the same as a system that understands gravity, and the difference is not a matter of degree. It is a matter of kind. Newton did not merely retrieve the fact that apples fall. He discovered the law that made the falling predictable across every possible case, including cases nobody had ever observed, and that discovery was intelligence operating on knowledge, not knowledge by itself magically becoming intelligence through scale or polish.&lt;/p&gt;

&lt;p&gt;Let me be very concrete about where this distinction matters most, because abstract arguments always benefit from landing in a specific place. Medical diagnosis is one of the most important domains in which AI is currently being deployed, and it is a domain that clearly illustrates the difference between knowledge and intelligence. A language model trained on medical literature has access to an enormous amount of clinical knowledge, symptoms, lab values, disease mechanisms, treatment protocols, drug interactions, and outcomes from millions of cases. When given a patient presentation, it can produce a differential diagnosis that looks impressive and often includes the right answer. But producing a list of possibilities from pattern matching is not the same as understanding why this particular patient, with this particular history, this particular constellation of risk factors, and this particular presentation, is more likely to have one disease than another. That reasoning requires a causal model of how diseases produce symptoms, how risk factors modify probability, how the timeline of symptom onset constrains the diagnosis, and how the biological mechanisms interact. A language model does not have that causal model. It has the text that describes that causal model, and those are different things, and the difference can cost lives when the system is wrong in a case that falls outside the patterns it learned. Research published in The Lancet Digital Health has documented that even high-performing AI diagnostic systems show systematic failures on atypical presentations precisely because they are matching patterns rather than reasoning causally (2). The knowledge is there. The intelligence is not.&lt;/p&gt;

&lt;p&gt;I also want to say something about the problem of knowledge staleness, because it is another dimension of why knowledge alone is never enough. A language model's knowledge has a cutoff date, after which the world has changed but the model has not. That cutoff problem is often presented as a technical limitation that retrieval augmented generation can fix, and that framing misses the deeper issue. Even with perfect retrieval of up-to-date information, the problem of integration remains. New knowledge does not automatically incorporate itself into a coherent understanding of the world. A person who learns that a new treatment has been shown to work for a disease they know well immediately begins the process of updating their causal model, asking how this treatment interacts with the mechanisms they already understand, what it implies for patients they are currently treating, and what questions it raises about the underlying biology. That integration is intelligence at work, and no amount of retrieved text performs it automatically. The retrieved text is still just more raw material, and raw material without the capacity to process it intelligently is not intelligence. It is a pile. A very large, very well-organized pile, but still a pile, and piles do not think.&lt;/p&gt;

&lt;p&gt;I want to close this section with something personal, because I promised myself I would always connect these arguments to lived experience, and I think it matters here. When I was struggling to find work after technology destroyed the career path I had been building, which I described in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt;, one of the things I did was read a huge amount, about systems design, about machine learning, about distributed systems, about everything I thought I needed to know to get a job in the industry that had displaced me. I had knowledge. I was accumulating it as fast as I could. But knowledge without the intelligence to apply it in a context-sensitive, adaptive, creative way felt like carrying bricks without knowing how to build anything. I could answer trivia questions about distributed consensus algorithms and still have no idea how to think about the specific design problem sitting in front of me in an interview. The knowledge was necessary but nowhere near sufficient. What the interviewers were testing, whether they articulated it this way or not, was whether I could think with the knowledge, not just about it. Current AI systems have the same limitation at a structural level, and the people who build them know it, and the people who deploy them are slowly discovering it, and the people who use them will eventually notice it in the moments that matter most.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Tools Actually Are, and Why They Cannot Give Intelligence to Their User
&lt;/h2&gt;

&lt;p&gt;The second half of the title of this post is about tools, and I want to give tools their proper treatment, because tools are often presented as the solution to the intelligence gap I described in the previous section. The logic goes like this: even if a language model cannot reason causally by itself, if you give it access to a calculator, it can do arithmetic; if you give it access to a search engine, it can look up current information; if you give it access to a code interpreter, it can verify whether its reasoning is correct by running the code. This is the foundation of the agent paradigm that has dominated AI discourse for the past 2-3 years, where a language model sits at the center of a system that can call tools, take actions, observe results, and iterate. And I want to be honest that this paradigm has produced some genuinely impressive demonstrations. But impressive demonstrations are not intelligence, and I want to explain precisely why tools cannot bridge the gap, because the explanation is important and it is being almost entirely missed in the current conversation.&lt;/p&gt;

&lt;p&gt;Tools extend what a system can do, but they do not change what the system fundamentally is. This is a principle that applies to humans as well as to machines, but in humans, there is always an intelligent agent deciding which tool to use, when to use it, how to interpret what the tool returns, and what it means in the context of the overall problem. When I use a calculator, the calculator does not make me a mathematician. It removes the burden of arithmetic so my intelligence can focus on the mathematical structure of the problem. The intelligence is still mine. The tool serves it. When a language model uses a calculator, the situation is structurally different, because there is no intelligence directing the use of the tool from outside the statistical process. There is only the statistical process trying to decide which tool to call based on patterns in its training data about when people typically use calculators. That call may often be correct, and when it is correct, the output looks intelligent. But correct and intelligent are not the same thing, and the system's inability to tell the difference between cases where the tool is appropriate and cases where it looks appropriate but produces the wrong answer is the clearest symptom of the intelligence gap. A truly intelligent system knows why it is using a tool. A system without intelligence knows only when tools tend to be used.&lt;/p&gt;

&lt;p&gt;This distinction becomes catastrophically important in agentic settings, where a language model is given a long-horizon task, access to multiple tools, and the responsibility to plan a sequence of actions over time. Research from Stanford and elsewhere has shown that even the most capable language model agents fail systematically on tasks that require more than a few steps, that involve genuine novelty, or that require the agent to recognize when its current approach is wrong and to backtrack and try something different (3). These failures are not random. They follow predictable patterns tied to the statistical nature of the planning process. The agent tends to continue in the direction that looks most locally plausible rather than stepping back to reconsider the global strategy, because local plausibility is what next-token prediction optimizes for, and global strategic reasoning requires a kind of self-monitoring that statistical pattern matching does not naturally support. It can call a search engine and retrieve results, but whether it correctly identifies which retrieved result is actually relevant to the problem it is trying to solve, and why, and what it implies for the next step, is a matter of intelligence that the tool use does not provide. The tool returns a result. The system has to understand the result. And understanding is the thing that is missing.&lt;/p&gt;

&lt;p&gt;I want to talk about function calling specifically, because it is the technical mechanism through which most AI tool use happens, and it is worth being precise about what it actually does. When a language model is given a set of function specifications and a user query, it learns to recognize patterns in the query that suggest which function should be called with which arguments, and to format the output accordingly. This is a genuine and useful capability, but from a cognitive standpoint it is essentially a very sophisticated keyword matching system that has learned to translate natural language requests into structured API calls. The intelligence that designed the functions, that decided what functions should exist, that determined what the functions should do, that understood how they relate to each other and to the domain they serve, that intelligence came from humans, and it lives in the function specifications, the documentation, and the structure of the API. The language model is calling functions that humans built using human intelligence. It is not exercising its own intelligence to figure out what functions should exist. When the right function exists and the query is within the distribution of what the model has seen during training, it works well. When the right function does not exist, or when the query requires combining functions in a genuinely novel way, or when the functions return unexpected results that require creative interpretation, the system has no recourse except to produce statistically plausible-sounding text, which may or may not be correct, and which the system has no reliable way to verify. That is not intelligence augmented by tools. That is pattern matching that happens to have some tools attached to it.&lt;/p&gt;

&lt;p&gt;I also want to address the retrieval augmented generation paradigm specifically, because it is strongly associated with the idea that tools can fix the knowledge limitations I described in the previous section. Retrieval augmented generation works by embedding a query, searching a vector store for documents that are semantically similar to the query, and providing those documents as context to the language model so it can generate a response grounded in specific retrieved information rather than relying purely on its training data. This is genuinely useful in many practical settings, and I do not deny that it improves reliability compared to pure in-context generation. But it has two deep limitations. The first is that the retrieval step uses semantic similarity, meaning it retrieves documents that look like the query, not documents that are causally relevant to answering the question correctly. Good retrieval requires understanding what evidence is needed to answer a question, not just finding text that sounds related, and that understanding is a matter of intelligence that the embedding-based retrieval system does not possess. The second limitation is that even after retrieval, the model must synthesize the retrieved information with its prior knowledge, reason about what the retrieved documents actually say versus what they imply, identify contradictions between sources, and construct a coherent answer. That synthesis and reasoning is where intelligence is required, and the retrieval step does not provide it. Tools can find the books. They cannot read them intelligently.&lt;/p&gt;

&lt;p&gt;The environmental and systemic costs of the tool-augmented AI paradigm deserve attention here, because they are real and they are being paid by people who did not choose to pay them. Running an agentic AI system that makes dozens or hundreds of API calls in the course of handling a single user request consumes computational, network, and financial resources at a rate that is dramatically higher than running a simple retrieval system. Those costs are not inherent to providing the tool-augmented capability. They are inherent to the specific architecture of a statistical language model trying to simulate intelligence by calling tools repeatedly. In &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether.&lt;/a&gt;, I noted that the environmental cost of training large models has been extensively documented in research by Strubell and colleagues (4). The cost of inference at scale is less frequently discussed but equally real, and when that inference is being done repeatedly by a system that is trying to compensate for its lack of intelligence by calling tools over and over until it finds something that works, the waste is structural and not incidental. A truly intelligent system would figure out which tool to call, call it once, and understand the result. A statistical system trying to simulate intelligence uses tools the way someone trying to remember a phone number they have forgotten might keep guessing digits, iterating through plausible combinations, hoping eventually to land on the right one. That is not intelligence using tools. That is the absence of intelligence being masked by repeated tool use.&lt;/p&gt;

&lt;p&gt;I want to say something here that connects to my personal experience, because the tools question is deeply personal for me in a way that most people would not expect. When I was building the systems I described in my earlier posts, I spent years learning how to use sophisticated tools. Version control, profilers, debuggers, distributed tracing systems, load testers, all of the instruments of modern software engineering. And I learned something important that I think applies directly to the AI tool use question: a tool in the hands of someone who does not understand what they are trying to accomplish is not just useless. It is actively dangerous. It produces outputs that look like results, that can be formatted and reported and presented as evidence, but that are actually noise. It gives false confidence. I have seen junior engineers use profiling tools to identify bottlenecks and act on the results without understanding whether the identified bottleneck was actually the source of their performance problem, and the result was often that they optimized the wrong thing while the real problem remained untouched. The tool had done its job. The intelligence to interpret the tool's output was missing. That is exactly the situation with AI systems that call tools today. The tools do their jobs. The intelligence to interpret what those jobs actually mean for the problem at hand is not there.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Intelligence That Was Never Stored
&lt;/h2&gt;

&lt;p&gt;The previous two sections established that knowledge and tools are real, useful, and genuinely present in modern AI systems, but that neither of them is intelligence. This section is about what intelligence actually is, because you cannot argue that something is missing without describing what the missing thing is, and describing what intelligence is happens to be one of the hardest problems in all of cognitive science. I am going to try to describe it anyway, not because I have solved the hard problem of consciousness or because I have a complete theory of mind, but because I think there is enough common ground in cognitive science, neuroscience, and philosophy to make the case clearly enough to support the argument I am building. And I want to use simple words the whole time, because complex vocabulary is often a way of hiding from difficult ideas, and I do not want to hide from this one.&lt;/p&gt;

&lt;p&gt;Intelligence, at its most basic, is the capacity to build internal models of the world that are generative, meaning they can produce new predictions about situations that were never directly experienced, and adaptive, meaning they update when the predictions turn out to be wrong. It is the capacity to abstract from specific observations to general principles, and then to apply those principles to new specific situations that were not part of the original abstraction. It is the capacity to recognize when a current approach is failing and to switch strategies without being told to switch. It is the capacity to formulate questions that have never been asked, because recognizing that a question needs to be asked is itself a product of understanding the domain well enough to notice what is missing. It is, at a deeper level, the capacity to be genuinely surprised by the world, because surprise requires a model of what was expected, against which the actual outcome can be compared and found to be different. None of these capacities are properties of stored knowledge, and none of them are properties of tools. They are all properties of a process, an ongoing, dynamic, self-correcting process of building, testing, and revising models of reality. Cognitive scientists call this process model-based reasoning, and a large body of research establishes that it is distinct from the kind of pattern matching that characterizes both animal conditioning and artificial neural network learning (5).&lt;/p&gt;

&lt;p&gt;The most important property of intelligence, the one that most clearly separates it from knowledge retrieval, is what researchers call systematic compositionality. Humans can take a finite set of known concepts and combine them in an infinite number of novel ways to produce new thoughts. If you understand what a dog is, and you understand what a purple is, and you understand what a mountain is shaped like, you can immediately form a coherent mental image of a purple dog sitting on top of a mountain, even though no text you ever read described exactly that scene. You can then reason about what kind of behavior such a scene would cause if you encountered it, how you might photograph it, what the lighting would look like at sunset, what the dog's fur would feel like at altitude. None of that reasoning requires retrieving a stored description. It requires composing known concepts in a new configuration and then running your model of the world forward to produce new predictions. Language models fail at systematic compositionality in ways that have been carefully documented (6). When tested on tasks that require combining known concepts in configurations that differ from training examples, their performance drops dramatically, while human performance stays consistent because humans are using a compositional generative model rather than pattern matching to a database of seen combinations. This is not a data problem. It is not a scale problem. It is a structural problem: statistical pattern matching over tokens does not naturally produce compositional generative representations, and no amount of training data changes that structural fact.&lt;/p&gt;

&lt;p&gt;I also want to talk about what researchers call the frame problem, because it is one of the oldest and most stubborn problems in artificial intelligence, and it is directly relevant to why intelligence cannot be reduced to knowledge plus tools. The frame problem, originally identified in the context of symbolic AI and later shown to be equally relevant to connectionist systems, is the problem of knowing what changes when something happens and what does not change (7). When you push a coffee cup across a table, you do not have to explicitly reason about the fact that the color of the walls has not changed, or that the gravitational constant is still the same, or that the laws of physics still apply. You take all of that for granted because your intelligence has an implicit model of the world that marks what is relevant to the current action and what is not. A system without that implicit model has to either reason about everything explicitly, which is computationally intractable, or make assumptions about what is relevant based on statistical priors, which leads to systematic failures whenever the situation differs from the training distribution. Language models handle the frame problem statistically: they learn what kinds of things tend to change together in the text they were trained on, and they generate outputs consistent with those learned associations. This works within the training distribution and fails outside it, exactly as the research on compositional generalization predicts. Intelligence has a solution to the frame problem. Knowledge does not. Tools do not. Only the generative model of the world that intelligent beings build from direct engagement with reality provides a principled way to know what changes and what stays the same, because that model encodes the causal structure of the world rather than the statistical associations in descriptions of the world.&lt;/p&gt;

&lt;p&gt;Let me connect this to something I said in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, because that post was building toward the same point from a different direction. I argued there that equations encode mechanisms rather than descriptions, and that this encoding is what gives them their power to generate predictions in multiple modalities from a single compact representation. What I want to add here is that this property of equations is exactly what a generative model of the world needs. An intelligent system that has discovered the differential equation governing a physical process does not need to store examples of what that process looks like. It can generate new examples by running the equation forward, it can predict outcomes by running the equation, it can diagnose interventions by modifying variables in the equation, and it can check its predictions against new observations. That is intelligence at work, and it is specifically the kind of intelligence that knowledge retrieval and tool use cannot produce, because it requires the internal model to be mechanistic and generative rather than associative and retrieval-based. The equation is not stored knowledge. It is a compressed theory of how the world works in a specific domain, and theories are the products of intelligence, not the inputs to it.&lt;/p&gt;

&lt;p&gt;I know there will be people who say that large language models show signs of emerging reasoning ability, that they can solve novel math problems, draw analogies, and demonstrate knowledge transfer that suggests something more than pure pattern matching is happening. I take these claims seriously because the researchers making them are often serious people, and I want to engage with them honestly rather than dismissing them. But I think the evidence, when examined carefully, consistently shows that what looks like reasoning is in most cases very sophisticated pattern completion. The argument that I find most compelling comes from work on symbolic reasoning tasks by Marcus and colleagues, and separately from the work on large language model failures by Dziri and colleagues, both of which show that model performance on reasoning tasks is highly sensitive to surface features of the problem presentation in ways that true reasoning should not be (8). If a model had genuinely reasoned its way to an answer, rephrasing the problem in a different surface form should not change the answer, because the reasoning would be operating on the underlying structure rather than the surface pattern. But in experiment after experiment, that is exactly what happens: changing the surface form changes the answer, revealing that the model was matching to patterns in its training data rather than reasoning from a model of the problem's underlying structure. This is the fingerprint of pattern matching, not intelligence, and it is visible in the data if you know where to look.&lt;/p&gt;

&lt;p&gt;The neuroscientific perspective adds another layer of evidence that is worth considering here. Human intelligence is not just a property of the neocortex doing something that looks like language processing. It is distributed across a vast system that includes sensorimotor representations, embodied predictions, emotional signals that carry information about risk and relevance, episodic memory that preserves the specific context of past events, and a default mode network that keeps running simulations of past and future situations even when no external task is being performed (9). A language model trained on text interacts with none of this biological substrate. It receives text, processes it through transformer layers, and produces text. The richness of the representations available to the human mind, representations shaped by years of embodied experience in a physical world that pushes back, that has real consequences for wrong predictions, that produces pain when you touch something hot and satisfaction when you solve a real problem, none of that richness is present in the statistical weights of a language model. I said in &lt;a href="https://wiseai.dev/blogs/language-is-limited-asi-is-impossible" rel="noopener noreferrer"&gt;Language is Limited. ASI is Impossible.&lt;/a&gt; that the brain is not a text machine, and what I am adding here is that intelligence is not a text property. It is a property of a certain kind of dynamic, embodied, feedback-coupled engagement with a real world, and language models, by design, have none of that engagement. They have the text that humans produced from that engagement. That is not the same thing. That will never be the same thing regardless of how large the model gets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Retrieval Is Not Reasoning, No Matter How Fast It Is
&lt;/h2&gt;

&lt;p&gt;One of the most powerful illusions in the current AI conversation is the conflation of retrieval speed with reasoning ability. When a language model produces an answer in seconds that would take a human expert minutes or hours to formulate, the speed feels like evidence of superior intelligence. This reaction is understandable and I have felt it myself. But speed of retrieval is not the measure of intelligence, and this confusion is worth confronting directly because it shapes so many of the most popular claims about what current AI systems can do. A search engine can return millions of results in milliseconds, and nobody claims that the search engine is intelligent. The speed is a property of the indexing structure and the hardware, not of any reasoning process. Language models retrieve from a much richer index, one that stores not just documents but compressed patterns of association between ideas, and they produce their retrievals in fluent prose rather than ranked links, which makes the retrieval feel much more like thinking. But the underlying operation is still fundamentally retrieval, and the felt quality of the output tells us nothing about whether the process that produced it is intelligent.&lt;/p&gt;

&lt;p&gt;I want to be precise here because I think the argument requires precision to be convincing. When a human expert produces an answer to a complex question, the process involves more than retrieving a stored answer. It involves identifying which parts of the question are familiar and which are novel, constructing a representation of the question's structure that allows the relevant knowledge to be brought to bear, reasoning about how the relevant knowledge connects to the specific question, checking the emerging answer against constraints imposed by other things the expert knows, and often revising the answer as the reasoning process produces unexpected implications. All of that is intelligence operating on knowledge. The output may look like a stored answer because the expert has answered similar questions before, but the process that produced it is generative and adaptive, not purely retrievative. When a language model produces a similarly fluent answer to a similarly complex question, the process is much closer to pattern completion over learned associations. The model has not identified the novel aspects of the question and reasoned about them specifically. It has found the region of its learned representation space that is most consistent with the input and sampled from the distribution of outputs associated with that region. When the question is similar to things in the training data, this produces impressive-looking results. When the question is genuinely novel in its structure, the results degrade in characteristic ways that reveal the retrieval nature of the underlying process.&lt;/p&gt;

&lt;p&gt;The benchmark results that are most often cited as evidence of AI reasoning ability deserve scrutiny here, because they are consistently misinterpreted in the public conversation. When a language model achieves a high score on a standardized reasoning benchmark, the natural interpretation is that the model has learned to reason in the way the benchmark was designed to test. But performance on a standardized benchmark can be achieved either through genuine reasoning ability or through having seen enough examples of the benchmark during training to pattern-match to the test questions. Researchers have repeatedly shown, by constructing modified versions of popular benchmarks that preserve the underlying reasoning structure while changing the surface form, that model performance drops dramatically on the modified versions while human performance stays consistent (10). This is the smoking gun. If the model had learned to reason, it would transfer that reasoning to the modified version, because the structure is the same. The fact that it does not transfer reveals that the performance was achieved through pattern matching to the training distribution rather than through genuine reasoning ability. I described something similar in &lt;a href="https://wiseai.dev/blogs/rethinking-arc-agi" rel="noopener noreferrer"&gt;Rethinking ARC‑AGI&lt;/a&gt;, where I showed that ARC-AGI version 1 was undermined by brute-force search precisely because the evaluation was measuring something that could be achieved without genuine reasoning, and the same principle applies to most of the benchmarks currently used to evaluate AI progress.&lt;/p&gt;

&lt;p&gt;I also want to address chain-of-thought prompting specifically, because it is often presented as evidence that language models can reason when given the opportunity to show their work. Chain-of-thought prompting involves asking a model to produce a sequence of reasoning steps before giving a final answer, and it does improve performance on many tasks compared to asking only for the answer directly. I want to be careful here because I think the evidence is genuinely nuanced, and oversimplifying it would make my argument weaker rather than stronger. Chain-of-thought prompting improves performance on tasks within the training distribution because producing reasoning steps is itself a pattern that the model has learned to reproduce, and that pattern, when reproduced, tends to activate the regions of representation space that contain the right answer. It is, in a meaningful sense, a productive pattern to reproduce. But research by Lanham and colleagues has shown that the chain-of-thought steps produced by language models are often not causally connected to the final answer in the way that genuine reasoning steps would be (11). In experiments where the intermediate reasoning steps are deliberately made incorrect while keeping the problem the same, models often still produce the correct final answer, which reveals that they were not actually using the intermediate steps to compute the answer. They were producing plausible-sounding reasoning steps in parallel with pattern-matching to the final answer, and the two processes were not causally coupled. That is the opposite of reasoning. Reasoning is when the steps cause the conclusion. What chain-of-thought prompting often produces is a conclusion that was already determined by pattern matching, followed by a plausible post-hoc rationalization of that conclusion. That is not a small distinction. It is the whole game.&lt;/p&gt;

&lt;p&gt;The speed question connects back to the economics of AI deployment in a way that I think deserves to be made explicit. When organizations evaluate AI systems for deployment in high-stakes settings, one of the most frequently cited advantages is speed. The system can process a thousand documents in the time it takes a human analyst to read three. The system can generate a first draft of a legal brief in seconds compared to the hours a junior associate would need. The system can evaluate loan applications at a rate that would require a hundred human underwriters. All of these speed advantages are real, and they are economically valuable, and they are part of why organizations are deploying these systems at the scale they are. But the comparison is being made on the wrong dimension. The relevant question is not whether the AI system is faster than a human at retrieving and organizing information. The relevant question is whether the AI system's outputs are as reliable as a human expert's outputs in cases that require genuine reasoning rather than pattern matching within the training distribution. And the answer to that question, for currently deployed systems in currently deployed settings, is consistently no, as documented in legal, medical, financial, and scientific contexts where AI-assisted decisions have been compared against ground truth (12). The systems are fast and impressive within their training distribution. They are unreliable in novel situations that require reasoning. And the novel situations are exactly the ones where speed matters most and where errors are most costly.&lt;/p&gt;

&lt;p&gt;I want to bring this down to something concrete from my own experience building systems, because I think the abstract argument lands better when it is grounded. When I was working on systems that needed to make reliable decisions under uncertainty, the most important thing I learned was what engineers call the difference between a system that is right and a system that sounds right. A system that sounds right produces fluent, confident, well-formatted outputs that are often correct. A system that is right produces outputs that are provably connected to the information they were computed from, that degrade gracefully when that information is incomplete, and that signal uncertainty when the underlying computation cannot produce a reliable answer. The first kind of system is easy to demonstrate to stakeholders and hard to catch in its failures until the failures have real consequences. The second kind of system requires more design effort and more epistemic humility, and it produces outputs that look less impressive in a demo, but it is the system you actually want when real decisions depend on it. Language models as currently built are the first kind of system. They sound right. A genuinely intelligent system would be the second kind: right, in the sense of being connected to the truth by a traceable chain of reasoning, not just in the sense of producing outputs that match the expected surface form.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Benchmark Trap: How We Trained Ourselves to Celebrate the Wrong Thing
&lt;/h2&gt;

&lt;p&gt;One of the most reliable signs that a field has lost track of what it was trying to measure is when its benchmarks become the goal rather than the measure, and modern AI has been in this trap for long enough that most people inside the field have stopped noticing the trap. I wrote about this specifically in the context of ARC-AGI in &lt;a href="https://wiseai.dev/blogs/rethinking-arc-agi" rel="noopener noreferrer"&gt;Rethinking ARC‑AGI&lt;/a&gt;, where I documented how version one of that benchmark was undermined by brute force search and how even the improved version two still measures something narrower than the general reasoning ability it was intended to capture. But the ARC-AGI problem is only one instance of a much more general problem, which is that nearly every benchmark used to evaluate AI progress measures performance in ways that can be achieved through sophisticated pattern matching rather than genuine intelligence, and when a system achieves high performance on such a benchmark, we announce progress toward intelligence when we have only measured progress toward benchmark performance. Those two things are not always the same, and the history of AI benchmarks is the history of systems achieving high performance by learning the distribution of the benchmark rather than by learning the underlying capability the benchmark was intended to proxy.&lt;/p&gt;

&lt;p&gt;The history of this problem is long enough to be instructive. Language modeling benchmarks like GLUE and SuperGLUE were introduced to measure natural language understanding, and models quickly achieved human-level performance on them, leading to announcements of human-level language understanding. But when researchers probed the same models with carefully constructed probes designed to test whether the models were using the right kind of information to achieve their performance, they consistently found that models were using spurious correlations in the datasets, artifacts of data collection that correlated with the correct label but that had nothing to do with the linguistic understanding the benchmark was supposed to measure (13). The models had learned the benchmark distribution, not natural language understanding. This phenomenon is so well-documented and so frequently rediscovered that it has a name: Goodhart's Law, which states that when a measure becomes a target, it ceases to be a good measure. Every time a new benchmark is introduced, systems eventually optimize for the benchmark's specific distribution, and every time that happens, the community either acknowledges the limitation and moves to a harder benchmark or, more commonly, continues to cite the benchmark performance as evidence of the capability it was supposed to measure. The trap resets and closes again around a new target.&lt;/p&gt;

&lt;p&gt;What makes this particularly damaging to the public conversation about AI is that each benchmark milestone gets reported in the popular press as evidence of AI approaching or surpassing human intelligence in a specific domain, and those reports shape the expectations of the people who make decisions about AI policy, AI deployment, and AI investment. When a model achieves human-level performance on a medical licensing exam, headlines announce that AI is ready to be a doctor. When a model achieves expert-level performance on a bar exam, headlines announce that AI is ready to practice law. These headlines are not technically false in a narrow sense, but they are deeply misleading in a broader sense, because they interpret performance on a standardized test as evidence of the real-world capability that the test was designed to proxy, and the evidence consistently shows that the proxy relationship is weaker than it sounds. Research comparing AI system performance on medical licensing exams to AI system performance on actual clinical reasoning tasks has found that the high exam scores do not translate to reliable clinical reasoning, precisely because clinical reasoning requires adapting to novel patient presentations that differ from the training distribution, while exam performance can be largely achieved through pattern matching to the distribution of exam questions (14). The benchmark is the map. The capability is the territory. And when the map becomes the goal, the territory is forgotten.&lt;/p&gt;

&lt;p&gt;I want to talk about what a benchmark for intelligence would actually need to look like, because I think it is possible to design better evaluations and I want to be constructive rather than only critical. A benchmark for genuine intelligence, not performance on a fixed distribution of tasks, would need several properties that current benchmarks lack. First, it would need to use genuinely novel tasks, tasks that were designed specifically to lie outside any plausible training distribution, so that pattern matching to seen examples is definitionally impossible. Second, it would need to test transfer: after exposing the system to a novel domain long enough to learn its basic structure, can the system apply that structure to new problems that were not part of the introduction, the way a genuinely intelligent person would? Third, it would need to test self-diagnosis: can the system accurately identify when it does not know something and appropriately signal uncertainty rather than producing confident-sounding outputs that happen to be wrong? Fourth, it would need to test systematic compositionality: can the system combine known concepts in genuinely novel configurations and correctly predict the properties of the combination without having seen that specific combination during training? ARC-AGI was attempting to measure some of these properties, which is why it is more interesting than most benchmarks, but the execution requirements have proven easier to satisfy through non-intelligent means than the designers hoped. Designing a truly contamination-proof evaluation of genuine reasoning ability is hard, and the field mostly responds to that difficulty by lowering the bar rather than by doing the hard design work.&lt;/p&gt;

&lt;p&gt;The economic incentives behind benchmark culture deserve explicit attention, because they are the fuel that keeps the trap running. Companies that build AI systems have strong incentives to report benchmark performance because benchmark numbers are legible to investors, partners, and journalists in a way that nuanced capability descriptions are not. A statement that a model achieves 90 percent accuracy on a named benchmark communicates something that sounds precise and impressive and that can be compared directly to previous models, even if the benchmark is measuring something that is not actually what anyone cares about. The incentive to report on benchmarks and the incentive to optimize for benchmarks are the same incentive, operating at different stages of the development pipeline, and the result is that the entire industry moves in the direction of benchmark optimization rather than in the direction of genuine capability development, because benchmark optimization produces numbers that look good in press releases, while genuine capability development produces systems that work reliably in real-world deployments and are harder to reduce to a single number. I have seen this same dynamic in every technology-driven industry I have worked in or adjacent to, and it always produces the same eventually visible hollowness, the gap between the reported capabilities and the actual performance that users eventually discover and that researchers document and that the press eventually reports on, usually much later than the evidence warranted.&lt;/p&gt;

&lt;p&gt;I want to be fair to the researchers who design these benchmarks, because most of them know their limitations and say so clearly in their papers. The problem is not that benchmark designers are dishonest. The problem is that the gap between what a benchmark can measure and what the field needs to know gets systematically collapsed in the translation from paper to press release to public understanding. A paper that introduces a new benchmark typically includes extensive caveats about what the benchmark does and does not measure, which aspects of the task are most likely to be achieved through pattern matching, and what the results should and should not be taken to imply. Those caveats are in the paper. They are not in the press release. They are certainly not in the Twitter thread or the LinkedIn post that most people in the public actually read about the result. The field has a responsibility to close that gap, and so far it has mostly chosen not to, because closing the gap would require saying things that are more complicated and less impressive than the simple narrative of steadily increasing AI capabilities measured by steadily improving benchmark performance. I am saying those complicated things here because somebody has to say them plainly, and because I think the people who read these posts deserve the plain version more than they deserve the comfortable one.&lt;/p&gt;

&lt;p&gt;The right response to the benchmark trap is not to give up on evaluation. It is to design evaluations that measure what we actually care about, to interpret benchmark results with appropriate humility about their scope, and to resist the pressure to translate every number into a narrative about general AI capability that the number cannot support. This is easier said than done in an environment where every major AI lab is in competition for investor confidence and public attention. But the researchers who are doing this work honestly, who design their benchmarks carefully and interpret their results conservatively, are the ones whose work will look good in ten years, when the gap between the benchmark numbers and the real-world performance of the systems those numbers were supposed to represent has become impossible to ignore. I have enormous respect for that kind of scientific honesty because I know how hard it is to maintain when the surrounding culture rewards the opposite.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Real Intelligence Would Look Like, If We Built It
&lt;/h2&gt;

&lt;p&gt;I have spent several sections describing what intelligence is not, and I owe the reader something more than that. I owe a description of what real intelligence would look like if a machine actually had it, because the argument that current systems are not intelligent is only useful if there is a direction the field could go toward instead. I am not going to pretend I have a fully worked-out engineering plan for building a genuinely intelligent machine, because I do not, and anyone who claims to have such a plan should be viewed with extreme skepticism. What I do have is a set of properties that any system would need to have in order to earn the word intelligence without putting quotation marks around it, and I think describing those properties is valuable because it clarifies what the goal actually is and how far away we currently are from it.&lt;/p&gt;

&lt;p&gt;The first property a genuinely intelligent system would need is a generative causal model of the domain it operates in. Not a collection of associations between inputs and outputs. Not a set of retrieved documents that happen to be relevant to the query. A structural model that encodes the causal relationships between variables in the domain: how this causes that, why changing X changes Y in this specific way, what would happen if you intervened on Z. I argued in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt; that mathematical equations are the most compressed and honest form of such models for physical domains, and I stand by that argument. For non-physical domains, the structure of causal models is less obvious, but the principle is the same: a model that encodes mechanisms rather than associations is a model that can reason genuinely rather than retrieve approximately. Research on causal representation learning has made real progress toward systems that can learn causal structure from data, and I believe this line of work is more important for the future of real AI than all of the work on scaling language models combined (15). It is harder, it is slower, it is less impressive in demos, and it is the right direction.&lt;/p&gt;

&lt;p&gt;The second property is systematic compositionality, which I already described earlier in this post. A genuinely intelligent system should be able to take known concepts and combine them in novel ways to produce new thoughts, new predictions, and new understanding about situations that were never in its training data. This is a property that human intelligence has reliably, that current AI systems lack reliably, and that researchers have been trying to build into neural systems since the early days of connectionism. The fundamental problem is that the kind of representations needed for systematic compositionality are structured representations, where the parts maintain their identity when combined and where the combination respects the structure of both parts, but the kinds of representations that neural networks learn naturally are distributed representations, where information is spread across many neurons and where composition is achieved through learned associations rather than through structural operations. There are architectures designed to bridge this gap, including memory-augmented neural networks, slot-based representation systems, and neural symbolic hybrids, and the research on them is promising but not yet at the capability level needed to demonstrate genuinely systematic compositionality at scale (16). The problem is real and the research is real and the gap from current language models is real and large.&lt;/p&gt;

&lt;p&gt;The third property is honest uncertainty representation. A genuinely intelligent system would not produce fluent, confident-sounding outputs when it does not have a good basis for confidence. It would signal uncertainty in proportion to the actual uncertainty in its knowledge, flag when a question is outside the scope of what it can reliably answer, and defer to other sources when those sources are more reliable. This is a property that humans have imperfectly but that we recognize as epistemically virtuous when we see it, and that we correctly identify as a fault when it is absent. Calibrated confidence is not just a nice-to-have property of intelligence. It is a fundamental epistemic capability, the ability to know what you know and to know what you do not know, and Socrates was identified as the wisest man in Athens specifically because he had this capacity when everyone else lacked it. Current AI systems are systematically overconfident in ways that are well-documented (17). They produce detailed, confident answers on topics where no confident answer is warranted, and they do so because their training objective rewards producing answers that look good, not answers that accurately represent the system's epistemic state. Fixing this requires either changing the training objective fundamentally, which is technically difficult, or acknowledging that the systems are not intelligent in the sense that includes self-knowledge, which is philosophically important.&lt;/p&gt;

&lt;p&gt;The fourth property is genuine novelty generation, which is the capacity to produce something that was not in any sense present in the training data and that represents a real advance beyond what was already known. This is the hardest property to evaluate and the one most likely to be confused with sophisticated interpolation. Language models can produce outputs that seem novel because they combine elements in ways that have not been seen before, but the combination is still a function of the training data in a way that true novelty is not. When a human scientist makes a genuinely novel discovery, they are not interpolating between known states of the training distribution. They are finding a pattern in the world that nobody had found before, using a model of the world that they built from scratch through years of engagement with the domain. That kind of genuine novelty generation is what I described in &lt;a href="https://wiseai.dev/blogs/announcing-kevin-rs" rel="noopener noreferrer"&gt;I described in Announcing Kevin RS&lt;/a&gt;, where the goal of the project is to build systems that can discover things that are genuinely new rather than systems that can fluently describe things that are already known. The difference between those two goals is the difference between the past and the future of AI, and I do not think the field is taking that distinction seriously enough.&lt;/p&gt;

&lt;p&gt;The fifth property, which might be the most uncomfortable to say out loud, is the capacity for genuine motivation. A genuinely intelligent system would need some internal representation of what it is trying to do and why, some sense of the difference between making progress and not making progress, some capacity to care about the quality of its own understanding rather than just the quality of its outputs. I am not making a claim about consciousness here, because I do not think consciousness is required for intelligence in the functional sense I am describing. I am making a claim about the distinction between a system that optimizes an external objective and a system that has internalized a goal structure that it pursues on its own terms. Current AI systems optimize external objectives, specifically the objectives encoded in their training losses and reinforcement signals. They do not have internalized goals that they pursue intelligently in their own right. That is why they are fundamentally tools rather than agents, even when they are called agents in the marketing materials. A tool does what it is pointed at. An agent does what it cares about. Intelligence lives on the agent side of that distinction, and current systems are on the tool side regardless of how agentic their packaging appears.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Gap Between Capability and Understanding.
&lt;/h2&gt;

&lt;p&gt;The argument I have made in this post could sound like an academic debate about definitions, about what counts as intelligence versus capability versus understanding. I want to be very clear that it is not. The gap between genuine intelligence and the impressive-but-not-intelligent systems we are currently deploying at massive scale has consequences for real people in real situations, and I want to make those consequences concrete because they are the reason I am writing this post rather than keeping the argument inside the technical literature where it might be safely ignored by everyone who needs to hear it most. I described some of these consequences in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt;, but here I want to go deeper into the specific ways that the conflation of capability with intelligence is causing harm that could have been avoided if the people deploying these systems had been honest about what they actually had.&lt;/p&gt;

&lt;p&gt;The legal system is one of the most consequential domains in which AI systems are being deployed, and it is a domain in which the difference between genuine reasoning and sophisticated pattern matching is the difference between justice and injustice. Courts have already encountered cases where AI systems were used to assess recidivism risk in sentencing recommendations, and where the outputs of those systems, produced by pattern matching over historical criminal justice data that encodes decades of racially biased policing and prosecution practices, were treated as objective assessments of individual defendants' future behavior. The COMPAS system is the most documented example: researchers at ProPublica found that the system's outputs were systematically biased against Black defendants, assigning higher risk scores even when controlling for actual reoffending rates (18). The system was not intelligent. It was pattern matching, and the patterns it learned from were discriminatory patterns encoded in historical data. The people who deployed it and the judges who used its outputs treated it as if it were providing intelligent assessment, and that misclassification of pattern matching as intelligence produced documented injustice. That is not an academic consequence. It is a consequence measured in years of people's lives.&lt;/p&gt;

&lt;p&gt;The healthcare deployment of AI systems raises the same constellation of concerns at a scale that is only growing. AI-assisted diagnostic systems, clinical decision support tools, triage algorithms, and treatment recommendation engines are being deployed in hospitals and clinics around the world, and the evaluations presented in marketing materials typically measure performance on test sets drawn from the same distribution as the training data, which tells us almost nothing about how the systems perform on the actual patients who will be affected by their outputs. Patients who fall outside the demographic distribution of the training data, patients who present with atypical symptoms, patients whose conditions are rare or novel, these are exactly the patients for whom the difference between real reasoning and pattern matching matters most, and these are exactly the patients for whom well-documented AI diagnostic failures tend to be concentrated. Research has shown that AI diagnostic systems trained predominantly on images from lighter-skinned patients perform significantly worse on darker-skinned patients, not because of any intentional design choice but because skin tone correlates with diagnostic signal in ways that differ systematically across demographic groups, and pattern matching over a biased training distribution learns the biased pattern (19). A genuinely intelligent diagnostic system would reason from biological mechanisms rather than matching to training patterns, and its performance would not depend on whether the patient looks like the patients in the training set. No currently deployed system has that property.&lt;/p&gt;

&lt;p&gt;The educational technology sector is deploying AI at an even larger scale, and the consequences here are subtler but no less real. When AI writing assistants and question-answering systems are used extensively in educational settings, and students come to rely on them for the kind of thinking that education is supposed to develop, the students absorb the pattern: retrieve from the AI, not think it through yourself. This might be acceptable if the AI were actually reasoning, because then the student would at least be learning from genuine reasoning. But if the AI is pattern-matching and the student is copying the pattern-match output while bypassing the reasoning process, then the student is not learning to reason. They are learning to use a tool that produces outputs that look like the results of reasoning, which is a very different thing. Education is supposed to develop the capacity for intelligence in students. AI tools that substitute for that development rather than scaffolding it are doing the opposite of what education is for. And the teachers who deploy these tools with good intentions, wanting to save students time or to differentiate instruction, are often not informed about the difference between a tool that reasons and a tool that pattern-matches, because the marketing materials for these tools do not make that distinction clearly, and because the research literature on the difference is not easily accessible to non-specialist practitioners.&lt;/p&gt;

&lt;p&gt;The economic consequences of the intelligence gap play out where I discussed in &lt;a href="https://wiseai.dev/blogs/as-engineers-llms-should-pay-us-for-tokens-usage" rel="noopener noreferrer"&gt;As Engineers, LLMs should pay us for tokens usage&lt;/a&gt; and &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether.&lt;/a&gt;. When people and organizations are told that AI systems are intelligent, they make decisions about labor, skill development, and investment on the basis of that characterization. Jobs that require actual reasoning are restructured around AI systems that do not actually reason, leading to output quality that is worse than it would have been with a human expert but that is harder to audit because the AI-produced material is fluent and plausible enough to pass cursory review. The humans who were doing the reasoning are displaced. The AI systems do not have the intelligence to actually replace them. The organizations end up with lower quality outputs and higher reputational risk when the failures become visible, and nobody is accountable because the decision to deploy was based on benchmark numbers interpreted as intelligence claims rather than on honest assessment of what the systems could and could not actually do. This is not a hypothetical. It is already happening in software development, in legal services, in content production, in financial analysis, and in dozens of other domains where AI has been deployed on the basis of capability claims that conflated pattern matching with reasoning.&lt;/p&gt;

&lt;p&gt;The most dangerous long-term consequence of the intelligence gap is epistemic, and it is the consequence that I think about most. If a large fraction of the information that people encounter is produced by systems that pattern-match rather than reason, and if people cannot easily distinguish AI-produced content from content produced by genuine reasoning, then the epistemic environment degrades in a specific and serious way. People learn to accept fluent, confident-sounding content as a proxy for reliable content, because that has always been a reasonable heuristic when fluent, confident-sounding content was mostly produced by people who had to know what they were talking about to produce it. AI changes the base rate: now fluent, confident-sounding content can be produced in unlimited quantities by systems that know nothing in the honest sense of the word, and the traditional heuristic fails. I wrote about this specific dynamic in &lt;a href="https://wiseai.dev/blogs/llms-destroyed-the-internet-lmms-will-make-it-alive" rel="noopener noreferrer"&gt;LLMs destroyed the Internet. LMMs will make it alive.&lt;/a&gt; where I described how the mass deployment of content generation has already degraded the reliability of the web as an information environment. The intelligence gap is the root cause of that degradation, and it will not be fixed by better content moderation or by users becoming more skeptical. It will only be fixed by building systems that actually reason, so that fluent and confident can once again be reasonable proxies for reliable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where We Go From Here, and What Honest Progress Would Look Like
&lt;/h2&gt;

&lt;p&gt;I want to end this post by saying something constructive, because I am aware that I have spent a long time taking apart the current state of affairs, and taking things apart without pointing toward what a repaired version would look like is the habit of cynics, and I refuse to be a cynic even when cynicism is easy. I believe in the possibility of building genuinely intelligent machines. Not because current systems are almost there, but because I believe the scientific problems that need to be solved to build them are genuine problems that science can make progress on, and I believe the people working on those problems are making real progress, even if that progress is slower and less glamorous than the scaling-based progress that gets most of the attention. The right direction is not more scale applied to the current architecture. The right direction is fundamentally different architectures that prioritize the properties I described in the previous sections: causal generative models, systematic compositionality, calibrated uncertainty, genuine novelty generation, and something like internalized goals.&lt;/p&gt;

&lt;p&gt;Causal machine learning is one of the most important and underinvested directions in AI research. The work of Judea Pearl on causal inference and do-calculus has given the field a rigorous mathematical framework for reasoning about causation rather than correlation, and researchers are beginning to extend this framework in ways that could eventually be realized in learned systems (15). Bernhard Schölkopf's group has been developing the theory of causal representation learning, which aims to learn the causal variables and their structural relationships from observational data rather than requiring explicit experimental intervention (20). These directions are hard. They require theoretical innovation at least as much as they require computational scale. They produce results that are harder to demonstrate impressively in a short demo than language model capabilities. And they are the right direction, because they are building toward systems that can genuinely understand rather than systems that can fluently retrieve. The &lt;a href="https://github.com/wiseaidotdev/lmm" rel="noopener noreferrer"&gt;lmm project&lt;/a&gt; I described in &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept.&lt;/a&gt; is one concrete implementation that moves in this direction, with symbolic regression for equation discovery, physics simulation, and explicit causal reasoning. It is a proof of concept, not a complete system, and I say that honestly, but proof of concept matters because it proves that the alternative exists and is engineerable.&lt;/p&gt;

&lt;p&gt;Neurosymbolic AI is another direction worth watching, because it attempts to combine the pattern recognition strengths of neural networks with the structural reasoning strengths of symbolic systems. The symbolic AI tradition, which dominated the field before the deep learning wave and which was prematurely written off as obsolete when deep learning achieved its first impressive results, has real strengths in exactly the areas where neural networks are weakest: systematic compositionality, explicit causal reasoning, logical inference, and honest uncertainty representation. The neurosymbolic research agenda is trying to build hybrid architectures that get the best of both worlds, and while the engineering challenges are significant, the theoretical case for the approach is strong. Vaishak Belle and Gary Marcus have argued that neurosymbolic integration is the most promising route to systems that combine the learning capabilities of neural networks with the reasoning capabilities of symbolic systems (16). The research is real, the progress is real, and the goal is architecturally appropriate in a way that pure scaling of language models is not.&lt;/p&gt;

&lt;p&gt;Honest benchmark design is another place where the field could make genuine progress without waiting for fundamental architectural breakthroughs. If the research community committed to designing benchmarks that are genuinely contamination-proof, that test transfer rather than memorization, that test systematic compositionality rather than interpolation within the training distribution, and that include measures of calibration and uncertainty quantification alongside measures of accuracy, the field would at least have better information about how far current systems are from genuine intelligence, and that honest information is more valuable than inflated benchmark numbers in the long run. Some researchers are already moving in this direction: the BIG-Bench project, the HELM benchmark suite, and the various adversarial evaluation approaches that have been developed in recent years all represent genuine attempts to evaluate capability more honestly. These efforts deserve more support, more institutional recognition, and more attention in the press than they currently receive.&lt;/p&gt;

&lt;p&gt;For me personally, the path forward is the same thing it has always been, which is writing what I believe honestly and building what I can concretely. I believe that knowledge is not intelligence. I believe that tools extend reach but do not create understanding. I believe that genuine intelligence requires causal models of the world, systematic compositional reasoning, calibrated uncertainty, and something like goals that go beyond optimizing an external loss function. I believe that the gap between what current AI systems have and what genuine intelligence requires is large, is structural rather than merely a matter of scale, and matters enormously for every domain in which these systems are being deployed. And I believe that saying these things plainly, in simple words, is more useful than softening them into a more comfortable form, because the comfortable form has been available for years and has not moved the conversation in the direction that needs to be moved.&lt;/p&gt;

&lt;p&gt;The people who will build genuinely intelligent machines are not necessarily the people at the largest labs with the most compute. They are the people who are working on the right problems, building causal models rather than larger language models, developing theory of compositionality rather than collecting more training data, designing honest evaluations rather than finding better benchmarks to optimize for. Those people exist. Their work is real. It is getting funded less than it should be. It receives less attention than it deserves. And it is more important for the actual future of AI than almost anything currently being discussed in the mainstream AI conversation. I have enormous respect for those researchers because I know what it is like to be working on something real in an environment that rewards something flashier, and I know that the determination required to keep working on the right thing when the rewards keep flowing to the wrong thing is a form of intelligence in itself.&lt;/p&gt;

&lt;p&gt;The title of this post is a description of the actual situation, not a slogan. Every AI system that calls a function, searches a database, retrieves a document, or generates a response is doing something with knowledge and tools. None of them are doing the thing I have called intelligence in this post. That is where we are. The direction I have described is where we need to go. The gap between those two points is the most important unsolved problem in artificial intelligence, and honest acknowledgment of that gap is the necessary first step toward actually closing it. I do not know when it will be closed. I do not know whether I will be around to see it. But I know that it is worth working toward and worth being honest about, and those two things are enough to keep going.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Bender, E. M. et al., &lt;em&gt;On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?&lt;/em&gt;, &lt;a href="https://dl.acm.org/doi/10.1145/3442188.3445922" rel="noopener noreferrer"&gt;ACM FAccT 2021&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; Obermeyer, Z. et al., &lt;em&gt;Dissecting racial bias in an algorithm used to manage the health of populations&lt;/em&gt;, &lt;a href="https://www.science.org/doi/10.1126/science.aax2342" rel="noopener noreferrer"&gt;Science, 2019&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Kambhampati, S. et al., &lt;em&gt;LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2402.01817" rel="noopener noreferrer"&gt;arXiv:2402.01817&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; Strubell, E. et al., &lt;em&gt;Energy and Policy Considerations for Deep Learning in NLP&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1906.02243" rel="noopener noreferrer"&gt;arXiv:1906.02243&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Lake, B. M. et al., &lt;em&gt;Building Machines That Learn and Think Like People&lt;/em&gt;, &lt;a href="https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/article/building-machines-that-learn-and-think-like-people/A9535B1D745A0377E16C590E14B94993" rel="noopener noreferrer"&gt;Behavioral and Brain Sciences, 2017&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Dziri, N. et al., &lt;em&gt;Faith and Fate: Limits of Transformers on Compositionality&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2305.18654" rel="noopener noreferrer"&gt;arXiv:2305.18654&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; McCarthy, J. &amp;amp; Hayes, P. J., &lt;em&gt;Some Philosophical Problems from the Standpoint of Artificial Intelligence&lt;/em&gt;, &lt;a href="https://www-formal.stanford.edu/jmc/mcchay69.html" rel="noopener noreferrer"&gt;Machine Intelligence, 1969&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; Marcus, G., &lt;em&gt;Deep Learning: A Critical Appraisal&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1801.00631" rel="noopener noreferrer"&gt;arXiv:1801.00631&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-9"&gt;&lt;/span&gt;&lt;strong&gt;9.&lt;/strong&gt; Damasio, A., &lt;em&gt;Descartes' Error: Emotion, Reason, and the Human Brain&lt;/em&gt;, &lt;a href="https://www.ncbi.nlm.nih.gov/nlmcatalog/9505285" rel="noopener noreferrer"&gt;NIH National Library of Medicine, 1994&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-10"&gt;&lt;/span&gt;&lt;strong&gt;10.&lt;/strong&gt; McCoy, R. T. et al., &lt;em&gt;Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1902.01007" rel="noopener noreferrer"&gt;arXiv:1902.01007&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-11"&gt;&lt;/span&gt;&lt;strong&gt;11.&lt;/strong&gt; Lanham, T. et al., &lt;em&gt;Measuring Faithfulness in Chain-of-Thought Reasoning&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2307.13702" rel="noopener noreferrer"&gt;arXiv:2307.13702&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-12"&gt;&lt;/span&gt;&lt;strong&gt;12.&lt;/strong&gt; Cabitza, F. et al., &lt;em&gt;Unintended consequences of machine learning in medicine&lt;/em&gt;, &lt;a href="https://doi.org/10.1001/jama.2017.7797" rel="noopener noreferrer"&gt;JAMA, 2017&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-13"&gt;&lt;/span&gt;&lt;strong&gt;13.&lt;/strong&gt; Gururangan, S. et al., &lt;em&gt;Annotation Artifacts in Natural Language Inference Data&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1803.02324" rel="noopener noreferrer"&gt;arXiv:1803.02324&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-14"&gt;&lt;/span&gt;&lt;strong&gt;14.&lt;/strong&gt; Kanjee, Z. et al., &lt;em&gt;Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge&lt;/em&gt;, &lt;a href="https://doi.org/10.1001/jama.2023.8288" rel="noopener noreferrer"&gt;JAMA, 2023&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-15"&gt;&lt;/span&gt;&lt;strong&gt;15.&lt;/strong&gt; Pearl, J., &lt;em&gt;Causality: Models, Reasoning, and Inference&lt;/em&gt;, &lt;a href="https://doi.org/10.1017/CBO9780511803161" rel="noopener noreferrer"&gt;Cambridge University Press, 2009&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-16"&gt;&lt;/span&gt;&lt;strong&gt;16.&lt;/strong&gt; Belle, V. &amp;amp; Marcus, G., &lt;em&gt;The Future Is Neuro-Symbolic: Where Has It Been, and Where Is It Going?&lt;/em&gt;, &lt;a href="https://doi.org/10.1609/aaai.v40i48.42130" rel="noopener noreferrer"&gt;Proceedings of the AAAI Conference on Artificial Intelligence, 2026&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-17"&gt;&lt;/span&gt;&lt;strong&gt;17.&lt;/strong&gt; Kadavath, S. et al., &lt;em&gt;Language Models (Mostly) Know What They Know&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2207.05221" rel="noopener noreferrer"&gt;arXiv:2207.05221&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-18"&gt;&lt;/span&gt;&lt;strong&gt;18.&lt;/strong&gt; Angwin, J. et al., &lt;em&gt;Machine Bias&lt;/em&gt;, &lt;a href="https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing" rel="noopener noreferrer"&gt;ProPublica, 2016&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-19"&gt;&lt;/span&gt;&lt;strong&gt;19.&lt;/strong&gt; Groh, M. et al., &lt;em&gt;Evaluating Deep Neural Networks Trained on Clinical Images in Dermatology with the Fitzpatrick 17k Dataset&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2104.09957" rel="noopener noreferrer"&gt;arXiv:2104.09957&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-20"&gt;&lt;/span&gt;&lt;strong&gt;20.&lt;/strong&gt; Schölkopf, B. et al., &lt;em&gt;Toward Causal Representation Learning&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2102.11107" rel="noopener noreferrer"&gt;arXiv:2102.11107&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>lmm</category>
    </item>
    <item>
      <title>the penguins are already sentient. Your neural network is just a distraction.</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 18:49:59 +0000</pubDate>
      <link>https://dev.to/wiseai/the-penguins-are-already-sentient-your-neural-network-is-just-a-distraction-1kgj</link>
      <guid>https://dev.to/wiseai/the-penguins-are-already-sentient-your-neural-network-is-just-a-distraction-1kgj</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This post was originally published on &lt;a href="https://wiseai.dev/blogs/the-penguins-are-already-sentient-your-neural-network-is-just-a-distraction" rel="noopener noreferrer"&gt;the main website&lt;/a&gt; on &lt;a href="https://github.com/wiseaidotdev/blog/pull/17" rel="noopener noreferrer"&gt;Apr 18 2026&lt;/a&gt;. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account &amp;lt;3!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;I was watching a documentary a few months ago, I do not remember which one exactly because I watch a lot of them late at night when I cannot sleep, and there was a segment about emperor penguins in Antarctica, specifically about how they recognize one another's calls across a colony of thousands of birds in the middle of a blizzard. Each individual has a unique vocalization. Each partner in a mated pair learns the other's call with such precision that they can find each other in conditions where visibility is zero and the wind is loud enough to drown out almost any sound. They do this every year. The colony disperses, reassembles, and the bonds hold through conditions that would kill most mammals in hours. And I remember sitting there in the dark, watching this, and thinking: what exactly is the story we are telling ourselves about what intelligence is and where it lives? Because whatever that penguin is doing when it picks its mate's voice out of a screaming Antarctic storm is not nothing. It is something sophisticated, something persistent, something that cannot be reduced to reflex or accident or blind evolutionary wiring without doing serious violence to the word "intelligence". It is, by any honest standard, cognition. And yet the conversation about intelligence in AI circles almost never mentions it, because the conversation is entirely organized around building and scaling the kinds of structures that humans use, language, symbols, text prediction, and is almost entirely silent on the question of whether the structures that already exist in the living world around us might tell us something important about what intelligence actually is. That silence bothers me. In fact, it has been bothering me for long enough that I need to write about it, which is what this post is for.&lt;/p&gt;

&lt;p&gt;This connects to everything I have been building toward across the last several posts. In &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt;, I argued that language models are locked inside a symbolic cage, describing the world without ever touching it. In &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, I argued that equations encode reality in a way that sentences never can, that mathematical structure is the most compressed and honest representation of physical truth that humans have ever produced. In &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept. LMMs Eliminates it Altogether.&lt;/a&gt;, I argued that the training paradigm is not just ethically wrong but architecturally bankrupt, that consuming unconsented human creative work at billion-token scale is a moral choice dressed up as an engineering necessity. In &lt;a href="https://wiseai.dev/blogs/llms-destroyed-the-internet-lmms-will-make-it-alive" rel="noopener noreferrer"&gt;LLMs destroyed the Internet. LMMs will make it alive.&lt;/a&gt;, I argued that the mass deployment of language systems as content factories has hollowed out the authenticity of the web. The argument in this post is the one that goes underneath all of those, the one that asks why, despite all of this obvious evidence, we continue to treat a very narrow kind of human cognitive output as the definition of what intelligence is, while billions of minds that process reality, form bonds, navigate dynamic environments, and solve genuine survival problems are dismissed as mere biology and declared irrelevant to the AI conversation. I want to make the case that this dismissal is not just scientifically wrong. It is philosophically catastrophic, and it has led the entire field in a direction that produces impressive demos and genuine misunderstanding at the same time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What We Mean When We Say "Sentient" and Why the Answer Keeps Changing
&lt;/h2&gt;

&lt;p&gt;The word sentient comes from the Latin word for feeling, and in its original usage it simply meant capable of sensation, capable of experiencing something rather than merely processing it as a switch would process voltage. That is a humble definition, and by that definition the question of whether nonhuman animals are sentient has been essentially settled for well over a century of serious empirical research. Animals have pain receptors. Animals have nervous systems that process threatening stimuli and produce responses that look, functionally, indistinguishable from pain behavior. Animals form bonds, grieve, play, and show something that looks from the outside very much like preference, anticipation, and disappointment. But the word sentient has been quietly colonized over the centuries by a much more ambitious claim, which is that sentience requires not just feeling but the specific kind of reflective self-awareness that humans associate with language-mediated consciousness, and that anything short of the ability to say "I am experiencing this" does not fully count. That colonization is a philosophical error, and it is an error that has done significant damage both to our treatment of other animals and to our scientific understanding of mind, because it has caused us to systematically underestimate the cognitive sophistication of creatures whose intelligence takes forms that our language-centric frameworks are not well equipped to see.&lt;/p&gt;

&lt;p&gt;The Cambridge Declaration on Consciousness, signed in 2012 by a distinguished group of neuroscientists, was a rare moment of institutional clarity on this question, and I want to spend some time on it because it has not received the attention it deserves (1). The declaration stated unambiguously that nonhuman animals possess the neurological substrates that generate consciousness, and that the weight of evidence indicates that humans are not unique in possessing the biological equipment for conscious experience. This was not a fringe statement. It was signed at a symposium at the University of Cambridge and included researchers from some of the most respected institutions in cognitive neuroscience. The declaration covered mammals and birds explicitly, but also noted that the evidence for conscious experience extends to invertebrates like octopuses, which have a genuinely alien nervous system architecture that produces highly flexible, goal-directed behavior of a kind that is very difficult to explain without some form of unified experience. The scientific community has known this for over a decade, and the AI conversation has largely failed to engage with it, which is telling. If the machines we are building are supposed to be intelligent, and there are already billions of intelligent beings on this planet whose intelligence takes forms we have not yet fully understood, then the obvious question is whether we might learn something from looking more carefully at what those beings are actually doing. That question is almost never asked in the mainstream AI conversation, and I want to ask it here.&lt;/p&gt;

&lt;p&gt;I grew up in a village, as I described in &lt;a href="https://wiseai.dev/blogs/who-am-i" rel="noopener noreferrer"&gt;my first post&lt;/a&gt;, and I spent a significant portion of my childhood around animals in ways that people who grew up in cities often do not. I grew up around chickens, goats, dogs, cats, and occasionally the stray cats that wandered through the farming communities near where I lived, and I learned something from that proximity that no paper or textbook ever quite articulates clearly enough, which is that every single one of those animals had a distinct personality, a distinct set of preferences, a distinct way of engaging with the world that was not predictable from general statements about their species. The chicken that would let me pick her up was different from the chicken that would not. The dog that was afraid of strangers was different from the dog that greeted everyone at the gate. These were not abstractions. They were individuals, and their individuality was part of their daily engagement with the world around them, and that individuality is one of the markers of something that we should, if we are honest, call inner life. I am not making a mystical claim here. I am making an empirical observation about behavioral complexity, and I am noting that the same observations that would lead any careful scientist to attribute cognitive sophistication to those behaviors are observations that most people in the AI field seem entirely uninterested in, because those behaviors do not involve language models and therefore do not attract funding, prestige, or product launches.&lt;/p&gt;

&lt;p&gt;This is where I want to introduce a distinction that I think is genuinely important and that the AI field has almost entirely failed to make, which is the distinction between the form of intelligence and the fact of intelligence. What I mean is this: human linguistic intelligence takes a very specific form, one that is organized around symbolic reasoning, narrative construction, and the explicit manipulation of abstract concepts using language as the medium. That form of intelligence is real and powerful, and it has produced everything from philosophy to particle physics. But the fact of intelligence, the thing that makes cognition cognitive, is not tied to that specific form. Crows solve multi-step tool-use problems without language. Octopuses solve novel escape problems without vertebrate neural architecture. Bees perform abstract distance calculations and communicate them to their hive-mates through dance without a neocortex. Elephants remember the locations of distant water sources across decades of drought without GPS or digital memory. All of these are demonstrations of the fact of intelligence without the specific form that human-centric AI research has decided is the thing worth building. The preoccupation with language as the medium of intelligence is not a conclusion derived from careful study of what intelligence is. It is a starting assumption inherited from centuries of philosophy that privileged human cognition as the gold standard, and that assumption has baked itself into the foundations of the field in ways that most practitioners never stop to examine.&lt;/p&gt;

&lt;p&gt;The research on penguin cognition specifically is worth looking at carefully, because it has produced findings that should disturb anyone who is still operating with a dismissive picture of animal minds (2). Emperor penguins have demonstrated what researchers call episodic-like memory, meaning they do not just learn rules but form something resembling autobiographical records of specific events, specific encounters, specific outcomes, and can draw on those records to make decisions in novel situations. They demonstrate theory of mind precursors, meaning they track the informational states of other individuals in their colony in ways that suggest an understanding of what others know and do not know. They show evidence of social learning, meaning they modify their behavior based on observing the outcomes experienced by others rather than only on their own direct experience. And they do all of this in an environment that is arguably more extreme and more demanding than almost any human habitat, where the margin between success and catastrophe is measured in hours and where mistakes are fatal. That is not reflex. That is sophisticated, environmentally embedded intelligence operating at a level of generality and flexibility that any honest comparison with current AI systems must acknowledge is far more impressive than current AI systems in the domains that actually matter for survival.&lt;/p&gt;

&lt;p&gt;I want to make one more point in this section before moving on, and it is the point that I think connects all of this most directly to the argument I am building toward. When researchers study animal cognition, they consistently find that the intelligence they observe is deeply integrated with the animal's body, its environment, and its social relationships in ways that resist clean separation of computation, sensing, and action. A bird navigating by magnetic field is not running an algorithm on separately stored data. The sensing, the computing, and the acting are intertwined in a biological architecture that does not have clean hardware-software boundaries. An elephant navigating a remembered landscape is not retrieving a map from a database and then executing a pathfinding algorithm. The knowledge is distributed through behavioral, social, and physiological systems in ways that make the boundary between memory, sensation, and movement genuinely unclear. This integration, this embodied, environmental, social embeddedness of cognition, is what the AI field almost entirely ignores when it builds language models, because language models process symbols in a context-free way that bears essentially no relationship to how cognition actually works in living systems. And yet the AI field claims to be building toward something it calls intelligence, using an architecture that shares almost none of the structural features of the intelligences that already exist everywhere in the living world. That claim deserves to be examined with more skepticism than it currently receives.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Neural Network Is Not a Theory of Mind. It Is a Theory of Curve Fitting.
&lt;/h2&gt;

&lt;p&gt;I want to be fair to the field here, because fairness is something I try to practice even when I am critical, and I am about to be quite critical. Neural networks are remarkable engineering achievements. The fact that you can take a dataset, specify a loss function, run gradient descent for long enough, and produce a system that can classify images, play Go better than any human, fold proteins, and generate fluent text is genuinely astonishing. I do not want to minimize that. I have studied neural networks in college, I have built neural networks, and I have a genuine appreciation for the elegance of backpropagation and the surprising power of the universal approximation theorem. These are real achievements, and the people who developed them deserve real credit. But engineering achievement and scientific theory are different things, and the AI field has a persistent and troubling tendency to confuse them. The fact that a neural network can do something impressive does not tell you why the network can do it, and more importantly, it does not tell you whether the network is doing it in a way that resembles anything going on in biological cognition. Those are separate questions, and collapsing them has produced some of the worst thinking in the field.&lt;/p&gt;

&lt;p&gt;A neural network, at its mathematical core, is a parameterized function that maps inputs to outputs. The training process finds parameter values that minimize a loss function on a training dataset. The result is a function that generalizes, sometimes very well, to inputs outside the training set. That is it. That is the whole mechanism. Everything else, the emergent capabilities, the apparent reasoning, the fluent language generation, is an output property of that mechanism operating on very large amounts of data with very many parameters. The mechanism itself does not have beliefs, it does not have goals in any rich sense, it does not have a model of the world, and it does not have anything resembling the subjective experience that is the hallmark of sentience in the biological systems I described in the previous section. What it has is a very flexible statistical model of the patterns in its training distribution, and that statistical model can produce outputs that look, from the outside, like belief, goal-directedness, world modeling, and experience. The appearance is not the thing. The map is not the territory. And the AI field has been spending twenty years and several hundred billion dollars studying the map while largely ignoring the territory, and the territory is what the penguins live in.&lt;/p&gt;

&lt;p&gt;The specific claim I want to make, and I want to make it precisely so that it can be evaluated rather than vaguely so that it sounds impressive, is this: neural networks are optimized for input-output mapping, and optimizing for input-output mapping does not, in general, produce the internal representations that characterize biological cognition. Biological cognition is not primarily organized around input-output mapping. It is organized around building, maintaining, and updating models of the world that can be used for prediction, planning, and action across a wide variety of tasks that the organism has never encountered before. The difference between a system optimized for input-output mapping and a system with a genuine world model is the same as the difference between a lookup table and a physics engine: the lookup table can give you the right answer for inputs it has seen, but the physics engine can give you the right answer for inputs it has never seen, because the physics engine contains the actual structure that generates the answers rather than just the answers themselves. I argued in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt; that equations encode mechanisms rather than surfaces, and that distinction maps directly onto the distinction between neural networks and the kind of world models that biological cognition actually relies on. A penguin navigating a blizzard is running a world model, not a lookup table, and the world model is what makes the navigation work.&lt;/p&gt;

&lt;p&gt;This is not a new observation. The philosopher Jerry Fodor made a version of this argument in his book "The Modularity of Mind" in the early 1980s, noting that the kinds of central cognitive processes that make human intelligence general and flexible are precisely the kinds of processes that computational systems of the behaviorist-inspired variety have the most trouble capturing (3). The argument has been made more recently and more specifically by researchers like Gary Marcus and by the whole school of thought around compositional generalization, which asks whether neural networks can learn to apply rules to novel combinations of familiar inputs rather than just pattern-matching to training examples (4). The evidence is mixed, and the honest version is that standard neural networks do not compositionally generalize in the way that biological cognition does, and that this is a structural property of the architecture rather than something that will be fixed by training on more data. The experiments are clear, the replication rate is high, and the implication is one that the field has consistently found ways to avoid drawing, which is that the neural network architecture as currently practiced is not modeling what biological cognition actually does, even in the relatively simple case of compositional rule application. If it cannot do that well, the claim that it is approaching general intelligence deserves serious scrutiny.&lt;/p&gt;

&lt;p&gt;I also want to say something about what I call the scale fallacy, which is the reasoning that says: current neural networks fail at X, but if we scale them up with more data and more parameters, they will eventually succeed at X. This reasoning is sometimes correct and sometimes catastrophically wrong, and the field has an embarrassing tendency to apply it indiscriminately without asking whether the failure at X is a quantitative limitation or a qualitative one. If X is "generate more fluent text," then scaling probably helps, because fluency in text generation is a quantitative property that more data and more parameters can plausibly improve. But if X is "form a genuine world model that supports causal reasoning across novel domains," then scaling alone does not help, because forming a world model is not the objective that the network is being optimized for. You can train a curve fitter on infinite data and you will still have a curve fitter. You will never get a physics engine from a curve fitter by training it longer, because the objective is wrong. The objective optimizes for one thing, and the thing you want is a different thing, and no amount of the first thing gives you the second thing. This is not a pessimistic claim about the future of AI. It is a specific claim about the relationship between objectives and outcomes, and it is a claim that the field's most honest researchers have been making for years in papers that attract far fewer citations than the scaling papers because they deliver less comfortable news (5).&lt;/p&gt;

&lt;p&gt;The penguins are relevant here in a way that I want to be explicit about, because I am not just using them as an emotional hook. The penguin finding its mate in a blizzard is demonstrating something that current neural networks, despite their scale and sophistication, would have extreme difficulty matching in any meaningful sense. It is not that the computation exceeds the network's capacity. It is that the kind of computation being done is structurally different from what neural networks do. The penguin is running a real-time, dynamic, embodied recognition system that integrates acoustic pattern matching with spatial navigation with social memory with motivational state in a way that is seamlessly unified and operates with extreme reliability in conditions that would challenge any engineered system. The neural network, given a training dataset of penguin calls and a test set of noisy versions, could certainly learn to classify calls, and it would do so by finding statistical features that discriminate between classes in the training distribution. That is useful. But the penguin is not running a classifier. It is running a survival system, and the difference between a classifier and a survival system is not the magnitude of the computation. It is the kind of computation. It is the fact that the penguin's recognition system is integrated with everything else the penguin knows and needs and wants, in a way that makes the recognition not just accurate but alive, not just functionally correct but embedded in a continuous engagement with reality that the neural network, by its architecture, cannot replicate.&lt;/p&gt;

&lt;p&gt;Let me also make the philosophical point explicitly rather than letting it hide inside the technical argument, because the technical argument can always be dismissed as a practical limitation that will eventually be overcome, while the philosophical point cuts deeper. The neural network does not have anything at stake when it classifies an input. It has no survival interest in the outcome. It has no mate to find. It has no colony to return to. It has no past that informs its present or future that it is trying to secure. It is operating a function with no inside to it, no orientation toward the world, no stake in anything. Whether the classification is right or wrong is a matter of the loss value on a metric, not a matter of survival or death or reunion or loss. That difference, the difference between a system with something at stake and a system with nothing at stake, is not an engineering detail that can be fixed by adjusting the hyperparameters. It is the difference between sentience and computation, and I am not prepared to believe that the gap can be closed by gradient descent alone, any more than I am prepared to believe that a map can be turned into a territory by making it more accurate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Embodied Intelligence Looks Like When You Actually Look at It
&lt;/h2&gt;

&lt;p&gt;The AI field talks about embodied intelligence a lot, mostly in the context of robotics, and mostly in a way that reduces "embodiment" to the fact that the robot has sensors and actuators connected to a neural network. That is an impoverished definition of embodiment, and I want to spend some time explaining what embodiment actually means in the context of real cognition, because the real version is much more interesting and much more instructive for anyone trying to build systems that genuinely understand the world. Real embodiment is not the fact that a system has inputs and outputs from the environment. It is the fact that the system's knowledge, its memory, its representations of the world, are organized by and inseparable from its physical capabilities, its evolutionary history, and its ongoing engagement with a specific kind of environment. A bat's echolocation system does not just give the bat access to acoustic information. It gives the bat a bat-shaped understanding of the world, organized around the specific capabilities and needs of a bat body in a bat environment. The bat's acoustic world model is not separable from the bat's life, and that inseparability is not a limitation. It is the source of the system's power, because it means the model is precisely tuned to the situation in which the bat actually operates.&lt;/p&gt;

&lt;p&gt;Research on animal navigation provides some of the most compelling evidence for this kind of embodied, integrated intelligence, and I want to spend some time on it because it is directly relevant to the argument I am making about what intelligence actually is when you look at it carefully. Clark's nutcrackers are birds that cache tens of thousands of seeds in thousands of locations across a landscape and then retrieve them months later with an accuracy that is astonishing by any standard (6). They do this without GPS, without a written map, without language, and without anything that resembles the cognitive tools that human chauvinism would suggest are necessary for sophisticated spatial reasoning. The hippocampal volume of seed-caching birds is proportionally larger than in non-caching birds, which suggests that the brain structures used for spatial memory are specifically adapted to this cognitive demand, meaning that the bird's brain architecture is shaped by and tuned to the specific cognitive challenges its lifestyle presents. That is embodied intelligence in the deep sense: the architecture of the system is structured by the architecture of the problem it evolved to solve, not by a general-purpose optimization scheme applied to a training distribution. Understanding that distinction matters enormously for anyone who is seriously trying to understand what intelligence is, as opposed to what impressive-looking outputs intelligence can produce.&lt;/p&gt;

&lt;p&gt;Honeybee cognition is another area that should disturb anyone operating with confident assumptions about the cognitive prerequisites for sophisticated behavior (7). Bees perform the distance-transformed waggle dance to communicate the direction and distance of a food source to their hive-mates, accounting for the angle of the sun, the time of day, the distance, and even the quality of the source on a numerical scale. This is not a simple signal. It is an abstract spatial encoding that other bees can decode and use to navigate to a location they have never visited, using information they received through a physical performance by another bee. That is a form of symbolic communication, it uses a learned code, it encodes abstract spatial relationships, and it works even when the communicating bee is indoors and cannot see the sun, meaning the dance references a representation of the sun's position rather than the sun's actual position. If I described that capability in a neural network architecture, people would call it a breakthrough. In a bee, people call it "just instinct," and the dismissal is so automatic and so culturally comfortable that most people who use it have never stopped to ask what exactly they mean. I know what instinct means. I have read the papers. Instinct means neurologically determined behavior. So does reading, for most people below a certain age. The distinction between instinct and intelligence that we think we are drawing when we say "just instinct" is largely a distinction between familiar and unfamiliar computational substrates, and that is not a meaningful distinction for anyone trying to understand what cognition actually is.&lt;/p&gt;

&lt;p&gt;I want to connect the embodied intelligence argument directly to the &lt;a href="https://github.com/wiseaidotdev/lmm" rel="noopener noreferrer"&gt;lmm project&lt;/a&gt; here, because the lmm architecture is my attempt to build toward this kind of embodied, grounded intelligence in a way that does not depend on the training paradigm I critiqued in my last post. The lmm system includes a perception layer that converts raw bytes and sensor streams into normalized tensors, a physics simulation layer that models dynamic systems using actual differential equations, a causal reasoning layer that maintains explicit structural causal models of the dependencies between variables, and a consciousness loop that ties these together by running a continuous cycle of perceiving, encoding, predicting, and acting. This architecture is closer in spirit to embodied cognition than to the standard language model architecture, not because it is biological, but because it is organized around engaging with the structure of physical reality rather than around predicting the next token in a text sequence. When you run &lt;code&gt;lmm consciousness --lookahead 5&lt;/code&gt;, the system performs one tick of this full loop: it takes raw input, converts it to a tensor, runs a world model prediction, evaluates the prediction against the actual state, and plans an action based on the discrepancy. The output is a state vector and a mean prediction error, both of which are observable, verifiable, and grounded in the structure of the input rather than in the statistical patterns of a training corpus. That is not biological intelligence. But it is a step toward the right kind of intelligence, in a direction that the training-based paradigm cannot go, because it is organized around the world rather than around text about the world.&lt;/p&gt;

&lt;p&gt;The field calculus capability of lmm is also worth thinking about in this context, because it represents a form of spatial and structural reasoning that is closer to what embodied cognition actually does than anything in a standard language model. When you run &lt;code&gt;lmm field --size 8 --operation gradient&lt;/code&gt;, the system computes the gradient of a scalar field using central differences, which is a numerical implementation of the mathematical operation that describes how quantities change across space. That operation is at the heart of everything from physical simulation to navigation to the neural population codes that actual animal brains use to represent space. A gradient is not a piece of text. It is a structure, a mathematical object that lives in the geometry of a field, and computing it correctly requires genuine engagement with that structure rather than statistical pattern matching against descriptions of gradients from a training corpus. The result, something like &lt;code&gt;[1.0, 2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 13.0]&lt;/code&gt; for the gradient of &lt;code&gt;x²&lt;/code&gt;, is a number sequence that can be verified against the exact analytical derivative, which is &lt;code&gt;2x&lt;/code&gt;. That verification is possible because the computation is transparent and the ground truth is accessible. No language model can offer that kind of verification, because language model outputs are not computations in the relevant sense. They are predictions of token sequences that describe computations, which is a different thing entirely. The embodied cognition of the penguin navigates reality. The lmm field operator computes reality. The language model narrates reality. Those are three structurally different activities, and only two of them constitute genuine engagement with the world.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Distraction Is Working. Let Me Show You How.
&lt;/h2&gt;

&lt;p&gt;I want to talk about attention, specifically about where the field's attention is pointed and what the costs of that pointing are. The AI field right now is experiencing what I can only describe as a collective hallucination about what it is building and how important it is. The rhetoric around large language models has reached a level of confidence and self-congratulation that is genuinely strange to anyone who has read the history of AI carefully, because the history of AI is a history of premature confidence followed by painful corrections, and the confidence is currently at levels I have not seen since the 1960s symbolic AI days when some of the most brilliant people in the world confidently predicted that general intelligence was ten to twenty years away (8). Those predictions were wrong, and the reason they were wrong was not that the researchers were stupid. They were not. The reason they were wrong was that they were distracted by the impressive capabilities of the systems they were building from asking honest questions about whether those systems were actually doing what the researchers thought they were doing. The distraction is happening again, at much larger scale, with much more money behind it, and with much more riding on the outcome.&lt;/p&gt;

&lt;p&gt;The distraction takes a specific form that I want to name clearly so that it can be recognized. It goes like this: a new capability appears in a scaled-up language model, the capability was not explicitly trained for, and people call this "emergence". The emergence is treated as evidence that scaling is a path to general intelligence, because look, the model can do something new that nobody taught it to do. The problem with this reasoning is that it conflates statistical emergence with genuine cognitive emergence. When a statistical model trained on text discovers that certain text patterns cluster together in ways that allow it to produce what looks like arithmetic, that is statistical emergence, the discovery of surface patterns that co-occur with arithmetic in text. It is not the same as genuinely learning the rules of arithmetic, and the empirical evidence shows exactly this: language models perform much better on arithmetic problems that appear frequently in their training distribution than on arithmetically equivalent problems presented in forms that appear rarely, which is exactly what you would expect from a system that learned statistical patterns about arithmetic rather than arithmetic itself (9). That is the distraction working in real time. The capability looks real from the outside, and the appearance creates confidence, and the confidence attracts resources, and the resources deepen the bet, and the whole cycle continues without anyone pausing to ask whether the appearance is the thing or just a very good imitation of the thing.&lt;/p&gt;

&lt;p&gt;The cost of the distraction is not just computational or financial, although those costs are enormous. The cost that I find most troubling is the cost to our understanding of intelligence, because every year spent building and scaling systems that imitate the surface of intelligence without engaging its deep structure is a year not spent trying to understand what intelligence actually is. The research program on animal cognition that I described in the previous sections is not well funded, not prestigious, not connected to major product launches, and not the subject of breathless press coverage. It is slow, careful, empirical work done by researchers who care more about understanding than about demos, and it is being systematically outcompeted for attention and resources by a paradigm that is very good at producing things that look impressive and very reluctant to ask whether they are genuinely intelligent. I described in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt; what it feels like to work on something real and have it invisible because it is not dressed up in the right narrative, and the researchers studying animal cognition live that experience constantly. Their subjects are genuinely intelligent in ways that matter. Their findings are genuinely important for anyone who cares about the nature of mind. And they are being crowded out by a conversation that has decided in advance what intelligence looks like and is building systems to match that predetermined picture.&lt;/p&gt;

&lt;p&gt;I want to make a specific claim about what the distraction has cost in terms of scientific progress, because vague claims about opportunity cost are easy to dismiss and specific claims are harder. The specific claim is this: if the resources invested in scaling language models over the last decade had been partially redirected toward understanding the computational principles of animal cognition, toward building formal mathematical models of navigation, social learning, episodic memory, and multi-sensory integration in biological systems, and toward implementing those models in engineered systems that could be tested and refined, we would today have a much clearer scientific picture of what intelligence actually is and a much more principled basis for building artificial systems that instantiate it. That is a claim about what would have happened under a different allocation of resources, and it is inherently speculative, but it is no more speculative than the claim that scaling language models will eventually lead to general intelligence. The difference is that the second claim is the one being funded and celebrated, while the first claim is the one that has the biological evidence on its side. The penguins, the bees, the nutcrackers, the octopuses: these are existence proofs of a kind of intelligence that does not pass through text, and existence proofs are the most powerful kind of evidence in any scientific argument, because they show that the thing is possible rather than merely arguing that it might be.&lt;/p&gt;

&lt;p&gt;The way the distraction sustains itself is worth understanding, because it is not sustained by stupidity or bad faith on the part of individual researchers, most of whom are genuinely curious and genuinely talented. It is sustained by an incentive structure that rewards impressive demos over careful science, market share over mechanistic understanding, and confidence over honesty. I wrote about this in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt;, where I described how the technology industry's incentive structure consistently produces outcomes that are good for the organizations at the center of the field and bad for the people at its margins, and the same dynamic applies to the scientific question of what intelligence is. The organizations at the center of the AI field have enormous incentives to believe that scaling language models is the path to general intelligence, because they have bet enormous sums on that belief, and anyone who challenges it from within faces strong institutional pressure to continue in the current direction. The challenge has to come from outside the current incentive structure, which is exactly what the work on animal cognition represents, and why I think it deserves much more attention than the AI conversation is currently giving it.&lt;/p&gt;

&lt;p&gt;The distraction also has a racial and geographic dimension that I want to name because I think it is usually left out of the conversation, and leaving it out makes the picture incomplete. The AI field is geographically concentrated, demographically uniform, and culturally shaped by a specific tradition that privileges certain kinds of cognitive performance, specifically the kinds associated with academic achievement within Western educational systems, as the gold standard of intelligence. That tradition has a long history of underestimating the intelligence of people who succeed through other means, through spatial navigation, through social intelligence, through craft and embodied skill, through forms of reasoning that are not well captured by standardized tests or academic publications. The same cultural predisposition that produces a dismissive attitude toward non-Western forms of intelligence also produces a dismissive attitude toward non-human forms of intelligence, and both dismissals serve the same function of maintaining a comfortable hierarchy with the AI researcher at the top. I am not saying this to accuse anyone of racism or malice. I am saying it because the cultural assumptions that a research community brings to its work shape what the community sees and what it misses, and a field that has systematically missed the intelligence of billions of non-human animals while spending hundreds of billions of dollars on systems that imitate a very specific kind of human cognitive output is a field that would benefit from examining its assumptions more carefully.&lt;/p&gt;

&lt;h2&gt;
  
  
  lmm as the Alternative Architecture: Building With the World Instead of About It
&lt;/h2&gt;

&lt;p&gt;I want to spend this section connecting the philosophical argument I have been making to the specific engineering choices in the lmm project, because I think the connection is real and important rather than just rhetorical. The lmm project is not presented as an imitation of animal cognition. It is presented as an alternative to the language model paradigm, built around mathematical structure and physical simulation rather than around text prediction and gradient descent. But the reason I brought animal cognition into this post is that animal cognition, honestly examined, points toward exactly the kind of architecture that lmm is trying to instantiate: an architecture organized around world modeling, causal reasoning, and embodied engagement with physical reality rather than around surface pattern matching in a high-dimensional statistical space. The connection between the penguin and the lmm is not that the lmm can navigate a blizzard. It is that both the penguin and the lmm are organized around contact with reality rather than around descriptions of reality, and that organizational principle is the thing that distinguishes intelligence from imitation.&lt;/p&gt;

&lt;p&gt;The symbolic regression capability of lmm is the most direct implementation of the "learn the mechanism, not the surface" principle that I keep arguing for. When you run &lt;code&gt;lmm discover --iterations 200&lt;/code&gt;, the system takes a set of observations, in the default case data points from a linear process, and runs a genetic programming search to find the symbolic equation that best explains those observations. The output is something like &lt;code&gt;(x + (1.002465056833142 + x))&lt;/code&gt;, which is the system's discovered approximation of the underlying law &lt;code&gt;2x + 1&lt;/code&gt;. The method used to find this equation is genetic programming: a population of candidate expressions is initialized, each one is evaluated for how well it fits the data, better-fitting expressions are selected and recombined and mutated, and the process iterates until either the fitness converges or the iteration budget is exhausted. This is superficially similar to neural network training in that both involve iterative optimization on data, but the difference is fundamental: the neural network optimizes parameters in a fixed architecture toward minimizing a loss, producing a black-box function; the genetic programming optimizes the structure of a symbolic expression toward explaining the data, producing a human-readable equation. The output of one is opaque. The output of the other is transparent. And transparency, as I argued in &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept&lt;/a&gt;, is not a cosmetic feature. It is the property that makes outcomes verifiable, and verification is the foundation of honest science.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Discover the governing equation from synthetic linear data&lt;/span&gt;
lmm discover &lt;span class="nt"&gt;--iterations&lt;/span&gt; 200
&lt;span class="c"&gt;# Output: Discovered equation: (x + (1.002465056833142 + x))&lt;/span&gt;
&lt;span class="c"&gt;# This approximates 2x + 1, discoverable in ~200 GP iterations&lt;/span&gt;
&lt;span class="c"&gt;# No training corpus, no gradient descent, no unconsented data&lt;/span&gt;

&lt;span class="c"&gt;# For more complex patterns, more iterations help&lt;/span&gt;
lmm discover &lt;span class="nt"&gt;--iterations&lt;/span&gt; 500 &lt;span class="nt"&gt;--data-path&lt;/span&gt; ./my_observations.csv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The physics simulation capability directly instantiates the idea of a world model, which is what I argued animal cognition relies on rather than a lookup table of trained associations. When you run &lt;code&gt;lmm physics --model lorenz --steps 500 --step-size 0.01&lt;/code&gt;, the system integrates the Lorenz chaotic attractor equations forward in time using the Runge-Kutta fourth-order method, producing the exact trajectory of a chaotic dynamical system from its governing equations. That trajectory is not predicted from training data. It is computed from the differential equations that describe the Lorenz system, and those equations are not learned from examples. They are specified from physical first principles and then used to generate predictions. The difference between generating predictions from learned statistical patterns and computing predictions from explicit physical laws is the difference between the curve fitter and the physics engine that I described earlier in this post. The curve fitter can only interpolate within what it has seen. The physics engine can extrapolate to states it has never simulated, because it has the generating structure, not just examples of outputs. This extrapolation capability is the thing that animal cognition has that enables animals to navigate novel environments, solve novel problems, and make decisions in situations that have no precedent in their experience.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# lmm uses known physical laws, not training data, to predict the future&lt;/span&gt;
lmm physics &lt;span class="nt"&gt;--model&lt;/span&gt; lorenz &lt;span class="nt"&gt;--steps&lt;/span&gt; 500 &lt;span class="nt"&gt;--step-size&lt;/span&gt; 0.01
&lt;span class="c"&gt;# =&amp;gt; Lorenz: 500 steps. Final xyz: [-8.900..., -7.413..., 29.311...]&lt;/span&gt;

lmm physics &lt;span class="nt"&gt;--model&lt;/span&gt; sir &lt;span class="nt"&gt;--steps&lt;/span&gt; 1000 &lt;span class="nt"&gt;--step-size&lt;/span&gt; 0.5
&lt;span class="c"&gt;# =&amp;gt; SIR: 1000 steps. Final [S,I,R]: [58.797..., 7.649e-15, 941.202...]&lt;/span&gt;

lmm physics &lt;span class="nt"&gt;--model&lt;/span&gt; pendulum &lt;span class="nt"&gt;--steps&lt;/span&gt; 300 &lt;span class="nt"&gt;--step-size&lt;/span&gt; 0.005

&lt;span class="c"&gt;# The equations are the world model. No training. No hallucination.&lt;/span&gt;
&lt;span class="c"&gt;# If the equations are right, the predictions are right. Always. Verifiably.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The causal reasoning layer of lmm is perhaps the most important piece for the argument I have been making about the gap between animal cognition and neural network processing, because causality is the specific kind of reasoning that animal survival most requires and that neural network pattern matching most systematically fails to provide. When a penguin learns that a specific behavior leads to food in some contexts and does not in others, it is not just learning a stimulus-response association. It is building a causal model that includes the conditions under which the association holds, and that model allows the penguin to behave appropriately in novel situations that share the relevant causal structure but not the specific sensory context. Neural networks trained to classify or predict do not, in general, learn causal models. They learn correlational patterns, and correlational patterns break down in exactly the novel situations that causal models handle correctly. The lmm causal module builds an explicit structural causal model and supports the &lt;code&gt;do(X=v)&lt;/code&gt; intervention operator from Judea Pearl's do-calculus, which allows you to ask not just "what is correlated with what" but "what would happen if I changed this variable". That is the question that causal reasoners ask, that animal cognition answers, and that neural networks are systematically unable to address without explicit causal structure.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Build a causal model and reason about interventions&lt;/span&gt;
&lt;span class="c"&gt;# The SCM is y = 2*x, z = y + 1&lt;/span&gt;
lmm causal &lt;span class="nt"&gt;--intervene-node&lt;/span&gt; x &lt;span class="nt"&gt;--intervene-value&lt;/span&gt; 10.0
&lt;span class="c"&gt;# Before intervention: x=Some(3.0), y=Some(6.0), z=Some(7.0)&lt;/span&gt;
&lt;span class="c"&gt;# After do(x=10): x=Some(10.0), y=Some(20.0), z=Some(21.0)&lt;/span&gt;

&lt;span class="c"&gt;# A language model given the same question would produce a&lt;/span&gt;
&lt;span class="c"&gt;# plausible-sounding description of the math from training patterns.&lt;/span&gt;
&lt;span class="c"&gt;# lmm computes the actual causal propagation from the explicit structure.&lt;/span&gt;
&lt;span class="c"&gt;# The penguin navigating food sources makes exactly this kind of inference.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I want to be transparent about the obvious gap, which is that lmm currently works with mathematical functions and explicit causal graphs, while animal cognition operates on raw sensory streams in a messy, partially observable, physically complex world. The lmm perception layer is the beginning of bridging that gap: it accepts raw byte streams and converts them to normalized tensors, and the consciousness loop is designed to be the integration point where perception, prediction, and action come together. But the current implementation is a proof of concept for the architecture, not a finished system, and the distance between a proof of concept and the navigation capability of an emperor penguin is significant. I am not claiming otherwise. The argument I am making is not that lmm already matches animal cognition. The argument is that lmm is organized around the right principles, toward physical structure rather than statistical text, toward explicit causal models rather than correlational patterns, toward equation discovery rather than parameter fitting, and that organizing around the right principles is the prerequisite for making genuine progress toward the kind of intelligence that the penguins already have. The specific implementation will improve. The principles are the thing that matters, and the principles are right. One of those principles, the one I want to spend the next section on because it is the most surprising and the most directly comparable to what neural networks do, is what I call stochastic determinism, and it is the principle that lets lmm produce unique, varied, natural-sounding text on every run without ever being trained on a single human-authored sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stochastic Determinism: The Honest Equivalent of Neural Network Output
&lt;/h2&gt;

&lt;p&gt;There is a specific objection to the lmm approach to text generation that I hear most often from people who have worked with language models, and I want to address it here because it is a reasonable objection that deserves a real answer rather than a dismissal. The objection goes like this: a neural language model, when generating text, introduces temperature-controlled randomness into its sampling process, which is what makes each output different from the last and what gives generated text its feeling of natural variety, and without a similar mechanism any training-free system will produce identical, robotic, repetitive output every time, which will be immediately recognizable as non-natural and therefore useless for practical applications. That objection is correct about neural language models. Temperature sampling is the mechanism that gives them variety, and without variety the output is mechanical and immediately distinguishable from human writing. But the objection assumes that variety requires randomness over a learned probability distribution, which is exactly the assumption that lmm challenges. The lmm project implements what I am calling stochastic determinism, a two-layer architecture in which the underlying generation is completely deterministic, traceable, and mathematically grounded, and the surface variation is introduced by a separate, explicit, auditable synonym replacement layer rather than by sampling from an opaque learned distribution.&lt;/p&gt;

&lt;p&gt;The way this works in practice is worth explaining carefully, because the architecture is more elegant than it might sound. When you run &lt;code&gt;lmm predict --text "Wise AI built the first LMM"&lt;/code&gt;, the system runs genetic programming on the context words to discover a trajectory equation that describes how word identity changes with position, discovers a rhythm equation describing how word length evolves, and uses those equations together with a curated vocabulary mapping and a syntactic Subject-Verb-Object sentence structure to produce a deterministic text continuation. That continuation is the same every time for the same input, which is a feature rather than a bug, because it means the system's reasoning is reproducible and auditable. But when you add the &lt;code&gt;--stochastic&lt;/code&gt; flag, a second layer activates: the &lt;code&gt;StochasticEnhancer&lt;/code&gt;, which draws from a built-in synonym bank to replace eligible words in the deterministic output with contextually appropriate alternatives, at a replacement rate controlled by the &lt;code&gt;--probability&lt;/code&gt; parameter. The mathematical structure of the sentence, the equation-derived skeleton, stays completely fixed. The specific surface words vary across runs according to the synonym selections. The result is an output that is unique on every run, reads with natural variety, and yet is grounded in a determinate mathematical computation that can be inspected, reproduced, and verified at any time by disabling the stochastic layer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Deterministic base output - same every time, fully reproducible&lt;/span&gt;
lmm predict &lt;span class="nt"&gt;--text&lt;/span&gt; &lt;span class="s2"&gt;"Wise AI built the first LMM"&lt;/span&gt;
&lt;span class="c"&gt;# Output: "Wise AI built the first LMM in the true law often long time"&lt;/span&gt;
&lt;span class="c"&gt;#         and a open path of an old scope is the solid order.&lt;/span&gt;

&lt;span class="c"&gt;# Stochastic mode - unique output every run, same mathematical skeleton&lt;/span&gt;
lmm predict &lt;span class="nt"&gt;--text&lt;/span&gt; &lt;span class="s2"&gt;"Wise AI built the first LMM"&lt;/span&gt; &lt;span class="nt"&gt;--stochastic&lt;/span&gt; &lt;span class="nt"&gt;--probability&lt;/span&gt; 0.4
&lt;span class="c"&gt;# Run 1: "Wise AI built the first LMM in the genuine principle often extended time"&lt;/span&gt;
&lt;span class="c"&gt;#        and a accessible route of an ancient domain is the firm sequence.&lt;/span&gt;
&lt;span class="c"&gt;# Run 2: "Wise AI built the first LMM in the real law frequently long duration"&lt;/span&gt;
&lt;span class="c"&gt;#        and a open trajectory of an old range is the stable order.&lt;/span&gt;

&lt;span class="c"&gt;# Single sentence generation with stochastic variation&lt;/span&gt;
lmm sentence &lt;span class="nt"&gt;--text&lt;/span&gt; &lt;span class="s2"&gt;"Mathematics is the language of the universe"&lt;/span&gt; &lt;span class="nt"&gt;--stochastic&lt;/span&gt;
&lt;span class="c"&gt;# Run 1: Cognition enables the dynamic significance of the world.&lt;/span&gt;
&lt;span class="c"&gt;# Run 2: Analysis facilitates the continuous meaning of the cosmos.&lt;/span&gt;
&lt;span class="c"&gt;# The equation-derived structure is identical. Only synonyms differ.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The contrast with neural language model temperature sampling is in the specific thing that is being randomized, and that contrast is the whole point. When a language model applies temperature to its softmax distribution and samples a token, the randomness is operating over a learned probability distribution across the entire vocabulary, and the distribution itself is an opaque artifact of the training process. You cannot inspect the distribution and understand why certain tokens were assigned certain probabilities, because those probabilities are the accumulated result of gradient descent over billions of training examples, and the individual contributions of those examples have been averaged and compressed beyond human comprehension. The randomness is real but its source is invisible. When lmm's &lt;code&gt;StochasticEnhancer&lt;/code&gt; replaces a word with a synonym, the randomness is operating over a curated, human-readable synonym bank that maps each eligible word to a set of alternatives with known semantic relationships. You can inspect the synonym bank, understand exactly which words are candidates for replacement, understand what semantic category each replacement belongs to, and reproduce any specific output by seeding the random number generator with a fixed value. The randomness is real but its source is completely visible, completely auditable, and completely separable from the deterministic mathematical computation that generates the underlying structure.&lt;/p&gt;

&lt;p&gt;This distinction matters for reasons that go beyond technical transparency, and I want to make those reasons explicit because they connect directly to the moral argument I have been developing across several posts. The key insight is that lmm separates what I would call the epistemic layer from the aesthetic layer of text generation. The epistemic layer, the part that determines the meaning, the structure, the mathematical relationships encoded in the output, is completely deterministic and completely auditable. The aesthetic layer, the part that determines the specific surface words used to express those relationships, introduces controlled randomness to produce natural variety. This separation means that the epistemic content of lmm's output is never contaminated by its aesthetic variability: you can always recover the deterministic spine of any stochastic output by disabling the stochastic layer, and the spine is exactly what you can verify, examine, and trust. Neural language models do not have this separation. For them, the epistemic and aesthetic layers are entangled in the same learned distribution, which is why it is so difficult to verify any specific claim a language model makes. You cannot turn off the temperature and ask "what does the model actually believe about this," because the model does not have beliefs in a form that is separable from its sampling behavior. The model's output is always a sample from an opaque distribution, and the distribution is the model, and the model is an artifact of an uncheckable training process.&lt;/p&gt;

&lt;p&gt;I also want to address why I call this stochastic determinism rather than just controlled randomness or probabilistic output, because the naming matters for understanding what is philosophically new here. Stochastic determinism means that the system is deterministic at the level of its reasoning and stochastic at the level of its expression. The reasoning, the equation discovery, the physics simulation, the causal propagation, all of these are fully deterministic computations that produce the same result for the same input every time. The expression, the specific words chosen to represent the output of those computations, varies according to an explicit and auditable probability structure. This is actually a much better model of how expert human communication works than temperature sampling over a learned distribution is. When I write these posts, the ideas I am trying to express are determined by my thinking, which is grounded in specific arguments, specific evidence, specific logical relationships that I have worked out carefully. The specific words I choose to express those ideas vary across drafts and revisions, and that variation is what makes my writing feel like my writing rather than like a lookup table. The ideas are deterministic. The phrasing is stochastic. Lmm implements exactly this architecture: deterministic ideas grounded in mathematics, stochastic phrasing grounded in a synonym structure that can be inspected and verified.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Paragraph generation - stochastic determinism at scale&lt;/span&gt;
&lt;span class="c"&gt;# Same seed, same mathematical structure, different surface words each run&lt;/span&gt;
lmm paragraph &lt;span class="nt"&gt;--text&lt;/span&gt; &lt;span class="s2"&gt;"Equations reveal hidden truths about nature"&lt;/span&gt; &lt;span class="nt"&gt;--sentences&lt;/span&gt; 6 &lt;span class="nt"&gt;--stochastic&lt;/span&gt;
&lt;span class="c"&gt;# Run 1: Simulation manifests the continuous symmetry of the truths.&lt;/span&gt;
&lt;span class="c"&gt;#        The symmetric wavelength connects infinity. Entropy remains...&lt;/span&gt;
&lt;span class="c"&gt;# Run 2: Modeling reveals the persistent balance of the realities.&lt;/span&gt;
&lt;span class="c"&gt;#        The balanced frequency links unbounded space. Randomness sustains...&lt;/span&gt;

&lt;span class="c"&gt;# The full generation pipeline without any training&lt;/span&gt;
lmm encode &lt;span class="nt"&gt;--text&lt;/span&gt; &lt;span class="s2"&gt;"The mathematical universe"&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt;   &lt;span class="c"&gt;# encode to equation&lt;/span&gt;
lmm decode &lt;span class="nt"&gt;--equation&lt;/span&gt; &lt;span class="s2"&gt;"..."&lt;/span&gt;                     &lt;span class="c"&gt;# decode back perfectly&lt;/span&gt;
lmm predict &lt;span class="nt"&gt;--text&lt;/span&gt; &lt;span class="s2"&gt;"..."&lt;/span&gt; &lt;span class="nt"&gt;--stochastic&lt;/span&gt;           &lt;span class="c"&gt;# continue stochastically&lt;/span&gt;
lmm ask &lt;span class="nt"&gt;--prompt&lt;/span&gt; &lt;span class="s2"&gt;"What is LMM?"&lt;/span&gt; &lt;span class="nt"&gt;--stochastic&lt;/span&gt;    &lt;span class="c"&gt;# Q&amp;amp;A without retrieval&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The advantages of this architecture over the neural network approach extend beyond the obvious ones of transparency and auditability, and I want to spend some time on the less obvious advantages because they are the ones that matter most for the long-term potential of the lmm project. The first non-obvious advantage is composability: because the deterministic layer and the stochastic layer are explicitly separated, it is possible to improve each independently without the improvements interfering with each other. You can make the genetic programming more powerful, allowing it to discover more complex equations from noisier data, without touching the synonym bank at all. You can expand the synonym bank, improving the variety and naturalness of the stochastic layer, without touching the equation discovery at all. In a language model, this kind of independent improvement is impossible, because the model's capabilities, its factual knowledge, its linguistic fluency, its reasoning ability, and its stochastic output behavior are all baked together in the same parameter matrix, which means you cannot improve one without risk of degrading the others. The architectural cleanness of the lmm approach is not just aesthetically pleasing. It is the property that makes systematic, directed improvement possible rather than the empirical, emergent, unpredictable improvement that scaling produces.&lt;/p&gt;

&lt;p&gt;The second non-obvious advantage is what I call zero hallucination by architecture rather than zero hallucination by alignment training. Language models hallucinate because their output is a sample from a distribution that was learned from text, and text contains errors, fabrications, and confident-sounding falsehoods, which means the learned distribution assigns non-zero probability to outputs that are factually wrong, and temperature sampling can land on those wrong outputs. The engineering response to this has been alignment training and RLHF, which attempt to shift the distribution away from commonly hallucinated outputs by rewarding correct outputs in a second training phase. That is a patch on a structural problem: you are trying to fix a system that does not know the difference between true and false by training it to mimic a preference for truth, and the mimic is only as good as the preference data, which is limited, biased, and never complete. The lmm system does not hallucinate in this sense because its epistemic layer is not sampling from a learned distribution at all. The causal inference produces results that are computed from an explicit causal graph. The physics simulation produces results that are computed from explicit differential equations. The symbolic regression produces equations that fit the actual data. If any of these computations are wrong, the error is traceable to a specific input, a specific equation, a specific causal assumption that can be examined and corrected. The stochastic layer introduces only synonym variation, not factual variation, which means the surface words change but the facts encoded in the sentence structure remain constant across runs. That is a fundamentally different and more honest error mode.&lt;/p&gt;

&lt;p&gt;The third advantage is the one I think has the most long-term significance, which is that stochastic determinism is the architecture that enables genuinely personalized outputs without privacy violations. When a language model is fine-tuned to sound like a specific person or to serve a specific user's preferences, the personalization is embedded into the model's weights, which means the model has encoded something about the target person into parameters that cannot be easily inspected, reverted, or isolated from the rest of the model's behavior. This is a privacy concern in addition to an architectural one. The lmm approach, because the stochastic layer is a separately specified and auditable synonym bank, could in principle support personalization by providing user-specific synonym preferences, domain-specific terminology banks, or style-specific structural templates that modify the expression layer without touching the reasoning layer at all. Your preferred vocabulary is stored in an explicit file that you can inspect, modify, revoke, and delete. It is not baked into an opaque parameter matrix that you cannot audit or remove. That is the difference between a tool that respects your agency over your own cognitive preferences and a tool that absorbs your preferences into its own body and uses them in ways you cannot fully see or control. For the same reasons I argued in &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept&lt;/a&gt; about training on creative work, personalization through opaque parameter absorption is a different and lesser kind of respect for the person being personalized than explicit, inspectable, revocable preference storage is.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Philosophical Stakes Are Higher Than Anyone Is Admitting
&lt;/h2&gt;

&lt;p&gt;The question of whether penguins are sentient is not just a question about penguins. It is a question about what sentience is, and the answer you give to that question shapes both how you treat the living beings around you and how you design the artificial systems you are building. If sentience requires human-like language-mediated reflective self-awareness, then you have a relatively clear design target and a very convenient excuse to ignore the welfare of the rest of the animal kingdom. If sentience is something more fundamental, something like the capacity for subjective experience and directed engagement with the world that underlies many different cognitive architectures, then you have a much harder design problem, a much richer set of existence proofs to learn from, and a much larger moral circle to contend with. I believe the second answer is the right one, and I believe the scientific evidence increasingly supports it, and I believe the AI field's failure to engage with that evidence is both a scientific failure and a moral one.&lt;/p&gt;

&lt;p&gt;The moral dimension is the one I find most troubling to articulate, because making ethical arguments in a technical conversation is a reliable way to be dismissed as impractical or sentimental, and I am neither. But I wrote in &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;Training Is an Evil Concept&lt;/a&gt; about the moral costs of the training paradigm as currently practiced, specifically about the extraction of value from human creative workers without consent, and I want to extend that moral argument here to include the broader question of what it means to build systems we call intelligent while ignoring the intelligence that already surrounds us. There is something ethically confused about a civilization that will spend hundreds of billions of dollars to simulate intelligence in silicon while systematically destroying the habitats of billions of beings that already instantiate the very kind of intelligence we claim to be trying to build. The climate crisis is reducing penguin populations (10). Habitat destruction is eliminating the seed-caching birds whose spatial memory we have barely begun to understand. Ocean acidification is threatening marine invertebrates whose distributed neural architectures we have not yet fully mapped. We are losing the existence proofs faster than we are reading them, and we are doing so while spending our energy and capital on a paradigm that is less like those existence proofs than any serious theory of their intelligence would recommend.&lt;/p&gt;

&lt;p&gt;The philosophical stakes also include the question of what we are building toward, which is something I have addressed in pieces across many posts but want to say directly here. The goal of building artificial general intelligence, if it is an honest goal rather than a marketing goal, should be to understand and instantiate the kind of general intelligence that enables a system to engage with the world flexibly, adaptively, and genuinely, across a wide range of novel situations without collapsing into confusion or confabulation. That goal is exactly what animal cognition demonstrates, at various levels of sophistication, across an enormous range of species and environments. The emperor penguin demonstrates it in the extreme conditions of Antarctica. The Clark's nutcracker demonstrates it in spatial memory. The honeybee demonstrates it in abstract symbolic communication. The octopus demonstrates it in distributed neural computation. None of these systems passed a text comprehension benchmark. None of them can write an essay. None of them can fine-tune a model or run backpropagation. But all of them are doing something that the most powerful language models in existence are not doing, which is engaging with the actual structure of physical reality in a way that is flexible, robust, and alive. If AGI research took that observation seriously, it would look very different from how it currently looks.&lt;/p&gt;

&lt;p&gt;I also want to say something about consciousness specifically, because it is the word that sits in the background of every argument I have been making and that I have been approaching carefully rather than carelessly, because carelessly deployed it becomes a conversation stopper. The question of what consciousness is, whether it requires specific biological substrates, whether it can exist in systems that lack the specific features of mammalian brains, and what the relationship is between intelligence and subjective experience, is genuinely hard, and I am not going to pretend I have answers that the philosophers and neuroscientists who have spent decades on these questions have not found (11). What I will say is this: the dismissal of animal sentience has historically been motivated less by the evidence and more by convenience, by the convenience of being able to use animals as resources without the moral complications that come with treating them as subjects. I am worried that the same convenience is at work in the AI field's dismissal of animal cognition as irrelevant to the question of what intelligence is, because taking animal cognition seriously would complicate the story that the current paradigm is on the right track, and complications are inconvenient when you have already made very large bets.&lt;/p&gt;

&lt;p&gt;The stakes for getting this right are not just philosophical or even just about the welfare of animals. They are about whether we understand intelligence well enough to build systems that can genuinely help humans with the problems that will define the coming century: climate modeling, pandemic prediction, materials discovery, protein engineering, ecological management. These are problems that require genuine understanding of complex physical systems, not fluent text about complex physical systems. They require the kind of reasoning that says "if we change this variable, here is how the system will respond," which is causal reasoning, not correlational pattern matching. They require the kind of generalization that extrapolates from known physical laws to novel situations, not the kind that interpolates within a training distribution. The penguin does its version of all of this every year, in Antarctic conditions, without electricity. The techniques it uses, embodied world modeling, causal environmental reasoning, robustly integrated multi-sensory processing, are exactly the techniques that the lmm project is attempting to build toward, and the techniques that the language model paradigm is structurally unable to instantiate. Getting this right matters. We do not have the luxury of being distracted by impressive demos when the actual problems require something more.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Field Gets Wrong About Cognition, and What Getting It Right Would Look Like
&lt;/h2&gt;

&lt;p&gt;I want to be constructive in this section rather than just critical, because pure criticism without a positive vision is the kind of writing I find most frustrating to read, and I do not want to produce it. I have been critical across seven sections now, and I think the criticism is valid, but I also think the positive vision is where the real energy should go. The positive vision, as I have been building toward it across this post and across all the previous posts, is this: intelligence is a property of systems that build and maintain accurate models of the world and use those models to navigate the world flexibly and adaptively. The world is mathematical in its deep structure, meaning it is organized by differential equations, causal laws, conserved quantities, and symmetries that can be discovered, encoded, and used for prediction. Animal cognition instantiates this kind of intelligence in biological form. The lmm project is an attempt to instantiate it in mathematical form. And the language model paradigm, whatever its surface capabilities, is not moving toward this kind of intelligence, because it is not organized around building world models, discovering physical laws, or reasoning causally about the structure of reality.&lt;/p&gt;

&lt;p&gt;Getting it right would look like a research program that takes the following things seriously simultaneously: the mathematical structure of the physical world as the primary target of representation, the empirical study of animal cognition as the richest source of existence proofs for the kind of intelligence we want to build, the development of architectures that learn from physical observations rather than from human creative expression, the construction of explicit causal models rather than opaque correlational parameters, and the commitment to transparency and verifiability as design principles rather than as optional features. None of this requires abandoning everything that has been learned in neural network research. Genetic algorithms informed modern neural architecture search. Reinforcement learning connects to the optimal control theory that underlies animal navigation. Attention mechanisms have genuine computational affinities with the selective attentional processes studied in animal cognition. The knowledge from the training-based paradigm is not worthless. The direction is wrong, and the moral costs are real, and the fundamental architecture is misaligned with what genuine intelligence requires, but the specific technical insights can be carried forward into the correct direction. The engineers who work on this problem are not the enemy. The paradigm is the problem, and paradigms can be changed.&lt;/p&gt;

&lt;p&gt;What it would look like from the outside is a field that celebrates the discovery of governing equations from data as much as it currently celebrates the scaling of parameter counts. A field where a paper saying "we built a system that discovered a physical law from noisy observations and predicted novel phenomenon X with that law" receives as much attention as a paper saying "we scaled our language model by a factor of ten and observed capability Y emerge". A field where the welfare and cognitive sophistication of the animals that already have the intelligence we claim to want to build is treated as relevant data rather than as a distraction from the real work. A field where "my system computed this" and "my system narrated this" are recognized as claims of completely different epistemic weight rather than being evaluated by the same surface appearance criteria. A field where a verifiable equation is worth more than a confident sentence, because a verifiable equation can be proven wrong while a confident sentence can only be doubted. That is the epistemically honest position, and it is the position that physics has always occupied, and it is the position that the current AI field has largely abandoned in its rush toward products that impress rather than toward understanding that grounds.&lt;/p&gt;

&lt;p&gt;The lmm project is my contribution toward making this field exist, and I want to be honest about how small a contribution it currently is. The codebase is one person's work, implemented in Rust, with capabilities that are impressive as proofs of concept but modest as production systems. The symbolic regression discovers equations from small datasets in a way that scales poorly to high-dimensional chaos. The physics simulations are limited to the models that have been explicitly implemented. The causal reasoning module handles small, explicitly specified graphs rather than automatically inferred complex causal structures. These are real limitations, and the gap between lmm and the kind of mathematical intelligence that could genuinely help with climate modeling or pandemic prediction is significant. I know that. But the limitations of a proof of concept are not evidence against the concept. They are evidence for the need to invest in developing the concept, and the concept, the direction that lmm points toward, is the correct direction. Physics over statistics, equations over parameters, causal structure over correlational patterns, world models over token prediction. Those are not just design choices. They are the difference between building toward genuine intelligence and building toward a very impressive distraction.&lt;/p&gt;

&lt;p&gt;I want to close this section by coming back to the penguins, because they are where this post started and they are the honest benchmark against which current AI systems should be measured. An emperor penguin that finds its mate in a blizzard has solved a problem that involves pattern recognition under severe noise, spatial navigation in a physically demanding environment, integration of multiple sensory modalities, maintenance of a long-term social memory, motivation strong enough to survive months of Antarctic winter, and action that is coordinated with the behavior of thousands of other individuals in the colony. This is not a simple problem. This is a hard problem solved by a system that evolved over millions of years to solve exactly this kind of problem, and studying how it solves this problem would tell us things about intelligence that a billion training examples of penguins in text cannot. The field should want to understand this. The field that claims to be building toward general intelligence should be deeply curious about every existence proof of general intelligence in the world around it. And the fact that the field is largely not curious, largely not funding that curiosity, and largely not building toward architectures that could instantiate what the penguin does is the clearest evidence I know that the field is distracted, and that the distraction is profound.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Post-Distraction Field Would Actually Build
&lt;/h2&gt;

&lt;p&gt;I want to end by being specific rather than rhetorical, because specificity is what I respect and vagueness is what I distrust, and I have been vague enough in this post already. If the field took seriously the argument I have been making, if it took seriously the evidence from animal cognition, if it committed to building systems organized around physical structure and causal reasoning rather than around text prediction and parameter scaling, what would it actually build? I want to try to answer that question concretely, drawing on both the existing animal cognition research and the existing lmm architecture, to give a sense of what the right direction looks like when made specific.&lt;/p&gt;

&lt;p&gt;It would build perception systems that convert raw sensory streams into mathematical representations grounded in physical reality rather than into token sequences grounded in language statistics. The lmm perception layer is the beginning of this: raw bytes are converted to normalized tensors, which are mathematical objects that can be operated on by the symbolic and physical reasoning layers that follow. Extending this to richer sensory streams, acoustic, visual, proprioceptive, thermal, chemical, and building the invariant representations that allow the same physical object or event to be recognized across variations in perspective, distance, and context is a research program with deep roots in computational neuroscience and a clear path from existing lmm architecture toward much more capable systems. The honeybee's encoding of spatial direction in the waggle dance is a specific existence proof of how a compressed, mathematical representation of spatial reality can be communicated between individuals, and understanding the computational principles of that encoding is directly relevant to building better perception systems.&lt;/p&gt;

&lt;p&gt;It would build physics-grounded world models that can predict the future state of a system from its current state and a mathematical description of the forces and constraints that govern it, and that can do this not just for the specific systems where human-derived equations exist but for novel systems encountered in the real world. The lmm physics simulation layer handles several well-understood physical systems, from harmonic oscillators to SIR epidemic models, and the genetic programming symbolic regression layer can discover equations for novel systems from observational data. The integration of these two capabilities, using symbolic regression to discover new physics and physics simulation to generate world-model predictions, is the architecture of a system that can learn about the physical world from observing it rather than from reading about it, and that distinction, as I have been arguing throughout this post, is the fundamental one. Clark's nutcracker's spatial memory system is an existence proof of a world model that is precise enough to locate tens of thousands of specific cached items across months and miles, and the computational principles of that system are waiting to be understood and instantiated.&lt;/p&gt;

&lt;p&gt;It would build causal reasoning systems that can automatically infer causal structure from observational data and experimental interventions, rather than requiring causal structure to be specified in advance by human domain experts. The lmm causal module currently handles explicitly specified structural causal models with the do-calculus intervention operator, which is mathematically rigorous and produces the right answers for the specified models. The hard and important next step is automatic causal discovery, algorithms that can infer the structure of the causal model from observational and interventional data, which is an active area of research with results from computational methods like PC, FCI, and GES that have not yet been fully integrated into the framework I am advocating. The animal that knows "changing X will cause Y to change" without having that relationship spoon-fed to it by a human experimenter is demonstrating exactly the capability that automatic causal discovery would implement, and the research on how animals learn causal structure from their environments, through play, exploration, and social observation, is directly relevant to making automatic causal discovery work at scale.&lt;/p&gt;

&lt;p&gt;It would build integration layers that tie these capabilities together into a unified system that perceives, models, discovers, reasons, and acts in a single continuous loop rather than passing data between separate specialized modules that do not share a common representational framework. This is the role of the lmm consciousness loop, which is designed to be the integration point where perception output becomes world model input and world model prediction becomes action plan. The architecture is correct, but the current implementation is limited to taking raw bytes, running a simple world model prediction, and evaluating prediction error. Making this genuinely powerful requires much deeper integration between the perception layer, the symbolic regression layer, the physics simulation layer, and the causal reasoning layer, so that what the perception layer sees informs what equations the symbolic regression layer searches for, which informs what the physics simulation layer uses as its governing equations, which informs what the causal reasoning layer treats as the structure of the world. That is the architecture of embodied cognition as animal cognition research reveals it, and it is the architecture that the lmm project is trying to move toward, one capability at a time.&lt;/p&gt;

&lt;p&gt;It would, finally, take seriously the moral implications of the evidence about animal sentience, not as a sentimental addendum to the technical program but as an integral part of the research agenda. If the existence proofs of genuine intelligence that we need to study in order to build genuine artificial intelligence are the same beings whose habitats we are destroying, whose welfare we are ignoring, and whose cognitive sophistication we are systematically underestimating, then the research program is in a genuine ethical contradiction with itself. A field that claims to value intelligence while disrespecting the intelligent beings that already exist is not a field that I trust to build systems that respect the intelligence of the humans who use them. The moral circle and the epistemic horizon expand together, and a field that refuses to expand either is a field that will keep producing impressive distractions rather than genuine understanding. The penguins are already sentient. They are already demonstrating the kind of embodied, causal, physically grounded intelligence that the AI field cannot yet build. They deserve our study, our respect, and our honest acknowledgment that they are ahead of us in ways we have barely begun to admit.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Low, P. et al., &lt;em&gt;The Cambridge Declaration on Consciousness&lt;/em&gt;, &lt;a href="https://fcmconference.org/img/CambridgeDeclarationOnConsciousness.pdf" rel="noopener noreferrer"&gt;Francis Crick Memorial Conference, Cambridge, 2012&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; Robisson, P., Aubin, T. &amp;amp; Brémond, J.C., &lt;em&gt;Individuality in the Voice of the Emperor Penguin Aptenodytes forsteri: Adaptation to a Noisy Environment&lt;/em&gt;, &lt;a href="https://doi.org/10.1111/j.1439-0310.1993.tb00445.x" rel="noopener noreferrer"&gt;Ethology 94, 279–290, 1993&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; Fodor, J.A., &lt;em&gt;The Modularity of Mind&lt;/em&gt;, &lt;a href="https://mitpress.mit.edu/9780262560252/the-modularity-of-mind/" rel="noopener noreferrer"&gt;MIT Press, 1983&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; Marcus, G., &lt;em&gt;The Algebraic Mind: Integrating Connectionism and Cognitive Science&lt;/em&gt;, &lt;a href="https://mitpress.mit.edu/9780262632683/" rel="noopener noreferrer"&gt;MIT Press, 2001&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Marcus, G. &amp;amp; Davis, E., &lt;em&gt;Rebooting AI: Building Artificial Intelligence We Can Trust&lt;/em&gt;, &lt;a href="https://www.worldcat.org/isbn/9781524748258" rel="noopener noreferrer"&gt;Pantheon Books, 2019, ISBN 978-1-524-74825-8&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Balda, R.P. &amp;amp; Kamil, A.C., &lt;em&gt;Long-term Spatial Memory in Clark's Nutcracker, Nucifraga columbiana&lt;/em&gt;, &lt;a href="https://doi.org/10.1016/S0003-3472(05)80302-1" rel="noopener noreferrer"&gt;Animal Behaviour 44(4), 761–769, 1992&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Riley, J.R. et al., &lt;em&gt;The Flight Paths of Honeybees Recruited by the Waggle Dance&lt;/em&gt;, &lt;a href="https://doi.org/10.1038/nature03526" rel="noopener noreferrer"&gt;Nature, 2005&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; McCorduck, P., &lt;em&gt;Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence&lt;/em&gt;, &lt;a href="https://archive.org/details/machineswhothink00mcco" rel="noopener noreferrer"&gt;W.H. Freeman, 1979, ISBN 978-0-716-71072-1&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-9"&gt;&lt;/span&gt;&lt;strong&gt;9.&lt;/strong&gt; Razeghi, Y. &amp;amp; Logan, R.L., &lt;em&gt;Impact of Pretraining Term Frequencies on Few-Shot Numerical Reasoning&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2202.07206" rel="noopener noreferrer"&gt;arXiv:2202.07206&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-10"&gt;&lt;/span&gt;&lt;strong&gt;10.&lt;/strong&gt; Trathan, P.N. et al., &lt;em&gt;Penguins and Climate Change&lt;/em&gt;, &lt;a href="https://doi.org/10.1098/rstb.2014.0217" rel="noopener noreferrer"&gt;Philosophical Transactions of the Royal Society B, 370(1669), 2015&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-11"&gt;&lt;/span&gt;&lt;strong&gt;11.&lt;/strong&gt; Nagel, T., &lt;em&gt;What Is It Like to Be a Bat?&lt;/em&gt;, &lt;a href="https://doi.org/10.2307/2183914" rel="noopener noreferrer"&gt;The Philosophical Review, 1974&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>lmm</category>
    </item>
    <item>
      <title>Training Is an Evil Concept. LMMs Eliminates it Altogether.</title>
      <dc:creator>Mahmoud Harmouch</dc:creator>
      <pubDate>Fri, 21 Aug 2026 18:46:50 +0000</pubDate>
      <link>https://dev.to/wiseai/training-is-an-evil-concept-lmms-eliminates-it-altogether-15ej</link>
      <guid>https://dev.to/wiseai/training-is-an-evil-concept-lmms-eliminates-it-altogether-15ej</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This post was originally published on &lt;a href="https://wiseai.dev/blogs/training-is-an-evil-concept-lmms-eliminates-it-altogether" rel="noopener noreferrer"&gt;the main website&lt;/a&gt; on &lt;a href="https://github.com/wiseaidotdev/blog/pull/16" rel="noopener noreferrer"&gt;Apr 16 2026&lt;/a&gt;. I am reposting it here for SEO reasons and enabling humble bumble discussions with the DEV community. Feel free to engage with this post and i am available to respond during weekends. Sorry about the spam posting all the blogs in one day. I forgor about my dev account &amp;lt;3!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hey everyone 👋,&lt;/p&gt;

&lt;p&gt;In my last few posts, I have been building a case, one piece at a time, that the direction most of the AI industry is moving in is not the direction that will produce genuine intelligence. In &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt;, I argued that language models are trapped inside a symbolic cage, that they can describe the world without ever touching it, and that the transition from text-prediction to mathematical perception is the most important shift happening in AI right now. In &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, I argued that equations are not tools for homework but the most compressed and honest representations of reality that humans have ever produced, and that any system built around equations inherits their multimodal power for free. In &lt;a href="https://wiseai.dev/blogs/llms-destroyed-the-internet-lmms-will-make-it-alive" rel="noopener noreferrer"&gt;LLMs destroyed the Internet. LMMs will make it alive.&lt;/a&gt;, I argued that the mass deployment of language models as content factories has quietly dissolved the authenticity that made the web worth using, and that only grounded intelligence tied to reality can reverse that damage. Each of those posts was a different face of the same underlying argument, which is that the current paradigm is built on a foundation that looks impressive from the outside and is rotten from the inside. And in this post I want to say the thing that connects all of those faces, the thing that I have been circling around for months without quite naming directly, because I was not sure I had earned the right to say it yet. The thing is this: training, as it is currently practiced and celebrated in the AI industry, is not a neutral engineering choice. It is a moral choice that most of the people making it have not examined honestly, and the consequences of that unexamined choice are visible everywhere the technology has touched, from the web I described in my last post to the lives of the people whose work was consumed to build these systems, to the lives of the engineers whose labor funds the entire enterprise while they are simultaneously told to be grateful. I have spent years trying to understand why brilliant people build systems with predictable harms and then seem genuinely surprised when the harms arrive, and I think the answer is that nobody forced them to sit with the question of what training actually is and what it actually does to the world. This post is my attempt to force that conversation, at least for the people who read me, and to connect it honestly to everything I believe about where intelligence should be heading.&lt;/p&gt;

&lt;p&gt;I want to start by being careful about the word "evil", because it is a word that generates heat rather than light if it is thrown carelessly, and generating heat without light is exactly the failure mode I am trying to avoid. I am not saying that every engineer who has ever trained a neural network is a bad person. I am not claiming that training is demonic in some metaphysical sense. I am using the word precisely, in the old-fashioned sense that is most useful here, to mean a systemic practice that causes harm in ways that its practitioners could have foreseen if they had chosen to look, and that the choice not to look was itself a moral failure rather than an innocent oversight. The harm is not hypothetical. It is documented in court filings, in academic research, in the testimonies of writers and artists and coders whose work was ingested without consent, in the documented degradation of the web I wrote about last time, in the bias that lives inside every model that learned from biased data and then confidently produced biased outputs for millions of users who trusted it. Evil in this sense does not require malice. It only requires the systematic application of power in ways that impose costs on the powerless while delivering benefits to the powerful, and the refusal to examine that asymmetry honestly. By that standard, the training paradigm as currently practiced is not merely unfortunate. It is genuinely evil, and I think saying so clearly is more useful than softening it into a policy concern, because soft language about AI ethics has been available for years and has changed almost nothing, and I am no longer interested in language that changes nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Training Is Not Learning. It Is Extraction Under a Different Name.
&lt;/h2&gt;

&lt;p&gt;The word "training" does the most extraordinary rhetorical work in the AI conversation, and I want to start by pulling it apart, because the choice of that word is not innocent. Training, as a concept borrowed from human education and behavioral science, implies a relationship between a learner and a teacher, a process that involves consent, care, structure, and the learner's eventual autonomous capability. When an organization says it is "training" a model, the word evokes that framework, which is designed to feel benign, because we all understand that training people and animals is a normal and often good thing. But that evocation is a misdirection, and the misdirection matters because it shapes how billions of people think about what is happening when a large language model is built. What actually happens when a modern language model is trained is not a teaching relationship. It is an extraction process. Data is collected from sources, often at massive scale, often without the knowledge or consent of the people who produced it. That data is consumed by an optimization process that extracts statistical regularities from it, regularities that capture the patterns, the styles, the facts, the errors, and the private details that were embedded in the source documents. The result is a system that has absorbed the structure of human knowledge and human expression without any of the humans who produced that knowledge or expression being asked, informed, compensated, or even acknowledged. An honest name for this process would not be "training." An honest name would be "extraction", or perhaps "consumption", because the relationship is fundamentally one of taking, not teaching. The reason the industry reached for "training" instead is exactly the same reason that extraction industries reach for friendly language in every domain, because friendly language reduces resistance and resistance is expensive. I am not interested in reducing resistance to something that should be resisted. I want to call it what it is.&lt;/p&gt;

&lt;p&gt;The legal system is slowly beginning to agree with this framing, awkwardly and incompletely, as legal systems always engage with new technology, but the direction of the argument is becoming visible. Authors, visual artists, musicians, and software engineers have filed lawsuits in multiple jurisdictions arguing that using their work without consent or compensation to train AI systems constitutes copyright infringement and other legal wrongs (1). The U.S. Copyright Office has produced a multi-part series of reports specifically addressing the question of whether training on copyrighted material is legally permissible, and the answer is not a clean yes. The reports describe training as raising complex questions about reproduction, transformation, and fair use that the current legal framework was not designed to handle, and they note that the matter is being litigated across dozens of active cases (2). The European Union's AI Act and the broader European regulatory framework impose transparency requirements on AI systems, including requirements related to training data disclosure, precisely because legislators recognized that what goes into a training run is not a private technical detail but a public interest question with real consequences for real people (3). The fact that courts and legislatures are taking this seriously should be significant to anyone who is still inclined to treat data collection for training as a morally neutral act. The law is a lagging indicator of public morality, not a leading one, which means by the time courts reach a settled conclusion about whether training on unconsented data is wrong, the harm will have been done at a scale that makes remediation essentially impossible. The time to grapple with the moral reality is before the verdict, not after, and the moral reality is that taking things from people without asking is wrong even when the things being taken are abstract, and even when the taking is technically possible, and even when the result of the taking is impressively useful to third parties.&lt;/p&gt;

&lt;p&gt;I want to connect this to something I said in &lt;a href="https://wiseai.dev/blogs/as-engineers-llms-should-pay-us-for-tokens-usage" rel="noopener noreferrer"&gt;As Engineers, LLMs should pay us for tokens usage&lt;/a&gt;, because I argued there that the value extracted from engineers who generate and share code, documentation, forum answers, and technical writing ends up enriching the companies that train on it without flowing back to the people who produced it. That argument is one specific instance of a much larger structural problem that training creates at every layer of the creative and intellectual economy. Writers who spent years developing distinctive voices have found those voices imitated at scale by systems trained on their work, without attribution, without permission, and without royalties, and the imitation is good enough to produce outputs that compete directly with them in the market for writing. Visual artists have found their unique styles synthesized and combined in ways that would be obviously infringing if a human artist had copied them directly, but that are treated as novel production because the copying happened inside a training run rather than inside a sketchbook. Researchers have found their papers consumed, their findings abstracted, and their academic labor transformed into model weights that are then sold as a product, with none of the value returned to the universities and funding bodies that supported the original research. Software engineers have found their open-source code, released under licenses that require attribution and in some cases financial compensation for commercial use, incorporated into training datasets and then used to build coding assistants that compete directly with the engineers in the job market. Each of these is an instance of the same structural dynamic: training is a mechanism for capturing the surplus value produced by creative and intellectual labor and concentrating it in the organizations that can afford the compute to run the training pipeline, and the mechanism is designed to be opaque enough that the people being extracted from cannot easily trace the connection between their work and the system's capability. I find it hard to describe that dynamic in any terms other than exploitation, and I do not think softening the language serves anyone except the people doing the extracting.&lt;/p&gt;

&lt;p&gt;The bias problem adds another dimension to the moral case against training as currently practiced, and it is a dimension that the technical community has been aware of since at least the early 2010s but has consistently underweighted in its actual design and deployment decisions. The fundamental reality is that when a model is trained on data produced by humans, it does not learn some idealized abstraction of human knowledge. It learns the specific human knowledge that was represented in the specific training dataset, including its demographic imbalances, its historical prejudices, its cultural blind spots, its overrepresentation of certain languages and communities and underrepresentation of others, and the cumulative biases that arise from centuries of unequal access to the tools of written expression. Research published by groups at major universities has consistently shown that large language models trained on web-scale data reproduce and sometimes amplify the social biases present in that data, producing outputs that systematically associate certain demographic groups with negative attributes, that perform worse on languages and dialects with less representation in training data, and that encode occupational and social stereotypes that have been empirically documented in both word embeddings and generated text (4). This is not a surface-level problem that can be fixed with a few rules added to the fine-tuning stage. It is a structural consequence of training on data that reflects the unequal world that produced it, and the only thorough solutions require either fundamentally different training data, which raises its own consent and collection questions, or fundamentally different architectures that do not encode the world's biases by absorbing its text. The reason this matters morally is that the outputs of biased systems are not distributed equally. They fall hardest on the communities that were least represented in the training data and most marginalized in the society that produced it, which means the people who were underrepresented in the input are the people who pay the highest price for the model's errors in the output. That is the structure of systemic harm, and it is the structure that training, as currently practiced, reliably produces.&lt;/p&gt;

&lt;p&gt;The issue of memorization deserves its own careful attention, because it represents the training paradigm's most direct collision with individual privacy, and privacy is one of the clearest moral principles in any framework of respect for persons. Research has shown that large language models trained on web-scale corpora can memorize and reproduce verbatim fragments of their training data, including fragments that contain personally identifiable information, including fragments of private communications that were exposed through data breaches and then ingested into training datasets, including fragments of copyrighted text, and including fragments of content that individuals have since deleted or corrected (5). The existence of this memorization is not speculative. It has been demonstrated empirically by researchers who were able to extract training data from deployed models through systematic prompting. What it means is that information you produced and shared in a specific context, with specific expectations about who would read it and what would happen to it, may have been absorbed into a model and may be reproducible by anyone who knows the right prompt to use. That is a violation of contextual integrity, the principle that information flows appropriately when they match the norms of the context in which the information was originally shared (6). A message you sent in a private group, a blog post you wrote and later deleted, a forum answer you gave before you understood how the internet worked, may be living inside a language model and waiting to be retrieved. The industry's response to this has generally been to acknowledge that memorization exists and then proceed without changing the fundamental approach, because the fundamental approach is the source of the capability, and capability is the source of the revenue, and revenue is the thing the industry is actually organized around. I have seen this same logic applied to my own career, as I described in &lt;a href="https://wiseai.dev/blogs/technology-has-destroyed-my-livelihood" rel="noopener noreferrer"&gt;Technology Has Destroyed My Livelihood&lt;/a&gt;, where the comfort of those who benefit from the system is routinely prioritized over the safety of those who are harmed by it. The pattern is tiresome and familiar and it is the pattern that training, as currently practiced, extends into the domain of artificial intelligence.&lt;/p&gt;

&lt;p&gt;Let me also say something about the environmental cost of training, because it is a dimension of the moral argument that I have not covered in previous posts and that I think deserves to be connected to the rest of the case. Training large language models requires enormous amounts of computational resources, which require enormous amounts of electricity, which produce significant carbon emissions and generate significant quantities of electronic waste. Research published in 2019 estimated that training a single large natural language processing model produced carbon dioxide emissions comparable to the lifetime carbon footprint of several passenger cars (7). Since then, the models have become dramatically larger, the training runs have become longer, and the number of organizations conducting these runs has grown substantially. The environmental cost of the current paradigm is real, it falls disproportionately on communities near data centers and power plants, and it is a cost that the people and communities most harmed by the climate crisis are absorbing so that a small number of technology companies can claim to have built impressive demos. I do not bring this up to claim that AI should never use electricity, because that would be absurd. I bring it up because it is one more dimension along which the cost of training is externalized onto people who did not choose to bear it and are not compensated for doing so. The pattern of externalizing cost while internalizing benefit is the defining feature of the training paradigm as a moral system, and the environmental case fits that pattern as clearly as the copyright case and the privacy case and the bias case. When I say training is evil, I mean it is a system that reliably concentrates benefits in a small number of hands while distributing costs across a much larger number of people who had little or no say in the arrangement, and that is a description of systemic injustice regardless of the technical sophistication of the mechanism that produces it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Training Actually Optimizes for, and Why That Is the Problem
&lt;/h2&gt;

&lt;p&gt;I want to go deeper into the technical argument here, because I think the moral case I have been making is actually stronger when it is connected to what training actually does at the mathematical level, rather than only at the policy and ethical level. The reason I care about the mechanics is the same reason I have argued across multiple posts that the specific architecture of intelligence matters morally and not just practically, because the architecture determines what the system is capable of knowing and what it will systematically miss, and those omissions have consequences for real people. In &lt;a href="https://wiseai.dev/blogs/language-is-limited-asi-is-impossible" rel="noopener noreferrer"&gt;Language is Limited. ASI is Impossible.&lt;/a&gt;, I argued that language models cannot reach general intelligence because the thing they learn from, namely text, is a symbolic representation of reality rather than reality itself, and learning from symbols about the world is categorically different from learning from the world. Here I want to make a related and sharper point, which is that the specific optimization objective used in training large language models, namely predicting the next token given the previous context, is not an objective that selects for truth or for grounding in the world. It is an objective that selects for plausibility given the training distribution, and plausibility given the training distribution is a property of text, not a property of the external world, and the gap between those two properties is where most of the model's failures live.&lt;/p&gt;

&lt;p&gt;When a language model is trained on next-token prediction, it learns which token sequences are most likely to occur in text of the kind that appeared in its training data. That is a sophisticated and useful thing to learn. But it is explicitly not a thing that teaches the model which sequences are most likely to be true, or most likely to be physically grounded, or most likely to be causally connected to any observable state of the world. The model that says "the Earth is approximately 4.5 billion years old" is not accessing a geological database and retrieving a verified fact. It is producing a token sequence that is highly likely given the patterns in billions of documents about geology, most of which happen to agree on that number, and the fact that the answer is correct is a coincidence of the training distribution rather than a consequence of the optimization objective. The same model, asked about a topic where the training distribution is confused, contradictory, or dominated by misinformation, will produce plausible-sounding text that reflects the confusion without any mechanism to flag its own uncertainty or defer to a more reliable source. This is not a configuration problem. This is not something that more compute or more data will fix. It is the direct and predictable consequence of training an objective that optimizes for plausibility rather than truth, and the distinction between plausibility and truth is the entire problem that the scientific method was invented to solve. We spent centuries developing tools, mathematics, experiment, replication, peer review, that could distinguish what seems true from what is true, and the training paradigm casually discards most of those tools in favor of a statistical proxy that is fast, cheap, and impressively wrong in exactly the cases where being impressively wrong is most dangerous.&lt;/p&gt;

&lt;p&gt;This is where I want to connect to what I argued in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, because that post was not just about equations being pretty. It was about what it means for a representation to be grounded, and grounding is exactly what next-token prediction is not. An equation is grounded in reality because it encodes a mechanism that generates verifiable predictions, and verifiable predictions are predictions that can be tested against observations and confirmed or refuted. The training paradigm produces representations that encode patterns in text, and text patterns cannot be tested against observations in any rigorous way, because text patterns are not predictions about the physical world, they are predictions about what kind of text tends to follow other kinds of text in the corpora that humans have produced. When I said in that post that equations are multimodal by default, I meant that mathematical structure derives all its modalities from a single grounded source, and that grounding is what makes the outputs trustworthy in a way that language model outputs are not. The point I want to make here is the negative of the same claim: training on text produces representations that are multimodal in surface appearance, because the training data contained descriptions of many modalities, but they are not multimodal in ground truth, because the training data was not grounded in any of those modalities at the mechanism level. A model that has read a million descriptions of how springs work is not a model that understands springs. It is a model that understands how people write about springs, which is a very different thing, and the difference is exactly what training on text cannot bridge.&lt;/p&gt;

&lt;p&gt;The environmental and resource dimensions of training also connect directly to this optimization argument, in a way that I think is underappreciated. Because next-token prediction is a statistical objective applied to massive corpora, the way to improve performance under this objective is to train on more data with more compute, and the relationship between scale and performance has been empirically observed to follow specific scaling laws, meaning that the benefits of additional scale are real and quantifiable (8). This has created a perverse incentive structure where the primary engineering lever for improving AI systems is spending more money on computation, which means the organizations with the most resources can build the best systems, which means the economics of AI concentrate in favor of the largest institutions, which means the people setting the direction of the field are the people who are most invested in the current paradigm continuing to be the right one. The training paradigm has made itself self-reinforcing not because it is the best possible approach to building intelligence, but because it happens to scale with money in a way that is visible and measurable, and visible measurable progress with money is the thing that attracts more money. The alignment between the training paradigm's scaling properties and the incentive structure of venture-backed technology companies is not a coincidence. It is the mechanism by which a methodologically questionable approach has become the defining paradigm of an entire industry, and I think understanding that mechanism is necessary to understanding why the paradigm has persisted despite its documented costs.&lt;/p&gt;

&lt;p&gt;I want to be honest about the steelman of the training paradigm, because honesty requires engaging with the strongest version of the opposing view rather than the weakest one. The strongest defense of training as it is currently practiced is something like this: the alternatives, whatever they might be, have not produced systems of comparable capability, and capability is what is needed to actually help people, and failing to help people by maintaining theoretical purity is its own kind of moral failure. That is not a stupid argument. It is the argument I would make if I were trying to defend the current approach, and it has real force. The systems produced by training, whatever their ethical costs, have genuinely helped some people in some domains: they have accelerated drug discovery research, they have made programming assistance available to people who could not otherwise afford expert developers, they have translated languages and summarized texts and answered questions in ways that have real value for real users. I acknowledge all of that, and I do not want to be the kind of critic who treats every benefit of the technology as invisible. But the steelman has a crucial hidden premise, which is that the current paradigm is the only path to capability, and that premise is not established. It is assumed, because it is convenient, and convenient assumptions are the most dangerous kind. The history of technology is full of paradigms that seemed inevitable until they were replaced by something better, and "we have not yet found a viable alternative" is not the same as "no viable alternative exists." The moral costs of training at scale are real and documented. The claim that they are unavoidable is not established. And the refusal to take that distinction seriously is the thing that most angers me about the current conversation.&lt;/p&gt;

&lt;p&gt;The fine-tuning process deserves its own examination, because it is often presented as the answer to training's ethical problems, and it is not. Fine-tuning, whether through reinforcement learning from human feedback or through other supervised adjustment processes, is designed to adjust a pre-trained model's behavior toward outputs that human evaluators prefer. That sounds like an improvement over raw training on internet data, and in some surface ways it is. But fine-tuning has its own moral complexities that have been well documented. The annotators who provide the human feedback that drives RLHF are often poorly compensated workers in low-income countries who are asked to evaluate disturbing, violent, or traumatic content as part of their work, and the conditions under which they perform that work have been the subject of investigative reporting that should disturb anyone paying attention (9). The fine-tuning process extracts value from their labor, under conditions that no organization in a wealthy country would consider acceptable for their own employees, in order to make a product more palatable to users in those wealthy countries. That is a moral cost that is structurally identical to the moral cost of unconsented data collection, just located in a different part of the pipeline. The consistent pattern across the entire training and fine-tuning process is that costs are externalized to people with less power and less visibility, while benefits are concentrated in organizations with more power and more visibility. Fine-tuning does not fix training's moral problem. It perpetuates the structure of the moral problem at a different stage.&lt;/p&gt;

&lt;h2&gt;
  
  
  LMM: The Proof That Pure Mathematics Can Replace Training Entirely
&lt;/h2&gt;

&lt;p&gt;I have been critical across several sections and I owe the reader something more than criticism: I owe them a demonstration that the alternative is real, not theoretical, not a future aspiration dressed up in confident language, but actual running code that somebody has actually built and that actually works without training. I want to talk about &lt;a href="https://github.com/wiseaidotdev/lmm" rel="noopener noreferrer"&gt;lmm&lt;/a&gt;, which is the project I have been building quietly alongside these blog posts, because it is the most concrete answer I have to the objection that training is necessary. I want to be honest about what lmm is and is not, because overselling it would undermine the entire argument I am making about epistemic honesty. It is not a production system. It is not a replacement for GPT-5 in the applications where GPT-5 is currently used. It is a proof of concept, and the concept it proves is specific and important: that a system can perceive the world, discover mathematical structure within it, reason causally about it, and even generate coherent language output, without training on a single human-authored document, without a gradient descent step across a corpus of unconsented creative work, and without any of the ethical costs that the training paradigm makes unavoidable. The system is implemented entirely in Rust, which matters because Rust's type system and ownership model make it possible to write a verifiable, auditable system whose behavior can be reasoned about from first principles, which is exactly the property that trained neural networks systematically lack. The architecture is organized around five layers: perception, which converts raw input into tensors; symbolic regression, which discovers governing equations from data using genetic programming; physics simulation, which models dynamic systems using differential equation integrators; causal reasoning, which constructs structural causal models and applies do-calculus interventions; and cognition, which ties these together into a perceive-encode-predict-act loop that resembles the structure of conscious engagement with the world without depending on the statistical average of everything ever written about it.&lt;/p&gt;

&lt;p&gt;The symbolic regression system is the heart of lmm and the demonstration I want to spend time on, because it is the specific capability that replaces what training does in language models while doing so in a fundamentally different and more honest way. What lmm's symbolic regression does is take a set of data points, which might be measurements of a physical phenomenon, a time series of sensor readings, a sequence of observations from any domain, and search for a symbolic mathematical expression that fits the data. The search is done through genetic programming, which means a population of candidate expressions is evolved over multiple generations, with expressions that fit the data better surviving and those that fit worse being replaced, and the result is a compact symbolic equation that captures the structure of the data in human-readable, verifiable, falsifiable form. The crucial difference from training is what the output represents. When a language model is trained, the output is billions of floating-point parameters whose relationship to the training data is opaque, untraceable, and irreversible, which is why memorization and bias are structural problems rather than configuration bugs. When lmm performs symbolic regression, the output is an equation, something like &lt;code&gt;(95.09 - cos(x))&lt;/code&gt; or &lt;code&gt;(x + (1.002 + x))&lt;/code&gt;, which is a representation that anyone can read, verify against the data, and reason about mathematically. That transparency is not cosmetic. It is the property that makes the output trustworthy in a way that trained model outputs cannot be, because it exposes the system's reasoning in a form that invites challenge and correction rather than hiding it in a parameter space that is beyond human comprehension.&lt;/p&gt;

&lt;p&gt;Let me give you a concrete example from lmm's own documentation, because concrete examples are always more honest than abstract principles. The system includes a command called &lt;code&gt;encode&lt;/code&gt; that encodes any text as a symbolic mathematical equation using genetic programming. When you run &lt;code&gt;lmm encode --text "The Pharaohs encoded reality in mathematics."&lt;/code&gt;, the system treats the text as a sequence of byte values indexed by position, runs genetic programming to find a symbolic equation &lt;code&gt;f(x)&lt;/code&gt; that approximates those byte values, stores the equation along with integer residuals that capture the approximation error, and produces a lossless encoding of the original text in the form of mathematics. The round-trip is verified to be perfect: you can run &lt;code&gt;lmm decode&lt;/code&gt; with the equation and residuals and recover the original text exactly. Now I want you to compare that with what a language model does when it "encodes" a piece of text. The language model embeds the text in a high-dimensional vector space, represents it as a linear combination of learned basis vectors derived from billions of documents, and produces a geometric point in a space that has no interpretable relationship to either the original text or the physical world. The lmm encoding is transparent and verifiable. The language model embedding is opaque and untraceable. Both are forms of compression, but only one of them is a form of honest compression, compression that shows its work and invites you to verify it. That difference is the entire argument in more concrete form than anything I have written in this paragraph.&lt;/p&gt;

&lt;p&gt;The physics simulation layer of lmm demonstrates a different dimension of what training-free intelligence can look like at the level of world modeling. The system includes implementations of several fundamental physical models: a harmonic oscillator governed by Hooke's law, the Lorenz chaotic attractor which produces the famous butterfly-shaped strange attractor from just three coupled differential equations, a nonlinear pendulum, a SIR epidemic model for disease spread, and an N-body gravitational system. Each of these models is not trained. It is formulated mathematically from first principles, implemented as a set of differential equations, and integrated forward in time using numerical methods including the Euler method, standard fourth-order Runge-Kutta, the adaptive RK45 method, and symplectic leapfrog integration for Hamiltonian systems. When you run &lt;code&gt;lmm physics --model lorenz --steps 500&lt;/code&gt;, you get the exact trajectory of the Lorenz attractor computed from the differential equations, a trajectory that is entirely determined by the equations and initial conditions, entirely transparent, entirely verifiable, and entirely training-free. A language model asked to predict the trajectory of a Lorenz attractor would produce a plausible-sounding description of chaos theory and might even produce numbers that look roughly right, but those numbers would be interpolations from its training distribution rather than computations from the actual equations, and the difference matters the moment you need to use the numbers for anything that requires them to actually be correct. The lmm approach does not just tell you about the Lorenz attractor. It computes it, which is the fundamental distinction between description and understanding that I have been trying to articulate across all of these posts.&lt;/p&gt;

&lt;p&gt;The causal reasoning layer is perhaps the most philosophically significant part of lmm for the argument I am making, because causality is precisely the thing that next-token prediction training cannot learn. There is a well-documented theorem in causal inference that statistical associations, no matter how thoroughly measured, cannot by themselves identify causal relationships, and that causal knowledge requires either controlled experimentation or theoretical commitment to a causal model (12). What this means for trained language models is that despite their ability to produce fluent text about cause and effect, they do not have access to causal structure in any rigorous sense. They have access to patterns of co-occurrence in text written by humans who had causal understanding, which is not the same thing. The lmm system, by contrast, implements structural causal models with explicit do-calculus intervention support. You can specify a causal graph, ask what happens when you intervene on a variable by setting it to a specific value, and the system computes the downstream effects by propagating the intervention through the causal structure rather than by pattern-matching to previous text about what usually happens. When you run &lt;code&gt;lmm causal --intervene-node x --intervene-value 10.0&lt;/code&gt; on a three-node causal model where &lt;code&gt;y = 2 * x&lt;/code&gt; and &lt;code&gt;z = y + 1&lt;/code&gt;, the system tells you that setting x to 10 causes y to become 20 and z to become 21, and this is not a guess or a plausible extrapolation from training data about causal relationships. It is a computation from an explicitly specified and verifiable causal structure. That is the difference between knowing that something causes something and being able to describe the general concept of causality in confident-sounding sentences.&lt;/p&gt;

&lt;p&gt;The text generation capability of lmm deserves special attention because text generation is the specific domain where the comparison to language model training is most stark and most revealing about what the two paradigms are actually doing. The &lt;code&gt;predict&lt;/code&gt; command generates text continuation without any training on human-authored corpora. Instead, it runs genetic programming on the context words to discover a trajectory equation describing how word identity changes with position, discovers a rhythm equation describing how word length changes with position, and then uses these equations together with a vocabulary mapping and a syntactic Subject-Verb-Object loop to produce coherent English sentences. The output of &lt;code&gt;lmm predict --text "Wise AI built the first LMM"&lt;/code&gt; looks like: "Wise AI built the first LMM in the true law often long time and a open path of an old scope is the solid order." That is not grammatically perfect. It is not fluent in the way that GPT-5 outputs are fluent. But it is honest in a way that GPT-5 outputs are not, because every word in that output is traceable to a mathematical computation, to a gene-programmed equation evaluated at a specific position, mapped through a vocabulary structure to a specific word, with no dependence on any human's writing that was absorbed without consent. The fluency of language model outputs is purchased at the price of the ethical costs I have been describing in this entire post. The relative imperfection of lmm's outputs is the cost of honesty, and I find that trade worth making, particularly because it is a proof of concept that can be improved with research investment rather than a fundamental limitation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Mathematical Alternative Is Not Just a Technical Preference
&lt;/h2&gt;

&lt;p&gt;I have been building toward this section across every post I have written about AI, and I want to try to say it as carefully and as precisely as I can, because the argument I am about to make sounds idealistic and I want to separate the idealism from the actual substance. The argument is this: the training paradigm is not just ethically problematic. It is architecturally limited in ways that are not fixable by scaling, and those architectural limitations are what make the ethical costs genuinely wasteful rather than merely unfortunate. If training at scale were producing genuinely intelligent systems that could reason causally about the world, model physical reality, and revise their beliefs based on evidence, the ethical costs would at least be buying something profound. What they are actually buying, as I argued in &lt;a href="https://wiseai.dev/blogs/llms-are-usefull-lmms-will-break-reality" rel="noopener noreferrer"&gt;LLMs are Useful. LMMs will Break Reality&lt;/a&gt;, is a system that can mimic the surface of intelligence without possessing its substance, and that is a terrible return on the human and environmental and creative capital being consumed to produce it. Large Mathematical Models, or what I prefer to call systems grounded in mathematical structure and physical simulation, eliminate training not as an aspirational possibility but as a demonstrated engineering fact, and the lmm project I described in the previous section is the proof. A system can perceive, encode, discover structure, simulate dynamics, reason causally, and generate language without training on a single token of human-authored text, and the resulting system is more transparent, more verifiable, and more respectful of the people it might eventually serve, because its reasoning is made of equations rather than of absorbed human expression.&lt;/p&gt;

&lt;p&gt;Let me try to be specific about what "mathematical grounding" means in this context, because I want to avoid the vagueness that infects most discussions of AI alternatives. In &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt;, I described how an equation like the wave equation encodes a mechanism, not a surface description, and how that mechanism allows the equation to generate outputs in any modality that the mechanism is relevant to. The wave equation was not learned from reading descriptions of waves. It was discovered by mathematicians and physicists who formulated it from first principles and then tested it against observations. That process of formulation and testing is fundamentally different from the process of training on descriptions, because it involves contact with reality at every step. When a system is built around discovering such structures from data rather than predicting text about such structures, the output of the system is a compact mathematical representation that can be verified against new observations, falsified if wrong, and refined based on evidence. The lmm system does exactly this: when given data from a linear process, its genetic programming engine discovers something like &lt;code&gt;(x + (1.002 + x))&lt;/code&gt;, which is an approximation of the underlying &lt;code&gt;2x + 1&lt;/code&gt; law, and that discovery is not borrowed from any human's writing about linear functions. It is inferred from the data by an algorithm that knows only how to evaluate mathematical expressions, not how to pattern-match to descriptions of mathematical expressions. That distinction between inferring and pattern-matching, between real discovery and sophisticated imitation of discovery, is the architectural distinction that the entire moral argument rests on.&lt;/p&gt;

&lt;p&gt;The question of what LMMs eliminate is worth stating as precisely as the project itself states it. What the lmm project eliminates is the specific dependency on consuming human creative expression at scale as the primary input to intelligence. The system learns from physical observations, from data measurements, from mathematical relationships inferred by symbolic regression, and those inputs are fundamentally different in their ethical character from the unconsented creative work that language model training consumes. Measurements of a pendulum's trajectory, readings from an epidemic simulation, positions of N bodies in a gravitational field, these are not the creative labor of specific individuals who were not asked for permission. They are data about how the physical world behaves, and learning about the physical world from physical observations is the oldest and most honest form of inquiry that humans have ever practiced. It is called science, and science has an established ethical framework for data collection that includes consent, anonymization, review, and attribution precisely because those practices matter. The lmm approach to intelligence is closer to the scientific framework than to the extraction framework of language model training, and that proximity is not accidental. It reflects a deliberate design choice to build intelligence that is accountable to physical reality rather than accountable only to the statistical distribution of human text.&lt;/p&gt;

&lt;p&gt;I want to address the obvious objection here, which is that lmm is currently less capable than GPT-5 in the domains where GPT-5 is most used, and that the capability gap is large enough to make the ethical argument feel like a luxury concern. That objection is partly true and I want to be honest about it. The lmm system produces text that is less fluent, answers questions with less apparent breadth, and handles the wide variety of natural language tasks that users expect from AI assistants with much less apparent smoothness than a trained language model does. I acknowledge all of that, and I am not going to pretend that a proof of concept is a production system. But I want to push back on the hidden premise that greater capability automatically justifies greater ethical cost, because that premise leads to an infinitely regressing justification: whatever capability the current paradigm produces can always be used to justify the costs that produced it, regardless of what those costs actually are. The lmm system demonstrates that capability without training is possible in principle. The question of how to scale that principle, how to extend it to broader domains, how to make it competitive with trained systems in the domains that matter most to people, is a research question that has not been seriously funded or staffed. It is not a question that has been answered and failed. It is a question that the field has mostly chosen not to ask, because asking it seriously would require confronting the possibility that the training paradigm is not necessary, and confronting that possibility is uncomfortable for everyone who has invested their careers in developing it.&lt;/p&gt;

&lt;p&gt;I also want to connect this to the argument from &lt;a href="https://wiseai.dev/blogs/llms-destroyed-the-internet-lmms-will-make-it-alive" rel="noopener noreferrer"&gt;LLMs destroyed the Internet. LMMs will make it alive.&lt;/a&gt; about what the internet loses when it becomes populated by systems trained on human expression rather than grounded in physical reality. The lmm project produces outputs that carry a different kind of evidence in them, evidence not of statistical averaging over human creative work but of mathematical computation over physical structure. When the system encodes a text as a symbolic equation and then decodes it back perfectly, the equation is a real discovery, a real mathematical object that captures something true about the byte-level structure of that particular text. When the system simulates a Lorenz attractor and reports the trajectory, that trajectory is computed from the actual equations that govern chaotic dynamical systems, not approximated from patterns in physics textbooks. That relationship between output and reality is the relationship that the early web's best content had, the relationship of genuine engagement with actual problems rather than sophisticated imitation of such engagement. Building intelligence from equations rather than from text is building intelligence that can restore that relationship, one mathematical discovery at a time, and the restoration is not just philosophical. It is architectural, implemented deliberately, and demonstrably achievable with the technology that exists right now.&lt;/p&gt;

&lt;p&gt;Let me also say what I think is genuinely hard about the transition I am describing, because I do not want to be the person who criticizes the current paradigm without acknowledging the difficulty of replacing it. Mathematical modeling of complex phenomena is genuinely more difficult than statistical imitation of text. Symbolic regression and physics-informed machine learning are active research areas with genuine open problems. The domains where mathematical grounding works most naturally are the domains where we already have good physical theories, and there are enormous domains of human experience and practical importance where we do not have those theories. Language understanding itself, which is perhaps the most practically important domain for AI, does not have a clean mathematical theory in the same sense that fluid dynamics or electromagnetism does, and it is not obvious that it ever will. I acknowledge all of this. The lmm project's current text generation produces output that is less fluent than language models, because fluency in the sense of smooth-sounding prose is a property of statistical averaging over human writing, and removing the averaging removes some of the smoothness. The argument I am making is not that the mathematical path abolishes that difficulty. The argument is that the moral costs of the training paradigm are real and serious enough to justify serious investment in alternatives, even difficult ones, and that the current allocation of research effort and capital, overwhelmingly toward scaling the training paradigm rather than toward alternatives, is not justified by necessity but by inertia, and inertia is not a moral defense.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Consent Problem Is Not Solved by Better Licensing
&lt;/h2&gt;

&lt;p&gt;I want to spend a section on the consent problem specifically, because it is the part of the training ethics debate that I think receives the most inadequate treatment, both from the industry and from most critics. The industry's response to consent concerns has typically been one of two things: either arguing that training data use falls under fair use or other existing legal exemptions, which is a legal argument rather than a moral one and tells us nothing about whether the practice is right rather than merely legal; or proposing licensing frameworks that would allow creators to opt in or opt out of having their work used for training, which treats consent as a commercial transaction rather than a moral foundation. Both responses miss the point in the same way, which is that they treat the consent question as a problem to be managed rather than a principle to be respected, and the difference between management and respect is the difference between compliance and ethics. The lmm approach sidesteps this entire problem not by solving the consent question within the training paradigm but by eliminating the dependency on human creative expression that makes the consent question arise in the first place.&lt;/p&gt;

&lt;p&gt;Consent as a moral principle is not primarily a contractual matter. It is the recognition that persons are not means to others' ends but ends in themselves, which is the foundation of every serious framework of human rights and dignity that has been developed in the post-Enlightenment tradition. When an organization trains a model on creative work produced by a person, it is treating that person's creative labor as raw material for a process that the organization controls and benefits from, and the person is reduced to the role of an input rather than recognized as an agent who gets to decide whether they want their work to serve this particular purpose. That reduction is wrong even when the creative work is technically accessible, even when the training does not copy the work verbatim, and even when the output system produces content that does not obviously resemble the specific person's work. The wrongness does not require legal infringement to be real. It requires only the structural treatment of a person's creative expression as a resource to be consumed rather than a contribution to be respected. A system like lmm that learns from physical measurements rather than from creative expression does not face this structural problem, because physical measurements are not the creative contributions of specific individuals who have rights and interests in how they are used. The SIR epidemic model integrated by a Runge-Kutta solver is not the intellectual property of a specific writer who did not consent to its use. It is a mathematical description of a biological mechanism, and learning from that description is learning from the world rather than from the people who wrote about the world, and that moral distinction is the entire point.&lt;/p&gt;

&lt;p&gt;I described in &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt; what it feels like to have the contributions you make not acknowledged, to put genuine effort into something and find that the effort disappears without leaving any trace on the world. I know that feeling in a personal way, and I think it is close enough to what creators experience when their work is consumed and transformed without acknowledgment that I am willing to use my own experience as evidence of the moral stakes. When a writer produces a body of work over years, each piece is an expression of something particular and personal, a way of engaging with the world that is irreducibly theirs. When that body of work is ingested into a training pipeline without the writer's knowledge and used to build a system that can then produce similar-sounding text on demand at a cost that makes the writer's own production economically non-viable, something real and important has been taken. It is not merely the market value that has been taken, although that too. It is the recognition that the work was the expression of a specific person with a specific life getting specific things from their engagement with specific ideas, and that recognition is what the training paradigm systematically fails to provide. That failure of recognition is the moral failure that licensing frameworks cannot repair, because recognition is not a contractual matter. It is a matter of how you conceptualize the people whose work you are using, and the training paradigm's conceptualization is one of resources rather than persons. Building intelligence from equations rather than from expression is one way to build a system that does not need to fail at recognition, because it does not depend on human expression in the first place.&lt;/p&gt;

&lt;p&gt;The consent problem also extends to the outputs of trained systems in ways that the licensing discussion has not fully addressed. When I use a language model and it produces text, I am often unable to know whether the text reflects patterns absorbed from specific sources in its training data or a genuine synthesis of diverse influences, and that unknowability is itself a violation of the contextual norms that govern honest communication. In any other context, producing text that closely resembles another person's work without acknowledgment would be considered plagiarism, and plagiarism is wrong not primarily because it is illegal but because it misrepresents the authorship and provenance of the work. The training paradigm creates a system that can produce such resemblances at scale, systematically, without any mechanism for tracking or acknowledging the specific sources of the patterns it is reproducing, and then positions the output as the product of the AI system. That misrepresentation is not incidental. The UNESCO Recommendation on the Ethics of AI specifically emphasizes transparency as a fundamental principle, including transparency about the origins and processes that produce AI outputs (10). Training as currently practiced cannot satisfy that principle. An lmm output, by contrast, is always traceable to its mathematical source: any sentence generated by the &lt;code&gt;predict&lt;/code&gt; command can be traced to the trajectory equation, the rhythm equation, the vocabulary mapping, and the positional rules that produced it, and none of those sources are anyone's intellectual property because none of them are anyone's creative expression.&lt;/p&gt;

&lt;p&gt;I want to be honest about one more dimension of the consent problem that I have not yet addressed, which is the consent of future people rather than only current and past creators. The training corpora used for large language models typically include a large proportion of text produced by people who are no longer alive, text from historical figures, from classical authors, from early internet users who could not have imagined the use to which their words would be put. The dead cannot give or withhold consent in any active sense, and the rights that govern posthumous use of creative work vary enormously across jurisdictions and traditions. But the use of historical human expression to train systems that then shape the information environment of living people is not morally neutral simply because the original authors are not present to object. Their expressions were produced in specific contexts, for specific purposes, and with specific expectations about the contexts in which they would be received and used, and using them to train statistical pattern matchers that then influence how a billion people understand the world is a transformation of context so radical that the original authors could not have contemplated it, let alone consented to it. The moral principle is the same one I cited from contextual integrity, that information flows appropriately when they match the norms of the context in which the information was originally produced (6). The lmm approach to intelligence avoids this entire historical dimension of the consent problem, because a system that learns from the trajectory of a pendulum or the spread of an epidemic is not appropriating the creative expression of any historical person. It is reading the book of nature rather than the books of the dead, and there is a fundamental moral difference between those two reading practices.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Memorization Reveals About the Training Paradigm's Soul
&lt;/h2&gt;

&lt;p&gt;I want to spend some time on memorization specifically, because I think it is the most revealing pathology of the training paradigm, the symptom that most directly shows what the system fundamentally is as opposed to what it is claimed to be. I mentioned memorization earlier in a legal and privacy context, but I want to go deeper here because memorization is not just an embarrassing bug. It is a window into the mechanics of what training actually produces and what the relationship is between the training data and the trained model. If training were truly a process of abstraction and learning, as the word "training" implies, we would expect the relationship between the training data and the model's outputs to be indirect and transformed, the way a student who has studied many books can discuss ideas from those books in their own words without being able to reproduce the books verbatim. The empirical fact that trained models can reproduce verbatim fragments of training data is evidence that what training actually produces is not pure abstraction but something closer to compressed storage with pattern matching on top, and that evidence is directly relevant to the moral case because it reveals that the relationship between the training data and the model is more extractive than transformative. The lmm encode-decode cycle is instructive by contrast: when you encode text and then decode it, you are explicitly performing lossless compression and recovery through mathematical structure, and the system does not pretend otherwise. It tells you exactly what equation it found, exactly what the residuals are, and exactly what the round-trip recovery looks like. That transparency is not a feature added for PR reasons. It is the natural state of a system that is built from equations rather than from absorbing human expression without acknowledgment.&lt;/p&gt;

&lt;p&gt;The research on memorization in trained models is worth engaging with carefully. A study specifically examining non-adversarial reproduction, meaning reproduction that occurs during normal model use rather than through deliberate extraction attacks, found that significant fractions of a model's outputs can match internet content verbatim when the prompts are similar to content that appeared in the training data (5). This is not a theoretical possibility. It is a documented empirical reality that occurs in ordinary use of current deployed models. The implications are striking. If you ask a language model to explain a concept and the model happens to have seen a good explanation of that concept during training, the model may reproduce that explanation or large fragments of it, without attributing the source, without the user knowing that they are reading someone else's words, and without the original author having consented to this use. The user believes they are receiving AI-generated synthesis. They may in fact be receiving a fragment of a specific human being's writing, laundered through a statistical process that removed the attribution while keeping enough of the content to be legally and morally questionable. That relationship between inputs and outputs is not the relationship that is advertised when AI systems are presented as creative and generative. It is the relationship of sophisticated storage and retrieval that happens to be opaque enough to escape the moral frameworks that govern explicit copying. An lmm system cannot memorize and reproduce your blog post without your consent because it does not ingest your blog post in the first place. It ingests the physical world, and the physical world belongs to no one and to everyone equally.&lt;/p&gt;

&lt;p&gt;The connection to what I argued in &lt;a href="https://wiseai.dev/blogs/llms-destroyed-the-internet-lmms-will-make-it-alive" rel="noopener noreferrer"&gt;LLMs destroyed the Internet. LMMs will make it alive.&lt;/a&gt; about the loss of the web's authenticity is direct and important. One of the things that made the early web valuable was that its content was traceable, which meant that the provenance of information was in principle recoverable. If you found a forum post that solved your problem, you could see who wrote it and when. If you found an article that made an extraordinary claim, you could trace the claim to its source and evaluate whether the source was reliable. The training paradigm systematically destroys traceability by creating systems that absorb attributable information and produce unattributable output, thereby breaking the provenance chain that made honest information exchange possible. The web that is increasingly populated by model outputs trained on unconsented human expression is less alive partly because its content has been detached from the specific human experiences and identities that made it meaningful and trustworthy in the first place. Memorization makes this visible in an extreme form, but it is the same process that occurs whenever training absorbs human expression and recombines it into outputs that appear to be generated fresh while actually being derived from specific human contributions that went unacknowledged. The lmm approach to the web would look different: instead of outputs that might or might not be reproducing someone's work in a way that nobody can trace, you would have outputs that are demonstrably mathematical, demonstrably grounded in physical structure, and demonstrably not derived from anyone's creative expression, because the derivation chain is entirely transparent and consists entirely of equations.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Scale Argument Is a Moral Red Herring
&lt;/h2&gt;

&lt;p&gt;The most sophisticated defense of the training paradigm that I encounter is what I think of as the scale argument, and it goes roughly like this: yes, there are ethical concerns about training, but the scale of the benefits produced justifies the costs, because the systems trained on massive corpora can help billions of people in ways that no alternative approach could currently match, and the aggregate benefit to humanity is large enough to outweigh the costs to the specific individuals whose work was consumed without consent. This is a consequentialist argument, and consequentialist arguments have real force, and I want to engage with it seriously rather than dismissing it, because dismissing strong arguments is the habit of people who are more interested in being right than in being honest. But the scale argument assumes something that the existence of lmm directly challenges: it assumes that training-based capability is the only path to AI utility. A system that can simulate epidemics, discover governing equations from data, reason causally about interventions, and generate coherent language output without training is a counterexample to that assumption, and counterexamples matter even when they are imperfect, because an imperfect counterexample refutes the claim of impossibility.&lt;/p&gt;

&lt;p&gt;The first problem with the scale argument is empirical. It assumes that the aggregate benefit of current AI systems is large and clearly positive, and that assumption is much less secure than it sounds when stated confidently. The benefits are real in some domains, drug discovery research, programming assistance, language translation, accessibility tools, and I do not deny them. But the costs are also real and substantial, and the empirical measurement of net benefit across an entire society is an extraordinarily difficult problem that nobody has solved. The evidence I cited in &lt;a href="https://wiseai.dev/blogs/llms-destroyed-the-internet-lmms-will-make-it-alive" rel="noopener noreferrer"&gt;LLMs destroyed the Internet. LMMs will make it alive.&lt;/a&gt; about the degradation of the web's information quality represents a broad diffuse harm that is very difficult to quantify but is clearly large in scale. The economic displacement of creative workers represents a concentrated harm to a specific class of people that is also difficult to quantify but is clearly real and ongoing. The bias harms documented in research fall heaviest on already-marginalized communities and represent systematic disadvantage that compounds over time. The privacy violations from memorization potentially affect anyone whose data was absorbed. The environmental costs are global and transgenerational. The scale argument needs to show that the benefits outweigh all of these costs summed together, and making that case requires an honest engagement with the full cost ledger that the AI industry has consistently refused to produce.&lt;/p&gt;

&lt;p&gt;The second problem with the scale argument is philosophical. Even if we could establish that the aggregate benefits outweigh the aggregate costs, which I doubt and which nobody has demonstrated, the consequentialist reasoning fails to respect the separateness of persons, which is one of the most fundamental insights of serious moral philosophy. The fact that a large aggregate benefit exists does not justify taking something from a specific person without their consent, because the person is not the aggregate. The writer whose work was consumed without permission is not made whole by the observation that the model has benefited many people, because they are a separate individual with their own interests and rights that cannot be traded away to produce benefits for others without their participation in the exchange. This is the insight that rights-based frameworks in ethics are designed to protect, and it is the insight that consequentialist arguments in favor of the training paradigm consistently violate. UNESCO's ethics framework is explicitly rights-based rather than purely consequentialist for exactly this reason, maintaining that fundamental human rights cannot be overridden by aggregate calculations of benefit however large the calculation appears (10). The training paradigm, as currently justified by its practitioners, routinely overrides the rights of specific individuals in favor of aggregate benefit claims, and that is a moral framework that has historically been used to justify a very wide range of abuses, and I am not comfortable with it.&lt;/p&gt;

&lt;p&gt;The third problem with the scale argument is that it is not static, and the lmm project makes this visible in a way that pure theory cannot. The people who use the scale argument are implicitly claiming that the current paradigm, with its documented costs, is the only path to the documented benefits. But lmm demonstrates that at least some of those benefits, equation discovery, physics simulation, causal reasoning, language generation, are achievable without training on human-authored text. The claim that training-based capability cannot be achieved through alternative means has not been seriously tested, because the field has been so strongly oriented toward scaling the training paradigm that alternatives have been chronically underfunded and understaffed. If a substantial fraction of the resources currently invested in training larger and larger language models on more and more unconsented data were redirected toward developing systems like lmm, toward genetic programming for equation discovery, toward physics-informed modeling, toward explainable causal inference, we do not actually know what the resulting capability would look like after five or ten years of sustained investment. The moral case for the training paradigm has borrowed its strength from a counterfactual that the field has not been willing to seriously invest in testing, and that is not a robust moral foundation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What We Lose When We Stop Asking Why, and What We Gain When We Start
&lt;/h2&gt;

&lt;p&gt;I want to close with the argument that I think is the most fundamental, the argument that is not primarily about rights or bias or privacy or economics but about something I can only call the soul of an inquiry. One of the most distinctive things about the training paradigm is that it is explicitly designed not to explain itself. The goal of training is to produce a model that generates good outputs, and the measure of goodness is the training objective, and as long as the training objective is satisfied, the internal mechanisms of the model are not required to be interpretable, causal, or grounded in any theory of the phenomenon being modeled. This is not an accidental property. It is the deliberate design choice of a paradigm that prioritizes empirical performance over theoretical understanding, and that choice has philosophical consequences that go beyond efficiency or interpretability in the narrow technical sense. When we build systems that work without understanding why they work, we are betting our technological future on a black box, and black box bets are only reliable within the envelope of the training distribution. Outside that envelope, the system has no principled way to recognize that it is outside its reliable range, and so it generates plausible-sounding outputs regardless of whether those outputs are trustworthy. The lmm system is the opposite: every output traces back to an equation, every equation can be evaluated for fit against the data it was discovered from, every simulation can be verified against known physical behavior, and every causal inference can be checked against the explicit causal graph it was derived from. That is not a black box. That is a system that answers the question "why" in the only way that deserves the name of an answer, by showing the mathematical mechanism that produced the output.&lt;/p&gt;

&lt;p&gt;The contrast with mathematical discovery is profound and worth dwelling on. When mathematicians and physicists discover the laws that govern physical phenomena, they are not just producing accurate predictions. They are producing understanding that can be extended, generalized, and applied to situations that were never in the training dataset, because the understanding is expressed in terms of mechanisms rather than patterns, and mechanisms can be reasoned about in ways that patterns cannot. Newton's laws were not learned from a dataset of planetary observations in the sense that a neural network is trained on data. They were formulated as structural relationships that could be derived from more fundamental principles and extended to phenomena that Newton never observed. That capacity for principled extension beyond the training distribution is what gives theoretical understanding its power, and it is what the training paradigm systematically sacrifices in favor of empirical performance within the training distribution. The lmm system embodies the opposite design philosophy: its symbolic regression engine does not memorize data points. It searches for an equation that captures the structure beneath the data points, and that equation, once found, can in principle be extended to new data points that were never part of the discovery process. That is generalization in the true sense of the word, not interpolation within a training distribution, but principled extension from discovered structure to new observations. I argued in &lt;a href="https://wiseai.dev/blogs/mathematical-equations-are-multimodal-by-default" rel="noopener noreferrer"&gt;Mathematical Equations are Multimodal by default&lt;/a&gt; that equations encode mechanisms rather than surfaces, and the lmm project is the implementation of that argument in executable Rust code.&lt;/p&gt;

&lt;p&gt;I also want to connect this to something in the personal posts, particularly &lt;a href="https://wiseai.dev/blogs/an-empty-life-filled-with-constant-suffering" rel="noopener noreferrer"&gt;An Empty Life Filled With Constant Suffering&lt;/a&gt;, where I talked about how hollow it feels to produce things that are effective without being meaningful, to do the right thing technically while something deeper is missing. I think there is an analogy to what I am describing about the training paradigm. A system that works without understanding why it works has a certain hollowness to it, a certain mechanical efficiency that is not the same as genuine comprehension, and building civilization on top of systems that are efficient but not comprehending is a bet that depends entirely on the distribution of our future problems staying close to the distribution of past data. If the future contains novel challenges, and it always does, then the systems we have built on pattern matching without understanding will be the wrong tools, and the cost of having built them, the moral cost in extracted labor and eroded rights and degraded information ecosystems, will have been paid for something that was not adequate to the moment when it mattered most. The lmm system is my answer to that hollow feeling, an attempt to build something whose mechanisms I can actually see and reason about and extend, something whose relationship to the world is direct and mathematical rather than indirect and statistical, something that asks for nothing from human creators because it is busy learning from the physical world instead. It is not finished. It is not production-ready. It is the beginning of a direction, and a direction is what matters when the current position is wrong.&lt;/p&gt;

&lt;p&gt;There is an important caveat I want to be honest about, which is that lmm currently learns from physical data and mathematical structure, but the domains where it can function are much narrower than the domains where trained language models are used. The simulation and equation discovery capabilities are genuinely training-free, but extending them to the full range of human knowledge and practical need is a research program that will take years of serious investment. I am not claiming that lmm solves the problem completely. I am claiming that it proves the problem is solvable, that intelligence without training on human creative expression is not a contradiction in terms but an achievable engineering target, and that the existence of a working proof of concept changes the moral conversation about training from "it is unfortunately necessary" to "it is a choice, and the choice can be made differently". That shift in framing is small, but it is important, because necessary evils and unnecessary evils require different responses. A harm that is truly unavoidable calls for mitigation and management, which is what most AI ethics frameworks try to provide. A harm that is avoidable but chosen calls for something stronger, which is refusal, and the existence of lmm is my attempt to make that refusal concrete rather than merely rhetorical. I am building the alternative because I believe that building something is more honest than only arguing for it.&lt;/p&gt;

&lt;p&gt;I do not know where the transition from the training paradigm to something better will come from or how fast it will arrive. I am building what I can in &lt;a href="https://github.com/wiseaidotdev/lmm" rel="noopener noreferrer"&gt;lmm&lt;/a&gt;, and I am writing what I believe in these posts, and I am watching the research that points in the right direction with more hope than I usually admit to. But I am not optimistic that the transition will happen quickly or that it will happen for the right reasons rather than because the training paradigm eventually hits limitations that force the field to look elsewhere. What I do know is that the arguments I have been making across these posts are not arguments for inaction or for despair. They are arguments for a specific direction, toward intelligence grounded in the physical world rather than in the statistical surface of human expression, toward systems that can be verified rather than systems that can only be trusted, toward a relationship between AI and human knowledge that is built on recognition and respect rather than on extraction and consumption. That direction is harder. It requires more intellectual honesty. It requires admitting that some of the most impressive systems ever built are built on a foundation that is morally questionable and architecturally limited. I think admitting hard things is the only way forward that I can respect, and so I am admitting them here, and building what I can alongside the admission, for whatever both are worth.&lt;/p&gt;

&lt;p&gt;Till next time 👋!&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;span id="ref-1"&gt;&lt;/span&gt;&lt;strong&gt;1.&lt;/strong&gt; Grynbaum, M. &amp;amp; Mac, R., &lt;em&gt;The Times Sues OpenAI and Microsoft Over A.I. Use of Copyrighted Work&lt;/em&gt;, &lt;a href="https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html" rel="noopener noreferrer"&gt;New York Times, 2023&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-2"&gt;&lt;/span&gt;&lt;strong&gt;2.&lt;/strong&gt; U.S. Copyright Office, &lt;em&gt;Copyright and Artificial Intelligence, Part 3: Generative AI Training&lt;/em&gt;, &lt;a href="https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf" rel="noopener noreferrer"&gt;copyright.gov, 2025&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-3"&gt;&lt;/span&gt;&lt;strong&gt;3.&lt;/strong&gt; European Parliament, &lt;em&gt;EU AI Act&lt;/em&gt;, &lt;a href="https://www.europarl.europa.eu/topics/en/article/20230601STO93804/eu-ai-act-first-regulation-on-artificial-intelligence" rel="noopener noreferrer"&gt;europarl.europa.eu, 2024&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-4"&gt;&lt;/span&gt;&lt;strong&gt;4.&lt;/strong&gt; Gallegos, I. et al., &lt;em&gt;Bias and Fairness in Large Language Models: A Survey&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2309.00770" rel="noopener noreferrer"&gt;arXiv:2309.00770&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-5"&gt;&lt;/span&gt;&lt;strong&gt;5.&lt;/strong&gt; Ippolito, D. et al., &lt;em&gt;Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2210.17546" rel="noopener noreferrer"&gt;arXiv:2210.17546&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-6"&gt;&lt;/span&gt;&lt;strong&gt;6.&lt;/strong&gt; Nissenbaum, H., &lt;em&gt;Privacy as Contextual Integrity&lt;/em&gt;, &lt;a href="https://digitalcommons.law.uw.edu/wlr/vol79/iss1/10/" rel="noopener noreferrer"&gt;Washington Law Review, 2004&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-7"&gt;&lt;/span&gt;&lt;strong&gt;7.&lt;/strong&gt; Strubell, E. et al., &lt;em&gt;Energy and Policy Considerations for Deep Learning in NLP&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/1906.02629" rel="noopener noreferrer"&gt;arXiv:1906.02629&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-8"&gt;&lt;/span&gt;&lt;strong&gt;8.&lt;/strong&gt; Hoffmann, J. et al., &lt;em&gt;Training Compute-Optimal Large Language Models&lt;/em&gt;, &lt;a href="https://arxiv.org/abs/2203.15556" rel="noopener noreferrer"&gt;arXiv:2203.15556&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-9"&gt;&lt;/span&gt;&lt;strong&gt;9.&lt;/strong&gt; Perrigo, B., &lt;em&gt;Exclusive: The $2 Per Hour Workers Who Made ChatGPT Safer&lt;/em&gt;, &lt;a href="https://time.com/6247678/openai-chatgpt-kenya-workers/" rel="noopener noreferrer"&gt;TIME Magazine, 2023&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-10"&gt;&lt;/span&gt;&lt;strong&gt;10.&lt;/strong&gt; UNESCO, &lt;em&gt;Recommendation on the Ethics of Artificial Intelligence&lt;/em&gt;, &lt;a href="https://www.unesco.org/en/artificial-intelligence/recommendation-ethics" rel="noopener noreferrer"&gt;UNESCO, 2021&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-11"&gt;&lt;/span&gt;&lt;strong&gt;11.&lt;/strong&gt; See &lt;a href="https://wiseai.dev/blogs/language-is-limited-asi-is-impossible" rel="noopener noreferrer"&gt;Language is Limited. ASI is Impossible.&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;span id="ref-12"&gt;&lt;/span&gt;&lt;strong&gt;12.&lt;/strong&gt; Pearl, J. &amp;amp; Mackenzie, D., &lt;em&gt;The Book of Why: The New Science of Cause and Effect&lt;/em&gt;, &lt;a href="https://www.basicbooks.com/titles/judea-pearl/the-book-of-why/9780465097609/" rel="noopener noreferrer"&gt;Basic Books, 2018&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>lmm</category>
      <category>agi</category>
    </item>
  </channel>
</rss>
