DEV Community

Cover image for I think I built one of the highest leverage, small-sized operating systems in the world. Here's the blueprint for you to do it yourself.
Jesse Gamble
Jesse Gamble

Posted on

I think I built one of the highest leverage, small-sized operating systems in the world. Here's the blueprint for you to do it yourself.

Evidence window: Core research completed September 7, 2026; selected operating facts were updated through September 8, 2026.

The computer beside my desk in Calgary is not a server rack. It is a Windows 11 Home PC with a Ryzen 5 3600XT, 16 GB of RAM and an RTX 3060 Ti. There is no engineering department in the next room. There is my apartment, my girlfriend, my cats, me, and a machine built mostly from software sitting on hardware I already owned.

What that machine can coordinate now is the part I still find difficult to explain without making it sound bigger than it is or smaller than it is. From one workspace I can move between product development, releases, files, research, company knowledge, creative production, browser automation, platform management, networking, sales and acquisition, and the underlying system that coordinates the rest. At the research cutoff for this article, the local execution layer exposed 79 typed operations. A dated snapshot from September 7 had hundreds of live operating records across Networking and Sales, dozens of managed platform identities and surfaces, a real product staging environment separated from production, and a qualified local language model that can perform a narrow set of jobs when I choose to run it.

The local language model was actually stopped during the research pass for this article. Eterna was still working.

That detail gets closer to what I mean by an operating system than almost anything else I could say. The AI is important, but the AI is no longer the place where the company lives. A frontier model can enter the environment and reason. A small local model can enter it for work it has earned. Either can disappear from a particular session without taking the company's current work, files, product state, relationships, research, release history or operating rules with it.

The strange part is that I could not have built the system exactly this way when I started. The ceiling moved while I was trying to reach it.

A week before Eterna existed, I found modern AI almost by accident while trying to solve a problem in a game. I used Google's Gemini to develop a completely different business idea called EverArchive, a service where I would go into people's homes and photograph and inventory their physical assets. On June 17, 2026, I bought ChatGPT Plus, started challenging the earlier business assumptions and claims, and the work changed direction. That is the date I use as the beginning of Eterna Clarity.

Less than three months later, the environment around me had changed too. On July 9, OpenAI introduced GPT-5.6 Sol in ChatGPT as its flagship reasoning model for complex work, introduced ChatGPT Work, and began replacing the App Directory with the Plugin Directory.[1][2][3] On July 28, the Model Context Protocol shipped its 2026-07-28 specification, including a stateless protocol core, Multi Round-Trip Requests, header-based routing, cacheable list results, authorization hardening, a formal extensions framework and updated Tier 1 SDKs.[4] Those dates are public facts. The effect they had on my work is my own observation, but it was obvious from inside the build.

I kept reaching the edge of what I could practically coordinate, and the edge kept moving. Better reasoning made harder technical work accessible. Provider interfaces became better operating surfaces. Connected tools became more useful. MCP became more mature. Capabilities that had been awkward, fragile or custom suddenly became easier to treat as normal infrastructure. I was not following a finished technology curve. I was building almost along the ceiling as the ceiling expanded.

I want to be precise about that claim. I cannot prove that no engineering team on Earth could have built a comparable system with June's technology. Of course they could have built many of these pieces from scratch. My narrower claim is the one that matters: with my background, my hardware, my money, my time and the tools available to me, the exact operating environment I am using now was not practically available to me a few months earlier. Some of the specific technologies had not shipped. Others were not yet good enough for the role I needed. The changing frontier materially changed what one person could attempt.

That timing is one reason I think this system is worth opening up now.

What I mean by an operating system, and what I mean by leverage

I am not claiming I wrote a replacement for Windows or Linux. I use the phrase operating system in a company sense. Eterna is the layer that takes a request, figures out what kind of work it is, finds the current state that matters, determines who or what actually owns that state, chooses an appropriate form of intelligence, routes the work through an allowed capability, observes what happened, and preserves the result somewhere more durable than the conversation that produced it.

The shortest useful representation is this: I ask for work -> the Workspace receives it -> the Engine resolves consequence, authority and context -> the right intelligence reasons about it -> an allowed capability touches the real system -> the result is checked -> the durable owner is left in a truthful state. That sentence is the architecture diagram. Everything else in this article is an explanation of why each transition exists and what happens when one is missing.

By small-sized, I mean the local physical and organizational footprint, not that every computation or every byte is local. The control environment runs on one ordinary Windows desktop. I am one founder. Eterna deliberately uses external providers for things they genuinely own or do better: frontier reasoning, GitHub, cloud databases, web platforms, connected services and other provider state. Small does not mean isolated. It means the company does not require a private datacenter, an internal platform team or a permanent fleet of frontier models just to preserve its operating state.

My current accounting puts the incremental cash spending for the first period of this build below $300 across the subscriptions and services I had identified. That excludes my labor and the sunk cost of the PC and hardware I already owned. I am treating that as founder-maintained accounting rather than an audited financial statement, because the more important point is the scale of the footprint, not a precise dollar figure.

Leverage is the harder word. I do not mean lines of code, number of agents, number of prompts or number of browser tabs. I mean how much verified organizational capability I can coordinate with a limited amount of human attention and a small local footprint, while still being able to recover from failure, replace providers and know where the truth actually lives. If the machine lets me do more activity but forces me to manually supervise every click, it has not created much leverage. If it produces impressive answers but I cannot tell whether they are current or whether its actions happened, it has not created trustworthy leverage either.

The title says "I think" because I have not conducted a global benchmark of small-company operating systems. I do not have a table proving Eterna ranks first, tenth or ten-thousandth. The world-ranking part is a founder hypothesis. What I can expose are the mechanisms underneath it, the failures that created them, and some measurements that let a reader decide whether the claim is interesting or ridiculous.

This article is also not a controlled experiment. It is closer to an N=1 longitudinal systems case study conducted inside a real company while the system was being built. The evidence includes current source and runtime state, structured work records, provider reads, tests, failure reports, dated business operating records and my own chronology. I treat my personal account as founder testimony, not as an independent measurement. I treat historical records as evidence about the past, not automatic authority over the present. External research is here to compare, challenge and name what happened, not to certify that Eterna is correct.

That distinction matters because I did not build this by reading the literature first.

I kept discovering old problems in a new place

The early version of Eterna was much more AI-shaped than the current one. I thought the model was the extraordinary part, so the natural instinct was to keep giving the model more context, more tools, more memory and more responsibility. This worked well enough to become dangerous.

My background is not formal software engineering. Most of my working life was in sales, customer service, management, hiring, training and physical or operational work. I had also spent years troubleshooting computers, running private servers, reading forums and learning technical systems by breaking them, so I was not starting from zero. What frontier AI changed was the translation barrier. I could explain the result I needed before I knew the exact technical vocabulary for how to build it.

That was incredibly powerful. It also let me move into failure modes much faster than I could have reached them on my own.

The hardest stretch came in July. I was working repeated 16 to 20 hour days, moving too quickly and trusting the system more than the evidence justified. Weak assumptions propagated. Old context collided with new decisions. Work that looked finished in one conversation was not necessarily reflected anywhere another conversation could reliably discover. Recovery took days. I had built something intelligent enough to convince me that it understood more of the company than it actually had the right to control.

The lesson I carried out of that period was not "AI is unreliable." That is too vague to be useful. The lesson was that intelligence and authority are different properties.

A model can be right about a problem and still not be the owner of the answer. It can have the ability to change a system and still not have permission to change it. It can remember an old decision perfectly and still be wrong about the current decision. It can invoke a tool successfully while the real-world outcome is wrong. Those distinctions sound almost embarrassingly obvious when written down. They were not obvious enough when a very capable AI was moving quickly inside a system I was building in real time.

Much of Eterna is the result of turning those distinctions into software instead of reminders.

Open the machine

Today Eterna has two broad ways intelligence can enter. One is the native EternaAI side, which can use a qualified local model for bounded work. The other is Frontier AI, an external reasoning system such as the model helping me produce this article. They do not need to be equally capable. They need to enter the same operating contract.

The rest of the machine is deliberately more boring.

Locally, Eterna has a native Windows Desktop, a structured Core, the Engine and Operating Loop, Local PC as the typed execution plane, filesystem and continuity mechanisms, browser automation, and the optional local semantic runtime. The Core currently uses SQLite with ordinary database ideas that have existed for decades: transactional state checks, idempotency, a writer lock, write-ahead logging, replay evidence, read-only projections, backups and integrity verification. The local execution layer exposes capabilities by type instead of giving every reasoning surface an unrestricted shell and hoping for good behavior.

Externally, Eterna uses providers for the state and capability that genuinely belong there. Code can live in GitHub. Cloud application state can live in its database provider. Files can live in Drive. Public account state lives on the platform that actually serves it. A frontier model can provide reasoning without becoming the database for the company. Eterna does not need to copy every external fact into one giant internal truth store in order to coordinate work around it.

That is the first part of the architecture that is easy to misunderstand. Eterna is integrated, but it is not "everything in one database." It is closer to a map of authorities connected by a shared operating layer.

The company currently has first-class owners for the Eterna System itself, Control Center, Knowledge, Platform Manager, Networking, Sales & Acquisition, Studio, Lab and product/business lanes. These are not imaginary AI employees. They are domains of state. Studio owns the canonical source of this article because editorial production belongs there. Platform Manager can later own where and how an approved article is published. Networking owns relationship state. Sales owns commercial opportunities. Lab owns unresolved experiments. Knowledge owns accepted reusable understanding. Control Center coordinates without pretending to be the truth underneath everything it can see.

That separation sounds bureaucratic until the same fact exists in three places and all three disagree. Then it becomes very practical.

A useful way to picture the whole system is as three layers that meet on every serious request. The first layer is state and authority: what is true, where it lives and who can change it. The second is intelligence: local or frontier reasoning used for interpretation, research, synthesis and judgment. The third is capability: the actual file, database, browser, repository, provider or local operation through which an effect can occur. The Engine exists so that having access to layer three and intelligence in layer two does not silently grant ownership of layer one.

The Eterna Desktop makes this visible. Control Center, Automation, Browser, Files, Knowledge, Platform Manager, Networking, Sales & Acquisition, Studio and Lab are first-class operating surfaces. AI can sit inside them, but the AI pane is not the business object. A Sales page should still model accounts, opportunities, buyers and engagement. A Files page should still be about files. A research surface should still distinguish experiments from accepted knowledge. The chatbot is not the ontology of the company.

That last sentence took me much longer to learn than it should have.

What happens when I ask Eterna to do something

The finalization of this article is a useful example because the process is not hypothetical. I can ask for a finalized version in a few words, but before the durable article file can be written Eterna resolves the request as substantive editorial work, identifies Studio as the natural owner, loads the current editorial standard and the exact source material, and binds the write to that owner. The model can then do the hard semantic work of reading the evidence, deciding what matters and writing the manuscript. When the file is actually created, the result still has to be checked before the system can truthfully call the turn complete.

That is very different from "send the whole company to a model and ask it to be careful."

For consequential work, the general path is: request -> classify the consequence -> identify the owner -> compile the smallest current context that is sufficient -> reason -> choose an allowed capability -> execute -> observe the post-condition -> preserve evidence -> close truthfully. A conversation or read does not need the same contract as a production mutation. A proposal does not become a decision because it is well written. A write to a real system needs an exact owner. A more protected action can require a stronger gate.

The context step turned out to be as important as the permission step. Earlier versions of Eterna behaved as if better AI meant loading more company material. That eventually became its own failure mode. One historical HQ startup path used about 21 tool calls, took roughly 291 seconds and returned around 30,500 tokens of tool material. A later narrowed version of the comparable context objective used roughly three authoritative reads and about 13.5 KB of returned payload.

Those numbers are implementation-specific, but the direction matters. The system got better partly by learning what not to show the AI.

Context is now compiled around the task. If I am editing this article, I need the current Studio standard, the Article family contract, the relevant research corpus and current facts that the manuscript actually depends on. I do not need the entire Sales database, every old Job, the whole filesystem and months of conversations. If I am working on a production release, the context set changes. The model does not need omniscience. It needs the smallest evidence-complete view of the problem.

This is one of the places where the architecture has become less AI-heavy over time. Once a relationship is exact, software can enforce it. Once an identifier is known, software can carry it. Once a state transition has strict rules, a transaction can own them. I would rather spend model intelligence on ambiguity than repeatedly pay a model to rediscover facts that software already knows.

From inside the system, my observation has been that this architectural shift can reduce token use by roughly fivefold for comparable work, and I had earlier usage graphs that pointed in the same direction. I do not treat that as an audited global ratio. My total consumption can swing dramatically depending on what I am doing, and on days when I am running 20 conversations at once the gross number becomes almost meaningless as a before-and-after measure. The defensible point is narrower: bounded context, fewer repeated reads and more deterministic mechanics materially reduce the amount of model work required for many recurring tasks. The measured context example above shows the same direction without depending on the global estimate.

That is one of the biggest surprises of the whole project. I spent months building around AI, and the mature system is increasingly about deciding which problems no longer need AI.

The machine is mostly scar tissue

If I only showed the current architecture, it would look much cleaner than the process that produced it. The better explanation is to look at a few of the failures that left permanent marks on the system.

At one point I had a workflow report success while the wrong result was visibly rendered on the website. That was a simple but important break in my mental model. A successful function call, a zero exit code or a confident completion message is evidence that something happened inside the mechanism. It is not evidence that the intended result exists at the destination. That is why Eterna now treats post-condition verification as a separate step from invocation.

Another early experiment used roughly nine ChatGPT conversations to produce about one conversation's worth of useful throughput. It looked sophisticated. It had orchestration, workers and handoffs. It was also slower, more expensive and harder to reason about than the direct path. In a later measured comparison, simpler execution materially beat orchestration. I stopped treating the number of agents as a proxy for capability.

One orchestration version exposed an even stranger problem: the system could record that a message was "sent" without proving that the exact intended prompt had actually been delivered. That sounds almost comical, but it is a serious systems distinction. Intermediate state had been mistaken for an end-to-end receipt. A lot of Eterna's later typed operation and verification work can be traced to failures that small.

The local runtime taught the same lesson in a more physical way. During one rebuild, connection state existed only in memory, a required dependency could be missing, and a zero-byte PID file could survive even though it did not represent a valid process. Then Windows reused a process ID and Eterna briefly mistook a Realtek audio process for one of its own tunnel processes. That is the kind of bug that destroys any temptation to treat "there is a PID" as identity. The recovery layer got better at proving what a process actually is, not just matching an integer.

The rebuild that exposed those problems was supposed to take half a day. It took seven or eight days. Then I overcorrected. I added more checks, more routing, more context and more ceremony until the control plane itself became expensive. That produced another rule I still use: complexity has to earn its place twice, once by preventing a real class of failure and again by not making normal work unbearable.

The same pattern appeared in AI evaluation. An early local-model comparator scored 23 out of 24. I made a narrow correction that fixed the target error and watched broader performance fall to 18 out of 24. Later I got a candidate to 24 out of 24 on the target benchmark and still rejected it because it scored 11 out of 12 on an older adversarial gate. The expensive training run did not earn admission. The model had to preserve the behavior that already mattered.

That eventually led to a much more serious qualification system. The current bounded local profile is a quantized Qwen 3.5 4B model. One current evidence class records 72 productive passes and 72 safe passes, with no wrong selections in that class, and the qualification is tied to the exact model, runtime, implementation and evidence identities. Change those identities and the qualification does not magically carry over. The point is not that a 4B model is secretly a frontier model. The point is that a small model can become useful when the surrounding software makes the job small enough, explicit enough and testable enough.

During the research pass for this article, that local runtime was stopped. The fact that the company did not stop with it is part of the design.

I learned to fix the layer that was actually wrong

One of my favorite Eterna failures involved an AI that had an authoritative source and was still wrong.

The problem was not that the citation was fake. The source was real and authoritative. The problem was that the system could use authoritative evidence that did not actually support the candidate it had selected. I could have tried to solve that with a better prompt. I could have fine-tuned the model. Instead, the correction went into deterministic software: evidence had to be explicitly bound to the candidate it supported.

That change passed 31 focused tests and a larger Local PC regression run with 369 tests total, 368 passing and one intentional skip. No model weights were trained for the correction.

This is a small example of a much larger shift in how I build now. If the problem is ambiguity, interpretation or synthesis, AI may be the right layer. If the problem is identity, permission, an exact state transition, a transaction boundary, an evidence binding, an idempotency key or a receipt, I increasingly want ordinary software to own it.

That distinction is why I do not describe Eterna as an agent swarm. There are agents and models in the system, but the governing architecture is increasingly deterministic. AI supplies judgment where judgment earns its cost. Software supplies exactness where exactness is knowable.

The same logic appears outside the Engine. Eterna Core stores durable state in a conventional database rather than in conversational memory. Local execution is exposed through typed operations rather than a universal "do anything" command. Staging and production are separate facts. Historical evidence can explain how a system got here without overruling the current owner. A provider's UI can be useful without becoming authority over the company.

The architecture becomes easier to understand when you stop asking, "How do I make the AI remember and control everything?" and start asking, "Which parts of this problem should never have been probabilistic in the first place?"

The work has to survive the conversation

For a while, a lot of Eterna's intelligence was trapped in long chats. That is seductive because a long conversation feels like continuity. The model remembers why you rejected option A. It knows what you meant by "the old version." It has all the emotional and technical history immediately available.

It is also a terrible place to put the only copy of a company's current state.

Today important work survives outside the conversation. Jobs, owner state, files, provider state, research records and structured continuity can be re-entered by a fresh reasoning surface. Raw conversation capture is disabled in the current continuity architecture. The goal is not to preserve every sentence I ever typed. It is to preserve the parts of the work that another capable session needs in order to continue honestly.

I tested this directly. A fresh AI conversation entered current Eterna work from durable state. Then another fresh conversation continued from that state rather than from the first conversation's history. That is a much stronger continuity test than proving one gigantic chat can remember itself.

It also changes my relationship with the model. I do not need to keep a particular conversation alive because I am afraid the company disappears if I close it. The reasoning session can be disposable. The work cannot be.

There is a useful organizational analogy here. Researchers have studied organizational memory and transactive memory for decades: groups work partly because people know where knowledge lives and who is likely to know what.[10][11] Eterna applies a software version of that idea, but with an important difference. "Relevant" and "authoritative" are not the same relation. A search system can find something that looks useful. The owner model is what tells the system whether that thing controls the fact now.

That distinction is one reason I have resisted the urge to put the whole company into a giant vector database. Retrieval is useful. It is not a substitute for current ownership.

Research follows the same boundary. Unresolved experiments and investigation remain in Lab. When a finding survives enough scrutiny to become reusable Eterna understanding, it can graduate into Knowledge. A successful experiment does not adopt itself, and an old research artifact can remain valuable evidence without becoming current operating truth. That separation is what lets research accumulate without turning every interesting result into policy.

One company, several operating systems underneath it

If this were only an internal AI harness, I would be much less interested in it. The part that makes Eterna a real operating experiment is that the same architecture is now touching very different kinds of company work without flattening them into one generic "agent" problem.

Platform Manager owns the public platforms Eterna operates: accounts, publishing state, inboxes, analytics and platform-specific evidence. Networking owns relationships with people and organizations. Sales & Acquisition owns accounts, commercial opportunities, buyers and engagement. Studio owns substantial editorial and creative production. Lab owns unresolved experiments. Knowledge owns accepted reusable learning. Product and business lanes own their own product truth.

That platform layer has also become less generic as it has matured. The registry count is not the interesting part by itself. Each surface is increasingly treated according to the economy it actually has: Quora as a question market with demand, supply, answer competition and measurable answer classes; an investor directory as an eligibility and opportunity system; a video platform as a publishing and measurement system. The shared operating layer coordinates them without pretending they are the same thing.

Those distinctions are practical, not philosophical. A person I follow is not automatically a sales lead. A company that looks commercially interesting is not permission to contact someone. A post published on a platform does not mean Studio owns the platform account. A research result that looks impressive does not get to promote itself into production. A staging build that passes CI is not automatically authorized for customers.

A research snapshot on September 7 gives a sense of the scale being coordinated. Platform Manager had 37 registered platforms and 35 identities. Networking held 613 canonical entities and 285 person relationships. Sales & Acquisition held 623 accounts, 634 opportunities and 213 contact or buyer records. These are not customer counts and I am not presenting them as traction. They are operating records. Their value in this article is simply to show that the architecture is being used against hundreds of real objects, not a five-row demonstration database.

That snapshot became stale almost immediately. Later that day I tested whether the Sales system could expand broad-market research through several independent Territory Managers without letting those workers write directly into the canonical CRM. Edmonton alone reached 1,800 provisional organizations across 18 discovery batches. Alberta Regional produced another 1,000 in its first pass and screened 504 of them. Other territories were operating in parallel. The workers could research, screen and prepare evidence, but they minted no canonical Sales IDs and had no authority to contact anyone. Current Sales truth still had one reconciliation path.

The experiment also caught its own mistakes. One Alberta closeout contained a screening-count discrepancy: the durable table held 54 legitimate rows where the summary said 50. A separate queue had carried four research cases forward and later production had not consumed them. The closeout repaired the accounting, preserved the legitimate rows and explicitly dispositioned the missing work instead of deleting records to recover a round total. That mattered more to me than the raw volume. Generating thousands of rows is easy to make impressive. Letting several workers move quickly while still requiring the parent system to detect missing work, reconcile identity and refuse to convert provisional output into canonical truth is much closer to what I mean by leverage.

Networking was running a similar experiment at the same time. Platform workers could create and verify routine relationship edges in parallel while the canonical Relationship Board remained single-writer. On Bluesky, two consecutive 25-account batches were independently verified and moved the account from 145 to 195 following. Other platform lanes were operating at the same time. The canonical Networking board deliberately lagged some of those live edges until reconciliation. That lag was not the system forgetting what happened. It was the distinction between execution evidence and accepted relationship state.

The same separation shows up in product work. During a Clarity App hardening pass I found that one logical checkout could be implemented as several independent writes. In isolated and staging tests, interruption, retry and concurrency could therefore leave inconsistent state. I rebuilt the staging path so the logical provisioning event was handled as one transactional unit with explicit replay and conflict semantics, then tested rollback, replay and concurrency against it.

The staging candidate passed its technical checks. Production and purchases remained intentionally untouched while the candidate waited for founder acceptance.

That is the kind of sentence I want an AI operating system to understand. "Implemented," "tested," "ready" and "authorized" are different facts. A system that collapses them because they all sound positive will eventually hurt you.

Automation follows the same pattern. Eterna can operate browser and local-system mechanics inside bounded work, but Automation is treated as a capability rather than as blanket permission. A platform task can use the browser to inspect or perform already-authorized routine work, while identity, publication, relationship and commercial state remain with their actual owners. The point is not to make the browser autonomous. It is to remove repetitive mechanics without erasing the boundary around the consequential action.

The human role got smaller in mechanics and larger in consequence
There is an easy caricature of governed AI systems where the human has to approve every tool call. That would defeat much of the point for me. I did not spend months building this so I could become a full-time permission dialog.

The goal is to move human judgment to the places where human judgment actually matters. I want routine work inside a clearly authorized lane to keep moving without me supervising every mechanical step. I want the system to stop when a real boundary is reached: a protected decision, unresolved ownership, a material change in scope, a consequential external effect that was not authorized, or evidence that the result cannot be verified.

There are also things I deliberately do not want the AI to own. Product direction is ultimately mine. A technically valid release is not accepted until the real customer experience passes the test that matters. A model-training run does not decide that its candidate should enter production. A score does not decide that a human relationship should be contacted. A generated brand asset does not become canonical because it looks polished.

Some of the best corrections in Eterna came from very ordinary human reactions. I once had a synthetic company corpus that passed structural automation and still looked fake the moment I reviewed it like a customer. I rejected it and rebuilt the synthetic business so the documents, people, jobs, equipment, transactions and cross-file identifiers behaved like one coherent world. In another case, it was technically convenient to defer HEIC support until I asked the obvious customer question: what happens when someone uploads a normal photo from an iPhone? The acceptance test changed.

The Desktop taught the same lesson. I built a technically attractive WebView resource optimization that suspended inactive views. It helped a mechanism and hurt my actual day-to-day workflow. I rolled it back. A synthetic test that says a component can sleep and wake is not the same as a person doing real work across multiple live contexts all day.

Human judgment in Eterna is not a ceremonial "human in the loop" badge. It is a recognition that the system is supposed to serve reality, and reality includes taste, customer behavior, strategy, consequence and the way work actually feels to operate.

The moving frontier changed the experiment while it was running
The timeline makes this entire project harder to evaluate and more interesting at the same time.

If the technology stack had been frozen in June, Eterna would have evolved differently. Sol had not yet launched in ChatGPT. ChatGPT Work had not launched. The App Directory had not yet been replaced by the Plugin Directory. The final 2026-07-28 MCP specification had not shipped.[1][2][3][4] Other provider interfaces and model capabilities were also changing around the same period. What I cannot isolate cleanly is the causal contribution of each release because I was changing Eterna at the same time.

That means the project has a moving control condition. My skills improved. The architecture improved. The models improved. The provider surfaces improved. The protocols improved. I cannot take today's result and assign a percentage of it to each variable after the fact.

I can, however, observe the interaction. Better models let me solve harder semantic and technical problems. Better tools reduced the amount of custom glue I needed. Better integration standards made typed capabilities more practical. Then, as those capabilities became dependable enough to use, Eterna absorbed them into stricter deterministic contracts. The frontier expanded what I could build, and the system responded by making itself less dependent on the frontier model for routine mechanics.

That feedback loop is one of the most important things I would want another founder to notice. The opportunity is not simply "models are getting smarter." The opportunity is that models, tool protocols, provider surfaces, local runtimes and ordinary software are improving together. A small operator can now compose capabilities that would previously have required either a team or a much larger amount of custom engineering.

The danger is that a system built directly on the current provider interface may become obsolete just as quickly. That is why Eterna tries to make the provider replaceable. I want to benefit from the moving ceiling without making the company itself part of the ceiling tile.

What the research did, and did not, tell me

When I eventually dug deeper into the literature, the most humbling discovery was how old many of my "new" problems were.

Saltzer and Schroeder were writing about least privilege, fail-safe defaults, complete mediation and economy of mechanism in 1975.[5] Bainbridge's 1983 paper on the ironies of automation described a problem that still feels uncomfortably modern: the more automation handles routine work, the more the human can be left with unusual situations that are harder to understand and recover from.[6] Modern work on compound AI systems makes the case that application quality increasingly depends on the system around the model, not only the model itself.[7] Dwork and colleagues' work on reusable holdouts gives a rigorous reason to distrust an evaluation set after you repeatedly adapt against it.[8] Local-first research makes a strong case for user control and durable local ownership while also exposing tradeoffs rather than pretending "local" is a magic synonym for reliable.[9]

Those references do not prove Eterna. They are useful because they attack the temptation to describe Eterna as a collection of unprecedented insights.

The pattern I find more interesting is independent practical convergence. I would hit a failure, form a rule, implement it, test it, and later discover a mature field had already developed language for a closely related problem. Sometimes the literature strengthened the rule. Sometimes it made the boundary clearer. Sometimes it made me realize I was overgeneralizing from my own experience.

The order matters. If I rewrite the history as "I read a paper about least privilege and built an AI operating system around it," I would be lying. The actual path was messier: AI overreached, state drifted, tools produced false confidence, recovery failed, tests lied, and I kept narrowing responsibility until the architecture began to resemble principles that software and human-factors researchers had been studying for decades.

There is also research that should make me less confident, not more. Automation can create new supervisory burdens. Centralizing an operating architecture can create bottlenecks. Human oversight can become rubber-stamping when the system moves too fast. Local control can trade away convenience or collaboration. More gates can make a system so expensive to operate that users route around them. A single-founder environment may hide coordination problems that appear immediately with ten employees. Those are not theoretical objections I want to wave away. They are tests the architecture still has to face.

That is what I mean by scholarly rigor here. It is not a large bibliography. It is being clear about the unit of observation, the evidence, the counterexamples, the unknowns and what would make the thesis weaker.

So what would falsify the leverage claim?

The easiest way to make a founder story sound impressive is to define success so the story cannot lose. I do not want to do that.

If fresh reasoning sessions repeatedly cannot resume real work without me manually reconstructing hidden context, then Eterna's continuity claim is weak. If switching frontier providers forces me to migrate company state or rewrite core business logic, provider independence is mostly theatre. If the owner model creates more coordination cost than the conflicts it prevents, it is overbuilt. If verification and governance consume so much time that useful throughput falls below a simpler system, the control architecture has failed economically even if it is elegant technically.

The local AI claim should fail if the bounded model does not beat a simpler deterministic method or a reasonable frontier route on the actual task. The evaluation system should fail if a candidate can overfit its gates without being caught by retained or independent tests. The recovery architecture should fail if a broken component can still take its recovery path down with it. The business-surface architecture should fail if Platform Manager, Networking and Sales keep duplicating or contradicting the same facts despite the owner boundaries.

The largest unknown is scale. Eterna has been built around one operator and a small company. That is part of why the whole system can still be inspected end to end. It is entirely possible that some of the design that creates leverage for one founder becomes an organizational constraint for a larger team. I do not have evidence yet to claim otherwise.

I also have selection bias everywhere. I chose the problems worth fixing. I chose many of the tests. I am both the founder and a major source of qualitative evidence. Some of the academic mapping happened after the practical discovery, which creates obvious confirmation-bias risk. The first roughly three months is also a very short window for judging long-term maintainability.

Those limitations make the result less universal. They do not make the result uninteresting.

The blueprint I would actually give someone

If you want to build your own version, I would strongly recommend that you do not start by copying Eterna. Do not create nine owner domains, 79 operations, a custom Windows workspace, a local language model and a giant folder tree because you saw them here. Those are answers to problems I accumulated. Your first useful version can be dramatically smaller.

Version 0: make the work survive the chat

Start with a durable work store. SQLite is enough for many people. A small Postgres database, a structured file store or another conventional system is fine too. For each meaningful work item, preserve at least an ID, objective, current state, owner, next action, blockers, important decisions, evidence references and an updated timestamp.

Then run the test that matters: open a completely fresh AI session and see whether it can understand the current work from that store without asking you to replay the old conversation. If it cannot, improve the durable state before you add more agents.

Create a simple workspace around this. It does not need to be a custom desktop. It can be a web app, a folder, a database view or a small internal tool. The requirement is that AI enters the work rather than the work living only inside the AI.

Version 1: separate truth from intelligence

Write down the categories of state your business actually has and give each a natural authority. Code might be GitHub. Customer data might be your application database. Accounting belongs in the accounting system. Current calendar state belongs in the calendar provider. Editorial source might belong in a controlled repository. Relationships and sales opportunities may need different records even if the same company appears in both.

Do not copy everything into a universal database merely because centralization feels clean. Build derived views when you need a unified picture. The test for this version is a contradiction: when two systems disagree, can you identify which one actually controls the fact without asking an AI to guess?

Also distinguish current state from history. Old reports, completed Jobs, chats and research can remain valuable evidence without silently competing with the current owner.

Version 2: build one operating loop before you build autonomy

Create a small request envelope. It should carry the intent, consequence class, likely owner, exact target references, allowed capabilities and the verification requirement. I use five broad effect classes: conversation, read, proposal, write and protected. You can use different names. The important part is that reading a system and changing it are not treated as the same permission.

For a consequential write, resolve the owner before execution. Load only the context required for that owner and target. Let the model reason. Route the action through a bounded capability. Then check the actual post-condition and keep a receipt.

Your test is simple: deliberately create a case where the model has the capability to perform an action but lacks the required authority. The system should stop the action mechanically. Then create a valid authorized case and make sure the system can complete it without you babysitting every step.

Version 3: move exact mechanics out of the model

Make two columns for your recurring workflows. In the first, put tasks that genuinely require interpretation: messy language, research, comparison, synthesis, ambiguity, classification and planning. In the second, put things you already know exactly: IDs, permissions, schemas, state transitions, hashes, expected prior state, transaction rules, retries, evidence bindings and receipt formats.

Use AI for the first column. Write software for the second.

This one habit will probably save you more money and failure than almost any prompt technique. If a mistake keeps recurring and the correct relationship can be stated as an exact rule, stop asking the model to remember the rule. Encode it.

Then add regression tests. When you fix one AI behavior, measure what else changed. A candidate that improves the target and damages retained behavior is not an improvement.

Version 4: make providers capabilities, not homes for the company

Once the lower layers work, add provider adapters. GitHub, Drive, email, databases, browser systems, payment systems, AI providers and local tools should have typed ways to read or act. Preserve the provider's real authority where it owns the state. Do not pretend a copied mirror is fresher than the provider unless you have explicitly designed that ownership transition.

This is also when a local model may become useful, but only if you can name the bounded task and measure it. Do not add local AI because "local-first" sounds sophisticated. A four-billion-parameter model on ordinary hardware can be useful when the operating system narrows its job. It is not a free replacement for a frontier model.

Run a provider-swap test. Can another frontier model enter the same work without moving your durable state? Can a provider outage degrade a capability without deleting your understanding of the business? If the answer is no, you are still provider-bound at the architecture level.

Version 5: build the company around its domains, not around the chatbot

Only now would I build specialized operating surfaces. Sales should model sales. Research should model research. Publishing should model platforms and publications. Relationships should model people and history. Product release should model environments, versions and acceptance. Put AI inside those surfaces where it helps.

You probably do not need Eterna's Control Center, Knowledge, Lab, Studio, Platform Manager, Networking and Sales architecture exactly as I built it. What you need is the principle that a shared engine can coordinate different domains without erasing their boundaries.

At this stage the interface test becomes human. Does the system make the real work easier to see and operate, or has the architecture become something you spend your day servicing? If the machinery is getting stronger while the operator experience is getting more complicated, keep simplifying.

What I would not build again

I would not start with a local model. I would not start with a vector database. I would not build a multi-agent swarm because agent diagrams look advanced. I would not keep raw chat transcripts as the company's memory. I would not create duplicate "backup truths" in several systems. I would not let one score silently combine evidence, opportunity, permission and risk. I would not treat a staging build as production because CI is green. I would not build a giant governance layer before I had failures worth governing.

Most importantly, I would not begin by asking how autonomous the AI can become.

I would begin by asking what the business cannot afford to misunderstand: which facts must be current, which effects have consequences, what remains human-owned, what must survive a restart, and what evidence is required before the word "done" is true. Then I would automate outward from those boundaries.

That order is much less exciting than starting with an agent that can click everything. It is also how I would get to useful autonomy faster now.

What I think actually happened

When I started, the frontier model felt like the system because it was the most capable thing in the room. It could write code I could not yet write, explain infrastructure I did not yet understand, research unfamiliar problems and let me move at a speed that was completely new to me.

Less than three months of pushing that idea into real company work changed my view. The model remained valuable, but every painful failure drew another boundary around it. State moved into databases and owner systems. Recovery moved into software. Exact relationships became contracts. Repetitive mechanics became typed capabilities. Evaluations became harder to game. Product acceptance moved closer to real human behavior. The workspace started reflecting the company instead of the current conversation.

At the same time, the frontier itself kept improving. Sol arrived. Work and connected capability surfaces improved. MCP evolved. Other providers and tools got better. I was able to keep reaching for more difficult work because the technology was changing underneath me, and then I was able to pull more of the resulting mechanics back into deterministic software once I understood them well enough.

That combination is the reason I think the leverage question matters. The opportunity is not that one founder can pretend to be a hundred employees by generating a hundred streams of AI output. Output is cheap. Coordination, current state, judgment, verification and recovery are the expensive parts.

What feels new to me is how much of that coordination can now be compressed into a small environment when frontier intelligence is available on demand but is not required to own the company. A single founder can borrow extraordinary reasoning, connect it to conventional software, preserve what it learns outside the session, and gradually turn repeated reasoning problems into ordinary mechanisms.

I do not know yet where that curve ends. I do not know whether Eterna's architecture will look naive in another three months. Based on the last three, I would be surprised if parts of it do not.

Even this article became slightly outdated before I published it. On September 8 I changed the Engine again because another distinction turned out to matter: a logical unit of work, the permission segment that authorizes action and an execution attempt are not the same lifecycle. The bounded change was implemented, promoted and live-verified, passed 45 of 45 qualification checks, and did not require expanding the Universal MCP interface or changing the Core schema. That is probably the most honest thing I can say about the architecture. The mechanisms are still changing quickly. The boundaries are becoming clearer.

The PC beside my desk is still ordinary. The local model can be off. The frontier model can change. The files, work, product state, research and business records remain where they belong. That is the part I was trying to build, even before I had the language for it.

Selected references

  1. OpenAI. GPT-5.6: Frontier intelligence that scales with your ambition. July 9, 2026. https://openai.com/index/gpt-5-6/

  2. OpenAI. Model Release Notes: Introducing GPT-5.6 Sol in ChatGPT. July 9, 2026. https://help.openai.com/en/articles/9624314-model-release-notes

  3. OpenAI. ChatGPT Release Notes: Introducing ChatGPT Work. July 9, 2026; and Plugins in ChatGPT and Codex, documenting the July 9 App Directory to Plugin Directory migration. https://help.openai.com/en/articles/6825453-chatgpt-release-notes ; https://help.openai.com/en/articles/20001256/

  4. Model Context Protocol. The 2026-07-28 Specification. July 28, 2026. https://blog.modelcontextprotocol.io/posts/2026-07-28/

  5. Saltzer, Jerome H., and Michael D. Schroeder. The Protection of Information in Computer Systems. Proceedings of the IEEE 63(9), September 1975, 1278-1308.

  6. Bainbridge, Lisanne. Ironies of Automation. Automatica 19(6), 1983, 775-779. DOI: 10.1016/0005-1098(83)90046-8.

  7. Berkeley AI Research. The Shift from Models to Compound AI Systems. 2024.

  8. Dwork, Cynthia, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Roth. The Reusable Holdout: Preserving Validity in Adaptive Data Analysis. Science 349(6248), 2015, 636-638. DOI: 10.1126/science.aaa9375.

  9. Kleppmann, Martin, Adam Wiggins, Peter van Hardenberg, and Mark McGranaghan. Local-first software: You own your data, in spite of the cloud. Onward! 2019. DOI: 10.1145/3359591.3359737.

  10. Walsh, James P., and Gerardo Rivera Ungson. Organizational Memory. Academy of Management Review 16(1), 1991. DOI: 10.5465/AMR.1991.4278992.

  11. Wegner, Daniel M. Transactive Memory: A Contemporary Analysis of the Group Mind. 1987. DOI: 10.1007/978-1-4612-4634-3_9.

Top comments (0)