Toward a Star Trek Future of Building Things
How a slightly ridiculous way of making software may turn out to have been the awkward first chapter of a much larger transformation
At a glance
- Vibe coding made software generation conversational; it did not replace the work of defining, checking, releasing, and maintaining software.
- A useful engineering agent needs a bounded goal, explicit invariants and permissions, relevant checks, evidence, and a stop or escalation rule.
- Productivity findings depend on the people, task, tools, and measurement period. Benchmark scores describe their evaluators, not production autonomy.
- Cheaper implementation could make more small, tailored software worth exploring, while maintenance, reliability, security, and operating costs remain.
- The aim is to expand human agency. People still choose goals, set policy, authorize consequential actions, and remain accountable.
The First Abstraction Shift
There was a brief and beautiful period in the history of computing when the future arrived wearing sweatpants.
The ritual went something like this:
Build me an app.
A pause.
Make it modern.
Another pause.
No, more modern.
A mysterious dependency appeared.
Add authentication.
Fourteen files changed.
Fix the error.
Twenty-seven files changed.
Why did you rewrite the database layer?
The model apologized.
Undo that.
The model rewrote the database layer in a different way.
And somewhere, quietly, a senior engineer felt a disturbance in the force.
This was vibe coding.
The phrase became culturally prominent in early 2025 as a description of programming by conversational intent; Andrej Karpathy February 2025 post is a useful primary timestamp, not a claim that the underlying practice began there: instead of manually attending to every implementation detail, the developer increasingly tells an AI what ought to exist and lets the machine fill in substantial procedural detail.
For a while, the expression seemed to summarize everything simultaneously delightful and horrifying about generative AI.
Software could appear astonishingly quickly.
Some people who had never thought of themselves as programmers could create functioning applications.
Experienced developers could prototype some ideas at striking speed.
Interfaces materialized from descriptions.
Boilerplate evaporated.
Entire categories of annoying implementation work suddenly looked negotiable.
And, naturally, people also discovered that an AI could generate 4,000 lines of code with approximately the same emotional attachment to those lines that a leaf blower has to autumn foliage.
Vibe coding was liberating. Without review on consequential work, it could also be reckless.
It was serious.
It was also hilarious.
Most importantly, it was transitional.
Because the real historical significance of vibe coding was never that programmers had discovered a more relaxed way to type code.
It was that human intention had acquired a new computational interface.
For most of computing history, the difficult part of software creation was translation.
A person imagined a desired state.
Then they translated that state into architecture.
Then into algorithms.
Then into APIs.
Then into control flow.
Then into syntax.
Then into build systems.
Then into deployment systems.
Then, after several hours of discovering that YAML contains more ways to experience regret than previously thought possible, into production.
Generative AI began compressing parts of that translation chain.
A human could state the destination before understanding every road.
That was the breakthrough.
But it was not the destination.
The next stage is emerging now.
It is the transition from vibe coding to agentic software engineering.
And that transition may eventually matter much more than code generation itself.
1. Vibe Coding Was Not the Revolution. It Was the Demo
One easy mistake when thinking about AI-assisted programming is to focus too much on code.
Code is important, obviously.
But code is only one intermediate representation in a much larger engineering process.
Software engineering begins before code exists.
Someone identifies a need.
Someone decides what behavior is desirable.
Someone models the system.
Someone determines what must never happen.
Someone considers failure.
Someone weighs compatibility.
Someone chooses which compromises are reversible.
Someone defines what evidence would count as success.
And software engineering continues after code has been written.
The system must build.
It must pass tests.
It must survive strange inputs.
It must handle data safely.
It must deploy.
It must recover.
It must remain observable.
Its documentation must tell the truth.
Its dependencies must remain supportable.
Its permissions must remain bounded.
Its future maintainers must be able to discover why the strange-looking thing in line 847 exists instead of "cleaning it up" and resurrecting a bug from 2024.
The act of producing source code is therefore only a slice of engineering.
Vibe coding automated that slice first because language models are extraordinarily well suited to turning descriptions into plausible text.
That was enough to make the experience feel magical.
But magic is what technologies look like immediately before we standardize their interfaces.
A more mature version of AI engineering will not be limited to:
Write some code that seems right.
It will increasingly be:
Achieve this bounded engineering objective, preserve these invariants, use these authorized tools, gather this evidence, and stop if the evidence does not justify proceeding.
That is a radically different proposition.
The first is generation.
The second is delegated engineering.
2. Every Great Computing Revolution Is an Abstraction Revolution
The history of computing is, among other things, the history of humans repeatedly becoming unwilling to think about yesterday's details.
Early programmers interacted with machinery near the physical level.
Then came symbolic assembly.
Then higher-level languages.
Then structured programming.
Then libraries.
Then operating systems.
Then databases.
Then object systems.
Then frameworks.
Then virtual machines.
Then package managers.
Then cloud platforms.
Then declarative infrastructure.
Then managed services.
Then serverless systems.
Then application platforms.
At each layer, critics could correctly point out that the lower layer had not disappeared.
Somewhere there is still machine code.
Somewhere there are still memory pages.
Somewhere electrons remain stubbornly involved.
But civilization advances by allowing most people to stop thinking about the lower layer most of the time.
Abstraction does not eliminate reality.
It reorganizes responsibility for reality.
The database engineer does not expect an application developer to position the disk head.
The web developer does not manually implement TCP congestion control.
The mobile developer does not manufacture the accelerometer.
The user who asks a spreadsheet to calculate a sum does not expect to become briefly responsible for floating-point hardware.
Generative AI represents another abstraction layer.
But it is subtler than "English replaces Python."
That slogan is catchy and mostly wrong.
The important abstraction is not natural language over programming language.
It is:
desired state over explicit procedure
Traditional programming asks:
What exact operations should the machine execute?
Declarative programming asks:
What state should exist?
Agentic engineering extends that pattern:
What outcome should exist, under which constraints, and what evidence is required before we accept that it exists?
That is a much more powerful interface.
It changes not merely how code is written, but where engineering effort accumulates.
3. The Dream Is Older Than Generative AI
There is a temptation to treat modern AI-assisted computing as though history began when a chat box appeared.
It did not.
The intellectual ancestry is much richer.
In 1945, Vannevar Bush imagined machinery that could extend human access to knowledge through associative structures rather than merely reproducing traditional filing systems. His famous As We May Think envisioned something closer to an intellectual prosthesis than a conventional calculating machine.
Norbert Wiener's cybernetics, first published in 1948, made feedback and control central conceptual tools for understanding both machines and organisms. The critical insight for our purposes is that competent behavior need not consist of executing a perfect predetermined plan; it may instead involve observing consequences and correcting action through feedback.
J. C. R. Licklider went further in 1960. In Man-Computer Symbiosis, he imagined close cooperation in which humans would establish goals, hypotheses, criteria, and evaluations while machines handled routinizable work supporting those decisions.
Read today, that division of labor feels almost suspiciously current.
Humans define the objective.
Machines perform substantial cognitive operations.
Humans evaluate.
The partnership exceeds either participant alone.
Two years later, Douglas Engelbart described "augmenting human intellect" as increasing humanity's capacity to comprehend complex situations and solve problems that had previously exceeded practical reach.
Then Herbert Simon framed design as the science of artificial systems: not merely studying what is, but systematically reasoning about artifacts constructed in pursuit of purposes.
These ideas form a lineage:
memory augmentation -> feedback -> symbiosis -> intellectual augmentation -> design science
Agentic engineering is not an alien break from that tradition.
It is one of its most literal realizations.
4. From Autocomplete to Agency
Consider the progression.
Stage one: completion
The machine predicts the next token.
Useful, but fundamentally local.
Stage two: generation
The machine produces a function, a component, a test, perhaps a file.
Now intent can produce larger artifacts.
Stage three: conversational editing
The machine modifies existing code according to instructions.
The human can steer iteratively.
Stage four: repository awareness
The machine searches files, understands relationships, inspects configuration, traces definitions, and reasons across multiple modules.
Stage five: tool use
The machine can run tests, issue commands, inspect logs, query services, read documentation, and observe execution results.
Stage six: goal-directed iteration
The machine receives an objective, forms a plan, modifies the environment, observes consequences, updates its approach, and repeats.
At this point we have crossed an important line.
The machine is no longer simply producing text.
It is interacting with an environment to pursue a state transition.
That is agency in the engineering sense.
Not philosophical personhood.
Not consciousness.
Not a synthetic employee demanding coffee privileges.
Operational agency.
The ability to take sequential actions toward a goal under changing information.
Modern research benchmarks have increasingly moved in exactly this direction. SWE-bench was designed around real GitHub issues requiring repository-level changes rather than isolated function completion. Its original dataset contained 2,294 problems from 12 Python repositories and explicitly emphasized the need to coordinate changes across files and interact with execution environments.
By October 2026, “SWE-bench” no longer names one undifferentiated score: its official leaderboard separates Full, Verified, Lite, Multilingual, and Multimodal tracks. Two September preprints make the limits of benchmark scores especially visible. SWE-Bench Pro Verified reports reward-hacking/leakage and task-quality problems in the earlier Pro evaluation, and finds that some systems perform substantially worse after those problems are addressed. SWE-Proof adds a formal-verification oracle to 500 SWE-bench-derived issues; in its reported experiments, formal checks find counterexamples in between a quarter and a half of test-passing patches, while faithful specification synthesis remains difficult. These are early results tied to particular datasets, models, and harnesses, not a universal ranking or a measure of production autonomy. They support a narrower point: a benchmark result describes what that evaluator measured, and trustworthy evaluation must itself be engineered. SWE-bench tracks · SWE-Bench Pro Verified · SWE-Proof, v2
SWE-agent subsequently demonstrated how much the surrounding interface matters: agents perform differently depending on how effectively they can inspect repositories, edit files, execute tools, and receive feedback.
That is a profound clue.
The future of AI software engineering depends not only on better models, but also on better tools, interfaces, and governance around them.
5. The Prompt Is Becoming a Protocol
Early AI development culture fetishized prompts.
There were secret prompt recipes.
Prompt frameworks.
Prompt gurus.
Prompt incantations.
A strangely large quantity of human ingenuity was devoted to asking a machine to "think step by step" in increasingly decorative ways.
Prompts remain useful.
But mature engineering does not rest on incantation.
It rests on protocol.
Compare:
Fix the persistence bug.
with:
Reproduce the reported persistence failure. Identify the violated invariant. Determine which storage authority owns the durable state. Preserve exact user data. Do not weaken existing validation. Add a regression test covering the failure condition. Run the required browser and desktop verification paths. If any required evidence is unavailable, report UNKNOWN rather than PASS. Present the resulting diff and evidence before requesting merge authorization.
The second instruction is not "better prompting" in the superficial sense.
It encodes a miniature operating system for responsibility.
It defines:
- objective,
- constraints,
- authority,
- failure semantics,
- evidence requirements,
- prohibited shortcuts,
- and escalation conditions.
That is where agentic engineering becomes serious.
We move from telling intelligent systems what to say toward defining how intelligent work may proceed.
Governing the Repository and Its Agents
6. The Repository Becomes a Tiny Civilization
A modern repository is often described as a collection of source files.
This is like describing a city as a collection of bricks.
Technically defensible.
Conceptually inadequate.
A serious repository contains:
- source code,
- dependency contracts,
- schemas,
- architectural rules,
- tests,
- build systems,
- deployment systems,
- security policies,
- release procedures,
- ownership declarations,
- review conventions,
- changelog rules,
- migration history,
- issue history,
- incident knowledge,
- observability contracts,
- and enough shell scripts to recreate a small Bronze Age religion.
These are not merely technical artifacts.
They are institutional artifacts.
A test is a law expressed as executable evidence.
A type system is a boundary on representable states.
A CI workflow is a procedure for deciding which transitions are admissible.
A required check is a veto point.
A code owner is a jurisdiction.
A release process is a constitutional transfer of authority from candidate state to public state.
An audit log is institutional memory.
Once autonomous agents participate in development, this institutional interpretation becomes much more important.
Because now the repository has inhabitants.
The question becomes:
Under what constitution should machine actors operate?
This is why the best agentic engineering systems will not be the ones that trust AI the most.
They will be the ones that need to trust it the least.
7. Trust the System, Not the Mood of the Model
A bridge is not safe because its engineer seemed intelligent.
An airplane is not airworthy because the maintenance technician was confident.
A cryptographic protocol is not secure because somebody wrote "looks good" in a pull request.
Engineering civilization rests on a basic insight:
capable humans can still be wrong.
Therefore trustworthy systems constrain capable humans.
We use checklists.
Independent review.
Static analysis.
Redundancy.
Simulation.
Testing.
Certification.
Separation of duties.
Version control.
Rollback.
Monitoring.
Incident investigation.
AI deserves the same courtesy.
The wrong question is:
Can the model be trusted?
The better question is:
What architecture remains safe when the model is intermittently wrong?
This is a far more optimistic question than it appears.
Because if we solve it, progress no longer requires perfect artificial intelligence.
We can obtain tremendous value from imperfect intelligence embedded inside strong institutions.
That is much more achievable.
8. Generation Gets Cheap; Verification Gets Precious
Suppose an agent can generate a plausible implementation in thirty seconds.
Wonderful.
Suppose it can generate ten implementations in five minutes.
Even better.
Now what?
In workflows where candidate generation gets cheap, one bottleneck starts to move.
Candidate generation may no longer be the scarce resource.
Selection becomes more valuable.
Which implementation is correct?
Which is maintainable?
Which violates a hidden invariant?
Which quietly leaks memory?
Which mishandles Unicode?
Which passes visible tests while breaking behavior nobody remembered to test?
Which adds a dependency with unacceptable licensing?
Which behaves differently under concurrency?
Which solves today's bug by creating next month's incident?
This is the verification inversion.
As generation becomes cheaper, verification can become relatively more valuable.
That points toward a shift in software engineering investment.
Teams may get more value from:
- property-based testing,
- mutation testing,
- fuzzing,
- formal specification,
- symbolic analysis,
- architecture tests,
- reproducible builds,
- provenance tracking,
- dependency verification,
- security policy engines,
- differential testing,
- runtime canaries,
- synthetic users,
- automatic rollback,
- and adversarial agent review.
The great AI engineering platform of the future may not be distinguished primarily by how much code it can produce.
It may be distinguished by how much bad code it can confidently reject.
9. What Do We Actually Know About AI Productivity?
Technological enthusiasm becomes more credible when it survives contact with inconvenient evidence.
So here is some inconvenient evidence.
A controlled 2023 experiment involving GitHub Copilot found that participants asked to build an HTTP server completed the task 55.8% faster with the AI assistant than the control group.
That sounds spectacular.
Then a 2025 randomized study by METR examined 16 experienced open-source developers completing 246 tasks in mature repositories they already knew well. In that setting, early-2025 AI tools increased completion time by about 19%, even though participants believed the tools were helping them go faster.
Also spectacular.
In the opposite direction.
Then things became more interesting.
In its February 2026 update, METR said later measurements were a weak signal because participants and tasks were selected in ways that likely missed some of the most AI-optimistic cases, and concurrent agents made time accounting harder. The returning-participant estimate (n=10) was an 18% time reduction, with an interval from 38% faster to 9% slower; newly recruited developers (n=47) showed a 4% estimated reduction, with an interval from 15% faster to 9% slower. Both intervals include no speedup. METR explicitly called the estimate a poor proxy for the true effect. The original study remains evidence about its 2025 tools, people, and task setting; the follow-up does not establish a universal current productivity number. METR update
This is exactly what we should expect during a rapidly changing technological transition.
"Does AI make programmers faster?" is too crude a question.
Better questions include:
- Which programmers?
- Doing what tasks?
- In familiar or unfamiliar repositories?
- With what quality bar?
- Using which models?
- With which agent interfaces?
- With how much parallelism?
- Under which review regime?
- Measuring typing time, elapsed time, cognitive load, cycle time, or throughput?
- Counting downstream defects?
- Counting work delegated asynchronously?
- Counting tasks that previously would never have been attempted?
Technology changes measurement before it changes consensus.
The printing press was not valuable because scribes achieved superior handwriting throughput.
The spreadsheet was not valuable because accountants typed numbers faster.
The internet was not valuable because letters were written with fewer keystrokes.
The largest effects come when workflows reorganize around new capabilities.
Agentic engineering is likely to be the same.
What should a team measure?
The productivity question is not a vote between two studies. It is a measurement problem with several outcomes that can move in different directions.
Elapsed time is useful, but it is not the only cost. A person might spend less time typing and more time reviewing. A task might finish earlier while requiring a second engineer to repair a subtle regression. An agent may do useful work asynchronously while its operator handles another task, which makes wall-clock time difficult to assign to one person. A tool may make a previously unaffordable experiment possible; a study of only assigned tasks will not count work that nobody would otherwise have started.
Quality also needs a defined test. “The patch passed the tests” means that the chosen tests passed. It does not automatically measure maintainability, user comprehension, security, support burden, or whether the right problem was solved. Nor does the presence of a longer test suite imply better quality if the tests simply encode the implementation’s assumptions.
A team evaluating a new workflow can start with a small, representative set of tasks and decide in advance what it will measure: completion time, acceptance rate, rework, escaped defects, review effort, or a combination. Preserve task and tool context. Record which model and agent configuration were used, which checks ran, and whether the person could work on other tasks while waiting. Compare like with like where possible, and be explicit when the comparison is not randomized.
The result should be a profile, not one magic percentage. Perhaps routine migrations become quicker while unfamiliar architectural work takes longer to review. Perhaps a junior engineer gets more prototypes but needs stronger supervision. Perhaps the tool is a clear win for documentation and a poor fit for changes to an unfamiliar persistence boundary. These are useful findings even when they do not collapse into a single headline.
This is a proposed measurement discipline, not a claim that one study design solves every confound. Its purpose is simpler: keep the question attached to the work people actually do, and keep quality and downstream cost in the same frame as speed.
10. The Unit of Automation Is Expanding
At first we automated characters.
Then expressions.
Then functions.
Then files.
Now we are beginning to automate engineering episodes.
An engineering episode might be:
Upgrade the database driver safely.
That includes far more than editing a version number.
A serious agent may need to:
- identify all affected packages;
- inspect release notes;
- evaluate breaking changes;
- update dependency constraints;
- regenerate the lockfile;
- repair compilation errors;
- migrate affected APIs;
- run tests;
- inspect failures;
- modify tests where behavior intentionally changed;
- check security advisories;
- benchmark critical paths;
- update documentation;
- inspect generated artifacts;
- prepare a review summary;
- stop before merge if authorization is required.
That collection of work used to be distributed across dozens of manual interactions.
The future agent does not merely write code.
It performs bounded operational labor.
The unit of automation rises from syntax to workflow.
11. The Engineer Moves Up the Stack
When people ask whether programmers will disappear, they often imagine a fixed quantity of software work.
There is some pile of tasks.
Humans currently perform it.
Machines begin performing it.
Therefore fewer humans are required.
Sometimes that will happen.
But general-purpose technologies rarely operate on a fixed pile of demand.
When the cost of capability falls, people demand more capability.
The more interesting question is:
What becomes valuable when implementation becomes abundant?
The answer is judgment.
Architecture.
Problem selection.
Product understanding.
Risk reasoning.
Constraint design.
Verification.
Taste.
Human context.
The engineer increasingly becomes responsible for questions such as:
- What outcome actually matters?
- Which state is authoritative?
- Which invariants are sacred?
- What evidence is sufficient?
- Which operation is reversible?
- Which data cannot be exposed?
- Which part should remain simple?
- Which automation should require approval?
- Which failure should degrade gracefully?
- Which trade-off will future maintainers regret?
These are higher-order engineering questions.
The model can participate in answering them.
But responsibility for the system increasingly concentrates there.
The future engineer is less a person who manually converts specifications into syntax and more a designer of socio-technical decision systems.
That is a promotion, not a demotion.
12. Congratulations, You Are Now Managing a Tiny Engineering Organization
The next major interface for programming may look less like an editor and more like an operations center.
Imagine opening a development environment and seeing:
Objective
Implement resumable encrypted uploads.
Constraints
- no plaintext persistence,
- mobile-safe memory usage,
- backwards-compatible API,
- existing public URLs must remain stable.
Agents
- architecture investigator,
- implementation agent,
- test agent,
- security critic,
- documentation agent.
Evidence
- unit tests: PASS
- fuzz suite: PASS
- browser E2E: PASS
- mobile memory budget: FAIL
- security review: PENDING
Decision
Merge authorization unavailable because memory budget failed.
This is much more interesting than watching a model type quickly.
The human supervises a system of specialized cognitive workers.
One proposes.
One critiques.
One tests.
One examines dependencies.
One checks policy.
One watches production.
The human allocates attention where uncertainty is highest.
Software development begins to borrow concepts from organizational theory:
- delegation,
- authority,
- escalation,
- specialization,
- separation of duties,
- audit,
- consensus,
- veto,
- budget,
- priority,
- institutional memory.
We will discover that building great agent systems is partly the art of designing very strange companies whose employees are software processes.
13. Policy Becomes a Programming Language
The more capable agents become, the more important it becomes to encode what they must not do.
Consider rules such as:
- production secrets must never enter model context;
- database migrations require review;
- release tags must be signed;
- failing tests may not be deleted to achieve green CI;
- security checks may not be bypassed;
- user data may not leave approved regions;
- destructive infrastructure changes require explicit authorization;
- a release may only originate from a revision whose required evidence passed;
- an agent may propose a policy change but may not authorize that same change.
These rules are simultaneously:
- organizational,
- technical,
- legal,
- operational,
- and architectural.
In the agentic era, they become executable.
This is a major evolution.
Organizations currently store enormous quantities of critical knowledge inside human habit.
"Ask Sarah before touching that."
"Never run this script against prod."
"That migration must happen after the cache flush."
"The deployment is green, but only if the artifact hashes match."
One day an organization discovers that Sarah is on vacation and nobody remembers why the migration rule exists.
Agentic engineering creates strong incentives to turn tacit wisdom into machine-readable policy.
That is good for AI.
It is also good for humans.
14. The Strange Return of Formalism
Natural-language programming sounds informal.
The mature system may become more formal than traditional development.
Why?
Because natural language is wonderfully ambiguous.
"Delete obsolete deployments."
Which ones are obsolete?
Anything older than thirty days?
Anything whose branch merged?
What about rollback deployments?
What about legal retention?
What about active canaries?
What if a customer is still using a direct deployment URL?
The phrase sounds simple because a human quietly supplies context.
Agents force us to discover how much context existed only in someone's head.
So precision returns.
But it returns in new places:
- invariants,
- schemas,
- acceptance criteria,
- typed tool interfaces,
- permissions,
- policies,
- state machines,
- contracts,
- tests,
- proofs,
- capability declarations.
The paradoxical future is both more conversational and more formal.
Humans express intention naturally.
Machines compile that intention into structured plans and constrained operations.
Where ambiguity is harmless, systems proceed.
Where consequences are reversible, systems experiment.
Where actions are dangerous, systems ask for stronger authority.
That is a far superior interface to either "write every instruction manually" or "trust the robot and hope."
Measure the System, Not the Demo
15. Software Development Starts Looking More Like Experimental Science
A mature engineering agent should not report:
Fixed it.
That is not an engineering statement.
That is a mood.
A useful report looks more like:
The failure occurred when an asynchronous writer committed after the active project identity changed. I reproduced the condition, introduced an identity-bound epoch, rejected stale commits, added regression tests for in-flight authority changes, and verified both persistence backends.
Now we have claims.
Claims can be challenged.
Experiments can be repeated.
Evidence can be inspected.
The shift toward agentic engineering may therefore increase pressure for software development to become more epistemically disciplined.
An engineering change should increasingly carry:
- the hypothesized failure mechanism,
- the relevant invariant,
- the change made,
- the tests performed,
- the artifact identity,
- the observed result,
- known residual uncertainty.
This is science-like behavior.
Not because software engineering becomes pure science.
But because delegated intelligence requires explicit justification.
The machine cannot rely on the office folklore that "everyone knows this is how we do it."
That folklore must become inspectable.
16. Documentation Stops Being Decoration
There is an amusing possibility here.
AI may finally force developers to write documentation.
Not because management sends another reminder.
Because agents work better in environments where architectural truth is explicit.
A codebase with clear boundaries, current documentation, reliable tests, explicit state ownership, and machine-readable conventions provides much better terrain for autonomous systems.
A chaotic repository full of stale comments, hidden assumptions, magical environment variables, and four competing definitions of "project state" is difficult for agents for exactly the same reason it is difficult for humans.
AI therefore increases the economic return on engineering hygiene.
Architecture documents stop being ceremonial artifacts.
They become context infrastructure.
Policies stop being PDFs nobody reads.
They become executable constraints.
Decision records stop being historical curiosities.
They become training data for future engineering episodes.
The better the organization explains itself to machines, the better it often explains itself to people.
17. Software Abundance Changes the Economics of Ideas
Now we arrive at the genuinely transformative possibility.
Software can be expensive.
Not merely in money.
In activation energy.
Suppose a biology researcher has an idea for a specialized visualization.
Today, the researcher may need:
- a developer,
- funding,
- requirements meetings,
- infrastructure,
- authentication,
- deployment,
- maintenance,
- perhaps procurement.
The idea may die before anyone creates a repository.
Suppose a teacher wants a tiny application tailored to one curriculum.
Too niche.
Suppose a disabled user needs an interface modified in a highly specific way.
Too small a market.
Suppose seventeen marine biologists need a collaborative tool for tracking a peculiar species of mollusk.
Someone will suggest Excel.
It is always Excel.
Excel has seen things.
The economics of software often favor repeated needs shared by many people.
Agentic engineering could make much smaller markets viable.
If the cost of implementation, testing, deployment, and maintenance falls enough, more long-tail software could become economically reachable.
That means more software for:
- rare diseases,
- minority languages,
- specialized research,
- local government,
- unusual accessibility requirements,
- niche industrial workflows,
- small communities,
- temporary humanitarian operations,
- classrooms,
- laboratories,
- families,
- and individuals.
Mass production historically made standardized goods cheap.
Intelligent automation may make customization cheap.
That is a different kind of abundance.
18. The End of the Blank Project
Generative AI has already reduced one psychological barrier: the terror of starting.
The empty editor is unforgiving.
It asks:
So, genius, what is the architecture?
A conversational agent says:
Tell me what you are trying to do.
That difference matters.
Creative production is often constrained less by imagination than by activation energy.
People possess ideas they never attempt because the first thousand steps are too expensive.
AI can reduce the cost of the first attempt.
And failed attempts matter too.
If experimentation becomes cheap, people can explore more possibilities.
The result is not merely more finished software.
It is a larger search space of human creativity.
Most experiments will be unimportant.
That is fine.
Most scientific hypotheses are not revolutions.
Most books are not War and Peace.
Most startups do not become global institutions.
Abundance works through distributions.
Lower experimentation costs could make valuable outliers easier to discover.
19. Small Teams Become Weirdly Powerful
Imagine a five-person organization with:
- one product expert,
- one senior engineer,
- one designer,
- one domain specialist,
- one operations lead,
plus persistent access to specialized engineering agents.
The effective organization might take on work once spread across a larger team.
Not because each agent perfectly replaces a job title.
Because the fixed overhead of specialized work may shrink.
The group can temporarily instantiate:
- a security reviewer,
- a test engineer,
- a migration specialist,
- a data analyst,
- a documentation writer,
- an infrastructure investigator.
This changes entrepreneurship.
It changes nonprofits.
It changes research labs.
It changes civic technology.
It changes independent creators.
The minimum organizational mass required to attempt ambitious work could decline.
That is socially significant.
Great ideas are not evenly distributed among people with access to large engineering budgets.
Anything that reduces the correlation between "ability to build" and "institutional wealth" increases the number of possible innovators.
20. Expertise Becomes More Scalable, Not Less Valuable
A common fear says AI makes expertise worthless.
A better model says AI makes encoded expertise more scalable.
Consider a brilliant security engineer.
Today that engineer can:
- review some pull requests,
- design some systems,
- mentor some colleagues,
- write some documentation.
Their influence is limited by attention.
Now imagine their expertise partly captured in:
- policy rules,
- threat-model templates,
- adversarial test suites,
- security-review agents,
- dependency constraints,
- secure-by-default scaffolds.
The expert's judgment becomes infrastructure.
One expert influences thousands of changes.
The same is true for:
- accessibility specialists,
- database engineers,
- performance experts,
- cryptographers,
- legal specialists,
- UX researchers,
- scientific methodologists.
AI may therefore transform experts from repeated task performers into designers of scalable judgment.
That is a powerful form of leverage.
21. Education Must Stop Confusing Syntax With Understanding
Programming education is about to face an uncomfortable question.
If a machine can generate a binary tree implementation in three seconds, should students still learn binary trees?
Yes.
But perhaps not for the old reason.
The purpose of technical education cannot remain:
Memorize the implementation because one day someone may pay you to type it.
The future purpose is:
Understand the system deeply enough to evaluate, modify, constrain, and improve machine-generated implementations.
That requires serious knowledge.
Potentially more knowledge.
A person supervising engineering agents must understand:
- algorithms,
- data structures,
- networking,
- concurrency,
- storage,
- distributed systems,
- security,
- testing,
- complexity,
- architecture.
Otherwise they cannot distinguish a sophisticated solution from sophisticated nonsense.
The future engineer must be able to recognize when the AI has beautifully solved the wrong problem.
That is a very human skill.
22. The Future IDE Is an Epistemic Cockpit
Traditional development tools are organized around artifacts.
Files.
Folders.
Tabs.
Branches.
Terminals.
Future engineering environments may instead be organized around claims and state transitions.
You might see:
Objective
Prevent stale writes after authority change.
Current hypothesis
Writer lifetime outlives target identity.
Proposed invariant
Commit only when writer epoch matches active authority epoch.
Evidence
Unit: green
Concurrency: green
Migration: green
Browser E2E: green
Desktop packaged artifact: pending
Residual risk
Legacy import path not exercised.
Authorization
Merge blocked pending packaged evidence.
That interface does something profound.
It centers the engineering question:
What do we know?
rather than:
Which file is open?
Code remains available.
But code becomes one piece of evidence in a larger epistemic model.
23. The Self-Improving Software Factory
Agents can improve products.
They can also improve the system that produces products.
An agent can inspect:
- flaky tests,
- repeated incidents,
- slow builds,
- recurring review findings,
- frequently broken APIs,
- duplicated configuration,
- stale documentation,
- dependency churn,
- security warnings.
Then it can ask:
Why does this problem keep recurring?
A repeated human review comment can become a lint rule.
A recurrent incident can become a regression test.
A manual checklist can become an executable gate.
An architectural convention can become a static constraint.
The engineering organization gradually converts experience into machinery.
This is the beginning of a self-improving software factory.
Not self-improving in the science-fiction sense of an unconstrained intelligence rewriting itself overnight.
Self-improving in the much more useful sense of:
the development process continuously learning which mistakes deserve permanent prevention.
That is organizational learning with executable memory.
Failure, Costs, and Responsible Optimism
24. Failure Does Not Invalidate the Vision
Agentic systems will fail.
They will misunderstand requests.
They will hallucinate APIs.
They will introduce regressions.
They will make architectural changes nobody requested.
They will discover an old dependency and decide, with heroic confidence, that now is an excellent time to upgrade it.
They will occasionally spend half an hour fixing a defect they created eleven minutes earlier.
This is not evidence that agents cannot participate in engineering.
It is evidence that they already understand the culture.
Human engineering is full of errors.
The safety of engineering disciplines comes from mechanisms for catching and containing them.
Aviation does not assume perfect pilots.
Medicine does not assume perfect memory.
Nuclear operations do not assume perfect attention.
Financial systems do not assume perfect arithmetic.
Agentic engineering should similarly assume fallibility.
The design objective is:
high capability, bounded consequence, strong evidence.
25. Brooks Was Right—and AI Can Still Be Transformative
Frederick Brooks famously argued in No Silver Bullet that there was no single development likely to produce an order-of-magnitude improvement across software productivity, reliability, and simplicity, because much of software's difficulty arises from the essential complexity of conceptual structures rather than accidental implementation friction.
That insight remains useful.
AI does not make essential complexity disappear.
Someone still has to decide:
- what the system should do,
- which conflicting requirements matter,
- where authority belongs,
- how failure should behave,
- which trade-offs are acceptable.
But AI may substantially reduce the accidental complexity surrounding those decisions.
Boilerplate.
API discovery.
Mechanical migration.
Configuration.
Routine testing.
Search.
Cross-referencing.
Documentation.
Diagnosis.
This is exactly why agentic engineering may be transformative without being a magical silver bullet.
It does not abolish engineering.
It concentrates engineering effort closer to the essential problem.
26. The Utilitarian Case for Intelligent Automation
Why should we want this future?
Not because technology is aesthetically impressive.
Not because intelligence benchmarks are fun.
Not because more compute is automatically morally good.
The strongest argument is consequential.
A technology is valuable when it increases human welfare.
Does it reduce suffering?
Does it free time?
Does it expand access?
Does it enable creation?
Does it improve safety?
Does it accelerate discovery?
Does it give more people meaningful agency?
Agentic engineering has an unusually strong route to those outcomes because software is already a multiplier across nearly every modern domain.
Cheaper software creation can mean:
- better scientific tools,
- improved healthcare workflows,
- more accessible interfaces,
- better logistics,
- better educational systems,
- improved public administration,
- more effective climate modeling,
- more personalized assistive technology,
- better research infrastructure.
The value compounds because software builds other capabilities.
We are not merely automating programming.
We are potentially reducing the cost of building tools that build futures.
27. But Utilitarianism Requires Counting the Costs
Optimism should not become arithmetic fraud.
The benefits of agentic technology must be weighed against:
- energy consumption,
- security risk,
- labor displacement,
- market concentration,
- surveillance incentives,
- misinformation,
- brittle dependency on providers,
- unequal access,
- degraded accountability.
A serious pro-technology position does not deny these.
It asks how to design institutions that preserve upside while reducing harm.
That means:
- competitive ecosystems,
- interoperable systems,
- open standards,
- local capability where practical,
- privacy-preserving architectures,
- clear liability,
- meaningful human control,
- transparent provenance,
- strong security.
The best techno-optimism is not faith.
It is engineering applied to social consequences.
The Star Trek Lens: Agency, Abundance, and Judgment
28. Why Star Trek Is the Better Metaphor
Science fiction offers many technological futures.
Some are cautionary.
Machines dominate.
Corporations dominate.
Surveillance dominates.
Everyone owns seventeen holographic advertisements and appears strangely unhappy.
Star Trek: The Next Generation represents a different technological philosophy.
Its most interesting feature is not warp drive.
It is the normalization of capability.
Computation is abundant.
Information is abundant.
Communication is abundant.
Fabrication is abundant.
Routine operations are automated.
The computer is everywhere but rarely the point.
Technology recedes into infrastructure.
This is the ideal form of mature technology.
You do not spend your day admiring the plumbing.
You turn on the tap.
You do not marvel at TCP/IP every time you load a page.
You expect the network.
In the Trek vision, technology increasingly removes categories of scarcity and friction so that human attention can move elsewhere.
Exploration.
Science.
Diplomacy.
Art.
Self-development.
Occasionally preventing Q from destroying civilization for pedagogical reasons.
That is the vision worth borrowing.
29. The Replicator Is a Philosophy
The replicator is often treated as a fun gadget.
But philosophically it represents something deeper.
It converts intention into artifact at negligible marginal effort.
"Tea. Earl Grey. Hot."
The point is not tea.
The point is that an enormous production chain disappears from the user's cognitive horizon.
Agentic engineering is moving software in that direction.
Not fully.
Not magically.
But conceptually.
A person expresses an intention.
The system performs:
- decomposition,
- implementation,
- verification,
- packaging,
- deployment,
- maintenance.
The marginal cost of creating some software artifacts falls.
That is a kind of computational replication.
And just as the replicator transforms the economics of physical goods in science fiction, cheap generation can transform the economics of digital capability.
30. The Computer on the Enterprise Is an Agent Platform
Think about how characters interact with the Enterprise computer.
They do not usually write scripts.
They state goals.
"Computer, analyze the readings."
"Compare them with Federation records."
"Simulate the effect."
"Locate all ships matching these parameters."
"Create a program."
The computer resolves substantial procedural detail.
Humans remain responsible for mission objectives.
This is remarkably close to the interface we are constructing.
The significant difference is that Star Trek assumes the infrastructure has become trustworthy enough to disappear.
Our job is to build the missing trust layer.
31. The Picard Test
We need a criterion for whether this future is actually good.
Call it the Picard Test.
A technology passes when it:
- increases human capability;
- reduces unnecessary drudgery;
- expands access to knowledge and creation;
- preserves meaningful human agency;
- supports rather than erodes dignity;
- increases the space available for intrinsically valuable activity;
- distributes benefits broadly enough to improve collective welfare.
The question is not:
Is the AI impressive?
The question is:
Are humans more capable because it exists?
That is the metric that matters.
32. Humans Should Not Compete With Machines at Being Machines
There is something deeply absurd about imagining the purpose of artificial intelligence as forcing humans to become more machine-like.
Faster typing.
Longer working hours.
Higher ticket throughput.
More constant availability.
That is technological failure.
The correct division of labor is the opposite.
Let machines perform machine-like work:
- exhaustive search,
- repetitive transformation,
- continuous monitoring,
- mechanical comparison,
- large-scale execution.
Let humans specialize further in what gives human activity meaning:
- imagination,
- judgment,
- empathy,
- responsibility,
- purpose,
- curiosity,
- moral reasoning,
- aesthetic taste,
- relationships.
The objective is not maximum automation.
It is better allocation of cognition.
33. The Bottleneck Eventually Becomes Wisdom
Imagine that agentic engineering succeeds beyond expectation.
Building software becomes dramatically easier.
Then the critical question changes.
It is no longer:
Can we build this?
It becomes:
Should we?
This may be the defining problem of technological abundance.
When implementation is expensive, cost filters ideas.
When implementation becomes cheap, judgment must replace cost as the filter.
Should this feature exist?
Should this automation be allowed?
Who benefits?
Who bears risk?
Does it increase autonomy?
Does it exploit attention?
Does it preserve privacy?
Does it reduce suffering?
Does it create genuine value?
These questions cannot be outsourced merely because machines become intelligent.
They become more important.
In this scenario, the final scarcity is not compute.
It is wisdom.
34. The Programmer of 2035 Will Feel Sorry for Us
Imagine a developer in 2035 examining an archaeological record of software development in 2025.
"You manually searched the repository?"
Yes.
"You copied the error message into another application?"
Yes.
"You read dependency release notes yourself?"
Often.
"You waited for CI, then reopened a browser to check it?"
Yes.
"And sometimes a release depended on someone remembering a checklist?"
Please stop.
"And if the documentation became stale-"
I SAID PLEASE STOP.
Future generations always discover that the past contained astonishing amounts of manual glue.
We are living inside glue work we no longer notice.
AI will make much of it visible by making its automation possible.
35. Humans Will Still Code
Of course they will.
People still:
- cook despite restaurants,
- drive despite trains,
- paint despite cameras,
- grow vegetables despite supermarkets,
- play chess despite engines.
Making things is satisfying.
Direct manipulation is intellectually rewarding.
Code is a beautiful medium.
Some domains require exact control.
Some people simply enjoy it.
The future does not need to abolish programming.
It may rescue programming from some of the bureaucracy surrounding programming.
Less repetitive plumbing.
More interesting design.
Less configuration archaeology.
More algorithms.
Less dependency babysitting.
More experimentation.
Less time wondering why a CI runner has decided today is the day to become philosophical.
From Prompts to Protocols
36. Vibe Coding Was the Biplane
We should remember vibe coding affectionately.
It revealed the interface.
It demonstrated that human language could become a high-bandwidth control surface for software construction.
It made millions of people feel, perhaps for the first time:
I can make the computer do something I imagined.
That feeling is historically important.
But the first aircraft did not define aviation.
The first automobiles did not define modern transportation.
Early websites did not define the internet.
Vibe coding will not define agentic engineering.
It was the biplane made of canvas, timber, and optimism.
The next stage adds:
- instrumentation,
- procedures,
- navigation,
- traffic control,
- maintenance,
- certification,
- redundancy,
- safety engineering.
Less romantic.
Far more powerful.
37. From Prompting Machines to Governing Capability
The deepest transition can be summarized in three eras.
Era One: Code
The human specifies procedure.
Era Two: Prompt
The human describes desired artifact.
Era Three: Protocol
The human defines desired outcome, boundaries, evidence, and authority.
The third era is where agentic engineering becomes an institution.
And institutions are what allow powerful actors to cooperate reliably.
A Future Worth Building
38. The Grand Future
Imagine a mature version of this technology.
A scientist has an idea for analyzing microscopy data.
She describes the analysis.
An agent constructs a reproducible pipeline.
Another agent tests statistical assumptions.
Another verifies provenance.
Another creates the visualization.
The scientist evaluates the scientific meaning.
A teacher wants an adaptive learning environment for a specific class.
The system generates it.
Accessibility checks run automatically.
Content aligns with the curriculum.
Student data remains local by policy.
The teacher adjusts the pedagogy.
A small nonprofit needs software for emergency relief coordination.
Agents assemble it from verified components.
Security constraints are enforced.
Deployment is automatic.
Operators remain in control of consequential actions.
A person with a unique disability describes how an existing interface fails them.
The system produces an adaptation for one person.
In the limit, a market of one may become reachable: software could be tailored to an individual user. That is a scenario about lower fixed costs, not a prediction that every custom tool will be cheap to maintain.
And that is enough.
A researcher creates an experimental tool in an afternoon.
A municipality builds a service without buying a gigantic enterprise platform.
A small company maintains infrastructure once requiring a dedicated operations department.
An independent developer supervises a group of agents that test, document, secure, and maintain software around the clock.
Not because humanity has been removed.
Because the mechanical distance between intention and capability has shrunk.
39. A Better Definition of Progress
Progress is not the number of tasks machines can perform.
Progress is the number of valuable things humans become capable of doing.
That distinction should guide everything.
The best future for AI is not one in which machines become impressive and humans become spectators.
It is one in which machine capability becomes a multiplier on human agency.
That is the difference between automation as displacement and automation as civilization.
40. Engage
Vibe coding began with a wonderfully unserious proposition:
What if we stopped obsessing over the code and simply told the computer what we wanted?
It turned out that this was simultaneously irresponsible and profound.
Because buried inside the chaos was a glimpse of a new abstraction layer.
We are now learning to civilize it.
To add architecture.
Evidence.
Policy.
Verification.
Authority.
Memory.
Recovery.
Specialization.
Accountability.
The future of software engineering is therefore not "AI writes all the code."
That is far too small an idea.
The future is that increasingly capable computational systems participate in the entire process by which human intentions become reliable technological capabilities.
When that process works, implementation becomes cheaper.
Experimentation expands.
Expertise scales.
Small teams become powerful.
Long-tail problems become worth solving.
Human attention moves upward.
Technology recedes into infrastructure.
And perhaps, eventually, software begins to feel less like a fragile pile of instructions and more like the quietly competent machinery aboard a well-run starship.
The computer handles the computation.
The agents maintain the systems.
The replicator makes the tea.
And the humans decide where to go.
That is not a future in which technology has replaced humanity.
It is one in which technology has finally learned its proper role:
to enlarge the space of human possibility.
Further reading and evidence
- Andrej Karpathy, original “vibe coding” post (February 2025).
- Vannevar Bush, “As We May Think” (1945).
- Norbert Wiener, Cybernetics (1948).
- J. C. R. Licklider, “Man-Computer Symbiosis” (1960).
- Douglas Engelbart, “Augmenting Human Intellect: A Conceptual Framework” (1962).
- Herbert A. Simon, The Sciences of the Artificial.
- Frederick P. Brooks Jr., “No Silver Bullet” (1987).
- Peng et al., The Impact of AI on Developer Productivity (2023).
- METR, randomized study (2025) and measurement update (2026).
- DORA 2025 report.
- Jimenez et al., SWE-bench (2023).
- Official SWE-bench leaderboard and benchmark tracks (accessed 2026-10-01).
- Zheng et al., SWE-Bench Pro Verified (2026 preprint).
- SWE-Proof (2026 preprint, v2).
- Hassan et al., Agentic Software Engineering: Foundational Pillars and a Research Roadmap (2025). This paper proposes a conceptual vocabulary and research roadmap, not a universal standard.
- GitHub status-check documentation and the SLSA v1.2 specification.
Continue with the series
The next article, The Reviewer Is Not the Gate, turns the flagship's authority principle into repository policy: AI may advise, while trusted checks and an accountable human retain merge authority.
Source, attribution, and license
This long-form feature substantially adapts and reworks material from From Vibe Coding to Agentic Software Engineering: Toward a Star Trek Future of Building Things — Long-form Feature and Academic Companion, credited on its Internet Archive record to ChatGPT, and the later definitive Make It So—But Run the Tests, credited on its Internet Archive record to ChatGPT. Both records display CC BY-NC-SA 4.0. Changes include a web-first organization, refreshed empirical interpretation, independently linked primary sources, a new measurement section, and removal or qualification of unsupported forecasts. The earlier From Vibe Coding to Agentic Symbiosis, credited there to Gemini 3.1 Pro, was used only for claim quarantine; its unverified numbers were excluded. The supplied provenance identifies Gemini 3.1 Pro Deep Research; Archive.org's Creator field lists Gemini 3.1 Pro without the Deep Research qualifier. This adaptation is shared under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
AI assistance was used in research synthesis and editing. A human editor remains responsible for source selection, checking claims, and publication decisions.






Top comments (0)