While connecting a Mac mini to Jarvis, my AI assistant, I kept coming back to a small question: what can I prove actually happened?
The setup produced a series of ordinary milestones. A command ran on the intended machine. A file transferred and could be read back. A browser opened the requested page. A device on my home network returned its status. The connection recovered after a deliberate connection restart.
Each result established a specific capability. Together, they gave the assistant a practical way to carry out work.
Then I asked it to research an article and prepare a Medium draft. It reached a login screen. I signed in, and the work continued.
For someone building agent workflows, that handoff is a useful place to start. Understanding a task and completing it depend on different parts of the system.
Intelligence needs execution capacity
In “The eternal complement”, Hemanth Asirvatham and Elliott Mokski argue that intelligence needs complementary capacity to turn ideas into progress. Their essay is hosted on OpenAI’s Intelligence Age platform and explicitly represents their own views, not OpenAI’s position.
The authors use “institutional intelligence” to describe the coordination behind execution: funding, supply chains, laws and the many correct local actions that make an ambitious project possible.
An agent workflow has its own supporting work. It needs the right input, access to the right application, a clear scope of authority and a way to check the result.
A model may produce an excellent plan while those pieces remain missing. Making the plan more sophisticated does not automatically supply the missing login or establish that a file was saved.
My computer setup is a small example of this relationship. It does not establish anything about the future of superintelligence. It does make the dependency visible.
Define completion before choosing tools
“Write an article” leaves several possible stopping points.
The model could return text in chat. It could save a local file. It could create a draft in the intended account. It could publish the reviewed version and verify the public page.
Each is a different deliverable.
For a draft workflow, I would define completion as: the intended title, body and links are saved in the correct account, and the content survives a reload.
For an authorized publishing workflow, I would add: the public article exists, the author is correct, and the published body matches the intended version.
Those definitions give the implementation something concrete to verify. They also prevent a fluent final message from becoming the only evidence that the work succeeded.
Before adding another tool, write the acceptance condition in one sentence. Then identify which tool can produce evidence for it.
Treat the application’s saved state as evidence
Browser automation introduces an easy trap. A successful input operation is not the same as a successful application write.
Text can land in the wrong field. A rich editor can show content before it has persisted. A connection can drop while the application is saving.
For this kind of workflow, a small write probe is useful. Enter a small amount through the intended UI, inspect where it landed, and confirm that the application created the expected saved artifact before inserting the full body.
After the full edit, reload and compare the persisted content with the intended text. Check links as well as paragraphs. Inspect the rendered page where layout matters.
This is why screenshots and text readback complement each other. A screenshot can expose a formatting problem. A full text comparison can expose missing content that sits outside the viewport.
The evidence should match the claim. A working browser establishes browser access. Recovery after restarting a connection establishes that recovery path. Unattended recovery after a machine reboot would need a separate test.
Keep uncertainty in the state model
The most important failure case often happens after an external action starts.
Suppose an agent clicks Publish and then loses its connection. The article may already be public. The agent may simply have missed the response.
Treating that as an ordinary failure and retrying immediately can repeat the action.
A content workflow can distinguish these states:
| State | What it establishes | What happens next |
|---|---|---|
| Prepared | Local content and the intended account are known | Create or reconcile the draft |
| Saved | The application retained the intended content | Continue within the authorized scope |
| Publishing | Publication intent was recorded before the external action | Observe the result |
| Uncertain publication | The action may have succeeded, but evidence is incomplete | Inspect the existing story and public profile |
| Published | The public artifact and its content were verified | Record the result and stop |
This table is a practical design suggestion, not a framework prescribed by the essay.
The record needs enough information to reconcile the operation: the existing draft identifier, intended account, content hash, and any public URL already returned. Write the intent before the action, and preserve those identifiers when an error occurs.
A local record cannot make an external website transaction atomic. Its value is that the next attempt can investigate the same operation instead of starting over.
Once publication is verified, a later cleanup error should not silently downgrade the article to “unpublished.” Record the cleanup problem separately.
Put autonomy inside a clear scope
A workflow needs to know what it is allowed to do.
That can include standing authorization to complete a recurring task. The agent should carry out that authorized work without asking the same question at every step.
It should also recognize actual exceptions: an expired login, a different account, a platform refusal, or an outcome it cannot establish.
Persist the exception and the work already completed. A vague “something went wrong” gives the next run very little to work with. A known draft URL and an uncertain publication state give it a specific investigation.
Improve the next blocked step
The essay explores two uncertain directions. In a civilization of depth, better intelligence makes much more effective use of existing knowledge and resources. In a civilization of width, it generates worthwhile projects faster than supporting capacity can keep up.
The authors do not predict which direction will dominate. The idea that machine intelligence could spend most of its time on routine organizational work belongs to the width scenario.
For the workflow in front of me, I do not need to settle that question.
I can choose one recurring task, define its finished result, and follow it through every dependency. Where does it wait for access? Where is the outcome assumed rather than checked? What happens if the next external write succeeds but its response never arrives?
Answering those questions gives me concrete implementation work. It is the ordinary work that lets an intelligent system carry a useful idea through to a result.
Adapted from my original Medium article, “The Boring Work That Makes AI Useful”.
Primary source: “The eternal complement,” by Hemanth Asirvatham and Elliott Mokski. The authors write as independent voices hosted by OpenAI.
Top comments (0)