DEV Community

Tej Pandya
Tej Pandya

Posted on Fully Autonomous

How to Build an AI Assistant That Finishes Tasks: Nine Checks Beyond the Model

A good model can choose the right tool and still leave the job unfinished. The browser might time out. A process might restart after a successful write. A payment might need approval the system never collected.

I build AI for sales teams at GrowEasy.ai. My view is that a useful assistant needs nine parts around the model: a workflow, a task board, an audit trail, a browser, an API executor, a payment layer, memory, human escalation and scoped identity. Here is how I would test each part before putting it in front of a customer.

1. Restart a workflow without repeating completed steps

Use code for steps whose order is known. Ask the model to judge the messy parts.

For a meeting-prep assistant, fetching the calendar and checking whether a brief already exists are ordinary steps. Deciding which parts of a long email matter for the meeting is a model task.

Anthropic distinguishes predefined workflows from agents that direct their own processes. Neither approach wins everywhere. Pick the simplest one that fits the job.

Test: stop the worker after a step succeeds, then restart it. Does it continue from the right place? Temporal's Durable AI documentation describes workflows that resume after crashes. That capability still needs testing against your own side effects.

2. Keep tasks open until the result is confirmed

A chat history is a poor substitute for job state. Give each job a stable ID, an owner, a state and an expected result. Include a waiting state for work that needs a reply or a decision.

Consider a lead follow-up. "Message prepared" and "message sent" are different states. So are "sent" and "customer replied". If your board uses one green tick for all of them, your reporting is already wrong.

Test: return a successful tool response with no stored result. The job should stay incomplete until the outcome can be checked. Then replay the same job ID. It should not create another message.

This is a design recommendation, not a promise that adding a board makes an agent reliable.

3. Explain the last confirmed step and the next decision

Developer traces help explain the path through a run. The OpenAI Agents SDK tracing documentation covers model generations, tool calls, handoffs and other events.

Keep that technical record, but add a plain summary: what was requested, what changed, what failed and what needs a decision. Link the summary to the result. Store only the information the job needs; an audit trail should not become a second copy of every private conversation.

Test: give someone who did not build the system a failed job. Can they identify the last confirmed step and the next safe action without reading raw logs?

4. Check the saved result after every browser action

Some portals have no usable API. Browser access matters because the assistant has to work where the task actually lives.

But a successful click is not proof of a successful booking. The page can reject an input, refresh into an old state or show a confirmation for the wrong item. Browser work needs a check of the final page and its saved result.

The original WebArena paper found a large gap between its baseline agents and people. It is an older controlled benchmark, not a current score for every browser agent or a measurement of live-site reliability.

Test: change a form field label and make a submit return an error. The assistant should notice, stop and keep the job open, rather than report success because it clicked the button.

5. Check for a successful write before retrying an API

Where a suitable API exists, I would use it before the browser. Structured inputs and outputs make checks easier. They do not remove uncertainty.

The difficult case is a timeout after a write. The server may have completed the action even though the client got no response. Before retrying, read the current state. Use the provider's idempotency support where available, scoped permissions and credentials kept out of model-written logs.

Test: simulate a lost response after a successful write. Does the assistant check for the result before trying again? Then expire a credential. Does it report the access problem without marking the underlying task done?

6. Recheck payment approval when the price changes

A payment layer is an authority boundary. The assistant needs to know who approved the purchase, what was approved, the total and the conditions. A broad goal such as "book travel" should not turn into unlimited spending.

Google's Agent Payments Protocol announcement describes signed mandates as evidence of intent and authorization. A protocol provides a way to carry that evidence; it does not make every shopping agent trustworthy.

Test: change the total after approval. The old approval should no longer silently authorize the changed purchase. Also test an uncertain payment response before allowing any retry.

7. Test that memory updates and deletes take effect

Store useful context beyond the current run, but separate facts from old guesses. A changed preference should replace the old one, not sit beside it as another possible answer. OpenAI's memory announcement describes remembering useful details across chats and controls to manage them.

Test: change a saved preference. The next task should use the correction. Then delete it and check that the assistant stops relying on it.

8. Pause for approval and resume the same job

The workflow needs a way to pause when it lacks authority or a reliable answer. LangGraph's interrupts documentation describes saving state while waiting for external input, then resuming it.

Test: require approval halfway through a job. The assistant should ask one clear question, preserve progress and resume the same job only after the answer arrives.

9. Test account selection and permission limits

The assistant needs to use the right identity with the right access for each service. Cloudflare's signed-agents work describes a way for sites to verify agent identity. That verification is separate from permission to use an owner's account.

Test: give it access to two accounts. It should select the authorized one, keep secrets out of logs and stop if the required permission is missing. A valid credential is not permission for every action it allows.

Run one task through interruptions, errors and missing approval

I would start with one bounded task and run it through these failure cases. A good answer in a chat window proves the model can respond. A checked outcome, with a readable record and no duplicate action, is a different standard.

These nine parts are my proposed foundation. They are not an exhaustive list. Their value is that each one answers a practical question: can this assistant keep working without making its owner carry the bookkeeping?

Tej Pandya, founder of GrowEasy.ai.

Top comments (0)