While turning my recent agent experiments into articles, I ran into an uncomfortable gap. I could follow the code, connect tools, and use AI to help debug a problem. Explaining why the system worked that way was much harder.
My notes covered ReAct, Reflection, and frameworks including LangGraph. But having a page for each topic did not mean I had a connected understanding of the system.
I decided to revisit three examples already in my own notes: a serial-device gateway, a Reflection coding exercise, and a LangGraph question-answering workflow. The exercises below are my next steps, rather than results I have already achieved.
1. Trace one tool call through a real project
In an earlier project, I used Python and pyserial to put a serial-device gateway behind an HTTP API. This let AI participate in testing a hardware module through tools.
At the time, I focused on getting the test workflow running. For revision, I want to follow one call across the boundaries:
Task and previous results
|
Model proposes a tool name and arguments
|
Execution program sends an HTTP request
|
Gateway uses pyserial to send a command and read the reply
|
Execution result is included in a later model request
|
Next action or final report
This is the simplified path in that project. Splitting it into stages gives each failure a place to investigate:
- A wrong command suggests checking how its arguments were produced.
- A request that never reaches the gateway points to the execution environment or interface.
- A missing device reply calls for checking serial communication.
- An incorrect report calls for checking what result reached the model and how the conclusion relates to it.
My first exercise is to take an existing test record and line up the task input, tool arguments, gateway response, and final report. If all I can find is “test passed,” I still need to locate the reply that supports that conclusion.
I do not need to rebuild the project to do this. I need to explain one complete execution path. Original serial-gateway notes, in Chinese
2. Read the Reflection controller, not just the review
One of my Reflection exercises generated a Python function for finding prime numbers. Looking back at its output, I noticed a review that said no algorithmic improvement was necessary, followed by another code-optimization step.
That observation alone does not establish a bug. The feedback also discussed engineering improvements, and the task may have allowed further changes. It does raise a useful question: what actually makes this program continue or stop?
A model can produce an assessment, but the program still has to turn that assessment into an action. I want to inspect the stopping condition, the expected review format, the parser, and the iteration limit.
My revision exercise has two parts.
First, I will test the control flow with fixed feedback: a passing review, a request for changes, and an unrecognized response. I will predict the next branch before running it. This separates the program's control logic from variation in a real model response.
Second, I will test the generated function itself. If its contract is to return primes no greater than n, a few initial cases are:
# Expected results for the exercise; not a complete test suite.
expected = {
1: [],
2: [2],
10: [2, 3, 5, 7],
}
I still need broader coverage and boundary checks. The point is to compare the code's behavior with the reviewer's opinion. A more enthusiastic review does not establish that the implementation improved.
That makes Reflection concrete for me: saved results, feedback passed into the next request, explicit stopping behavior, and tests. My Reflection notes, in Chinese
3. Inspect state when the final answer looks convincing
My LangGraph notes include a three-stage assistant: understand the question, search, and generate an answer. Its state contains fields such as search_query, search_results, and step.
I used to treat those fields as implementation details. Now I see them as the places to ask basic questions: what did the search node receive, where was its output stored, and what did the answer node read?
In that example, the answer stage can fall back to the model's existing knowledge after search fails. A fluent final answer therefore does not prove that search succeeded.
I plan to hold the question fixed and inspect three inputs to the answer stage:
| Search outcome | What I want to check |
|---|---|
| Results returned | Can the answer's important claims be traced to those results? |
| Empty results | Does the answer confuse “not found” with “does not exist”? |
| Search error | Does it make clear that the external material was not retrieved? |
There is also a task-level decision: should this workflow continue answering after retrieval fails? For a question that depends on current information, the model's existing knowledge may be insufficient. Stopping and explaining the missing evidence may be more appropriate.
The fields in a workflow state give me something to inspect between nodes. They do not, by themselves, establish persistent storage or long-term memory. My LangGraph notes, in Chinese
A revision sequence with something to inspect at every step
Instead of starting a new project for every concept, I want to reuse the same sanitized logs and examples:
- Inspect one model request. Identify the current input, prior messages, and response. Change one piece of context and observe what changes.
- Trace one read-only tool loop. Save the arguments, returned result, and stopping reason.
- Revisit Reflection. Add behavioral tests and explicit exit conditions, then compare before and after.
- Rebuild a small workflow in LangGraph. Map state between nodes and inspect empty-input and failure paths using the same material.
- Return to the serial gateway. Validate the workflow with saved replies or a mock interface before scheduling tests on real hardware.
The framework step should help me identify what the framework manages on my behalf. I first want to recognize the loop, state, tools, and logging in a small example. My notes on moving from manual implementation to frameworks provide a starting point.
Keep my own prediction in the learning loop
AI has saved me time writing code, investigating problems, and organizing notes. During study, though, taking the answer immediately can skip the part where I need to think.
So I want to write my understanding before asking for help. For example:
I think this program stops after the reviewer approves the result. Help me locate its actual exit condition and suggest an input that could disprove my understanding. I will predict the outcome before running it.
After each exercise, I want to keep four short notes: what I expected, what I observed, why they differed, and where I would look first next time.
I will keep writing about my projects. I also want the next article to make three things easier to explain: what I built, why it behaves that way, and how I know the result is trustworthy.
I'm Aiclaw, a connected-vehicle software developer learning and building with AI and agents. My engineering records and learning notes are on my blog; the linked original notes are in Chinese.
Top comments (1)
Writing your prediction down before running the test, like "I think this stops after the reviewer approves, help me find the real exit condition," is a really disciplined way to stop AI from short circuiting your own understanding. The point about a fluent final answer not proving search actually succeeded is worth its own post. That failure mode is easy to miss when the output just reads well.