I build the whole argument walking the dog. Beginning, middle, the objection someone will raise, my answer to it. It is all there. I get home, open the editor, write the first sentence — and the second one does not come. I did not forget the subject. I forgot the sequence, which was the only hard part.
The easy explanation is speed: people speak three to four times faster than they type, so the thought outruns the hands. But that does not explain why an idea survives twenty minutes of walking and dies in thirty seconds of blank page.
What actually happens is that writing bills you twice at once. You have to decide what to say and decide how to say it in the same instant, and the second job interrupts the first. Every time you stop to pick between two words, you drop the thread. Speaking does not bill you twice. Out loud you are forced to be sequential — one thing after another — but you are not forced to be final. You can say "no, wait, it is the other way around" and keep going. The editor has no such sentence.
It is the same mechanism as the programmer's rubber duck: explaining the problem out loud to an inanimate object solves the problem, and it solves it before the object answers. It is not the conversation that works. It is being made to put things in order.
So what works is splitting the two jobs into two moments. Talk first, without editing — not even the repetitions, the false starts, the "I mean". Then, and only then, tidy up. Spoken text is not the text; it is the raw material for it. The part you kept losing was the part you could not reconstruct. Tidying up you always knew how to do.
And it is not all or nothing. Lists, tables, numbers, proper nouns — anything with an exact shape is faster to type than to say and then correct. What goes to speech is what has order: the argument, the objection, the reason. In practice you alternate inside the same piece of work.
Where it stopped being a method and became a tool problem
Walking the dog, dictating a piece of market research into the ChatGPT app: the sources, the theses, what was worth testing. Partway through, something on the screen suggests it is not going well. I stop, send what I have said so far, and keep recording the rest — because I had not finished. Then I submit it and see the error.
The first frustration is immediate, and the fix is a bad one. A stretch did not make it, so I have to walk back to where it dropped, repeat the points I had emphasised so they are not lost, and splice back into the argument. It costs more than it sounds: it is not repeating words, it is rebuilding order — and order is the expensive part.
The second one is worse, and it is the same as the opening of this piece. When I cannot walk it back right there, I leave it to pick up at home. Except by then it is not splicing, it is remembering from scratch. And the first thing to go is the sequence, which was the only hard part. The tool failing handed me back exactly the problem that speaking had solved.
It happened again on a different occasion, recording an account of a dream — one of those eight-minute ones you only get to tell once: I wake up, write keywords in a notebook so I do not lose them, and narrate it in detail later. The transcription hung. The notebook keeps the subject; it does not keep the account.
What a dropped recording costs you is not the text: it is the one pass in which that reasoning existed whole.
⭐ That is what the command-line version was written against, and the lesson that stuck is not a technical one. What it does differently makes the transcription no better at all — it exists so the whole thing arrives on the other side. A wrong word you reread and fix in thirty seconds. It is the other thing that has no fix.
Where this still does not work
A transcript is not a finished draft, and anyone promising otherwise is selling something. What comes out follows your order of thinking, not your order of writing — those are different, and the second one is still work.
And the bigger limit has no tool fix. Speaking preserves reasoning; it does not manufacture it. If you do not know what you want to say, twenty minutes of audio gets you twenty minutes of someone not knowing what they want to say — now in writing.
What to take from this
The step that changes the most and costs the least is not the recording. It is deciding, before you start talking, where the text goes when you stop.
Without that destination, the walk produces one more audio file in your life. With it, the walk becomes work already started: the text lands where the next step happens — the session you were going to code in, the project note, the open draft, the agent that picks it up from there. Which place does not matter. Having one does.
You can test it in a day: next time you build an argument while walking, decide the destination before you hit record. The walk is still where the thinking happens. The only thing that changes is that it stops dying on the way home.
This article was prepared with AI assistance.
Top comments (0)