DEV Community

Cover image for Shipping practices from five AI engineering episodes
Conor Bronsdon
Conor Bronsdon

Posted on Originally published at chainofthought.show

Shipping practices from five AI engineering episodes

A demo is cheap. Shipping it means you can say what you reviewed, what you designed, what you measured, and who owns production when the model is wrong. These five episodes are the set I would hand a builder who already has a demo and needs a practice. Transcripts and chapters for every episode are on the Chain of Thought episodes page.

What do you review when an agent writes the code?

Anush Elangovan, AMD's VP of AI Software, describes an agent-first workflow at real scale, the case where AI writes most of the code. He runs 10 to 12 Claude Code agents in parallel, burns 6.5 billion tokens a week, and rewrote a 25-year-old Slurm replacement in Rust overnight. He argues that normal SDLC is dead. He treats testing as the new code review. He puts the rest in one line: software is just tokens. I would take the testing claim as the operating rule. Once that much code is machine-written, reading every diff will not keep up. Write tests that state the behavior you will accept, run them as the gate, and spend review time on those tests. The overnight rewrite is a reminder that the tokens show up before an old review process is ready. Treating software as tokens only works if you still own the definition of a passing test.

What should you design before you pick a model?

Aishwarya Srinivasan, Head of AI Developer Relations at Fireworks AI, argues that most AI agents are built backwards, starting with models instead of system architecture. The shift she describes runs from picking a model to architecting a system, and from prompt engineering to context engineering. In her account, a production agent needs careful orchestration of multiple models, memory systems, and tool calls, along with evaluation-driven development. Benchmarks alone will not save you. She also covers human-in-the-loop control of autonomy, LLM-as-judge checks, open source models such as DeepSeek v3, and compliance in regulated environments. Before I named a model I would write down the context the agent receives, the tools it may call, the memory that persists, and the eval that fails the build. A person stays in the loop wherever a bad action is expensive to undo. You can swap a model inside a system you have already specified.

How do you get from a vibe check to a shipped feature?

Malte Ubl, CTO of Vercel, describes agents as a new type of software for solving flexible tasks. The episode is about the engineering between a demo and a shipped feature. He points at Vercel's developer-first stack, including the AI SDK and AI Gateway, as a way to move from a quick proof of concept to a trusted, production-ready application. I would hold a team to that order. Name the flexible task in one sentence a user would recognize. Build the proof of concept on the developer surface you are willing to run in production, which is the job he gives that SDK and gateway. Treat the result as shipped when you trust it the way you trust the rest of the product. The vibe check ends when the task has to survive real users.

What should you look at before you trust a metric?

Hamel Husain, of Parlance Labs, says the first question he usually gets is which tools to use, and that it is the wrong question. The episode is about the mindset that separates engineers who ship reliable AI from engineers who chase metrics. He argues that vanity dashboards and a buffet of metrics create a false sense of security, and that this is no substitute for customized evals tailored to domain-specific risks. The discipline he describes is error analysis: look at the data yourself, find real-world failures, draw qualitative notes from logs, and turn those notes into quantitative guardrails. That is the experimentation loop he ties to taking a prototype into production. I would not add a metric until I could point at a failure I had read in our own logs. The first artifact is a short list of domain failures, each with an example and an eval that fails when that failure comes back. Tools can wait.

What does AI leave with the engineering team?

Charity Majors, co-founder and CTO of Honeycomb and the person behind charity.wtf, starts from her article "Generative AI is not going to build your engineering team for you." The episode is about production reliability, and about what AI fixes and what it leaves inside your hardest engineering problems. The question she and I take up is whether generative AI makes it easier to build, lead, and sustain a high-performing team. I would plan as if the article title is still true. Keep a team that can build the system, lead the work, and hold production reliability when the model is wrong or unavailable. AI can sit inside that team. Headcount, leadership, and the reliability of the system stay with the people who own it.

Checklist

  • Gate agent-written code with tests. Anush Elangovan calls normal SDLC dead and treats testing as the new code review.
  • Specify context, tools, memory, and evals before you commit to a model. Aishwarya Srinivasan's warning is that benchmarks alone will not save you.
  • Carry the flexible task from a proof of concept to a trusted production app on a stack you mean to keep, which is how Malte Ubl frames the path from vibe check to a shipped feature.
  • Read real failures before you trust a dashboard. Hamel Husain calls a buffet of metrics a false sense of security.
  • Staff and lead the team as if generative AI will not build it for you. Charity Majors ties that limit to production reliability and the hardest engineering problems.

More on this topic, with the related episodes, is on Chain of Thought. It draws on this episode.

Subscribe to the Chain of Thought newsletter for new episodes and write-ups like this one.

Drafted with AI assistance from the episode transcripts.

Top comments (0)