<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Valipireddy Kowshik</title>
    <description>The latest articles on DEV Community by Valipireddy Kowshik (@k0wsh1k_0x).</description>
    <link>https://dev.to/k0wsh1k_0x</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4125336%2F8808a527-e17c-475f-9cd0-d996d7e92fe7.jpg</url>
      <title>DEV Community: Valipireddy Kowshik</title>
      <link>https://dev.to/k0wsh1k_0x</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/k0wsh1k_0x"/>
    <language>en</language>
    <item>
      <title>Why I Walked Away From "Just" Full-Stack Dev to Build AI Agents</title>
      <dc:creator>Valipireddy Kowshik</dc:creator>
      <pubDate>Wed, 16 Sep 2026 09:05:32 +0000</pubDate>
      <link>https://dev.to/k0wsh1k_0x/why-i-walked-away-from-just-full-stack-dev-to-build-ai-agents-4127</link>
      <guid>https://dev.to/k0wsh1k_0x/why-i-walked-away-from-just-full-stack-dev-to-build-ai-agents-4127</guid>
      <description>&lt;h2&gt;
  
  
  The moment I realized I'd already built one
&lt;/h2&gt;

&lt;p&gt;A few months ago I was deep in the architecture of a voice-first AI technical interviewer — a platform that speaks to a candidate out loud, watches their keystrokes live over a WebSocket connection, nudges them with a hint when they're stuck, and at the end, outputs a structured hiring scorecard instead of a vague "went okay" note for a recruiter to guess at.&lt;/p&gt;

&lt;p&gt;I wasn't building a chatbot. I was building something that observed, decided, and acted, on its own, inside a real workflow that used to require a human sitting in the room. Somewhere in the middle of wiring up the WebSocket editor deltas and the finite state machine that decided when to intervene, it hit me: this &lt;em&gt;is&lt;/em&gt; an AI agent. I just hadn't been calling it that.&lt;/p&gt;

&lt;p&gt;That realization is the reason I now spend most of my time building AI agents for businesses, not because I abandoned web development, but because I finally had a name for the kind of engineering I'd already drifted toward.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I actually started
&lt;/h2&gt;

&lt;p&gt;For the last three-plus years I've been a full-stack and React Native engineer, not an "AI person." My days looked like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Shipping patient health flows at &lt;strong&gt;Tap Health&lt;/strong&gt; — glucose tracking, meal logging, multi-step onboarding — in React, Next.js, and React Native, and chasing down the unglamorous bugs that actually break trust: timezone mismatches, state sync issues, navigation deadlocks.&lt;/li&gt;
&lt;li&gt;Building doctor portals, HR platforms, and admin dashboards from scratch at &lt;strong&gt;ZarvisGenix&lt;/strong&gt;, wiring up role-based access control, real-time data sync, and secure APIs with Prisma and PostgreSQL.&lt;/li&gt;
&lt;li&gt;Sitting in rooms with investors and clinicians, translating what they actually needed into software that had to work correctly the first time, because in health tech, "close enough" isn't a shippable standard.
None of that reads like an AI career on paper. But there's a thread running through all of it: I was always the person who got handed the workflow nobody had automated yet, and asked to make it disappear into software.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The pattern I couldn't unsee
&lt;/h2&gt;

&lt;p&gt;At ZarvisGenix, I also built and integrated AI-powered speech-to-text pipelines, intelligent voice transcription that fed straight into automated workflows. It wasn't the headline feature of the project. It was infrastructure. But it was the first time I noticed something: the AI part of the system was never the hard part. The hard part was everything &lt;em&gt;around&lt;/em&gt; it, the authentication, the rate limiting, the data schema, the error handling for when a transcription came back malformed, the part where the "smart" feature actually had to survive contact with real users and real edge cases.&lt;/p&gt;

&lt;p&gt;That's a pattern I kept seeing repeat, project after project. Everyone wanted "AI" bolted onto their product. Almost nobody had thought through what happens when the AI is wrong, or slow, or needs to talk to four other systems to do its job. The gap wasn't a model problem. It was an engineering problem. And engineering problems are exactly what three years of production full-stack work had trained me to solve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why most "AI agents" on the market fall apart
&lt;/h2&gt;

&lt;p&gt;Once I started paying attention, I noticed the same failure mode everywhere, on Fiverr, in client requests, in demos I got shown by other teams: someone wraps a system prompt around GPT or Claude, calls it an "agent," and ships it. It works beautifully in the demo. It falls over the moment a real user asks something slightly off-script, or the business needs it to actually &lt;em&gt;do&lt;/em&gt; something instead of just reply, like check a calendar, update a CRM record, or hand off to a human at the right moment.&lt;/p&gt;

&lt;p&gt;That's not an agent. That's a chatbot wearing an agent's name tag.&lt;/p&gt;

&lt;p&gt;The AI Technical Interviewer project is the clearest proof I have that the difference is real. It didn't just generate text, it had to synthesize speech in under 300ms to feel human, track state deterministically so it never lost context mid-interview, evaluate code in a live sandbox, and produce a structured decision a hiring manager could actually act on. Every one of those requirements is a systems engineering problem wearing an AI costume. Get the engineering wrong and it doesn't matter how good the underlying model is.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq33dvqlywpafiyrdy5uf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq33dvqlywpafiyrdy5uf.png" alt="Custom AI agent gig cover" width="680" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I think this makes me a better fit than most agent builders
&lt;/h2&gt;

&lt;p&gt;Here's the honest version, not the sales pitch version: a lot of people entered the "AI agent developer" space over the last year with no software engineering background at all. They learned a no-code automation tool, connected a couple of APIs, and started calling themselves AI agent specialists. Some of that work is genuinely fine for simple use cases. But it tends to break exactly where it matters most, at scale, under load, or when a client's business actually depends on the thing not failing silently at 2am.&lt;/p&gt;

&lt;p&gt;I came at this from the opposite direction. I didn't start with prompts and work backward into engineering. I started as a production engineer, obsessive about frame budgets, latency, and deterministic state, and only later started pointing that discipline at AI systems. That means when I build an agent for a business, it's not a demo wearing a business's logo. It's:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Connected to real tools, CRMs, calendars, databases, not just answering in a vacuum&lt;/li&gt;
&lt;li&gt;Built with proper error handling for the moment the model gets something wrong, because it will&lt;/li&gt;
&lt;li&gt;Documented and handed off cleanly, so a client isn't permanently dependent on me to keep it running&lt;/li&gt;
&lt;li&gt;Designed around the actual workflow a business has, not a generic template I resell to everyone
## What I actually build now&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F20svycethjcd296eha63.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F20svycethjcd296eha63.png" alt="Custom AI agent gig gallery" width="680" height="680"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These days, most of my client work falls into one of these buckets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Lead qualification and booking agents&lt;/strong&gt; that engage a lead the moment they message, ask the right questions, and book directly into a calendar&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Knowledge-base agents&lt;/strong&gt; trained on a business's actual PDFs, docs, or website content, so answers are accurate instead of generic&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-channel agents&lt;/strong&gt; that work across a website, WhatsApp, and voice, from one underlying system&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Workflow automation&lt;/strong&gt; using n8n to connect the agent into the tools a business already runs on
None of this replaces the full-stack work. If anything, it's the same skill set pointed at a newer kind of problem. The React Native apps, the Next.js dashboards, the Prisma schemas, all of that is still the foundation. The agents just sit on top of it now.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  If you're weighing the same move
&lt;/h2&gt;

&lt;p&gt;If you're a web or full-stack developer looking at the AI agent space and wondering whether it's a real shift or just a rebrand, here's what I'd tell you: it's real, but only if you bring the engineering discipline with you. The market is already full of people who can wire up a chatbot in an afternoon. It is not full of people who can make that chatbot survive a real business's actual workflow. That gap is where the work is.&lt;/p&gt;

&lt;p&gt;I build custom AI agents for businesses using GPT and Claude, you can see exactly what that looks like &lt;a href="https://www.fiverr.com/kowshik_v_/build-a-custom-ai-agent-for-your-business-using-gpt-or-claude" rel="noopener noreferrer"&gt;on my Fiverr gig&lt;/a&gt;. If you're trying to figure out whether your business actually needs one, or you're a fellow developer thinking about making the same jump, I'm happy to talk through it in the comments.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>webdev</category>
      <category>automation</category>
    </item>
    <item>
      <title>Why I Moved Us Off Microservices Back to a Monolith (And What the Original Decision Got Wrong)</title>
      <dc:creator>Valipireddy Kowshik</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:47:09 +0000</pubDate>
      <link>https://dev.to/k0wsh1k_0x/why-i-moved-us-off-microservices-back-to-a-monolith-and-what-the-original-decision-got-wrong-2p9</link>
      <guid>https://dev.to/k0wsh1k_0x/why-i-moved-us-off-microservices-back-to-a-monolith-and-what-the-original-decision-got-wrong-2p9</guid>
      <description>&lt;p&gt;We split a perfectly working app into eight microservices because it was 2023 and that's what "scalable" teams did. Eighteen months later, I was the one writing the RFC to put most of it back together.&lt;/p&gt;

&lt;p&gt;Nobody likes admitting that. It feels like undoing someone else's work, or worse, undoing your own. But I'd rather write an honest postmortem than keep pretending the architecture was fine while three engineers spent half their week just keeping services talking to each other.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why we split it up in the first place
&lt;/h3&gt;

&lt;p&gt;The reasoning at the time sounded solid on paper. Independent deployability. Teams owning their own services. The ability to scale the busy parts of the app without scaling the whole thing. Every blog post about companies at our stage said some version of the same thing, and it was easy to believe we were behind if we hadn't already made the jump.&lt;/p&gt;

&lt;p&gt;What none of those blog posts mentioned, or maybe just assumed was obvious, was the size of team those benefits actually require. We had twelve engineers. Splitting the app into eight services meant most services had one, maybe one and a half people who actually understood them end to end. The "independent teams owning independent services" model doesn't work when the teams don't exist yet. We built the architecture for a company we hoped to become, not the one we were.&lt;/p&gt;

&lt;h3&gt;
  
  
  What actually broke
&lt;/h3&gt;

&lt;p&gt;The failures weren't dramatic. There was no single outage that made everyone say "okay, this was a mistake." It was slower and more corrosive than that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Debugging turned into archaeology.&lt;/strong&gt; A bug that used to mean stepping through one codebase now meant tracing a request across four services, three of which had inconsistent logging, to find out where the data actually went wrong. What used to take twenty minutes started taking half a day, and that was on a good day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Local development got painful.&lt;/strong&gt; Spinning up the full stack locally meant running eight services, each with its own environment variables, its own database, its own quirks about what port it wanted. New engineers took two weeks just to get a working local setup instead of two hours.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deployments needed careful choreography.&lt;/strong&gt; Services depended on each other in ways that weren't always obvious until deploy day. We'd ship a change to one service, break another, and spend the next hour figuring out which one actually needed to go out first. The "independent deployability" we split up for became its own coordination problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;We were paying an infrastructure tax with no matching benefit.&lt;/strong&gt; Eight services meant eight sets of CI pipelines, eight sets of monitoring dashboards, eight things that could each fail independently at 2am. We were paying operational complexity for scale problems we didn't have yet. Our actual traffic could have run comfortably on a single well built service.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nobody could hold the whole system in their head.&lt;/strong&gt; This one mattered more than any single technical issue. When something went wrong, there was no longer one person who could reason about the entire request path. Every incident started with a scramble to figure out who even understood the piece that broke.&lt;/p&gt;

&lt;h3&gt;
  
  
  The decision to go back
&lt;/h3&gt;

&lt;p&gt;I didn't propose the monolith because microservices are bad. They're not. They solve real problems for teams at a certain scale with certain organizational boundaries already in place. I proposed it because we'd adopted the pattern before we had the problem it solves, and we were paying the cost of that mismatch every single week.&lt;/p&gt;

&lt;p&gt;The RFC was blunt about it. We are not Amazon. We do not have twelve teams that each own a service end to end. We have twelve engineers who spend more time on service boundaries than on the product. Here is what a single, well organized codebase with clear internal module boundaries buys us back.&lt;/p&gt;

&lt;p&gt;It wasn't a popular conversation at first. Some of that split had taken real effort, and nobody wants to hear that effort didn't pay off. But once we framed it as "what does our actual team, at our actual size, need right now" instead of "what did we lose face on," it got easier to agree on.&lt;/p&gt;

&lt;h3&gt;
  
  
  What I'd tell myself eighteen months earlier
&lt;/h3&gt;

&lt;p&gt;Architecture decisions should follow the shape of your team and your actual traffic, not the shape of the team and traffic you're hoping to have in two years. If you can't clearly name which team owns which service, and mean it, you probably don't need separate services yet. You can build a monolith with clean internal boundaries, module by module, and split it later when a specific piece genuinely needs to scale or deploy independently. That split will be easier to do later, with real evidence behind it, than it is to undo now.&lt;/p&gt;

&lt;p&gt;We didn't fail because microservices are wrong for us forever. We failed because we borrowed an architecture built for a different sized team and assumed the benefits would show up automatically. They don't. They show up when the organizational structure that justifies them already exists.&lt;/p&gt;

&lt;p&gt;If your team is mid split right now and it's starting to feel heavier than it should, that feeling is worth listening to. It was for us, even if it took longer than I'd like to admit before I said it out loud.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>devops</category>
      <category>security</category>
    </item>
    <item>
      <title>What a Bad Project Taught Me (That No Great Project Ever Could)</title>
      <dc:creator>Valipireddy Kowshik</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:33:54 +0000</pubDate>
      <link>https://dev.to/k0wsh1k_0x/what-a-bad-project-taught-me-that-no-great-project-ever-could-ki7</link>
      <guid>https://dev.to/k0wsh1k_0x/what-a-bad-project-taught-me-that-no-great-project-ever-could-ki7</guid>
      <description>&lt;p&gt;Nobody writes a retrospective about the sprint that went smoothly. We move on, ship the next feature, and forget the details. But the project that hurt? That one stays with you line by line. It becomes the mental checklist you run every time someone says "let's just build it."&lt;/p&gt;

&lt;p&gt;I recently came off a project like that. It wasn't a technical disaster the code shipped, the app worked, the demo looked fine. What broke it was process: no real user research, no design principles anyone could point to, and a growing habit of treating AI-generated "research" as a substitute for actually talking to users. Here's what I took away from it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The pattern I kept seeing
&lt;/h3&gt;

&lt;p&gt;In the AI era, it's tempting for product managers to skip the slow, expensive parts of building a product. Why run five user interviews when you can ask an LLM to "research user needs for a fintech onboarding flow" and get a confident, well-formatted answer in ten seconds?&lt;/p&gt;

&lt;p&gt;The problem isn't that AI research is useless  it's a decent starting point for hypotheses. The problem is when it quietly replaces the step where you validate those hypotheses against real humans. On this project, feature decisions increasingly traced back to "the research says users want X," where "the research" meant a chat transcript, not a single user conversation. Nobody could tell me who "users" actually referred to.&lt;/p&gt;

&lt;p&gt;That's not research. That's a plausible-sounding guess wearing research's clothes.&lt;/p&gt;

&lt;h3&gt;
  
  
  What actually went wrong
&lt;/h3&gt;

&lt;p&gt;A few concrete anti-patterns emerged, and I think they're worth naming plainly:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;AI output was treated as ground truth, not a hypothesis.&lt;/strong&gt; Nobody followed up an AI-generated persona or use case with an actual interview, survey, or usability test. It went straight into the backlog as a requirement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;UI decisions had no design principles behind them.&lt;/strong&gt; There was no consistent spacing system, no defined interaction patterns, no accessibility baseline. Every screen was designed in isolation, so the app felt like five different products stitched together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feedback loops were one-directional.&lt;/strong&gt; Features shipped, nobody measured whether they solved the problem they claimed to solve, and the team moved on to the next AI-suggested feature.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speed was the only metric that mattered.&lt;/strong&gt; "We shipped it fast" became the substitute for "we shipped the right thing." Velocity was celebrated; outcomes were rarely discussed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Engineers were kept out of the "why."&lt;/strong&gt; We were handed tickets, not problems. When you don't understand the user problem you're solving, you can't push back when the solution doesn't fit it  you just build what's written.
None of this was one dramatic failure. It was a hundred small shortcuts that compounded until the product felt directionless, even though everyone individually was working hard.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Why bad projects teach you more than good ones
&lt;/h3&gt;

&lt;p&gt;On a healthy project, good decisions look invisible. Nobody points at the moment someone insisted on a usability test before committing to a flow  it just happens, and the project is better for it, quietly, forever.&lt;/p&gt;

&lt;p&gt;On a bad project, every missing safeguard becomes visible because you feel its absence. You don't need someone to explain why user research matters when you watch a feature ship, confuse every real user who touches it, and get silently reworked three sprints later. The lesson isn't taught to you it's &lt;em&gt;demonstrated&lt;/em&gt;, at cost, in real time.&lt;/p&gt;

&lt;p&gt;That's the uncomfortable value of a bad project: it turns abstract principles ("talk to your users," "have a design system," "measure outcomes, not output") into concrete, felt memories. You stop treating them as things senior engineers say in blog posts and start treating them as things you personally never want to relearn.&lt;/p&gt;

&lt;h3&gt;
  
  
  What I'm doing differently now
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AI research is a first draft, not a final answer.&lt;/strong&gt; If a PM shares an AI-generated set of user needs, my question is always: "Who did we validate this with?" If the answer is nobody, that's the next step, not a nice-to-have.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I ask for the "why" before I estimate the "how."&lt;/strong&gt; A ticket without a user problem attached doesn't get pointed. This isn't about gatekeeping  it's about being able to catch a bad idea before I've built it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Design principles get written down, even informally.&lt;/strong&gt; A shared doc of spacing, states, and interaction patterns saves so much rework it's not optional anymore.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I treat shipped ≠ solved.&lt;/strong&gt; After launch, I want to know if the feature actually moved the metric it was meant to move. If nobody's tracking it, I flag that as a gap, not a footnote.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I push back on speed as the only success metric.&lt;/strong&gt; Fast and wrong isn't a win  it's the same failure with better PR.
### The takeaway&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Good projects give you confidence. Bad projects give you conviction. Working on something that skipped user research and leaned on AI-generated assumptions instead didn't just teach me what &lt;em&gt;not&lt;/em&gt; to do it taught me how expensive it is to skip the boring, human parts of product development, even when the tools make skipping them feel effortless.&lt;/p&gt;

&lt;p&gt;AI can accelerate research. It can't replace the moment where a real user tells you your assumption was wrong. If your team is using AI to speed up thinking, make sure you're still checking that thinking against reality  because the app I worked on didn't have a technology problem. It had a "we stopped listening to users" problem, and no model can fix that for you.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Have you worked on a project where AI-generated research quietly replaced real user research? I'd love to hear how your team caught it  or didn't.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>discuss</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Why I Ditched "Just Let the LLM Handle It" for a State Machine (And Slept Better at Night)</title>
      <dc:creator>Valipireddy Kowshik</dc:creator>
      <pubDate>Tue, 15 Sep 2026 02:56:27 +0000</pubDate>
      <link>https://dev.to/k0wsh1k_0x/why-i-ditched-just-let-the-llm-handle-it-for-a-state-machine-and-slept-better-at-night-4i1p</link>
      <guid>https://dev.to/k0wsh1k_0x/why-i-ditched-just-let-the-llm-handle-it-for-a-state-machine-and-slept-better-at-night-4i1p</guid>
      <description>&lt;p&gt;When I started building my AI technical interviewer, I did what most people do: I threw a big system prompt at the LLM and told it to "act like an interviewer, ask coding questions, give hints when the candidate is stuck, and score them at the end."&lt;/p&gt;

&lt;p&gt;It worked... about 80% of the time.&lt;/p&gt;

&lt;p&gt;The other 20% is what this post is about.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with trusting an LLM to run your whole flow
&lt;/h2&gt;

&lt;p&gt;Here's what kept happening. Mid-interview, the model would:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Randomly decide the interview was "done" after one question&lt;/li&gt;
&lt;li&gt;Give away the answer while trying to give a "gentle hint"&lt;/li&gt;
&lt;li&gt;Forget it already asked a question and ask it again&lt;/li&gt;
&lt;li&gt;Jump straight to scoring before the candidate even submitted code&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this was a prompting skill issue. I rewrote that system prompt probably fifteen times. The real issue is structural: an LLM generating the &lt;em&gt;next thing to say&lt;/em&gt; has no actual memory of "what phase are we in," unless you spoon-feed it that context perfectly, every single turn, forever. And even then, it can just... decide to do something else. It's a language model, not a state tracker. Treating it like one is where things fall apart.&lt;/p&gt;

&lt;p&gt;For a casual chatbot, that's fine — mild chaos is charming. For a product where someone's actual hiring decision depends on the interview going through every stage correctly, "mild chaos" is a support ticket and an angry candidate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I did instead
&lt;/h2&gt;

&lt;p&gt;I pulled the &lt;em&gt;structure&lt;/em&gt; of the interview out of the LLM entirely and put it into a plain old finite state machine.&lt;/p&gt;

&lt;p&gt;Something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SETUP → GREETING → QUESTION_ASKED → CANDIDATE_CODING →
HINT_CHECK → EVALUATING → SCORING → DONE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The state machine — not the model — decides what phase we're in and what's allowed to happen next. The LLM only gets called &lt;em&gt;inside&lt;/em&gt; a state, to do the one thing that state needs: generate a question, generate a hint, or generate a scorecard. It never gets to decide "we're done now" or "let's skip to scoring." That's not its job anymore.&lt;/p&gt;

&lt;p&gt;Practically, this meant:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The LLM can't hallucinate a phase transition because it doesn't control phase transitions&lt;/li&gt;
&lt;li&gt;Hints only trigger off a real signal — 35+ seconds of no keystroke activity — not the model deciding "the candidate seems stuck"&lt;/li&gt;
&lt;li&gt;Scoring only fires after code is actually submitted and executed in the sandbox, not whenever the model feels like wrapping up&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The LLM still does all the hard, "actually intelligent" work — writing a good question, phrasing a helpful hint, writing a fair evaluation. It's just not allowed to drive the car anymore. It's a really good passenger with really good opinions.&lt;/p&gt;

&lt;h2&gt;
  
  
  A simplified look at the transition logic
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;InterviewState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;GREETING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;greeting&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;QUESTION_ASKED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;question_asked&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;CANDIDATE_CODING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;candidate_coding&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;HINT_CHECK&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hint_check&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;EVALUATING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evaluating&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;SCORING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scoring&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;DONE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;done&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;transition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current_state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# The FSM owns "what happens next" —
&lt;/span&gt;    &lt;span class="c1"&gt;# the LLM only fills in content within a state
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;current_state&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;InterviewState&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CANDIDATE_CODING&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;idle_35s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;InterviewState&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;HINT_CHECK&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;code_submitted&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;InterviewState&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;EVALUATING&lt;/span&gt;
    &lt;span class="c1"&gt;# ...deterministic, testable, no hallucinated jumps
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare that to letting the model implicitly track state through conversation history alone — there's no guarantee, no test coverage, and no way to catch a bad transition before it reaches the user.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that surprised me
&lt;/h2&gt;

&lt;p&gt;I expected this to make the product feel more robotic. It did the opposite.&lt;/p&gt;

&lt;p&gt;Because the &lt;em&gt;structure&lt;/em&gt; is now guaranteed, I could actually let the LLM be more creative and natural within each state, without worrying about it going off the rails. Constraining the skeleton let me loosen up the muscle. Counterintuitive, but it checks out — a lot of "unpredictable AI" complaints aren't really about the model being too creative, they're about the model having too much control over things it was never designed to control.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;If you're building something where an LLM is orchestrating a multi-step process — not just answering one-off questions — ask yourself: does the model actually need to decide &lt;em&gt;what happens next&lt;/em&gt;, or does it just need to generate good content &lt;em&gt;within&lt;/em&gt; a step someone else decides?&lt;/p&gt;

&lt;p&gt;Most of the time, in my experience, it's the second one. And the moment I stopped asking the LLM to be both the actor and the director, my "why did the interview just end after one question" bugs basically disappeared overnight.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I write about the messier, more practical side of building AI products — the stuff that doesn't make it into the demo. If you're curious, I built this exact system as a live product: &lt;a href="https://ai-interviewer-ten-delta.vercel.app/" rel="noopener noreferrer"&gt;AI Technical Interviewer&lt;/a&gt;. Happy to talk through the architecture more in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
