TL;DR
I recently completed another project from Udacity's Future AWS Agent Engineer Nanodegree Program, which I was able to take through the AWS ...
For further actions, you may consider blocking this person and/or reporting abuse
This reminds me of what Andrej Karpathy once said:
LLMs could become something analogous to a new operating system.
Ohh I haven't come across this one; now I want to read it π
Do you have the link to the Andrej Karpathy one by any chance? Sorry, your comment just made me curious π
Iβm glad you found it interesting π Itβs actually a talk by Andrej Karpathy.
Here is the link:
Andrej Karpathy: Software Is Changing (Again)
And the exact place where he said is at timestamp "9:08" π
Ohh yes, I got it wrong π I think I just mixed up the tweet and the talk because the idea in the tweet sounded so similar to what you were saying. Sorry, my curiosity got the better of me π
And I just watched it... wow, you were so on point with the 9:08 timestamp π Iβm going to watch the whole talk now. Thanks for sharing it!
Gave you a gem too π
π₯Ήπ My first gem, and from my sis too β€οΈ That makes it even more special.
Thank youu!
Hopefully more to come, especially for your posts! Wishing you all the success always, my bro β€οΈ
Thats the funny thing about the human mind. All the small pieces together look overwhelming, especially when something is new. But when you break it down into small parts, and be forgiving enough to slow down and give yourself time, you start understanding.
Sometimes the understanding goes a little to far and you get side tracked (at least I do) π but that's where all the cool finds usually hide out at. We can just call that part 'extra learning' π€£
This looks like quite a lofty feat to me! β¨οΈ
This happens to me too, Anna! Iβll start looking into one small thing and somehow end up 5 tabs later learning something completely different π I love the βextra learningβ way of looking at it though π
Also, I have to say, I LOVE the emojis you always use in your comments π They make the comments feel so much lighter and fun, both when you comment on my posts and around the community. So keep them coming, I love them π
I like it when other people use emojis too! Also, punctuation. If I dont see punctuation my brain struggles. π
The refund lesson is the whole thing: "the model said it happened" and "the system actually did it" are different claims, and treating them as one is where agents quietly break. Loved that you found the hardest parts were permissions, calculations, and monitoring β the boring, unglamorous scaffolding is exactly where real agents live or die. "Make the next problem smaller" is a keeper.
That was one of the biggest things I took away from this project too, James. I went in thinking the AI part would be the hardest, and then permissions, calculations, debugging, monitoring... π They kept showing up. And Iβm glad you liked βmake the next problem smallerβ because that became my little survival rule during this project π
Thanks for reading π
Thank you for such a thoughtful piece and reply! "Make the next problem smaller" is a keeper β genuinely enjoyed reading this. π
Great article! πΊ
I especially liked the distinction between βthe model says an action happenedβ and βthe system actually performed the action.β
That feels like one of the most important parts of building AI agents. The AI is only one piece, permissions, backend operations, monitoring, deployment, and normal software bugs still matter just as much.
Your point that βan AI application still contains normal software engineering problemsβ really stood out to me.
I havenβt built an AI agent myself yet, but Iβll definitely come back to this article when I do! πΈ
Thank you! And yes, that was one of the things that surprised me too. I went in thinking mostly about the AI part, but there was so much normal software stuff around it that I didn't expect π
Definitely come back to this when you build your first one, and I'd love to hear how it goes π
Hema, I really liked your point about the model saying an action happened versus the system actually doing it. π I've hit that in an agent I built myself, where the model reported success even though the backend call hadn't completed.
For a first agent, I'd check the actual result instead of trusting the confirmation text. A returned status or an actual record change tells you what really happened. Testing each capability separately, like you did, makes these failures much easier to spot.
Shubhra, yes, thatβs actually such a good example of what I was talking about! The βdoneβ message can sound so convincing, but checking what actually happened in the backend is a whole different thing π Thanks for sharing this from your own experience too π
Yes, I also think that when a problem is so big that we cannot understand it, understanding the outline first and then moving on to the details later is a good idea. Nice try! π
Yes, that helped me a lot with this project. Once I understood the bigger picture, the individual services started making much more sense instead of feeling like a huge list of AWS names π
Thanks for reading π
What do you recommend for someone building their 1st AI agent? Any tips, lessons, or maybe mistakes from your own projects that can help others? I love how DEV brings together people with different experience, so let's make the comments a small knowledge-sharing space. Even one small tip can help someone who just started π
Also, if anyone wants another take on this project, my classmate @earlgreyhot1701d wrote about her experience with it too! We worked on the same project, but our articles ended up being pretty different, so I thought some of you might enjoy reading hers as well π
Nova Adiutrix: My Second Agent Built My First Project's To-Do List
One small tip from my side: test things as you build them instead of waiting for the whole agent to be done. I found it much easier to catch what was actually wrong when I tested each capability separately. It saved me from having one huge βwhy is my agent not working?β problem at the end π
Congratulations on finishing another project! I especially admire the mindset youβve gained through this work. Thank you so much for sharing these insights with everyone for free β Iβve learned a great deal.
Iβm in a similar situation right now: new company, new project, and weβre building an AI agent OS. Your article addressed and eased my hidden anxiety perfectly. This is awesome π@hemapriya_kanagala
Thank you so much π It means a lot to hear that the article helped with something you were feeling too π I know youβre gonna do so well at the company and on the project too. Wishing you all the success with it @xulingfeng π
Thank you for your kind wishes π It means a lot. There are still many challenges ahead on the project, I will keep moving forward. Wish you more interesting discoveries in your own work and reading! π
One production lesson that stands out is the difference between deployment success and capability success. An agent can be deployed correctly while one tool, permission, or downstream dependency is still broken. At IT Path Solutions, weβve found that testing each capability as an independent contract makes this much easier to reason about: the agent shouldn't just prove that it can call a tool, but that the requested outcome actually occurred and can be verified. That also makes failures easier to localize because you can distinguish routing problems, permission failures, tool failures, and incorrect results instead of treating everything as βthe agent didn't work.β For production systems, that capability-level evidence can be much more valuable than a single successful deployment check.
This is actually why testing each capability separately helped me so much, Glen. When something failed, I could narrow it down instead of trying to figure out what was wrong with the whole agent. I ran into permission issues with a few of the capabilities myself, so I definitely saw how a deployment can be working while one part of the system is still broken. Thanks for sharing this π
the refund test is where i'd add a seventh case: the refund runs but the result never reaches the agent. retrying blindly can turn one refund into two; saying it failed can be wrong too. give the operation a stable request id and check the recorded outcome before retrying. your six capability checks prove the happy paths, that one checks whether the system can recover without guessing.
Ohh, I see what you mean! I hadn't really thought about the case where the refund goes through but the result doesn't make it back to the agent. Retrying without checking could definitely cause a problem. The stable request ID and checking the recorded outcome before retrying is a really good production consideration. Thanks for bringing this up π
glued together for what sounds like "six features" is the real story here. I hit the exact same wall wiring up Bedrock + Lambda + Knowledge Base last month β spent 4 hours just on IAM policies before the agent could even call one tool. The "what's the next thing I need to understand" framing is honestly the only way through a stack that wide without losing your mind.
Onizuka, 4 hours just on IAM π I can definitely relate to that part! I ran into permission issues too, and it made me realize how much of this project was figuring out what I needed to understand next rather than trying to understand everything at once π That mindset honestly helped me get through it. Thanks for sharing your experience π
Connecting refunds through AgentCore Gateway to Lambda is the useful split: the model can say a refund ran, but only the Lambda moves money.
The gap is a retry. If the first call may have already refunded and the agent asks again, you need a check before you charge that refuses a second create on that order.
When the gateway times out after Lambda may have run, do you block a fresh refund until reconcile says settled or failed?
For this project I was mainly focused on getting the required refund workflow working through Gateway and Lambda, but I hadn't thought about the timeout/retry case at that level. The possibility of Lambda completing the refund while the response gets lost definitely makes the retry part tricky. That's a really good production consideration to keep in mind for future projects. Thanks for bringing it up π
The shift from βinstructions the agent should rememberβ to βconstraints the system can continuously verifyβ is a useful distinction. One thing Iβd add is that ai:verify can become more than a feedback mechanism, it can act as a regression contract for agent behavior. If an architectural rule is encoded as a check, that check should ideally survive changes to prompts, models, and agent workflows. That makes the verification layer independent of how the agent was instructed to behave. It also gives teams something measurable: not just whether the agent followed AGENTS.md, but whether the resulting change satisfied the projectβs actual constraints. Over time, that could make instruction files much smaller while making the enforced behavior more durable.
What stood out to me is how you turned a huge AWS architecture into smaller, understandable problems. The βmake the next problem smallerβ mindset is especially relatable for anyone learning agents and cloud. The IAM and deployment lessons are great reminders that the AI is only one piece of the system.
As I mentioned in one of the comments earlier, βMake the next problem smallerβ became my little survival rule during this project, Jyanthi π It helped so much whenever everything started feeling like a lot. Thanks for reading π
Same experience - the model is the easy part. The wiring and the plan are the real work.
What helped me: write the whole system in a .md first - every service, what talks to what, what can fail - and have the agent poke holes in it before anything is deployed. "Six features that are really eleven services" shows up on paper instead of in the logs at 2am.
Great write-up.
The refunds section contains the piece's most production-relevant failure mode: "a model saying that something happened and a system actually performing that action are two different things" β which means each of your six evidenced tests should assert on the backend record, not the agent's final response. Shubhradev's comment proves the point: the model can report success while the Lambda call never completed, so a test that passes on the agent's narration is really just testing its confidence. The operating-the-system section gestures at the fix β those six capability checks belong on a schedule against failed requests and tool failures, or deploy-day evidence goes stale.
Yes, exactly! For this project, I tested each of the required capabilities end to end and made sure the actual workflows were working, but your point about production checks going beyond the agent's final response is really good. I hadn't thought as much about those checks becoming stale over time either. Definitely something I'll keep in mind for future projects π