DEV Community

Cover image for The 5 Layers Behind Every AI App
Blackwatch
Blackwatch

Posted on

The 5 Layers Behind Every AI App

It's been some time since the introduction of AI and now it has become an essential part of our life, isn't it?

As a developer I can vouch for this. My day now starts with AI and ends with AI too 😁.

AI gave us the power to be a one man army, to do things which typically needed a whole team in the past. Now we can do those things on our own. Just a subscription to an AI model and we are done.

But things aren't this easy. To use AI to its full extent, we need to understand how it works.

As a human, the best thing we can do is give a well defined prompt, informative enough for the AI to understand our requirement, and it's good to have some examples of what we expect.

Then comes one of the most important parts in this, iterations. Even AI can miss some things, or what if you change your mind? That's why it's good to iterate until we reach the goal we set in our mind.

To write an effective prompt I already wrote an article, Prompt Engineering: How to Actually Get What You Want from AI. You can refer to this to improve your prompts.

Now taking one step further. In this article we will dive into AI's layers and understand the stack behind it.

One quick clarification before we start. I am not talking about the layers inside a neural network, the input and hidden and output ones. I am talking about the layers of the stack that every AI product you use is built on top of.

So, majorly there are five layers behind the AI tools we use today. We will go top-down, starting with what you see and ending with what it all runs on.

  1. Gen AI Apps (Application Layer)
  2. Agent Layer
  3. Platform Layer
  4. Models Layer
  5. Infrastructure Layer

The AI stack: Gen AI Apps, Agent Layer, Platform Layer, Models, and Infrastructure, from the user at the top down to the hardware

Now one layer at a time.

1. Gen AI Apps or Application Layer

User centric apps which the user uses to interact and work with AI

Have you worked with ChatGPT, Gemini, DeepSeek, Claude etc? Then most of the time you must have used some sort of app or the official website.

Here is the thing though. You are not talking to the model directly. There is an app sitting in between, managing your chat history, rendering the response, handling your file uploads. The model itself has no idea what a conversation is.

That app is the Gen AI app, and it is the layer the user directly interacts with.

2. Agent Layer

This layer utilizes the capabilities of the model layer to perform more complex actions. Think of them as digital workers with a brain.

You specify a task to them, they gather information from the user, interact with the environment to gather more information, make decisions, and execute actions based on the info received.

They differ from simple AI in the fact that they can actually do the work. Like searching on the internet to look for the best restaurant, checking your schedule for free time, booking the table, and sending a mail or message to the other party for the invitation. Rather than just giving you the steps, it actually does the work.

This is the agent layer.

3. Platform Layer

This is one of the most important layers in AI. It acts as a bridge between the models and the layers above.

Here is the problem it solves. A trained model on its own is just a huge file of numbers sitting on some machine. It has no way for you to reach it, no idea who you are, and no way to handle a thousand people asking it things at the same time.

The platform layer is everything that turns that file into something an app can actually call.

So what does it actually do?

  • APIs — it gives you an endpoint to hit. This is the thing that makes a model usable from your code instead of just from a website.
  • Authentication — it checks who you are. This is where your API key comes in, and it is also how the provider knows which account to bill.
  • Rate limits and quotas — it decides how much you are allowed to use, and how fast. Anyone who has seen a 429 error while building with AI has met this part of the layer.
  • Routing and scaling — your request has to land on a machine that has a free GPU. The platform picks one, sends your request there, and spins up more machines when traffic goes up.
  • Monitoring and logs — it keeps the record. How many tokens you used, how long the request took, what failed and why.
  • Extra tooling — most platforms also give you the surrounding pieces, like vector databases for storing your own data, guardrails to filter unwanted output, and tools to evaluate or fine tune a model.

An example

Remember the restaurant agent from the last layer? Let's see what the platform does while that agent is working.

The agent decides it needs to think about your request, so it sends it to a model. That request hits an API endpoint. The platform checks your API key, confirms your account is allowed to use that model, and checks you haven't crossed your rate limit.

Then it finds a machine with a free GPU and sends your request there. The model does its thing and the answer comes back. On the way out, the platform counts the tokens you used, writes it to your usage log, and bills your account.

The agent gets its answer and moves on to the next step. All of that happened in between, and neither you nor the agent had to think about it.

Some platforms you may have heard of are Amazon Bedrock, Google Vertex AI and Microsoft Foundry, along with the APIs that the model providers run themselves.

If you have ever built a backend, this layer will feel familiar. It is the same set of concerns you already deal with every day, auth, rate limiting, scaling, logging. The only difference is that there is a model at the other end instead of a database.

4. Models

The elephant in the room, the brain of any AI. These give the power to AI to do the things that they are able to do nowadays, from simple conversations to generating videos.

There are different types of models as per the user's requirement.

  • LLMs (Large Language Models) — good for conversations or text based things. From translation to creating mails for you.
  • Diffusion Models — LLMs are good for text, meanwhile diffusion models are for generating new media. They start from random noise and refine it step by step until an image or a video comes out.

5. Infrastructure Layer

We talked about most of the things, but all these things need infrastructure to run.

That physical part comes under this layer, from CPUs to GPUs to RAM and networking.

We can blame this layer for the increasing RAM and SSD prices currently.

Conclusion

So the next time you type something into a chat box, remember how much is sitting underneath it. Your message passes through the app, the agent and the platform, and then it gets turned into actual math running on actual hardware.

Five layers. The model gets all the attention, but it is only one of them.

This is a small guide on the layers behind AI. I hope you found it helpful.

Until next time.

Peace ✌️

Top comments (0)