DEV Community

Cover image for The Real Post-Mortem: Serving Open LLMs on AWS (SageMaker vLLM vs. Bedrock Custom Models)
Rahul R
Rahul R

Posted on

The Real Post-Mortem: Serving Open LLMs on AWS (SageMaker vLLM vs. Bedrock Custom Models)

Running an open-source LLM on AWS sounds straightforward at first. Pick a model, deploy it, send requests, and you're done. But once you start thinking about production, things become more complicated. Where should the model run? Who manages the GPUs? How much control do you actually need? What happens when traffic increases?

Two approaches that caught my attention are Amazon SageMaker with vLLM and Amazon Bedrock with custom or supported models. Both can be useful, but they solve the problem at different levels.

The biggest lesson for me is that the decision isn't really about choosing between two AWS services. It's about deciding how much of the LLM infrastructure you want your team to own.

SageMaker + vLLM

SageMaker Multi-AZ Auto-Scaling Infrastructure

With SageMaker, you can deploy an open model and use vLLM as the inference engine. This gives you much more control over the serving layer. You can choose the model, configure the inference environment, select the underlying compute, and tune the serving configuration based on your workload.

The basic architecture looks like this:

Application
     |
     v
SageMaker Endpoint
     |
     v
    vLLM
     |
     v
  Open LLM
     |
     v
    GPU
Enter fullscreen mode Exit fullscreen mode

This approach becomes interesting when the model and inference performance are important parts of your application.

For example, your application can invoke a SageMaker endpoint using boto3:

import boto3
import json

runtime = boto3.client("sagemaker-runtime")

response = runtime.invoke_endpoint(
    EndpointName="my-vllm-endpoint",
    ContentType="application/json",
    Body=json.dumps({
        "inputs": "Explain Apache Spark in simple terms.",
        "parameters": {
            "temperature": 0.2,
            "max_new_tokens": 200
        }
    })
)

result = response["Body"].read().decode("utf-8")

print(result)
Enter fullscreen mode Exit fullscreen mode

The application only needs to know about the endpoint. The model and GPU infrastructure are behind it.

One of the reasons vLLM is interesting for LLM serving is its focus on efficient inference and request handling. Techniques such as continuous batching can help the server handle multiple requests efficiently instead of treating every request as an isolated workload.

Imagine five users sending requests at approximately the same time:

Request 1 ──┐
Request 2 ──┤
Request 3 ──┼──> vLLM ──> GPU
Request 4 ──┤
Request 5 ──┘
Enter fullscreen mode Exit fullscreen mode

This becomes increasingly important when your application moves from testing to real traffic.

Test vLLM before deploying to AWS

One thing I would recommend is testing the serving layer locally before introducing AWS infrastructure. If the model doesn't behave as expected locally, moving the problem to a GPU endpoint won't magically fix it.

A basic vLLM server can be started with:

pip install vllm
Enter fullscreen mode Exit fullscreen mode

Then:

vllm serve <your-model-name> \
    --host 0.0.0.0 \
    --port 8000
Enter fullscreen mode Exit fullscreen mode

If the server exposes an OpenAI-compatible API, you can test it with:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="dummy"
)

response = client.chat.completions.create(
    model="<your-model-name>",
    messages=[
        {
            "role": "user",
            "content": "Explain Spark partitioning in simple terms."
        }
    ],
    temperature=0.2,
    max_tokens=200
)

print(response.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

This gives you a useful development flow: validate the model and serving configuration first, then move the workload to your AWS environment.

But SageMaker + vLLM comes with responsibility

The extra control is also the downside.

Once you're responsible for the serving layer, you need to think about GPU selection, memory usage, model loading time, batching, concurrency, endpoint scaling, monitoring, deployment failures, and cost.

For example, if an endpoint suddenly becomes slow, there are many possible things to investigate:

High latency
    |
    +-- GPU utilization
    +-- GPU memory
    +-- Concurrent requests
    +-- Batch configuration
    +-- Model loading
    +-- Instance capacity
    +-- Token generation
Enter fullscreen mode Exit fullscreen mode

That's powerful when you have the expertise to tune these areas, but it also means more operational work for your team.

Amazon Bedrock

Amazon Bedrock Custom Model Import Control Plane

Now let's look at the other approach.

Amazon Bedrock provides a more managed way of working with foundation models. Instead of thinking primarily about GPU instances and model-serving infrastructure, your application interacts with a managed API.

The architecture becomes much simpler:

Application
     |
     v
Amazon Bedrock
     |
     v
Model
     |
     v
Response
Enter fullscreen mode Exit fullscreen mode

For example, a model can be invoked using the Bedrock Runtime API:

import boto3

client = boto3.client("bedrock-runtime")

response = client.converse(
    modelId="<model-id>",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "text": "Explain Apache Spark in simple terms."
                }
            ]
        }
    ],
    inferenceConfig={
        "maxTokens": 200,
        "temperature": 0.2
    }
)

answer = response["output"]["message"]["content"][0]["text"]

print(answer)
Enter fullscreen mode Exit fullscreen mode

Notice how different this feels from managing an inference endpoint. The application mainly cares about sending a request and receiving a response.

You aren't writing code to select a GPU, start vLLM, or manage the model server.

That's the main attraction of a managed approach.

A practical example

Imagine we're building an internal data engineering assistant.

A user asks:

"Why did my AWS Glue job fail with an out-of-memory error?"

With Bedrock, the application could send that question directly to the model:

response = client.converse(
    modelId="<model-id>",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "text": """
                    A Spark job failed with an out-of-memory error
                    during a Delta write.

                    Explain possible causes and debugging steps.
                    """
                }
            ]
        }
    ]
)

answer = response["output"]["message"]["content"][0]["text"]

print(answer)
Enter fullscreen mode Exit fullscreen mode

The same application could instead call a SageMaker endpoint hosting an open model:

import boto3
import json

client = boto3.client("sagemaker-runtime")

response = client.invoke_endpoint(
    EndpointName="llm-inference-endpoint",
    ContentType="application/json",
    Body=json.dumps({
        "inputs": """
        A Spark job failed with an out-of-memory error
        during a Delta write.

        Explain possible causes and debugging steps.
        """
    })
)

answer = response["Body"].read().decode("utf-8")

print(answer)
Enter fullscreen mode Exit fullscreen mode

The application problem is almost identical. What's different is the infrastructure underneath the API.

Control vs. simplicity

This is where I think the comparison becomes much clearer.

SageMaker + vLLM is closer to the infrastructure:

Application
    |
    v
SageMaker
    |
    v
vLLM
    |
    v
Open LLM
    |
    v
GPU
Enter fullscreen mode Exit fullscreen mode

Bedrock is closer to the application:

Application
    |
    v
Bedrock API
    |
    v
Model
Enter fullscreen mode Exit fullscreen mode

With SageMaker and vLLM, you get more control, but you also take on more responsibility. With Bedrock, you give up some infrastructure-level control in exchange for a more managed experience.

Neither approach is automatically better.

What about cost?

This is where I would avoid making a simple statement such as "Bedrock is cheaper" or "SageMaker is cheaper."

The answer depends heavily on your workload.

If you have high and predictable traffic, dedicated inference infrastructure can potentially be well utilized. If your application receives requests only occasionally, continuously running GPU capacity may be difficult to justify.

For example:

High predictable traffic

████████████████████████████
████████████████████████████
████████████████████████████
Enter fullscreen mode Exit fullscreen mode

Compared with:

Unpredictable traffic

██

          ███

                       ██

     █
Enter fullscreen mode Exit fullscreen mode

These workloads can lead to very different architectural decisions.

I would measure cost per request, cost per token, average latency, P95/P99 latency, throughput, GPU utilization, and error rate before deciding.

Don't forget engineering cost

There's another cost that is easy to overlook: engineering time.

Suppose your team spends several days troubleshooting GPU memory, model loading, batching, scaling, containers, and inference performance.

That work has a cost too.

Sometimes a managed service is worth using not because it is the cheapest infrastructure option, but because it removes operational work that your team doesn't need to own.

This is especially important for teams whose main product isn't an ML infrastructure platform.

When SageMaker + vLLM makes sense

I'd lean toward SageMaker + vLLM when the serving layer itself is important.

For example, you may have a specific open-source model that you need to run, significant inference traffic, strict latency requirements, or a team that already understands GPU and container workloads.

In that situation, the additional control can be worth the additional complexity.

You're not just consuming an LLM anymore. You're building part of an LLM serving platform.

When Bedrock makes more sense

I'd lean toward Bedrock when the main requirement is simply to add LLM capabilities to an application without taking ownership of the underlying inference infrastructure.

If your team doesn't want to spend time managing GPU workloads, a managed approach can significantly reduce the operational burden.

The question becomes:

"Do we need to operate an LLM platform, or do we simply need an LLM capability?"

That's an important distinction.

A simple decision framework

Before choosing an approach, I'd ask five questions.

Do we need deep control over inference? If yes, SageMaker + vLLM becomes more attractive.

Is GPU infrastructure already a skill within the team? If yes, the additional control may be easier to manage.

Is the model itself important to the product? If you're heavily dependent on a particular open model, more deployment control can be valuable.

What does our traffic look like? High, predictable traffic and occasional traffic can lead to very different cost and scaling decisions.

How much operational complexity are we willing to own? This may actually be the most important question.

A simple benchmark

Before making the final decision, I'd run a small load test rather than relying only on architecture diagrams.

For a SageMaker endpoint, you could start with something as simple as:

import time
import boto3
import json

client = boto3.client("sagemaker-runtime")

endpoint = "llm-inference-endpoint"

payload = {
    "inputs": "Explain partitioning in Apache Spark.",
    "parameters": {
        "max_new_tokens": 100
    }
}

start = time.time()

response = client.invoke_endpoint(
    EndpointName=endpoint,
    ContentType="application/json",
    Body=json.dumps(payload)
)

latency = time.time() - start

print(f"Latency: {latency:.2f} seconds")

print(
    response["Body"]
    .read()
    .decode("utf-8")
)
Enter fullscreen mode Exit fullscreen mode

For a proper benchmark, don't stop at one request. Send enough requests to understand the behavior of the system under realistic concurrency.

I'd collect at least:

Average latency
P50 latency
P95 latency
P99 latency
Requests/second
Error rate
Cost
GPU utilization
Enter fullscreen mode Exit fullscreen mode

That gives you something much more useful than simply saying that one service "feels faster."

Final comparison

Area SageMaker + vLLM Bedrock
Infrastructure control High Lower
Serving customization High More abstracted
GPU management More responsibility More managed
Operational complexity Higher Lower
Model flexibility High, depending on deployment Depends on supported models/options
Performance tuning More control More abstracted
Best fit Custom LLM serving Managed LLM integration

AWS model availability and deployment options change over time, so this should be treated as an architectural comparison rather than a permanent feature matrix.

What I would choose

If I were building a small internal application or proof of concept, I'd start with the simplest managed approach that satisfies the model requirements.

I wouldn't introduce GPU infrastructure just because I could.

But if inference became a core part of the product and I needed deep control over an open model, I'd seriously consider SageMaker + vLLM.

The workload should drive the architecture.

Not the other way around.

The real post-mortem

The biggest lesson isn't that SageMaker is better than Bedrock, or that Bedrock is better than SageMaker.

It's that every additional layer of control also creates another layer of responsibility.

SageMaker + vLLM gives you more control over how the model is served. Bedrock gives you more abstraction around the infrastructure.

So before deploying an open LLM on AWS, ask yourself:

"Do I actually need to manage the inference infrastructure, or do I just need an LLM API?"

That question can save a lot of engineering time.

And once the system is running, don't stop at asking whether the model works. Measure latency, throughput, utilization, token usage, cost, and reliability.

Because the architecture that looks best on a diagram isn't necessarily the architecture that performs best in production.

That's the real post-mortem.


Note: AWS model availability, model IDs, vLLM versions, supported deployment options, and Bedrock capabilities can change. Verify the current AWS documentation for your chosen model and AWS Region before using these examples in production.

AWS #SageMaker #AmazonBedrock #vLLM #LLM #GenerativeAI #DataEngineering #MachineLearning #AI

Top comments (0)