<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Andy Tan</title>
    <description>The latest articles on DEV Community by Andy Tan (@combo-andy).</description>
    <link>https://dev.to/combo-andy</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4009508%2F40758e2d-c679-4620-ac1b-2b09d5b0f313.png</url>
      <title>DEV Community: Andy Tan</title>
      <link>https://dev.to/combo-andy</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/combo-andy"/>
    <language>en</language>
    <item>
      <title>Installing the AWS CLI, kubectl and eksctl on macOS</title>
      <dc:creator>Andy Tan</dc:creator>
      <pubDate>Fri, 04 Sep 2026 14:25:22 +0000</pubDate>
      <link>https://dev.to/combo-andy/installing-the-aws-cli-kubectl-and-eksctl-on-macos-4779</link>
      <guid>https://dev.to/combo-andy/installing-the-aws-cli-kubectl-and-eksctl-on-macos-4779</guid>
      <description>&lt;h2&gt;
  
  
  Overview
&lt;/h2&gt;

&lt;p&gt;Working with Amazon Elastic Kubernetes Service (EKS) from a Mac takes three command-line tools, and they build on each other in the order below:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AWS CLI — authenticates you to AWS and is the tool every later step relies on for credentials and identity.&lt;/li&gt;
&lt;li&gt;kubectl — the standard Kubernetes client, used to talk to the API server of an EKS cluster once one exists.&lt;/li&gt;
&lt;li&gt;eksctl — a higher-level tool that creates and manages EKS clusters, driving CloudFormation underneath.
The versions used in the source walkthrough were Kubernetes 1.30 (kubectl v1.30.8-eks-aeac579) and eksctl 0.207.0. Check the AWS documentation for current releases before copying version-pinned URLs — the S3 paths below embed both a version and a build date, and they change.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  1. AWS CLI
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1.1  Create an access key&lt;/strong&gt;&lt;br&gt;
Before installing anything, create an access key for the IAM user you intend to work as. In the AWS Management Console this is under IAM → Users → your user → Security credentials → Create access key. The console shows you the Access Key ID and the Secret Access Key once; the secret cannot be retrieved again afterwards, so store it somewhere safe at that moment.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Image note: the original article illustrates this step with screenshots of the IAM console's access-key creation screens. Those images are the author's and are not reproduced here — see the original post if you want the visual reference.&lt;br&gt;
Security note: an IAM access key is a long-lived credential. Prefer a dedicated, least-privilege IAM user or, better, IAM Identity Center / SSO sessions where your organisation supports them, and rotate keys regularly.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;1.2  Install or update the CLI&lt;/strong&gt;&lt;br&gt;
AWS ships a signed macOS installer package. Download it and install it system-wide:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s2"&gt;"https://awscli.amazonaws.com/AWSCLIV2.pkg"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="s2"&gt;"AWSCLIV2.pkg"&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;installer &lt;span class="nt"&gt;-pkg&lt;/span&gt; AWSCLIV2.pkg &lt;span class="nt"&gt;-target&lt;/span&gt; /
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same two commands also perform an in-place upgrade of an existing AWS CLI v2 installation — there is no separate update path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.3  Confirm the shell can find it&lt;/strong&gt;&lt;br&gt;
Check that the binary resolved onto your $PATH and reports a version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;which aws
aws &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If which aws returns nothing, open a new terminal session so the shell re-reads its PATH, or add /usr/local/bin to it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.4  Configure credentials&lt;/strong&gt;&lt;br&gt;
Run the interactive configuration and supply the access key you created in step 1.1, along with a default region and output format:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;aws configure

AWS Access Key ID     [****************I66M]:
AWS Secret Access Key [****************o4pv]:
Default region name   [us-east-1]:
Default output format [json]:
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The bracketed values are the existing settings, shown masked; pressing Return keeps them. Answers are written to ~/.aws/credentials and ~/.aws/config.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.5  Verify who you are&lt;/strong&gt;&lt;br&gt;
This call returns the identity AWS resolves your credentials to, and is the quickest way to confirm the CLI is working:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;aws&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;sts&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;get-caller-identity&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"UserId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="s2"&gt;"AIDAxxxxxxxxxxxxxxxxx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"Account"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"123456789012"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"Arn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;     &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:iam::123456789012:user/your-user"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Make a note of this identity. The eksctl step further down requires that every command be run as the same IAM principal.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. kubectl
&lt;/h2&gt;

&lt;p&gt;AWS publishes kubectl binaries matched to each supported EKS Kubernetes version. The walkthrough targets Kubernetes 1.30.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.1  Download the binary&lt;/strong&gt;&lt;br&gt;
Fetch the binary for your platform from the Amazon EKS S3 bucket. The path encodes the Kubernetes version, the build date and the CPU architecture — use arm64 in place of amd64 on Apple Silicon:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-O&lt;/span&gt; https://s3.us-west-2.amazonaws.com/amazon-eks/1.30.8/2025-01-10/bin/darwin/amd64/kubectl
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;The source article's command line ends in kubectl.sha256, which downloads the checksum file rather than the binary itself. The checksum is useful — download it as well and verify with openssl sha1 -sha256 kubectl — but the binary is the file without the .sha256 suffix, as written above.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;2.2  Make it executable and put it on your PATH&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;chmod&lt;/span&gt; +x ./kubectl
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you already have another kubectl installed, the recommended approach is a per-user copy in $HOME/bin placed ahead of everything else on the PATH, so it takes precedence without disturbing the existing installation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="nv"&gt;$HOME&lt;/span&gt;/bin &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cp&lt;/span&gt; ./kubectl &lt;span class="nv"&gt;$HOME&lt;/span&gt;/bin/kubectl &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;PATH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;/bin:&lt;span class="nv"&gt;$PATH&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Persist that PATH entry in your shell's startup file so it survives new terminal windows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'export PATH=$HOME/bin:$PATH'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; ~/.bash_profile
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;~/.bash_profile only applies if your login shell is bash. macOS has defaulted to zsh since Catalina — in that case append the same line to ~/.zshrc instead. Check with echo $SHELL.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;2.3  Verify&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;kubectl version --client

Client Version: v1.30.8-eks-aeac579
Kustomize Version: v5.0.4-0.20230601165947-6ce0bf390ce3
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The -eks- suffix confirms you are running the AWS-published build rather than an upstream one.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. eksctl
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;3.1  Prerequisites&lt;/strong&gt;&lt;br&gt;
eksctl creates clusters by provisioning CloudFormation stacks, so the IAM principal you use needs permission to work with Amazon EKS IAM roles, service-linked roles, AWS CloudFormation, VPCs and the networking resources that go with them. A user without those permissions will fail partway through cluster creation rather than at the start.&lt;br&gt;
All steps must be carried out as the same user. Re-run the identity check if you are unsure which principal the CLI is currently using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws sts get-caller-identity
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3.2  Install via Homebrew&lt;/strong&gt;&lt;br&gt;
On macOS the maintained route is the Weaveworks Homebrew tap:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;brew tap weaveworks/tap
brew &lt;span class="nb"&gt;install &lt;/span&gt;weaveworks/tap/eksctl
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3.3  Verify&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;eksctl version

0.207.0
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Putting it together
&lt;/h2&gt;

&lt;p&gt;At this point all three tools are installed and the AWS CLI holds working credentials. The usual next step is to create a cluster with eksctl, which also writes the cluster's connection details into your kubeconfig so that kubectl can reach it. If you are attaching to a cluster that already exists, the AWS CLI can generate that kubeconfig entry directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws eks update-kubeconfig &lt;span class="nt"&gt;--region&lt;/span&gt; us-east-1 &lt;span class="nt"&gt;--name&lt;/span&gt; your-cluster-name
kubectl get nodes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A successful kubectl get nodes confirms that all three pieces — credentials, Kubernetes client, and cluster access — are wired together correctly.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>cli</category>
      <category>kubectl</category>
      <category>eksctl</category>
    </item>
    <item>
      <title>Fully Managed DeepSeek-R1 Arrives on Amazon Bedrock</title>
      <dc:creator>Andy Tan</dc:creator>
      <pubDate>Fri, 04 Sep 2026 14:06:12 +0000</pubDate>
      <link>https://dev.to/combo-andy/fully-managed-deepseek-r1-arrives-on-amazon-bedrock-496d</link>
      <guid>https://dev.to/combo-andy/fully-managed-deepseek-r1-arrives-on-amazon-bedrock-496d</guid>
      <description>&lt;p&gt;&lt;strong&gt;Abstract&lt;/strong&gt;&lt;br&gt;
DeepSeek-R1 is now officially available on Amazon Bedrock and can be accessed through Bedrock Marketplace and the Custom Model Import feature. With its robust security controls and reasoning capabilities, the model is already serving thousands of enterprise customers. Amazon Web Services recently added a serverless option, further simplifying deployment. This article explains how to use DeepSeek-R1 securely in Amazon Bedrock and provides a practical walkthrough.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzn74cvppmpdoczbq2tt5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzn74cvppmpdoczbq2tt5.png" alt=" " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ku3iobu45qm5zbi92qb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ku3iobu45qm5zbi92qb.png" alt=" " width="800" height="515"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Key Advantages: Fully Managed and Enterprise-Grade Security&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;No infrastructure to manage&lt;br&gt;
DeepSeek-R1 is offered as a fully managed service through Amazon Bedrock. Users do not need to operate the underlying infrastructure; they can integrate the model through a single API and quickly build generative AI applications.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Enterprise-grade security&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;Data privacy: User inputs and model outputs are not shared with third parties by default. Encryption at rest, encryption in transit, and fine-grained access control through IAM policies are supported.&lt;/li&gt;
&lt;li&gt;Compliance certifications: The service aligns with multiple industry security standards, supporting compliant AI deployment at scale.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt; Capabilities across multiple use cases
DeepSeek-R1 is available under the MIT open-source license. It is strong at complex reasoning, code generation, and natural-language understanding, making it suitable for intelligent decision support, software development, mathematical problem solving, data analysis, and knowledge management.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwu25mdhdue71onj36roa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwu25mdhdue71onj36roa.png" alt=" " width="800" height="413"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Deployment Considerations&lt;/strong&gt;&lt;br&gt;
When deploying DeepSeek-R1, pay particular attention to the following security practices:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Data security&lt;br&gt;
Use Amazon Bedrock's built-in monitoring and cost-control capabilities to keep data under control throughout the workflow and reduce the risk of sensitive-information exposure.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Responsible AI&lt;br&gt;
Use Amazon Bedrock Guardrails for content filtering and policy management:&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;Block harmful content, such as violence or biased language.&lt;/li&gt;
&lt;li&gt;Define custom filters for sensitive information, such as national identification and bank-card numbers.&lt;/li&gt;
&lt;li&gt;Use context-aware controls to reduce model hallucinations.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt; Model evaluation
Use Amazon Bedrock model-evaluation tools to select the best model with automated metrics (such as accuracy and robustness) or human evaluation (such as consistency with brand voice). You can validate performance with built-in or custom datasets.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;3. Practical Guide: From Access to Invocation&lt;/strong&gt;&lt;br&gt;
3.1 Enable model access&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sign in to the Amazon Bedrock console, open the "Model access" page, and request access to DeepSeek-R1.&lt;/li&gt;
&lt;li&gt;In "Playgrounds," choose Chat/Text mode and select DeepSeek-R1 as the model category to test it online.
Example prompt (to test reasoning):
PLAINTEXT
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A family has $5,000 to save for their vacation... (example from the original article)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;3.2 API invocation examples&lt;br&gt;
AWS CLI&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;aws bedrock-runtime invoke-model \
    --model-id us.deepseek-r1-v1:0 \
    --body "{\"messages\":[{\"role\":\"user\",\"content\":\"...\"}]}" \
    --region us-west-2 \
    invoke-model-output.txt

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PYTHON SDK&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;import boto3
client = boto3.client("bedrock-runtime", region_name="us-west-2")
response = client.converse(
    modelId="us.deepseek.r1-v1:0",
    messages=[{"role": "user", "content": [{"text": "Describe 'hello world'."}]}]
)
print(response["output"]["message"]["content"][0]["text"])

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;3.3 Configure Guardrails&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In the console, create protection rules under "Safeguards," including keyword filters, denied topics, and blocked-response templates.&lt;/li&gt;
&lt;li&gt;Validate the protections through repeated tests to ensure generated content complies with business policies.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;4. Supported Regions and Pricing&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Currently available regions: US East (N. Virginia), US East (Ohio), and US West (Oregon).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;5. Conclusion&lt;/strong&gt;&lt;br&gt;
DeepSeek-R1's fully managed capabilities, combined with Amazon Bedrock's security tools, give enterprises an efficient and dependable option for deploying AI. Developers can experience the model directly through the Bedrock console.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>aws</category>
      <category>deepseek</category>
      <category>bedrock</category>
    </item>
    <item>
      <title>Amazon Bedrock Deployment Guide: From Environment Setup to Production Operations</title>
      <dc:creator>Andy Tan</dc:creator>
      <pubDate>Tue, 30 Jun 2026 11:18:05 +0000</pubDate>
      <link>https://dev.to/combo-andy/amazon-bedrock-deployment-guide-from-environment-setup-to-production-operations-2hja</link>
      <guid>https://dev.to/combo-andy/amazon-bedrock-deployment-guide-from-environment-setup-to-production-operations-2hja</guid>
      <description>&lt;p&gt;Amazon Bedrock, AWS's fully managed service for foundation models, makes it much easier to build and deploy generative AI applications through a model-as-a-service (MaaS) approach. This guide outlines a structured deployment workflow that covers permissions, network architecture, model onboarding, API integration, and performance optimization, helping teams build AI services that are scalable, secure, and operationally reliable.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Core Benefits and Technical Context&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Organizations typically choose Amazon Bedrock for the following reasons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Resource isolation and elastic scalability: Dedicated compute capacity helps reduce contention with other workloads, while scaling policies can adjust capacity based on demand. Under the right conditions, this can improve cost efficiency significantly.&lt;/li&gt;
&lt;li&gt;Security and compliance: Bedrock integrates with AWS security controls such as VPC networking and IAM, helping organizations meet strict security and compliance requirements, including standards such as SOC 2 Type II, HIPAA, and GDPR.&lt;/li&gt;
&lt;li&gt;Operational simplicity: Because AWS manages the underlying infrastructure, teams can reduce deployment time and lower operational overhead compared with self-managed model serving stacks.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;Pre-Deployment Preparation&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;2.1 AWS Account and Permission Setup&lt;/p&gt;

&lt;p&gt;For better security, use a dedicated IAM user or role instead of the root account, and enable AWS CloudTrail for auditing and operational traceability.&lt;/p&gt;

&lt;p&gt;Example IAM policy (JSON):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2012-10-17"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Statement"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Effect"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Allow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"bedrock:*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"ec2:Describe*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"s3:GetObject"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"Resource"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;Note: In production environments, always follow the principle of least privilege and scope &lt;code&gt;Resource&lt;/code&gt; permissions as narrowly as possible.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;2.2 Local Environment Configuration&lt;/p&gt;

&lt;p&gt;Install and configure the AWS CLI (version 2.15 or later is recommended) so that you can manage resources from the command line.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws configure
&lt;span class="c"&gt;# Enter your Access Key ID, Secret Access Key, Region (for example, us-west-2), and preferred output format (such as json)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;2.3 Network and Storage Architecture&lt;/p&gt;

&lt;p&gt;A three-tier architecture is commonly recommended to support high availability and security:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Frontend layer: Use an Application Load Balancer (ALB), ideally protected by AWS WAF against common web threats.&lt;/li&gt;
&lt;li&gt;  Application layer: Deploy Bedrock-related application components across multiple Availability Zones (AZs) for resilience.&lt;/li&gt;
&lt;li&gt;  Data layer: Use Amazon S3 for model artifacts, logs, and intermediate data. Where appropriate, use VPC endpoints or PrivateLink to reduce public internet exposure.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;Model Deployment Workflow&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;3.1 Model Preparation and Conversion&lt;/p&gt;

&lt;p&gt;If you plan to work with a custom model such as DeepSeek-R1, prepare the model artifacts in a format compatible with your deployment pipeline, such as FP16 or FP8 where applicable.&lt;/p&gt;

&lt;p&gt;Example conversion code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;deepseek_r1.converter&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BedrockExporter&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;deepseek_r1_base.pt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;exporter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BedrockExporter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;framework&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;pytorch&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;output_path&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s3://model-bucket/deepseek/&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;precision&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;fp16&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;  &lt;span class="c1"&gt;# supports fp32/fp16/bf16
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;exporter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;convert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is generally recommended to package model artifacts as a &lt;code&gt;.tar.gz&lt;/code&gt; file and keep the package size below 50 GB.&lt;/p&gt;

&lt;p&gt;3.2 Deployment Through the Console or API&lt;/p&gt;

&lt;p&gt;You can deploy model-related resources through the Bedrock console or via API-driven automation.&lt;/p&gt;

&lt;p&gt;Example API workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;

&lt;span class="n"&gt;bedrock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;bedrock-runtime&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;us-west-2&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bedrock&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;deepseek-r1-prod&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_model_identifier&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;deepseek-ai/deepseek-r1-6b&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;inference_configuration&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;preferred_compute_type&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;gpu_t4&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;min_worker_count&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;max_worker_count&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;3.3 Auto Scaling Strategy&lt;/p&gt;

&lt;p&gt;To balance responsiveness and cost efficiency, define scaling rules such as the following:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Scale out when: Request queue depth exceeds 50, or latency rises above 2 seconds.&lt;/li&gt;
&lt;li&gt;  Scale in when: CPU utilization remains below 30% for 5 minutes.&lt;/li&gt;
&lt;li&gt;  Cooldown period: 300 seconds to avoid rapid scaling oscillation.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;API Integration Patterns&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;4.1 Basic Text Generation&lt;/p&gt;

&lt;p&gt;Use the &lt;code&gt;invoke_model&lt;/code&gt; API for synchronous inference requests.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;botocore.config&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Config&lt;/span&gt;

&lt;span class="n"&gt;bedrock_config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;max_attempts&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;mode&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;adaptive&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;read_timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;bedrock-runtime&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bedrock_config&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;invoke_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;modelId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;deepseek-r1-prod&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain the basic principles of quantum computing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;512&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;temperature&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.7&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;body&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;())[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;generation&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;4.2 Streaming Responses and Multi-Turn Conversations&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Streaming output: Use &lt;code&gt;invoke_model_with_stream&lt;/code&gt; to deliver responses incrementally and improve the user experience.&lt;/li&gt;
&lt;li&gt;  Conversation handling: Use Bedrock conversation-oriented APIs or your own session layer to preserve context for assistants, customer support bots, and similar use cases.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;4.3 Batch Processing Optimization&lt;/p&gt;

&lt;p&gt;For non-real-time workloads, dynamic batching can improve throughput substantially. A batch size of 32 to 64 requests is often a practical starting point.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Performance Optimization and Monitoring&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;5.1 Performance Tuning Approaches&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Model quantization: Moving from FP32 to FP16 or FP8 can reduce memory usage and improve inference speed.&lt;/li&gt;
&lt;li&gt;  Caching: Integrate ElastiCache Redis and apply an LRU strategy to frequently repeated queries.&lt;/li&gt;
&lt;li&gt;  Asynchronous processing: Route non-real-time requests through Amazon SQS to decouple frontend traffic from backend inference workloads.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;5.2 Example Benchmark Targets&lt;/p&gt;

&lt;p&gt;Metric  Test Method Target&lt;br&gt;
Time to First Token (TTFT)  Empty request test  &amp;lt; 800 ms&lt;br&gt;
Throughput  100 concurrent requests sustained for 5 minutes &amp;gt; 80 TPS&lt;br&gt;
Error rate  Measured across 1,000 consecutive requests  &amp;lt; 0.1%&lt;/p&gt;

&lt;p&gt;5.3 CloudWatch Monitoring and Alerts&lt;/p&gt;

&lt;p&gt;Set up alerts on key operational metrics such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  CPUUtilization: Above 85% for 5 minutes -&amp;gt; trigger an SNS notification and scale out automatically.&lt;/li&gt;
&lt;li&gt;  ModelLatency: P99 latency above 1000 ms -&amp;gt; investigate load levels or switch traffic to a backup endpoint.&lt;/li&gt;
&lt;li&gt;  Invocations 4xx: More than 10 per minute -&amp;gt; inspect client request formatting and permissions.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;Security, Compliance, and Cost Management&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;6.1 Data Protection&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Network isolation: Use VPC endpoint policies to restrict traffic to private subnets where appropriate.&lt;/li&gt;
&lt;li&gt;  Encryption: Use AWS KMS customer-managed keys (CMKs) to protect sensitive data.&lt;/li&gt;
&lt;li&gt;  Auditability: Log API metadata to support investigation, traceability, and compliance review.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;6.2 Cost Structure and Optimization&lt;/p&gt;

&lt;p&gt;Running a model such as DeepSeek-R1 on Bedrock may involve compute, storage, and data transfer costs.&lt;/p&gt;

&lt;p&gt;Optimization ideas include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Use Lambda@Edge where low-latency global access is needed.&lt;/li&gt;
&lt;li&gt;  Cache frequent requests to reduce unnecessary inference traffic.&lt;/li&gt;
&lt;li&gt;  Review utilization regularly and adjust Reserved Instances or Savings Plans where applicable.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;Troubleshooting&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Symptom Possible Cause  Recommended Action&lt;br&gt;
503 Service Unavailable Capacity overload   Increase &lt;code&gt;max_worker_count&lt;/code&gt; or enable auto scaling&lt;br&gt;
Garbled model output    Encoding mismatch   Verify that &lt;code&gt;Content-Type&lt;/code&gt; is &lt;code&gt;application/json&lt;/code&gt;&lt;br&gt;
Unstable latency    Network jitter  Consider AWS Direct Connect or review the network path&lt;br&gt;
Access Denied   Missing IAM permissions Check whether the IAM role includes &lt;code&gt;AmazonBedrockFullAccess&lt;/code&gt; or an equivalent custom policy&lt;/p&gt;

&lt;p&gt;By following the practices outlined above, teams can deploy AI capabilities on Amazon Bedrock in a way that is efficient, secure, and scalable, while accelerating integration into real business applications.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>aws</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
