DEV Community

confident_prep
confident_prep

Posted on Originally published at confidentprep.com

๐ŸŒก๏ธ What Do Temperature and Top-p Really Do? Test Them on Amazon Bedrock From the CLI (Hands-on)

๐ŸŽค The interview question

"What do temperature and top-p actually do, and when would you change them?"

It's an AIF-C01 favourite, and most answers stop at "temperature controls creativity". A stronger answer: you've seen the difference yourself. Let's send the same prompt to Amazon Bedrock at different settings from the AWS CLI. About 10 minutes, under one cent. ๐Ÿ‘‡

๐Ÿ‘‰ Flow: Low temperature โ†’ High temperature โ†’ Low vs high top-p โ†’ Hit the output ceiling


๐Ÿงฐ Before you start

  • ๐Ÿ’ป A current AWS CLI v2. Old builds have no bedrock-runtime converse; update with curl -fsSL https://awscli.amazonaws.com/v2/install.sh | bash.
  • ๐Ÿ‘ค Signed in (aws sts get-caller-identity works) as a user or role allowed bedrock:InvokeModel on the us.amazon.nova-micro-v1:0 inference profile and the amazon.nova-micro-v1:0 foundation model.
  • ๐ŸŒ Region: us-east-1. Amazon models like Nova are enabled by default, so there is no model-access form.
  • ๐Ÿ’ต Cost: Nova Micro on-demand is $0.000035 per 1K input tokens and $0.00014 per 1K output tokens (AWS price list, us-east-1). The ~15 calls below cost well under a cent. On the Free plan this comes out of your credits.

Set two variables once, so every command below pastes as-is:

M=us.amazon.nova-micro-v1:0
P='[{"role":"user","content":[{"text":"Suggest one name for a new coffee shop. Reply with the name only."}]}]'
Enter fullscreen mode Exit fullscreen mode

๐ŸงŠ Step 1: Low temperature, three runs

Why: temperature controls how widely the model samples from its next-word probabilities. Low means it picks the most likely wording almost every time. That's what you want for extraction and classification.

for i in 1 2 3; do
  aws bedrock-runtime converse --region us-east-1 --model-id $M --messages "$P" \
    --inference-config '{"temperature":0.1,"maxTokens":20}' \
    --query 'output.message.content[0].text' --output text
done
Enter fullscreen mode Exit fullscreen mode

โœ… Expected: the same name (or nearly) three times. AWS promises "more deterministic", not identical.

โš ๏ธ If you see AccessDeniedException โ†’ your identity lacks bedrock:InvokeModel on both ARNs above. Run aws sts get-caller-identity to check who you are, then fix that identity's policy.


๐Ÿ”ฅ Step 2: High temperature, same prompt

Why: high temperature lets less likely words through. Use it when variety is the point, such as subject lines or brainstorming.

for i in 1 2 3; do
  aws bedrock-runtime converse --region us-east-1 --model-id $M --messages "$P" \
    --inference-config '{"temperature":1,"maxTokens":20}' \
    --query 'output.message.content[0].text' --output text
done
Enter fullscreen mode Exit fullscreen mode

โœ… Expected: usually three different names. Sampling is random, so run it again if two match.

๐Ÿ’ก Interviewers listen for this: low temperature makes output consistent, not correct. A wrong answer at low temperature is wrong the same way every time. ๐Ÿ“š The full chapter explains why that matters for the exam.


๐ŸŽฏ Step 3: Top-p, low vs high

Why: top-p limits the candidate words to the smallest set whose probabilities add up to p. With 0.1 only the very top words are allowed; with 1.0 every word stays in. AWS's Nova guide says to change temperature or top-p, not both, so temperature stays at Nova's default here.

for p in 0.1 1.0; do
  echo "topP=$p"
  for i in 1 2 3; do
    aws bedrock-runtime converse --region us-east-1 --model-id $M --messages "$P" \
      --inference-config "{\"topP\":$p,\"maxTokens\":20}" \
      --query 'output.message.content[0].text' --output text
  done
done
Enter fullscreen mode Exit fullscreen mode

โœ… Expected: under topP=0.1, repeats or near-repeats; under topP=1.0, more variety.


๐Ÿงจ Break it on purpose

Set the output ceiling far too low and ask for a long answer:

aws bedrock-runtime converse --region us-east-1 --model-id $M \
  --messages '[{"role":"user","content":[{"text":"Explain what temperature does in a language model."}]}]' \
  --inference-config '{"maxTokens":5}' \
  --query '{text: output.message.content[0].text, stop: stopReason, out: usage.outputTokens}'
Enter fullscreen mode Exit fullscreen mode

โœ… Expected: a sentence cut off mid-way, "stop": "max_tokens" and "out": 5. maxTokens is a hard stop, not a request to be brief. If you want a short answer, ask for it in the prompt.


๐Ÿงน Cost and cleanup

  • ๐Ÿ’ต Under one cent in total, and every call above is capped at 20 output tokens or fewer
  • ๐Ÿงน Nothing to delete: Converse calls create no resources.

๐Ÿ’ฌ Say it in the interview

"Temperature and top-p both control how the model samples its next word. Low settings give consistent output, which suits extraction; high settings give variety, which suits brainstorming. I change one of them, not both, and I remember that low temperature makes a model consistent, not correct."


๐ŸŽฏ What you can now say (and the exam)

โœ… "Low temperature = consistent, not correct"
โœ… "High temperature or high top-p = more variety"
โœ… "Change temperature or top-p, not both"
โœ… "maxTokens truncates; stopReason: max_tokens is the tell"


๐Ÿ“š Go deeper

This is the hands-on cut of Chapter 8 of our free AWS AI Practitioner (AIF-C01) course: choosing a foundation model, inference parameters and agents, with practice questions.

๐Ÿ‘‰ Read the full chapter on Confident Prep

Top comments (0)