๐ค The interview question
"What do temperature and top-p actually do, and when would you change them?"
It's an AIF-C01 favourite, and most answers stop at "temperature controls creativity". A stronger answer: you've seen the difference yourself. Let's send the same prompt to Amazon Bedrock at different settings from the AWS CLI. About 10 minutes, under one cent. ๐
๐ Flow: Low temperature โ High temperature โ Low vs high top-p โ Hit the output ceiling
๐งฐ Before you start
- ๐ป A current AWS CLI v2. Old builds have no
bedrock-runtime converse; update withcurl -fsSL https://awscli.amazonaws.com/v2/install.sh | bash. - ๐ค Signed in (
aws sts get-caller-identityworks) as a user or role allowedbedrock:InvokeModelon theus.amazon.nova-micro-v1:0inference profile and theamazon.nova-micro-v1:0foundation model. - ๐ Region:
us-east-1. Amazon models like Nova are enabled by default, so there is no model-access form. - ๐ต Cost: Nova Micro on-demand is $0.000035 per 1K input tokens and $0.00014 per 1K output tokens (AWS price list, us-east-1). The ~15 calls below cost well under a cent. On the Free plan this comes out of your credits.
Set two variables once, so every command below pastes as-is:
M=us.amazon.nova-micro-v1:0
P='[{"role":"user","content":[{"text":"Suggest one name for a new coffee shop. Reply with the name only."}]}]'
๐ง Step 1: Low temperature, three runs
Why: temperature controls how widely the model samples from its next-word probabilities. Low means it picks the most likely wording almost every time. That's what you want for extraction and classification.
for i in 1 2 3; do
aws bedrock-runtime converse --region us-east-1 --model-id $M --messages "$P" \
--inference-config '{"temperature":0.1,"maxTokens":20}' \
--query 'output.message.content[0].text' --output text
done
โ Expected: the same name (or nearly) three times. AWS promises "more deterministic", not identical.
โ ๏ธ If you see AccessDeniedException โ your identity lacks bedrock:InvokeModel on both ARNs above. Run aws sts get-caller-identity to check who you are, then fix that identity's policy.
๐ฅ Step 2: High temperature, same prompt
Why: high temperature lets less likely words through. Use it when variety is the point, such as subject lines or brainstorming.
for i in 1 2 3; do
aws bedrock-runtime converse --region us-east-1 --model-id $M --messages "$P" \
--inference-config '{"temperature":1,"maxTokens":20}' \
--query 'output.message.content[0].text' --output text
done
โ Expected: usually three different names. Sampling is random, so run it again if two match.
๐ก Interviewers listen for this: low temperature makes output consistent, not correct. A wrong answer at low temperature is wrong the same way every time. ๐ The full chapter explains why that matters for the exam.
๐ฏ Step 3: Top-p, low vs high
Why: top-p limits the candidate words to the smallest set whose probabilities add up to p. With 0.1 only the very top words are allowed; with 1.0 every word stays in. AWS's Nova guide says to change temperature or top-p, not both, so temperature stays at Nova's default here.
for p in 0.1 1.0; do
echo "topP=$p"
for i in 1 2 3; do
aws bedrock-runtime converse --region us-east-1 --model-id $M --messages "$P" \
--inference-config "{\"topP\":$p,\"maxTokens\":20}" \
--query 'output.message.content[0].text' --output text
done
done
โ
Expected: under topP=0.1, repeats or near-repeats; under topP=1.0, more variety.
๐งจ Break it on purpose
Set the output ceiling far too low and ask for a long answer:
aws bedrock-runtime converse --region us-east-1 --model-id $M \
--messages '[{"role":"user","content":[{"text":"Explain what temperature does in a language model."}]}]' \
--inference-config '{"maxTokens":5}' \
--query '{text: output.message.content[0].text, stop: stopReason, out: usage.outputTokens}'
โ
Expected: a sentence cut off mid-way, "stop": "max_tokens" and "out": 5. maxTokens is a hard stop, not a request to be brief. If you want a short answer, ask for it in the prompt.
๐งน Cost and cleanup
- ๐ต Under one cent in total, and every call above is capped at 20 output tokens or fewer
- ๐งน Nothing to delete: Converse calls create no resources.
๐ฌ Say it in the interview
"Temperature and top-p both control how the model samples its next word. Low settings give consistent output, which suits extraction; high settings give variety, which suits brainstorming. I change one of them, not both, and I remember that low temperature makes a model consistent, not correct."
๐ฏ What you can now say (and the exam)
โ
"Low temperature = consistent, not correct"
โ
"High temperature or high top-p = more variety"
โ
"Change temperature or top-p, not both"
โ
"maxTokens truncates; stopReason: max_tokens is the tell"
๐ Go deeper
This is the hands-on cut of Chapter 8 of our free AWS AI Practitioner (AIF-C01) course: choosing a foundation model, inference parameters and agents, with practice questions.
Top comments (0)