A practical revision guide for choosing AWS AI and ML services, approaches, and controls
I made these notes to revise the choices the exam repeatedly asks you to make. They are less about memorising isolated definitions and more about spotting what a scenario is really asking for: the right level of model customisation, the appropriate managed service, or the trade-off behind a metric.
How I approached the exam
I found the exam broadly falls into three areas: baseline AWS-service knowledge, basic AI and ML concepts, and security fundamentals. If you have already passed AWS Certified Cloud Practitioner or a higher AWS certification, the AWS-services material should be familiar. Likewise, anyone who has studied security or passed security-focused exams should be comfortable with much of the security content.
That was my starting point: I already had baseline AWS and security knowledge, so I spent about two weeks concentrating on the AI material. If you are beginning from scratch, I would allow another two or three weeks to build those AWS and security foundations before focusing on the AI-specific topics.
Exam format
| Questions | 65 (50 scored and 15 unscored) |
|---|---|
| Time | 90 minutes |
| Pass mark | 700 / 1000, using compensatory scoring |
| Question types | multiple choice, multiple response, ordering, and matching |
Compensatory scoring means there is no minimum mark for each domain. One weaker topic will not, on its own, fail the exam. That said, ordering, matching, and multiple-response questions are all-or-nothing: getting four of five pairs right still scores zero. It is worth practising those formats, not just reading the content.
There is no penalty for guessing and an unanswered question is wrong. With roughly 80 seconds per question, a sensible approach is a quick first pass, flagging anything uncertain, then coming back to it.
Where the marks are
| Domain | Weight |
|---|---|
| Applications of foundation models | 28% |
| Fundamentals of generative AI | 24% |
| Fundamentals of AI and ML | 20% |
| Guidelines for responsible AI | 14% |
| Security, compliance and governance | 14% |
Over half of the exam is in generative AI and foundation-model territory, so that is where I would put most revision time.
The customisation ladder
This is probably the most useful idea to have clear. Ways of changing a foundation model sit on a rough ladder, from the simplest and cheapest to the most involved:
- Prompt engineering: adjust tone, format, length, persona, or audience.
- RAG / Amazon Bedrock Knowledge Bases: bring in your organisation's data, especially when it changes.
- Fine-tuning: teach a task, style, or domain terminology using labelled prompt/completion pairs.
- Continued pre-training: build domain knowledge from raw, unlabelled documents.
- Training from scratch: the greatest cost, but also the greatest control and security ownership.
My rule of thumb: start as low on this ladder as you can. If a question mentions cost, that is usually an invitation to choose the least complex option that actually meets the requirement.
The data clues are helpful too. Labelled data suggests fine-tuning. Unlabelled data suggests continued pre-training. If the information is proprietary or needs to stay current, think RAG. Fine-tuning does not keep a model's knowledge up to date, and it does not reduce inference cost.
Amazon Bedrock
| Feature | What it does |
|---|---|
| Knowledge Bases | Managed RAG, used to ground responses in your documents |
| Agents | Multi-step task orchestration, including calls to APIs and tools |
| Guardrails | Safety filtering for inputs and outputs |
| Model evaluation | Automatic evaluation for the lowest overhead, or a human workforce where style, tone, and preference matter |
| Advanced prompts | The place to provide examples to an Agent |
Bedrock pricing modes
| Mode | When to use it |
|---|---|
| On-Demand | Pay per token; suitable for variable or low volume where there is no commitment |
| Batch | Discounted bulk processing when latency is not important |
| Provisioned Throughput | For a steady, predictable rate; it is required to serve a custom model |
At inference time, token count drives cost, both input and output tokens. Temperature, training duration, and model size do not. Bedrock inputs and outputs are not shared with model providers or used to train base models.
Guardrails policies
The names are close enough that it pays to learn the distinctions rather than relying on the general idea.
| Policy | What it covers |
|---|---|
| Content filters | Hate, insults, violence, sexual content, misconduct, and prompt attacks |
| Denied topics | Avoid a whole subject area, such as legal or investment advice |
| Word filters | Block particular words or phrases |
| Sensitive information filters | PII in prompts or responses |
| Contextual grounding check | A response is not supported by the retrieved source documents |
Contextual grounding is an easy one to overlook. If a RAG application makes up details that are absent from its source material, that is a grounding issue, not a content-filter issue.
Guardrails work at runtime on inputs and outputs. They do not clean a fine-tuning data set; PII removal before fine-tuning happens earlier.
Prompt engineering
| Technique | What it is |
|---|---|
| Zero-shot | An instruction with no examples |
| One-shot / few-shot | Examples that establish a format or set of labels |
| Chain-of-thought | "Think step by step" for reasoning and multi-step logic |
| Prompt chaining | Split a large task into sequential sub-prompts |
| ReAct | Reasoning combined with tool or live-data calls |
| Negative prompts | Say what should be excluded; common in image generation |
| Role description in the prompt | The cheapest way to change tone or target audience |
Three risks are worth keeping separate. Prompt injection is an input vulnerability where user content manipulates model behaviour. Jailbreaking is an attempt to bypass safety controls and get harmful output. Prompt template extraction exposes configured system behaviour.
Inference parameters
Temperature controls randomness. At 0, results are deterministic and consistent; a higher value is more creative. Reducing temperature also reduces hallucination.
Top K is the number of candidate tokens considered at each step. Top P is the cumulative-probability cut-off. Max tokens limits response length, while the context window is the amount of text that fits into a single prompt.
If an application fails with long documents, it is likely exceeding the context window. That calls for a different model, not a different inference parameter.
Evaluation metrics
| Task | Metric |
|---|---|
| Translation | BLEU |
| Summarisation | ROUGE |
| Semantic similarity to a reference text | BERTScore |
| Balanced classification | F1 score |
| Minimising false positives | Precision |
| Minimising false negatives | Recall |
| Correct / total classified | Accuracy |
| Per-class right-and-wrong breakdown | Confusion matrix |
| Regression error | RMSE, MSE, R² |
| Production runtime efficiency | Average response time |
| Business impact of a chatbot | Cost per conversation, average handle time, CSAT |
Questions often describe precision and recall without naming them. If the concern is spending time reviewing items that turn out to be fine, it is precision. If missing even one case is unacceptable, it is recall. I translate the wording into false positives or false negatives first; the answer usually becomes obvious.
ML fundamentals
| Situation | Approach |
|---|---|
| Labelled data with known outputs | Supervised learning |
| Unlabelled data used to find groups or structure | Unsupervised learning |
| Learning from rewards or feedback | Reinforcement learning, including RLHF with human feedback |
| Adapting a pre-trained model to a related task | Transfer learning |
| Training without centralising data | Federated learning |
Useful algorithm matches:
| Need | Algorithm |
|---|---|
| Classify by nearest examples | k-NN |
| Segment customers into groups | K-means |
| Numeric prediction | Linear regression |
| Interpretable, adjustable weights | Logistic regression |
| Explain the internal decision path | Decision trees |
| Time-series forecasting | DeepAR or ARIMA |
| Generate synthetic data | GAN |
| Predict missing words | BERT-based models |
Overfitting means strong performance on training data but weak performance on new data. Counter it with more, and more varied, data; greater regularisation; earlier stopping; or a simpler model. Underfitting calls for more epochs, more features, or more capacity.
The lifecycle is: business goal, frame the ML problem, collect data, pre-process and explore, feature engineering, train and tune, evaluate, deploy, monitor. Compliance and regulatory requirements belong at the business-goal stage. Correlation matrices, summary statistics, and plots are exploratory data analysis.
SageMaker
| Feature | What it does |
|---|---|
| Clarify | Bias detection and explainability together |
| Model Cards | Standardised model documentation for audit |
| Model Monitor | Drift and quality degradation in production |
| Ground Truth | Human data labelling; Ground Truth Plus uses a managed workforce |
| Canvas | No-code model building for non-engineers |
| Data Wrangler | Visual data preparation |
| Feature Store | Sharing features across teams |
| JumpStart | Pre-built models and solutions |
| Model Registry | Model version management |
| Network isolation, plus a VPC with an S3 endpoint | Training and inference with no internet access |
SageMaker inference options
| Option | When to use it |
|---|---|
| Real-time | Low latency and steady traffic |
| Serverless | Intermittent or unpredictable traffic, with no infrastructure to manage |
| Asynchronous | Payloads up to 1 GB, processing up to an hour, and near-real-time results |
| Batch transform | Large offline data sets where results are not needed immediately |
The asynchronous limits above are documented service limits. They are a useful way of telling the four options apart, so learn the numbers rather than just the broad descriptions.
Matching a service to the data
Before choosing a service, check the input type. Several AWS services sound broadly similar but are meant for different modalities.
| Task | Service |
|---|---|
| Speech to text, call recordings, subtitles | Amazon Transcribe |
| Text to speech | Amazon Polly |
| Language translation | Amazon Translate |
| Sentiment, entities, toxicity, and PII in text | Amazon Comprehend |
| Medical entity extraction | Amazon Comprehend Medical |
| Clinical documentation from dictation | AWS HealthScribe |
| Documents and PDFs to structured text | Amazon Textract |
| Image and video analysis or moderation | Amazon Rekognition |
| Personalised recommendations | Amazon Personalize |
| Enterprise document search | Amazon Kendra |
| Conversational interfaces with intents | Amazon Lex |
| Sensitive-data discovery in S3 | Amazon Macie |
| Vector storage | OpenSearch (k-NN) or Aurora PostgreSQL with pgvector |
| Energy-efficient training hardware | EC2 Trn (Trainium); use EC2 Inf for inference |
Responsible AI
AWS documents eight dimensions: fairness, explainability, privacy and security, safety, controllability, veracity and robustness, governance, and transparency. A few come up particularly often in scenarios.
| Scenario | Dimension |
|---|---|
| Unrepresentative training data that disadvantages a group | Fairness |
| Users needing the rationale behind a decision | Explainability |
| Documenting purpose, limits, and lineage | Transparency |
| Anonymising personal data before training | Privacy and security |
| Preventing harmful output | Safety |
Bias vocabulary: sampling bias means the data does not represent the population. Measurement bias comes from an incorrect or proxy measurement. Observer bias is when labeller expectations leak in. Confirmation bias is seeking data that supports an existing belief.
Common remedies include collecting balanced, diverse data; measuring class imbalance and adapting training; and, for a biased fine-tuned model, adding diverse data and fine-tuning again. Human-in-the-loop review can deal with bias and toxicity at the post-processing stage.
One pair I make a point of separating: Model Cards document models you build, while AI Service Cards are AWS documentation for AWS AI services.
Security and governance
| Requirement | Service or control |
|---|---|
| Who called which API and unauthorised access attempts | CloudTrail |
| Metrics, logs, and alarms | CloudWatch |
| Resource-configuration compliance over time | AWS Config |
| Continuous audit evidence against a framework | AWS Audit Manager |
| Download SOC, ISO, and PCI compliance reports | AWS Artifact |
| Restrict which users can access which models | IAM policies and least privilege |
| Customer-managed encryption keys | AWS KMS |
| Reach Bedrock from a VPC without internet access | AWS PrivateLink |
| Record prompts and completions | Bedrock model invocation logging |
| Data must remain in-country | Data residency |
Under the shared responsibility model, the customer is responsible for its data in transit and at rest, access control, and the use case. AWS is responsible for the underlying infrastructure and patching.
The Generative AI Security Scoping Matrix goes from Scope 1, a consumer application, to Scope 5, a self-trained model. Security responsibility increases as you move up those scopes.
Reading the requirements
AWS wording is usually deliberate. These clues have helped me narrow options quickly:
- Cost qualifier: choose the cheapest approach that still meets the need, usually lower on the customisation ladder.
- Operational-overhead qualifier: favour a managed or serverless service instead of assembling components yourself.
- "Without managing infrastructure": serverless inference or a managed endpoint.
- "Real time" versus a scheduled frequency: real-time inference versus batch.
- "The company's own documents": RAG.
If two answers differ mainly by service name, go back to the data modality. If they differ by effort, ask whether the question actually needs the extra effort.
Last-minute checklist
Use the lowest rung of the customisation ladder that works. Labelled data means fine-tuning; unlabelled data means continued pre-training; changing data means RAG. BLEU is for translation, ROUGE for summarisation, BERTScore for similarity, and F1 for balanced classification. Temperature 0 is consistent. Inference cost is token-based. Clarify covers bias and explainability, Model Cards cover documentation, and Model Monitor catches drift. CloudTrail shows who did what; Artifact provides compliance reports. And do not leave anything blank.
Sources
These notes are based on the official AIF-C01 exam guide, the AWS Skill Builder official practice-question set, and AWS service documentation. The interpretations and revision reminders are my own.
Top comments (0)