Skip to content
BytePatterns

MLA-C02 · Domain 2: ML Model and Foundation Model (FM) Development · 24% of the exam

Task 2.1: Choose appropriate modeling approaches for ML and AI solutions.

Picking between an AWS AI service, a built-in algorithm, a custom model and a foundation model in Amazon Bedrock, choosing a fine-tuning strategy or a RAG pattern, and weighing accuracy against latency, training time and cost.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A lender receives thousands of scanned loan application forms each day. It wants to pull out fields such as "Applicant name" with their values, and the rows of an income table, so that a downstream model can score each application. Which AWS service fits this step with the least development effort?

  1. AAmazon Rekognition, detecting text in the scanned images
  2. BAmazon Comprehend, detecting entities in the scanned files
  3. CAmazon Textract, with forms and tables analysis
  4. DA custom object detection model trained on SageMaker AI
Show the answer and why
  • AAmazon Rekognition, detecting text in the scanned images

    Incorrect

    Rekognition is an image and video analysis service. Detecting words in an image does not link a field name to its value or rebuild a table.

  • BAmazon Comprehend, detecting entities in the scanned files

    Incorrect

    Comprehend finds entities such as names and dates in text. It does not turn a form layout into key-value pairs and table rows.

  • CAmazon Textract, with forms and tables analysis

    Correct

    Textract analysis returns form data as key-value pairs and extracts tables with their cells, which is exactly the structure the scoring model needs from scanned forms.

  • DA custom object detection model trained on SageMaker AI

    Incorrect

    A custom model could locate regions on a page, but it would need labeled data, training and extra logic to read and pair the text, which Textract already provides as a managed service.

When a managed AWS AI service already solves the task, it is usually the fastest and cheapest path. Document structure (forms, tables) is Textract's job; images in general are Rekognition's; plain text understanding is Comprehend's.

Question 2 · choose 1

A utility collects readings from thousands of smart meters. Nobody has labeled which readings are abnormal, and the team wants a SageMaker AI built-in algorithm that assigns an anomaly score to each reading. Which algorithm should the team use?

  1. ALinear Learner
  2. BDeepAR forecasting
  3. CRandom Cut Forest
  4. DObject2Vec
Show the answer and why
  • ALinear Learner

    Incorrect

    Linear Learner is a supervised algorithm for classification and regression. Without labels of normal and abnormal readings it has nothing to learn from.

  • BDeepAR forecasting

    Incorrect

    DeepAR forecasts future values of time series. It predicts what comes next rather than scoring how unusual each existing reading is.

  • CRandom Cut Forest

    Correct

    Random Cut Forest (RCF) is the built-in unsupervised algorithm for anomaly detection, such as spotting when a sensor sends abnormal readings, and it needs no labels.

  • DObject2Vec

    Incorrect

    Object2Vec learns embeddings of pairs of objects, for example to find similar support tickets. It is not an anomaly detector.

Match the problem type to the algorithm family: unlabeled anomaly scoring points to RCF (or IP Insights for IP addresses), labeled tabular prediction to XGBoost or Linear Learner, and forecasting to DeepAR.

Question 3 · choose 1

A support organization wants to route about 400,000 incoming emails a day into 12 fixed categories. It already has 60,000 emails labeled with the right category. The solution should be fully managed, cheap per email at this volume, and quick to build. Which approach fits best?

  1. ASend every email to a large foundation model in Amazon Bedrock with a classification prompt
  2. BTrain an Amazon Comprehend custom classifier on the labeled emails
  3. CWrite a custom PyTorch training script and host the model on a SageMaker AI endpoint
  4. DUse Amazon Comprehend built-in sentiment analysis to route the emails
Show the answer and why
  • ASend every email to a large foundation model in Amazon Bedrock with a classification prompt

    Incorrect

    A prompt-based classifier can work, but paying for a large model's input and output tokens on 400,000 emails a day costs far more than a trained classifier, and the labeled data would go unused.

  • BTrain an Amazon Comprehend custom classifier on the labeled emails

    Correct

    Comprehend custom classification trains a classifier on your own categories from labeled examples, then classifies documents in real time or in asynchronous jobs. It is managed and built for exactly this kind of routing.

  • CWrite a custom PyTorch training script and host the model on a SageMaker AI endpoint

    Incorrect

    This can reach good accuracy, but it means building, tuning and operating the model and its endpoint, which is more effort than the managed service the requirements point to.

  • DUse Amazon Comprehend built-in sentiment analysis to route the emails

    Incorrect

    Sentiment returns positive, negative, neutral or mixed. It cannot assign the organization's own 12 categories.

Weigh managed services, custom models and foundation models against volume, labels and effort. With good labels, fixed classes and very high volume, a trained managed classifier is usually cheaper per prediction than a general-purpose foundation model.

Question 4 · choose 1

Sales managers want to ask questions in plain English such as "What was total revenue by region last quarter?" The answers live in tables in an Amazon Redshift data warehouse, and the numbers must be computed exactly from those tables. Which retrieval pattern should the ML engineer choose?

  1. AA vector knowledge base built from nightly CSV exports of the tables
  2. BA GraphRAG knowledge base on Amazon Neptune Analytics
  3. CA knowledge base that uses Redshift as a structured data store
  4. DFine-tuning a foundation model on a snapshot of the sales tables
Show the answer and why
  • AA vector knowledge base built from nightly CSV exports of the tables

    Incorrect

    Vector retrieval returns the text chunks most similar to the question. It cannot sum and group millions of rows, so totals assembled from a few chunks would be wrong.

  • BA GraphRAG knowledge base on Amazon Neptune Analytics

    Incorrect

    GraphRAG connects entities and relationships across document chunks for multi-step reasoning over text. It does not run aggregate queries over warehouse tables.

  • CA knowledge base that uses Redshift as a structured data store

    Correct

    A knowledge base connected to a structured data store converts natural-language questions into queries for that store, such as SQL, retrieves the result and generates the answer, so the totals come straight from the tables.

  • DFine-tuning a foundation model on a snapshot of the sales tables

    Incorrect

    Fine-tuning teaches patterns from training data, not an exact, current copy of it. The model would produce stale or invented numbers instead of querying the warehouse.

Pick the RAG pattern from the shape of the data: unstructured text uses vector retrieval, connected entities across documents can use GraphRAG, and tables in a database or warehouse use a structured data store that turns questions into queries.

Question 5 · choose 1

A team wants a model in Amazon Bedrock to write SQL queries for its own schema. Writing ideal answers for thousands of questions would be slow and expensive, but each generated query can be checked automatically by running it against a test database and comparing the result. Which customization method fits best?

  1. AReinforcement fine-tuning with a Lambda reward function that runs each query
  2. BSupervised fine-tuning on question and ideal-query pairs written by the team
  3. CDistillation from a larger teacher model
  4. DA larger context window filled with the full schema
Show the answer and why
  • AReinforcement fine-tuning with a Lambda reward function that runs each query

    Correct

    Reinforcement fine-tuning learns from reward scores instead of labeled input-output pairs, and reward functions can be written in AWS Lambda. It suits tasks whose success can be verified programmatically and where labeled examples are expensive.

  • BSupervised fine-tuning on question and ideal-query pairs written by the team

    Incorrect

    Supervised fine-tuning needs labeled examples of the desired output, which is the slow and expensive work the team wants to avoid.

  • CDistillation from a larger teacher model

    Incorrect

    Distillation transfers a teacher's ability to a smaller student. It does not use the team's ability to verify each query by running it.

  • DA larger context window filled with the full schema

    Incorrect

    Including the schema in the prompt helps, but it does not change the model and does not use the automatic correctness check to improve it.

Supervised fine-tuning learns from labeled pairs, distillation from a teacher's answers, and reinforcement fine-tuning from a reward signal. When correctness can be checked by code but ideal answers are costly to write, rewards are the cheaper signal.

Question 6 · choose 1

A customer-facing assistant on Amazon Bedrock must answer as fast as possible during unpredictable business-hour peaks. The traffic does not justify reserving capacity around the clock, and the team accepts paying more per token for faster responses. Which inference option should the ML engineer use?

  1. AThe Priority service tier
  2. BThe Reserved service tier
  3. CThe Flex service tier
  4. DBatch inference jobs
Show the answer and why
  • AThe Priority service tier

    Correct

    Priority delivers the fastest response times for a price premium over standard on-demand pricing, needs no reservation, and is meant for customer-facing workloads that do not warrant 24x7 capacity.

  • BThe Reserved service tier

    Incorrect

    Reserved capacity is booked for one or three months with minimum tokens-per-minute commitments. That is the round-the-clock reservation the team does not want.

  • CThe Flex service tier

    Incorrect

    Flex gives a discount for workloads that can handle longer processing times. It trades speed for cost, the opposite of what is needed.

  • DBatch inference jobs

    Incorrect

    Batch inference processes prompts asynchronously in bulk. It cannot answer a live user in a conversation.

Bedrock offers Reserved, Priority, Standard and Flex tiers, plus batch inference. Choose by how much latency matters and whether capacity is needed all the time.

Question 7 · choose 1

A bank wants a chatbot, for its website and its phone line, that recognizes requests such as "check my balance" and asks follow-up questions to collect details like the account type. It wants a managed service built for conversational interfaces with voice and text. Which AWS service should the ML engineer start with?

  1. AAmazon Polly
  2. BAmazon Comprehend
  3. CAmazon Lex V2
  4. DAmazon Translate
Show the answer and why
  • AAmazon Polly

    Incorrect

    Polly converts text to speech; it does not understand user intents or manage a conversation.

  • BAmazon Comprehend

    Incorrect

    Comprehend analyzes text for entities, sentiment and similar insights; it is not a conversational bot service.

  • CAmazon Lex V2

    Correct

    Amazon Lex V2 is an AWS service for building conversational interfaces with voice and text, providing natural language understanding and automatic speech recognition for chatbots.

  • DAmazon Translate

    Incorrect

    Translate converts text between languages.

For intent-driven conversational interfaces across voice and text, start with a service built for that purpose before building one from models.

Question 8 · choose 1

A marketplace lets users upload product photos. The trust and safety team wants each image checked for inappropriate or offensive content before it is published, without training its own model. Which AWS capability fits?

  1. AAmazon Textract document analysis
  2. BAmazon Comprehend sentiment analysis of file names
  3. CA SageMaker AI k-means model on pixel values
  4. DAmazon Rekognition content moderation
Show the answer and why
  • AAmazon Textract document analysis

    Incorrect

    Textract extracts text and structure from documents; it does not judge whether images are offensive.

  • BAmazon Comprehend sentiment analysis of file names

    Incorrect

    File names say little about what an image shows.

  • CA SageMaker AI k-means model on pixel values

    Incorrect

    Clustering pixels groups similar images but does not identify inappropriate content, and it means building a model.

  • DAmazon Rekognition content moderation

    Correct

    Rekognition moderation APIs detect inappropriate, unwanted or offensive content in images and videos for use cases such as user-generated content.

Managed moderation services provide ready-made detection of unsafe visual content for user-generated media.

Question 9 · choose 1

A real estate firm wants to predict sale prices from 30 numeric and categorical property features in a table of 200,000 past sales. Someone suggests prompting a large language model in Amazon Bedrock with each row. What should the ML engineer recommend instead?

  1. AA tabular algorithm such as XGBoost
  2. BA text-to-image foundation model
  3. CThe k-means clustering algorithm
  4. DA Bedrock knowledge base over the sales table
Show the answer and why
  • AA tabular algorithm such as XGBoost

    Correct

    AWS maps tabular regression, such as estimating the value of a house, to built-in algorithms like XGBoost, LightGBM, CatBoost and Linear Learner, which learn directly from labeled tabular data.

  • BA text-to-image foundation model

    Incorrect

    Image generation does not predict numeric prices.

  • CThe k-means clustering algorithm

    Incorrect

    K-means groups similar records; it does not predict a numeric target.

  • DA Bedrock knowledge base over the sales table

    Incorrect

    Retrieval returns similar records to a prompt; it does not train a regression model on all 200,000 labeled sales.

Match the problem to the model family: labeled tabular regression is a job for classical supervised algorithms, not generative models.

Question 10 · choose 1

A help desk wants to find duplicate support tickets and route new tickets to the team that handled the most similar past tickets. It has many past ticket pairs marked as related or unrelated. Which SageMaker AI built-in algorithm fits this similarity problem?

  1. ADeepAR
  2. BObject2Vec
  3. CRandom Cut Forest
  4. DIP Insights
Show the answer and why
  • ADeepAR

    Incorrect

    DeepAR forecasts time series values.

  • BObject2Vec

    Correct

    Object2Vec learns embeddings that preserve the relationship between pairs of objects, which can be used to find nearest neighbors; AWS lists identifying duplicate support tickets as an example.

  • CRandom Cut Forest

    Incorrect

    RCF scores anomalies; it does not learn similarity between ticket pairs.

  • DIP Insights

    Incorrect

    IP Insights learns associations between IPv4 addresses and entities.

Pairwise similarity problems, such as duplicates or routing by resemblance, fit embedding algorithms such as Object2Vec.

Question 11 · choose 1

A publisher must assign one of 40 predefined subject categories to millions of short book descriptions, using its existing labeled examples. It wants a SageMaker AI built-in algorithm for text classification rather than a generative model. Which algorithm should the ML engineer use?

  1. ALatent Dirichlet Allocation (LDA)
  2. BSequence-to-Sequence algorithm
  3. CBlazingText text classification
  4. DK-means clustering algorithm
Show the answer and why
  • ALatent Dirichlet Allocation (LDA)

    Incorrect

    LDA discovers topics that are not known in advance; here the 40 categories are predefined and labeled.

  • BSequence-to-Sequence algorithm

    Incorrect

    Sequence-to-sequence models produce output sequences, such as translations or summaries, not category labels.

  • CBlazingText text classification

    Correct

    AWS maps text classification, assigning predefined categories to documents such as books by discipline, to the BlazingText algorithm and Text Classification - TensorFlow.

  • DK-means clustering algorithm

    Incorrect

    Clustering groups similar items without using the known labels.

Predefined labels mean supervised classification; unknown groupings mean topic modeling or clustering.

Practise domain 2 →Practise all domains →