Skip to content
BytePatterns

AIF-C01 · Domain 1: Fundamentals of AI and ML · 20% of the exam

Task 1.1: Explain basic AI concepts and terminologies.

The vocabulary the rest of the exam assumes: how AI, machine learning, deep learning, generative AI and agentic AI nest inside one another, the kinds of data and learning, and the ways a trained model is asked for predictions.

Study it

  • AI, ML, deep learning, generative AI and agentic AI: how the terms nest

    Partly covered by: What Is Machine Learning

  • Data and learning: labeled and unlabeled data, supervised, unsupervised and reinforcement learning

    Partly covered by: What Is Machine Learning

  • Training and inference: real-time, serverless, asynchronous and batch

    Partly covered by: Training vs Inference

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A media company runs a trained model that analyzes uploaded video files. Each file is about 500 MB and takes up to 20 minutes to process. Every upload must be queued for processing as soon as it arrives, uploads come in irregular bursts, and the company does not want to pay for idle instances between bursts. Which Amazon SageMaker AI inference option fits best?

  1. AAsynchronous inference
  2. BReal-time inference on a persistent endpoint
  3. CServerless inference
  4. DBatch transform jobs over the full dataset
Show the answer and why
  • AAsynchronous inference

    Correct

    Asynchronous inference queues requests, accepts payloads up to 1 GB with processing times up to one hour, and can scale the endpoint down to zero when there is nothing to process.

  • BReal-time inference on a persistent endpoint

    Incorrect

    Real-time inference is for low-latency online requests and supports payloads up to 25 MB with 60-second processing for regular responses, far below a 500 MB file that runs for 20 minutes.

  • CServerless inference

    Incorrect

    Serverless inference suits intermittent traffic, but it supports payloads only up to 4 MB and processing times up to 60 seconds.

  • DBatch transform jobs over the full dataset

    Incorrect

    Batch transform is for offline processing when a large dataset is available upfront. Here each file arrives on its own and must be queued the moment it is uploaded.

Large payloads, long processing times and bursty traffic with a wish to scale to zero are the profile SageMaker AI documents for asynchronous inference.

Question 2 · choose 1

Which statement correctly describes how artificial intelligence (AI), machine learning (ML), deep learning and generative AI relate to each other?

  1. ADeep learning is the broadest of the four fields, and both AI and ML are specialized techniques that sit inside deep learning.
  2. BML is one branch of AI, deep learning is a subset of ML, and generative AI is a type of AI that creates new content.
  3. CAI and ML are two names for the same thing, because every AI system has to learn its behavior from data.
  4. DGenerative AI is separate from ML, because generative models are written as fixed rules and are not trained on any data.
Show the answer and why
  • ADeep learning is the broadest of the four fields, and both AI and ML are specialized techniques that sit inside deep learning.

    Incorrect

    It is the other way round: AI is the umbrella term, and deep learning is a subset of machine learning.

  • BML is one branch of AI, deep learning is a subset of ML, and generative AI is a type of AI that creates new content.

    Correct

    AWS describes ML as one of many branches of AI, deep learning as a subset of ML, and generative AI as a subset of AI that creates new content such as conversations, stories, images and music.

  • CAI and ML are two names for the same thing, because every AI system has to learn its behavior from data.

    Incorrect

    Not all AI is ML. The field of AI also includes methods such as rule-based systems and search algorithms.

  • DGenerative AI is separate from ML, because generative models are written as fixed rules and are not trained on any data.

    Incorrect

    Generative AI grew out of advances in deep learning, and foundation models are ML models trained on a broad spectrum of data.

Think of nested circles: AI contains ML, ML contains deep learning, and generative AI is the part of AI, built on deep learning, that produces new content.

Question 3 · choose 1

A logistics company wants software that learns how to route its warehouse robots. No labeled examples of ideal routes exist. The software can try routes in a simulator, and each attempt receives a higher score when deliveries finish sooner. Which type of machine learning does this describe?

  1. ASupervised learning
  2. BUnsupervised learning
  3. CReinforcement learning
  4. DSemi-supervised learning
Show the answer and why
  • ASupervised learning

    Incorrect

    Supervised learning trains on input data paired with labeled outputs, and there are no labeled ideal routes to learn from here.

  • BUnsupervised learning

    Incorrect

    Unsupervised learning finds patterns and groups in unlabeled data. It does not learn from a score given to the actions it tries.

  • CReinforcement learning

    Correct

    Reinforcement learning learns by trial and error: actions that move toward the goal earn rewards, and the model learns the actions that maximize the cumulative reward.

  • DSemi-supervised learning

    Incorrect

    Semi-supervised learning combines a small labeled dataset with a large unlabeled one. There is no labeled data at all in this scenario.

A reward signal for actions taken in an environment, with no labeled answers, is the defining setup of reinforcement learning.

Question 4 · choose 2

A retailer is listing the data it could use for AI projects. Which of these are examples of unstructured data? (Choose TWO.)

  1. ARecorded audio of customer service calls
  2. BAn inventory listing stored in a relational database table
  3. CVideo footage from in-store security cameras
  4. DHR records with fields for name, start date and salary
  5. EMonthly sales figures per store kept in a spreadsheet table
Show the answer and why
  • ARecorded audio of customer service calls

    Correct

    Audio files do not fit neatly into rows and columns, so AWS lists them as unstructured data.

  • BAn inventory listing stored in a relational database table

    Incorrect

    Inventory listings with short text and numeric fields are a typical example of structured data, and relational databases are structured data stores.

  • CVideo footage from in-store security cameras

    Correct

    Video monitoring is one of the examples AWS gives of unstructured data.

  • DHR records with fields for name, start date and salary

    Incorrect

    Records made of short text, dates and numbers fit a table with one column per attribute, which makes them structured data.

  • EMonthly sales figures per store kept in a spreadsheet table

    Incorrect

    Sales figures are discrete numeric data in rows and columns, the kind of data AWS describes as structured.

Structured data fits a table of rows and columns. Audio, video, images and long free text do not, and they are where machine learning is most often applied.

Question 5 · choose 1

A company has a chat assistant that answers employee questions when asked. It now wants a system that is given a goal such as "onboard this new hire", decides the steps on its own, calls the HR and IT systems, and keeps working without a person prompting each step. What does the new requirement describe?

  1. AA workflow engine that runs a fixed, predefined sequence of steps
  2. BA generative AI model that writes text only in reply to each prompt
  3. CAn unsupervised model that groups past onboarding tickets
  4. DAgentic AI that plans steps and uses tools to reach the goal
Show the answer and why
  • AA workflow engine that runs a fixed, predefined sequence of steps

    Incorrect

    Traditional software follows predefined rules. The requirement is for a system that chooses its own next actions toward a goal.

  • BA generative AI model that writes text only in reply to each prompt

    Incorrect

    That is what the existing assistant already does. AWS contrasts this prompt-by-prompt guidance with agentic AI, which acts on its own.

  • CAn unsupervised model that groups past onboarding tickets

    Incorrect

    Grouping unlabeled records finds patterns in data. It does not plan steps or act on other systems.

  • DAgentic AI that plans steps and uses tools to reach the goal

    Correct

    Agentic AI is an autonomous system that acts independently to achieve goals; AI agents choose their next actions themselves and use tools and systems to complete tasks without constant human oversight.

The difference between a generative AI assistant and agentic AI is agency: an agent is given a goal, plans its own steps and acts through tools.

Question 6 · choose 1

Once a month, an insurer scores its entire customer database, about ten million records stored in Amazon S3, with a churn model. Nobody needs answers between runs, and the team does not want an endpoint running all month. Which Amazon SageMaker AI inference option fits best?

  1. AReal-time inference
  2. BServerless inference
  3. CAsynchronous inference
  4. DBatch transform
Show the answer and why
  • AReal-time inference

    Incorrect

    Real-time inference provides a persistent endpoint for low-latency online requests and sustained traffic, which this monthly job does not need.

  • BServerless inference

    Incorrect

    Serverless inference targets intermittent request traffic with payloads up to 4 MB and 60-second processing, not one large offline dataset.

  • CAsynchronous inference

    Incorrect

    Asynchronous inference queues individual requests with large payloads. The whole dataset is available upfront, which is the batch case.

  • DBatch transform

    Correct

    Batch transform is for offline processing when large amounts of data are available upfront and you do not need a persistent endpoint; it handles datasets that are gigabytes in size.

All the data at once, no one waiting, no endpoint wanted: that is batch inference, which SageMaker AI calls batch transform.

Question 7 · choose 1

A startup exposes a small text classification model through an API. Traffic is unpredictable, with long idle periods between bursts; each request is a few kilobytes and is answered in under a second; and the occasional cold start is acceptable. The startup does not want to pay for idle time. Which Amazon SageMaker AI inference option fits best?

  1. AServerless inference
  2. BReal-time inference on an always-on instance
  3. CBatch transform
  4. DAsynchronous inference
Show the answer and why
  • AServerless inference

    Correct

    Serverless inference suits intermittent or unpredictable traffic that can tolerate cold starts; it scales to zero when idle so you pay only for what you use, with payloads up to 4 MB.

  • BReal-time inference on an always-on instance

    Incorrect

    A real-time endpoint is persistent and backed by the instance type you choose, so it is billed while idle between bursts.

  • CBatch transform

    Incorrect

    Batch transform runs offline over a dataset that is available upfront; it does not serve individual API requests.

  • DAsynchronous inference

    Incorrect

    Asynchronous inference is for queued requests with large payloads and long processing times, not small requests answered in under a second.

Small payloads, short processing, spiky traffic and no budget for idle time is the profile SageMaker AI describes for serverless inference.

Question 8 · choose 1

A company has 50,000 past emails, each tagged by staff as "spam" or "not spam". It wants a model that tags new emails the same way. Which type of machine learning is this?

  1. AUnsupervised learning
  2. BSupervised learning
  3. CReinforcement learning
  4. DGenerative pre-training
Show the answer and why
  • AUnsupervised learning

    Incorrect

    Unsupervised learning works on unlabeled data to discover patterns. Here every email already has a label.

  • BSupervised learning

    Correct

    Supervised learning trains on inputs paired with labeled outputs, and AWS gives email spam classification as an example of it.

  • CReinforcement learning

    Incorrect

    Reinforcement learning learns from rewards for actions taken in an environment, not from a labeled dataset.

  • DGenerative pre-training

    Incorrect

    Foundation models are pre-trained on broad, generalized and unlabeled data. Learning to reproduce known labels is supervised learning.

Known answers attached to every example means supervised learning; tagging into two categories makes it binary classification.

Question 9 · choose 1

In machine learning, what is the difference between training and inference?

  1. ATraining and inference are the same step run on different hardware
  2. BInference adjusts the model's parameters, while training is the step that produces the outputs
  3. CTraining adjusts parameters from data; inference uses the trained model on new input
  4. DInference happens only once, before the model is deployed
Show the answer and why
  • ATraining and inference are the same step run on different hardware

    Incorrect

    They are different steps: training adjusts parameters, and inference uses the trained model to produce outputs.

  • BInference adjusts the model's parameters, while training is the step that produces the outputs

    Incorrect

    It is the other way round. Training adjusts parameters to reduce error; inference generates an output from an input.

  • CTraining adjusts parameters from data; inference uses the trained model on new input

    Correct

    During training a model adjusts its parameters to minimize the gap between predictions and known outcomes. Inference is getting predictions from the trained model for a given input.

  • DInference happens only once, before the model is deployed

    Incorrect

    Inference is how a deployed model is used; SageMaker AI provides deployment options to get predictions, or inferences, from trained models.

Training is the learning phase; inference is the using phase, and it runs every time the deployed model answers a request.

Question 10 · choose 2

Which of these tasks are examples of natural language processing (NLP)? (Choose TWO.)

  1. ADetecting whether product reviews express positive or negative feelings
  2. BIdentifying cars and pedestrians in traffic camera images
  3. CTranslating support articles from English into German
  4. DPredicting next month's electricity use from past meter readings
  5. EGrouping network traffic into types to spot security incidents
Show the answer and why
  • ADetecting whether product reviews express positive or negative feelings

    Correct

    NLP lets computers interpret human language, including understanding the intent or sentiment hidden in it.

  • BIdentifying cars and pedestrians in traffic camera images

    Incorrect

    Deriving information from images and video is computer vision.

  • CTranslating support articles from English into German

    Correct

    Machine translation is an NLP task; NLP research began with experiments in translating sentences between languages.

  • DPredicting next month's electricity use from past meter readings

    Incorrect

    Predicting a numeric value from past data is a regression or forecasting problem, not language processing.

  • EGrouping network traffic into types to spot security incidents

    Incorrect

    Grouping unlabeled records by similarity is clustering, an unsupervised learning technique.

NLP works on human language, written or spoken; images belong to computer vision, and numeric prediction or grouping are general ML techniques.

Practise domain 1 →Practise all domains →