Skip to content
BytePatterns

AIB-C01 · Domain 1: AI Fundamentals and Literacy · 24% of the exam

Task 1.1: Describe core AI concepts and define terminology.

The vocabulary a business leader needs: how AI, machine learning and generative AI relate, training and inference, structured and unstructured data, why data quality decides outcomes, and the global standards that give AI a shared language.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A retail bank trained a fraud model last quarter on two years of labeled card transactions. Since launch, the model scores every new card payment within milliseconds and flags suspicious ones for review. The program sponsor asks what to call the step the model performs on each new payment so that the budget line for it is named correctly. What is that step?

  1. ATraining, because the model updates its parameters every time it scores a new payment
  2. BData preprocessing, because each payment must be cleaned before it is scored
  3. CInference, which applies the trained model to new payments it has never seen
  4. DModel evaluation, because each score is tested against validation data
Show the answer and why
  • ATraining, because the model updates its parameters every time it scores a new payment

    Incorrect

    Training is the phase in which the algorithm learns patterns from historical examples and adjusts its parameters. Scoring a payment does not change the model; it uses the parameters it already has.

  • BData preprocessing, because each payment must be cleaned before it is scored

    Incorrect

    Preprocessing cleans and transforms raw data, mainly to prepare it for training. It is not the step in which the model produces a fraud score.

  • CInference, which applies the trained model to new payments it has never seen

    Correct

    Inference is the process of a trained model generating an output, here a fraud score, from a new input. It is the recurring production cost of a model, separate from the one-time training effort.

  • DModel evaluation, because each score is tested against validation data

    Incorrect

    Evaluation measures how well a model generalizes by testing it on a separate validation dataset before or after release. Scoring live payments is not evaluation; the true outcome is not known yet.

Training is where a model learns from historical data; inference is where the trained model is used to produce predictions on new data. A fraud model that scores each live payment is performing inference, which is why inference volume, not training, drives most of the ongoing cost of a model in production.

Question 2 · choose 1

A vendor demonstrates a claims-routing product to an insurer's operations director and calls it "generative AI". In the demo, the product sends each claim to a team using 120 if-then conditions that the vendor's consultants wrote by hand; the conditions change only when a consultant edits them, and the product never writes any new text. How should the director describe the product when comparing vendors?

  1. AAs machine learning, because the product makes routing decisions on its own
  2. BAs generative AI, because the product produces a different routing result for each claim
  3. CAs deep learning, because the 120 conditions work together like layers of a network
  4. DAs rule-based automation, which is neither machine learning nor generative AI
Show the answer and why
  • AAs machine learning, because the product makes routing decisions on its own

    Incorrect

    Making a decision automatically is not what defines machine learning. Machine learning derives its decision logic from data; this product's logic was written by people and changes only when people edit it.

  • BAs generative AI, because the product produces a different routing result for each claim

    Incorrect

    Generative AI creates new content such as text or images. Choosing one of a fixed set of teams from fixed conditions creates nothing new.

  • CAs deep learning, because the 120 conditions work together like layers of a network

    Incorrect

    Deep learning trains neural networks on data. A list of hand-written conditions is not a neural network and is not trained.

  • DAs rule-based automation, which is neither machine learning nor generative AI

    Correct

    Machine learning learns its rules from examples instead of being explicitly programmed, and generative AI creates new content. Fixed, hand-written conditions that never learn and never generate anything are traditional rule-based software.

Machine learning learns patterns from data without explicit instructions, and generative AI creates new content. A product whose behavior comes entirely from hand-written conditions is traditional automation, whatever the sales deck calls it. Knowing the difference lets a buyer compare like with like and ask the right questions about data, monitoring and cost.

Question 3 · choose 2

A logistics company is inventorying the data it could use for AI. Its data office wants to separate structured sources, which software can query by predefined fields, from unstructured sources, which need different processing before most AI systems can use them. Which sources are unstructured? (Choose TWO.)

  1. ARecorded phone calls between dispatchers and drivers about delays
  2. BThe fleet telemetry table of GPS position, speed and fuel level per truck
  3. CThe CRM table of accounts with ID, region and value
  4. DThe point-of-sale ledger of parcel shipments with date, weight and price
  5. EScanned PDF copies of signed customer contracts and amendments
Show the answer and why
  • ARecorded phone calls between dispatchers and drivers about delays

    Correct

    Audio recordings have no predefined schema of rows and columns. They are unstructured data and must be transcribed or otherwise processed before their content can be analyzed.

  • BThe fleet telemetry table of GPS position, speed and fuel level per truck

    Incorrect

    Telemetry stored as rows with fixed fields is structured data. AWS lists IoT sensor data and fleet telemetry among common structured sources.

  • CThe CRM table of accounts with ID, region and value

    Incorrect

    A CRM account table has defined attributes in rows and columns, which makes it structured data that SQL and BI tools can query directly.

  • DThe point-of-sale ledger of parcel shipments with date, weight and price

    Incorrect

    A ledger with a predefined schema of dates, weights and prices is structured data, readily used for reporting and traditional machine learning.

  • EScanned PDF copies of signed customer contracts and amendments

    Correct

    Free-form documents and images of documents do not follow a schema. Contracts in this form are unstructured data.

Structured data follows a predefined schema, typically tables of rows and columns, and is easy for software to query. Audio, images and free-form documents are unstructured: they hold much of the value generative AI can use, but they need extra processing to become usable and are harder to govern.

Question 4 · choose 1

A telecom company's data office declared its customer database "AI-ready" because it scores 97% on the company's standard data health check, which counts a field as complete and valid when it holds any allowed value. A model trained on the database to predict which business customers will cancel performs poorly. An audit finds that the billing system records every account as "retained" unless an agent changes it, and agents often skip that step, so many canceled accounts appear as retained in the outcome field the model learns from. The analytics director is planning the next attempt and wants to avoid a repeat. What should she require?

  1. ARaise the health check target to 99% for every field in the database before retraining
  2. BSwitch to a more advanced algorithm that is designed to tolerate noisy records
  3. CSet quality checks for this use case, such as outcome accuracy, and pass them before training
  4. DFix the billing default for new accounts, then retrain right away on the existing records
Show the answer and why
  • ARaise the health check target to 99% for every field in the database before retraining

    Incorrect

    A defaulted "retained" value is complete and valid, so a higher generic score would still pass the wrong outcomes. AWS advises tying data quality metrics to how the data will be used rather than to generic data health.

  • BSwitch to a more advanced algorithm that is designed to tolerate noisy records

    Incorrect

    A model learns the relationships in its training data, and noise in that data degrades accuracy. When many canceled accounts are recorded as retained, any algorithm learns from the same wrong outcomes.

  • CSet quality checks for this use case, such as outcome accuracy, and pass them before training

    Correct

    Data quality means different things for different uses. Working back from what the model must achieve, the team sets metrics and minimum thresholds, here the accuracy of the outcome field, and treats them as a gate before the data is used.

  • DFix the billing default for new accounts, then retrain right away on the existing records

    Incorrect

    Fixing the source protects future records, but retraining now still teaches the model from the historical outcomes that are wrong, so its predictions stay unreliable.

A model is only as good as the data it learns from, and a generic health score can hide the defect that matters most. Quality requirements should come from the use case: for a model that predicts cancellations, the recorded outcome must be right. Checking those requirements before training stops a known data problem from becoming a model problem.

Question 5 · choose 1

A software company sells an AI-based underwriting product to banks. Several bank customers now ask in their vendor questionnaires for independent evidence that the company runs a managed, continually improving process for developing and operating AI responsibly. Which standard should the company pursue certification against?

  1. AISO/IEC 23053, a global framework that gives organizations a shared vocabulary for AI systems
  2. BISO/IEC 42001, the certifiable standard for an AI management system
  3. CThe Responsible AI Lens, completed as a self-review
  4. DThe ISO/IEC 42001 certification that AWS holds for its own AI services
Show the answer and why
  • AISO/IEC 23053, a global framework that gives organizations a shared vocabulary for AI systems

    Incorrect

    ISO/IEC 23053 is one of the global frameworks that give teams a unified AI vocabulary, so that people mean the same things by the same terms. A shared vocabulary is not the managed AI process with independent certification that the questionnaires ask for.

  • BISO/IEC 42001, the certifiable standard for an AI management system

    Correct

    ISO/IEC 42001 specifies requirements for establishing, implementing, maintaining and continually improving an Artificial Intelligence Management System (AIMS), and accredited bodies certify organizations against it.

  • CThe Responsible AI Lens, completed as a self-review

    Incorrect

    The Responsible AI Lens is advisory best-practice guidance. AWS states that it should not be used as a compliance or assurance checklist, and a self-review is not independent evidence.

  • DThe ISO/IEC 42001 certification that AWS holds for its own AI services

    Incorrect

    AWS's certificate covers AWS's own AI management process for the services in its scope. Customers are not certified by association; the company needs its own certification.

Customers asking for independent evidence of a managed, improving AI process are describing an AI management system. ISO/IEC 42001 is the international standard for that system, and certification by an accredited body is the evidence they can rely on. Building on AWS can make the company's certification easier, but it does not transfer AWS's certificate.

Question 6 · choose 1

A sales director asks a general-purpose generative AI model about the company's new price list, published last week, and gets confident but outdated figures. She concludes the model is broken. What is the more accurate explanation?

  1. AThe model only knows what was in its training data
  2. BThe model is broken and must be replaced with a newer version
  3. CIt hides current prices on purpose
  4. DThe model retrieves prices from the internet but read an old page
Show the answer and why
  • AThe model only knows what was in its training data

    Correct

    Large language models are trained on large volumes of data and generate answers from what they learned; information outside their training data, such as a new internal price list, is unknown to them unless it is supplied.

  • BThe model is broken and must be replaced with a newer version

    Incorrect

    Nothing is broken; any model lacks information it was never given, so a newer model would have the same gap.

  • CIt hides current prices on purpose

    Incorrect

    Models do not hold back known facts for security; the price list was simply never part of what the model learned.

  • DThe model retrieves prices from the internet but read an old page

    Incorrect

    A plain model call does not search the internet; it generates text from its training.

Knowing where a model's knowledge comes from sets realistic expectations. For current or private information, the system must be given that information at question time, for example through retrieval.

Question 7 · choose 1

A company trains a model to screen job applicants using ten years of past hiring decisions. In that period, managers rarely hired people from certain backgrounds. What should the HR director expect if nothing else is done?

  1. AThe model will remove past bias automatically, because algorithms are neutral
  2. BThe model may repeat the past pattern, because it learns from history
  3. CThe model will ignore history and judge each applicant from scratch
  4. DBias is impossible because the data contains no names
Show the answer and why
  • AThe model will remove past bias automatically, because algorithms are neutral

    Incorrect

    An algorithm learns from the data it is given; it does not know which past patterns were unfair.

  • BThe model may repeat the past pattern, because it learns from history

    Correct

    A model learns the patterns in its training data. AWS guidance asks whether training data might inappropriately represent groups, because such data can carry unwanted bias into the system's decisions.

  • CThe model will ignore history and judge each applicant from scratch

    Incorrect

    A trained model's behavior comes from its history; it has no other basis for its predictions.

  • DBias is impossible because the data contains no names

    Incorrect

    Other fields can act as proxies for demographic attributes, so removing names does not remove bias.

Training on historical data means learning from historical decisions, good and bad. Leaders should expect past patterns to reappear unless the data and the system are examined and adjusted.

Question 8 · choose 1

A logistics company receives thousands of scanned delivery notes a week. Its reports run on a SQL database, and operations wants the shipment number, date, recipient and signature status of every note loaded into that database as exact field values. A vendor proposes a generative AI tool that writes a short summary of each note. Which AI capability fits the requirement?

  1. ADocument extraction that turns each note into structured fields
  2. BA generative model that writes a summary of each note
  3. CA forecasting model trained on past delivery volumes
  4. DSQL queries run directly against the scanned images
Show the answer and why
  • ADocument extraction that turns each note into structured fields

    Correct

    Intelligent document processing extracts data from documents and enters it in the format the target system needs, turning unstructured scans into the structured fields a SQL database can store and query.

  • BA generative model that writes a summary of each note

    Incorrect

    Generative AI creates new content such as text; a free-text summary is still unstructured and cannot be loaded as exact field values.

  • CA forecasting model trained on past delivery volumes

    Incorrect

    Forecasting predicts future trends from historical data; it does not read the fields on individual notes.

  • DSQL queries run directly against the scanned images

    Incorrect

    SQL works on structured data with a predefined schema; scanned images are unstructured, so their fields must be extracted before they can be queried.

Whether data is structured or unstructured decides which tools can use it. Scanned notes are unstructured; document extraction converts them into the structured fields that existing reports rely on, which a generated summary or a forecast cannot do.

Question 9 · choose 1

An online bookstore wants its home page to show each visitor the titles that visitor is likely to buy, without the visitor having to ask. It has years of purchase, browsing and search history for millions of visitors, and the marketing team admits it could never write rules by hand for every taste. Which approach fits?

  1. AA recommendation engine trained on visitors' past behavior
  2. BA bestseller list chosen by sales rules, shown to every visitor
  3. CA sales-trend dashboard that marketing reviews every month
  4. DAnomaly detection on visitors' purchase histories
Show the answer and why
  • AA recommendation engine trained on visitors' past behavior

    Correct

    Machine learning finds patterns in large amounts of historical data without explicit instructions, and retailers use it to recommend products based on previous purchases, browsing history and search patterns.

  • BA bestseller list chosen by sales rules, shown to every visitor

    Incorrect

    Hand-set rules give every visitor the same list; the bookstore needs predictions for each visitor, learned from history rather than written as rules.

  • CA sales-trend dashboard that marketing reviews every month

    Incorrect

    A dashboard describes past results for people to read; it makes no prediction for each visitor, while machine learning uses historical data to predict for every visitor automatically.

  • DAnomaly detection on visitors' purchase histories

    Incorrect

    Anomaly detection finds unusual patterns in data; it does not predict which titles each visitor is likely to buy.

A model trained on historical behavior can make a prediction for every visitor at the moment the page loads, which no set of hand-written rules can match at this scale. Recognizing this as a recommendation problem sets expectations for the data it needs.

Question 10 · choose 1

A hotel chain's BI dashboards report bookings, prices and star ratings well from its SQL database. Leadership now wants to know what guests think about specific parts of their stays, such as cleanliness, staff and noise, from 40,000 free-text reviews already collected, including comments like "not exactly spotless". Why can't the existing SQL reports answer this, and what is needed?

  1. AThe reports only need SQL counts of words such as "dirty" or "rude"
  2. BThe average star rating already stored answers the question
  3. CSentiment is in unstructured text, which needs language processing
  4. DStoring the reviews as a text column lets the dashboards chart them
Show the answer and why
  • AThe reports only need SQL counts of words such as "dirty" or "rude"

    Incorrect

    Counting listed words is a basic rule-based approach that misses context such as negation and sarcasm, which AWS notes machines struggle with; meaning in free text needs language processing.

  • BThe average star rating already stored answers the question

    Incorrect

    A star rating is structured data that gives one overall score; what guests say about cleanliness, staff or noise is in the review text, which structured reports cannot capture.

  • CSentiment is in unstructured text, which needs language processing

    Correct

    Structured databases track quantities such as prices but cannot capture sentiment, which comes in unstructured text; analyzing it requires techniques such as natural language processing and text mining.

  • DStoring the reviews as a text column lets the dashboards chart them

    Incorrect

    Storing text in a table does not interpret it; unstructured data lacks the order that conventional data-mining techniques need, so it must be analyzed with language processing first.

Much of what customers think is captured in unstructured text. Natural language processing turns reviews and comments into insights about specific topics that structured reports and simple word counts cannot provide.

Practise domain 1 →Practise all domains →