Skip to content
BytePatterns

AIB-C01 · Domain 4: Business Readiness, Leadership, and AI Transformation · 24% of the exam

Task 4.2: Establish data and infrastructure foundations for AI.

Data readiness, quality, access and silos, the data strategy, ownership and sharing rules AI depends on, and the technology foundations an initiative needs before it can grow.

Study it

  • Data readiness, silos, ownership and sharing

    Lesson: Data Readiness for AI

  • Technology foundations for AI

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A hospital group wants AI to predict patient discharge dates. The data it needs sits in seven separate systems run by different departments, each with its own access process, and analysts spend weeks assembling extracts for each request. What should the chief data officer prioritize?

  1. AA separate copy of the data for each AI project team, refreshed by hand
  2. BA more advanced prediction model that tolerates missing inputs
  3. CRestricting AI projects to one department's data
  4. DUnified, governed access to data across the seven systems
Show the answer and why
  • AA separate copy of the data for each AI project team, refreshed by hand

    Incorrect

    Ad hoc copies multiply the silos, add security exposure and leave each project to repeat the same assembly work.

  • BA more advanced prediction model that tolerates missing inputs

    Incorrect

    A model that tolerates gaps does not give the organization access to the data it needs; the bottleneck is fragmentation.

  • CRestricting AI projects to one department's data

    Incorrect

    Limiting projects to single silos avoids the problem rather than solving it, and discharge prediction needs data from several departments.

  • DUnified, governed access to data across the seven systems

    Correct

    AWS's generative AI guidance calls for breaking down data silos with a unified data system and unified access to data, independent of where it resides, while maintaining security and compliance.

Fragmented data environments and unclear ownership slow AI down. A unified, well-governed way to access data across systems lets many AI initiatives draw on the same foundation instead of each rebuilding its own extracts.

Question 2 · choose 1

A manufacturer's AI team keeps finding errors in the product master data its models depend on. Each time, engineering, sales and finance each say another department is responsible for the data. What should the company establish?

  1. AA rule that the AI team fixes any data errors it finds
  2. BA data lake that copies all of the product data into one central location
  3. CClear data ownership, with roles and standards for each data domain
  4. DAn annual data quality audit performed by an external firm
Show the answer and why
  • AA rule that the AI team fixes any data errors it finds

    Incorrect

    The AI team consumes the data but does not control how it is created. Fixing errors downstream leaves their cause in place.

  • BA data lake that copies all of the product data into one central location

    Incorrect

    Centralizing storage can improve access, but it does not decide who is accountable for the data's correctness.

  • CClear data ownership, with roles and standards for each data domain

    Correct

    Data governance determines roles, responsibilities and standards for data usage, including who can take what action on which data. Clear ownership gives each data domain someone accountable for its quality.

  • DAn annual data quality audit performed by an external firm

    Incorrect

    An audit can measure quality, but without owners there is still no one responsible for correcting and preventing errors.

AI depends on data that someone is accountable for. A governance model that assigns owners and defines roles, responsibilities and standards turns recurring disputes into clear accountability for data quality.

Question 3 · choose 1

An airline will run its first generative AI proofs of concept next quarter. The CIO asks what technology foundation must be in place before the teams start. What should the platform lead prepare?

  1. AA complete production platform with automated pipelines for every future use case
  2. BA self-managed GPU cluster to train the airline's own foundation model
  3. CNo foundation, letting each team use its own personal accounts and tools
  4. DA secure landing zone, developer permissions and a managed experimentation space
Show the answer and why
  • AA complete production platform with automated pipelines for every future use case

    Incorrect

    Production-grade pipelines come when solutions move to production. For proofs of concept AWS advises a minimal platform built on managed services.

  • BA self-managed GPU cluster to train the airline's own foundation model

    Incorrect

    Pretraining a foundation model is the costliest option and is not needed to run proofs of concept on existing models.

  • CNo foundation, letting each team use its own personal accounts and tools

    Incorrect

    Unmanaged personal accounts bypass security and governance and create shadow AI from the start.

  • DA secure landing zone, developer permissions and a managed experimentation space

    Correct

    AWS's maturity guidance lists, for the Experiment level, setting up foundation infrastructure such as a landing zone and the permissions developers need, and an environment for experimentation such as an Amazon Bedrock playground.

The infrastructure should match the stage. Proofs of concept need a secure, governed place to experiment quickly with managed services; production platforms and custom training infrastructure come later, when a use case has proven its value.

Question 4 · choose 1

A bank's analysts will start generative AI experiments next quarter, and each team plans to gather and label its own data in its own way. The steering group must be able to compare the experiments' results, and the experiments have to start on time. What data foundation should the data office set up first?

  1. AEach team's own data handling now, standardized after the experiments
  2. BPublic datasets in place of the bank's own data for the experiments
  3. CA structured approach to collecting, cleaning and labeling data
  4. DCleaning and labeling of every dataset in the bank before any experiment
Show the answer and why
  • AEach team's own data handling now, standardized after the experiments

    Incorrect

    Results built on differently collected and labeled data cannot be compared, and AWS places a common approach to data before the experiments, at the Envision stage.

  • BPublic datasets in place of the bank's own data for the experiments

    Incorrect

    Public data does not reflect the bank's own customers and processes, so results would say little about its use cases, and it leaves the teams without a common way to prepare the bank's data.

  • CA structured approach to collecting, cleaning and labeling data

    Correct

    AWS's data strategy guidance for the Envision stage is to define a structured approach to collecting, cleaning and labeling data for generative AI experiments.

  • DCleaning and labeling of every dataset in the bank before any experiment

    Incorrect

    Preparing all the bank's data first would delay the experiments; the guidance is a defined approach for the data the experiments need.

Data foundations start with consistent practices. A shared approach to collecting, cleaning and labeling data makes experiments comparable and their data reusable, without waiting for every dataset to be perfect.

Question 5 · choose 1

A retailer's AI teams each check data quality by hand before every project, and errors still slip through as data volumes grow. What foundation should the chief data officer invest in?

  1. AMore analysts to check data by hand before each project
  2. BAutomated validation and cleansing, monitored continuously
  3. CAccepting data errors, because models will learn around them
  4. DA one-time cleanup of all current data with no ongoing checks
Show the answer and why
  • AMore analysts to check data by hand before each project

    Incorrect

    Manual checks do not scale with growing data and still miss errors.

  • BAutomated validation and cleansing, monitored continuously

    Correct

    AWS's generative AI data guidance calls for automated data validation, cleansing processes and continuous monitoring to maintain data integrity at scale, under a data governance framework.

  • CAccepting data errors, because models will learn around them

    Incorrect

    Poor data leads to poor AI outcomes; models do not reliably correct bad inputs.

  • DA one-time cleanup of all current data with no ongoing checks

    Incorrect

    Quality degrades again without continuous validation.

Data quality is the bedrock of effective AI. Building validation and monitoring into pipelines keeps quality consistent as data and AI use grow.

Question 6 · choose 1

A pharmaceutical company's Experiment-stage team will test retrieval augmented generation on research reports and lab data. Which technology foundation does AWS's data strategy guidance list for this stage?

  1. AFederated data platforms serving every business unit at enterprise scale
  2. BA custom-trained foundation model on all company data
  3. CEarly data pipelines for both data types, plus vector databases for RAG
  4. DNo technology work, because experiments run on spreadsheets
Show the answer and why
  • AFederated data platforms serving every business unit at enterprise scale

    Incorrect

    Federated enterprise-scale data platforms are a Scale-stage investment.

  • BA custom-trained foundation model on all company data

    Incorrect

    Pretraining a model is unnecessary for RAG experiments on existing models.

  • CEarly data pipelines for both data types, plus vector databases for RAG

    Correct

    For the Experiment stage, AWS lists implementing early-stage data integration, including structured and unstructured data pipelines, and setting up vector databases for RAG experiments.

  • DNo technology work, because experiments run on spreadsheets

    Incorrect

    RAG experiments need data pipelines and a vector store, not spreadsheets.

Infrastructure should match the stage. RAG experiments need pipelines for the relevant data and a vector store; enterprise-scale platforms come later.

Question 7 · choose 1

A media company's on-premises storage was sized for last year's analytics. AI projects now bring unpredictable spikes in data volume and processing. Which infrastructure principle should guide the data foundation?

  1. AElastic infrastructure that scales with AI data workloads
  2. BFixed capacity sized to this year's average data load
  3. CBuying hardware for the largest imaginable future load now
  4. DLimiting AI projects to what current storage can hold
Show the answer and why
  • AElastic infrastructure that scales with AI data workloads

    Correct

    AWS's guidance lists scalable, elastic data infrastructure that can handle varying AI workloads and data volumes among the key success factors for generative AI data architecture.

  • BFixed capacity sized to this year's average data load

    Incorrect

    Fixed capacity sized to an average cannot absorb unpredictable spikes.

  • CBuying hardware for the largest imaginable future load now

    Incorrect

    Over-buying ties up capital in capacity that may sit idle.

  • DLimiting AI projects to what current storage can hold

    Incorrect

    Constraining projects to old capacity limits the value AI can deliver.

AI workloads are variable and growing. Elastic infrastructure lets capacity follow demand instead of forcing a choice between shortages and idle hardware.

Question 8 · choose 1

A retailer wants AI to personalize offers. Its online store, loyalty program and in-store payment system each keep customer records with their own IDs, spellings and addresses, so one shopper can appear as three different people, and nobody can see a shopper's full purchase history. The data team proposes copying all three systems into one data lake as they are. What should the chief data officer add before the data is used for personalization?

  1. AA separate personalization model trained on each system's data
  2. BUse of the loyalty data alone, as the largest of the three sources
  3. CReliance on the personalization model to learn which profiles match
  4. DMatching and linking the records into one view per customer
Show the answer and why
  • AA separate personalization model trained on each system's data

    Incorrect

    Each model would see only part of a shopper's history, which keeps the silo problem in the models instead of solving it in the data.

  • BUse of the loyalty data alone, as the largest of the three sources

    Incorrect

    Dropping the online and in-store records leaves the incomplete purchase history the retailer is trying to fix; AWS guidance is to break down data silos, not to work around them.

  • CReliance on the personalization model to learn which profiles match

    Incorrect

    AWS treats accurate, complete and consistent data across sources as the bedrock of effective AI; duplicate profiles fed to a model become poor inputs, not something it reliably corrects.

  • DMatching and linking the records into one view per customer

    Correct

    AWS Entity Resolution matches, links and enhances related records stored across applications and data stores to build a unified view of customer interactions; copying the systems into one place without linking keeps the duplicates.

Breaking down data silos is more than moving data into one store. When each system identifies customers differently, records must be matched and linked before AI can see one customer's full history.

Practise domain 4 →Practise all domains →