Skip to content
BytePatterns

AIP-C01 · Domain 1: Foundation Model Integration, Data Management, and Compliance · 31% of the exam

Task 1.3: Implement data validation and processing pipelines for FM consumption.

Getting data ready for a model: quality checks with AWS Glue Data Quality, processing text, images, audio and tables, formatting requests for each model and API, and cleaning input to improve answers.

Study it

  • Data quality and processing for model input: Glue Data Quality, Transcribe, Data Automation and request formats

    Partly covered by: Tokenization

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company loads product records from several suppliers into Amazon S3 with an AWS Glue ETL job, and the records are then ingested into a knowledge base. Some supplier files arrive with empty descriptions or duplicate SKUs, which leads to poor answers. The data engineers want declarative rules instead of custom code, a quality score for each run, and no data written to the target when the rules fail, with an alert when that happens. Which solution meets these requirements?

  1. AWrite a Lambda function that is triggered by each new S3 object and checks every record before the knowledge base ingests it
  2. BAdd a DQDL ruleset with AWS Glue Data Quality to the ETL job, fail it without loading the target, and publish the results
  3. CCall Amazon Comprehend on each description during the job and drop records whose entity list is empty
  4. DRun Amazon Macie discovery jobs on the landing bucket and stop the ingestion pipeline whenever Macie reports new findings
Show the answer and why
  • AWrite a Lambda function that is triggered by each new S3 object and checks every record before the knowledge base ingests it

    Incorrect

    A Lambda function can validate files, but it is custom code without declarative rules or a quality score, which the engineers want to avoid.

  • BAdd a DQDL ruleset with AWS Glue Data Quality to the ETL job, fail it without loading the target, and publish the results

    Correct

    Glue Data Quality evaluates DQDL rules inside the ETL job, reports a data quality score, can fail the job without loading data, and publishes metrics to CloudWatch and EventBridge for alerting.

  • CCall Amazon Comprehend on each description during the job and drop records whose entity list is empty

    Incorrect

    Comprehend extracts entities from text, which is useful for enrichment, but it is not a rule engine for empty fields or duplicate keys and adds per-record cost.

  • DRun Amazon Macie discovery jobs on the landing bucket and stop the ingestion pipeline whenever Macie reports new findings

    Incorrect

    Macie discovers sensitive data such as PII in S3. It does not check completeness or uniqueness of business fields.

Validation belongs before the data reaches the vector store. Glue Data Quality gives declarative DQDL rules, rule recommendations, a score per run, configurable job failure behavior and CloudWatch and EventBridge integration, so bad supplier files never reach the knowledge base.

Question 2 · choose 1

A contact center wants a live assistant that suggests answers to agents while a call is in progress. The assistant needs the customer's words as text within seconds, must know which party said each sentence, and must not pass card numbers or addresses to the foundation model. Which processing step should the developer put in front of the model?

  1. AAn Amazon Transcribe batch job that runs on the recording every five minutes during the call
  2. BAn Amazon Bedrock Data Automation project that processes each call recording after the call ends
  3. CAmazon Transcribe streaming transcription with speaker partitioning and PII redaction enabled
  4. DAmazon Comprehend DetectPiiEntities on the raw audio stream before it is sent to the model
Show the answer and why
  • AAn Amazon Transcribe batch job that runs on the recording every five minutes during the call

    Incorrect

    Batch transcription works on stored files, so suggestions would trail the conversation by minutes instead of seconds.

  • BAn Amazon Bedrock Data Automation project that processes each call recording after the call ends

    Incorrect

    Data Automation is a strong choice for extracting insights from stored audio, but processing after the call ends cannot help the agent during the conversation.

  • CAmazon Transcribe streaming transcription with speaker partitioning and PII redaction enabled

    Correct

    Streaming transcription returns text while the call is in progress, speaker partitioning labels each speaker, and streaming PII redaction replaces identified PII with a [PII] tag before the text reaches the model.

  • DAmazon Comprehend DetectPiiEntities on the raw audio stream before it is sent to the model

    Incorrect

    Comprehend analyzes text, not audio. The speech must be transcribed first, and speaker labels would still be missing.

Real-time requirements rule out post-call and batch processing. Transcribe streaming covers all three needs in one step: low-latency text, speaker labels and redaction of PII before the transcript is sent to the model.

Question 3 · choose 1

A developer prepares 40,000 support transcripts for summarization with Amazon Bedrock batch inference. The application code that will later call the same model interactively already builds Converse API request bodies, and the team wants to reuse that code to build the batch records. Each summary must be matched back to its transcript afterward. How should the developer prepare the input?

  1. AWrite JSONL lines with a recordId and a Converse-format modelInput, and set the job's model invocation type to Converse
  2. BWrite JSONL lines with only the provider-specific InvokeModel body and rely on line numbers to match the output
  3. CWrite one JSON array that holds all Converse request bodies in exactly the same order as the transcripts in the database
  4. DWrite CSV rows with a transcript column and a prompt column, and let batch inference build the request bodies
Show the answer and why
  • AWrite JSONL lines with a recordId and a Converse-format modelInput, and set the job's model invocation type to Converse

    Correct

    Each line of a batch input file has a recordId and a modelInput. With the Converse invocation type, modelInput uses the Converse request body, and the recordId lets results be matched even though output order is not guaranteed.

  • BWrite JSONL lines with only the provider-specific InvokeModel body and rely on line numbers to match the output

    Incorrect

    The InvokeModel format is allowed, but it would not reuse the Converse code, and line numbers are unreliable because the output order can differ from the input order.

  • CWrite one JSON array that holds all Converse request bodies in exactly the same order as the transcripts in the database

    Incorrect

    Batch inference expects JSONL with one record per line, and output order is not guaranteed to match input order, so position cannot be used for matching.

  • DWrite CSV rows with a transcript column and a prompt column, and let batch inference build the request bodies

    Incorrect

    Batch inference does not accept CSV input. Records must be JSON objects in JSONL files with a modelInput field.

Batch records are JSONL objects with recordId and modelInput, and the job's invocation type decides whether modelInput follows the InvokeModel body of the model or the Converse request format. Always carry your own recordId, because results are not returned in input order.

Question 4 · choose 1

Field service reports reach a summarization model after optical character recognition. The reports mix date formats such as 03/04/26 and 4 March 2026, and mix units such as psi and bar, so the model sometimes reports wrong intervals and pressures. The conversion rules are fixed and known, the output must be identical every time the same report is processed so that auditors can reproduce it, and the step should add as little cost as possible. What should the developer add before the model call?

  1. AA Lambda function that converts dates to ISO 8601 and units to one system with fixed rules before the prompt is built
  2. BAn Amazon Comprehend DetectEntities call that tags each date and quantity in the report
  3. CA higher temperature setting so that the model considers more ways to read each date and unit
  4. DA first Amazon Bedrock call that rewrites each report into a consistent format before the separate summarization call
Show the answer and why
  • AA Lambda function that converts dates to ISO 8601 and units to one system with fixed rules before the prompt is built

    Correct

    Known, fixed conversion rules are best applied deterministically in code. A Lambda function gives the same output for the same input, can be audited, and costs far less than a model call.

  • BAn Amazon Comprehend DetectEntities call that tags each date and quantity in the report

    Incorrect

    Entity detection finds and labels dates and quantities, but it does not convert them to a single format or unit.

  • CA higher temperature setting so that the model considers more ways to read each date and unit

    Incorrect

    A higher temperature makes the output more random, which works against both accuracy and reproducibility.

  • DA first Amazon Bedrock call that rewrites each report into a consistent format before the separate summarization call

    Incorrect

    A model can reformat text and is useful when the rules are fuzzy, but its output is not guaranteed to be identical across runs and it adds the cost of a second inference.

Improving input quality does not always need a model. When the rules are fixed and results must be reproducible and cheap, deterministic preprocessing in Lambda is the right tool; a model-based rewrite fits fuzzy cleanup where some variation is acceptable.

Question 5 · choose 1

Scanned loan application forms contain labeled fields such as "Applicant name" and "Annual income". Before a model drafts a decision memo, the pipeline must extract these field labels and their values as structured pairs. Which Amazon Textract feature fits?

  1. AAnalyzeExpense, built for invoices and receipts
  2. BDetectDocumentText, which returns lines and words
  3. CAnalyzeDocument with only the TABLES feature type
  4. DAnalyzeDocument with the FORMS feature type
Show the answer and why
  • AAnalyzeExpense, built for invoices and receipts

    Incorrect

    AnalyzeExpense targets invoices and receipts, not loan applications.

  • BDetectDocumentText, which returns lines and words

    Incorrect

    Text detection returns lines and words without linking labels to their values.

  • CAnalyzeDocument with only the TABLES feature type

    Incorrect

    Tables extracts table structure, not labeled form fields.

  • DAnalyzeDocument with the FORMS feature type

    Correct

    Textract extracts form data as key-value pairs when you request the FORMS feature type.

Choose the Textract feature by document structure: FORMS for labeled fields, TABLES for grids, and expense analysis for invoices and receipts.

Question 6 · choose 1

A support center's third-party phone system writes yesterday's recorded calls to Amazon S3 each night as single-channel audio files, about 20,000 calls. Each morning a model summarizes every call, and compliance also requires a verbatim text record of each call that shows which speaker said each part. Nothing needs to be processed during the call, and the team wants managed services without hosting any model. Which approach fits?

  1. ABuild an Amazon Lex V2 bot that processes each recording
  2. BRun Amazon Transcribe batch jobs with speaker diarization enabled
  3. CRun Amazon Transcribe Medical batch jobs on the recordings
  4. DSend each recording to Amazon Nova 2 Sonic for transcription
Show the answer and why
  • ABuild an Amazon Lex V2 bot that processes each recording

    Incorrect

    Lex builds conversational voice and text interfaces for live interactions. It is not a bulk transcription service for archived recordings.

  • BRun Amazon Transcribe batch jobs with speaker diarization enabled

    Correct

    Transcribe processes media files stored in Amazon S3 as batch jobs, and speaker diarization labels which speaker said each part of the transcript.

  • CRun Amazon Transcribe Medical batch jobs on the recordings

    Incorrect

    Transcribe Medical is designed for medical speech such as physician-dictated notes. Customer support calls are not its domain.

  • DSend each recording to Amazon Nova 2 Sonic for transcription

    Incorrect

    Nova 2 Sonic is a speech-to-speech model for real-time spoken conversations. It does not produce batch transcripts of archived files.

Stored recordings call for batch transcription, and the record's requirements decide the features: diarization when the transcript must say who spoke.

Question 7 · choose 1

Data analysts must clean and transform tabular customer feedback before it is used to fine-tune a model. They prefer a visual or natural-language interface over writing code, and they work in the current SageMaker Studio experience. Which tool should they use?

  1. AData Wrangler in SageMaker Studio Classic
  2. BAmazon Comprehend topic modeling
  3. CData preparation in SageMaker Canvas
  4. DA SageMaker AI real-time endpoint
Show the answer and why
  • AData Wrangler in SageMaker Studio Classic

    Incorrect

    Studio Classic is not available for new onboarding, and Data Wrangler has moved into Canvas.

  • BAmazon Comprehend topic modeling

    Incorrect

    Topic modeling is not available to new customers and does not clean tabular data.

  • CData preparation in SageMaker Canvas

    Correct

    Data Wrangler capabilities are now in SageMaker Canvas, which supports visual and natural language data preparation.

  • DA SageMaker AI real-time endpoint

    Incorrect

    Endpoints serve model inference, not data preparation.

For low-code data preparation, use the current Canvas experience instead of the Studio Classic tools that are closed to new onboarding.

Question 8 · choose 1

Each time a partner uploads a document to an S3 bucket, a pre-processing step must convert it to clean text before indexing. Uploads are irregular, and processing takes under a minute per file. What is the simplest event-driven design?

  1. AS3 event notifications that invoke a Lambda function
  2. BA nightly AWS Glue crawler on the bucket
  3. CAn EC2 instance that lists the bucket every minute
  4. DA SageMaker AI endpoint that the partner calls directly
Show the answer and why
  • AS3 event notifications that invoke a Lambda function

    Correct

    Amazon S3 can invoke a Lambda function when objects are created, which suits short, per-file processing.

  • BA nightly AWS Glue crawler on the bucket

    Incorrect

    A crawler catalogs data. It does not convert documents to text.

  • CAn EC2 instance that lists the bucket every minute

    Incorrect

    Polling on a server adds cost and management for irregular uploads.

  • DA SageMaker AI endpoint that the partner calls directly

    Incorrect

    Partners would need to change their process, and an endpoint is not an event trigger.

Event-driven processing with S3 notifications and Lambda fits short, irregular jobs without servers to manage.

Question 9 · choose 1

Chat transcripts from 40 web and mobile apps flow through an Amazon Data Firehose stream into the S3 prefix that a knowledge base ingests. Some apps send JSON records with a missing conversation ID or truncated text, which pollutes answers. Each record must be checked against the company's rules before it lands in that prefix, rejected records must be kept with the reason for later review, new records must arrive within minutes, and the team wants no servers or extra batch jobs. Which solution meets these requirements?

  1. AS3 event notifications that invoke a Lambda function to check each new object and delete bad records
  2. BA Firehose data transformation Lambda function that returns rejected records as ProcessingFailed
  3. CFirehose dynamic partitioning on the app ID so that each app writes to its own prefix
  4. DFirehose record format conversion to Apache Parquet with a schema from an AWS Glue table
Show the answer and why
  • AS3 event notifications that invoke a Lambda function to check each new object and delete bad records

    Incorrect

    S3 can invoke a Lambda function when objects are created, which suits per-file processing. Here the bad records would already sit in the ingestion prefix, and deleting them loses them for review.

  • BA Firehose data transformation Lambda function that returns rejected records as ProcessingFailed

    Correct

    Firehose treats records returned as ProcessingFailed as unsuccessfully processed and delivers them to the processing-failed folder with error details, while valid records continue to the destination. The function can log the rule each rejected record broke to CloudWatch Logs.

  • CFirehose dynamic partitioning on the app ID so that each app writes to its own prefix

    Incorrect

    Dynamic partitioning groups streaming data by keys within the records. It organizes records but does not check them.

  • DFirehose record format conversion to Apache Parquet with a schema from an AWS Glue table

    Incorrect

    Format conversion turns JSON into columnar files for analytics. It is not a rules check, and Parquet is not among the file formats a knowledge base ingests.

Validate streaming data in the stream itself: a transformation function can pass, drop or fail each record, and Firehose keeps failed records with their error details.

Practise domain 1 →Practise all domains →