Skip to content
BytePatterns

AIF-C01 · Domain 4: Guidelines for Responsible AI · 14% of the exam

Task 4.1: Explain the development of AI systems that are responsible.

The dimensions of responsible AI, the legal risks of generative AI, how bias and variance show up in data and models, and the AWS tools that detect and limit them.

Study it

  • The dimensions of responsible AI and Amazon Bedrock Guardrails

    Partly covered by: Guardrails

  • Bias, variance and datasets: subgroup analysis, bias metrics and human review

    Lesson coming

  • Legal risks of generative AI and sustainable model choice

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A bank is launching a generative AI assistant on Amazon Bedrock. The assistant must refuse to discuss investment advice, and it must mask customers' phone numbers and account numbers if they appear in prompts or responses. Which AWS capability provides both controls?

  1. AAmazon Bedrock Guardrails
  2. BAmazon Macie
  3. CAmazon Inspector
  4. DAmazon SageMaker Model Registry
Show the answer and why
  • AAmazon Bedrock Guardrails

    Correct

    Guardrails can block denied topics such as investment advice in prompts and responses, and their sensitive information filters can block or mask PII.

  • BAmazon Macie

    Incorrect

    Macie discovers sensitive data such as PII in Amazon S3. It does not sit between users and a model to filter conversations.

  • CAmazon Inspector

    Incorrect

    Inspector scans workloads for software vulnerabilities and unintended network exposure. It does not filter model prompts or responses.

  • DAmazon SageMaker Model Registry

    Incorrect

    Model Registry catalogs model versions and their approval status. It does not inspect conversations.

Denied topics and PII masking are both Guardrails safeguards, applied to user inputs and model responses alike.

Question 2 · choose 1

A team tests whether its model still gives correct answers when users send misspelled, unexpected or deliberately adversarial inputs. Which dimension of responsible AI, as AWS defines it, is the team assessing?

  1. AFairness
  2. BVeracity and robustness
  3. CTransparency
  4. DControllability
Show the answer and why
  • AFairness

    Incorrect

    Fairness is about considering the impacts on different groups of stakeholders, not about resisting unusual inputs.

  • BVeracity and robustness

    Correct

    AWS defines veracity and robustness as achieving correct system outputs, even with unexpected or adversarial inputs.

  • CTransparency

    Incorrect

    Transparency is about enabling stakeholders to make informed choices about their engagement with an AI system.

  • DControllability

    Incorrect

    Controllability means having mechanisms to monitor and steer AI system behavior.

AWS lists eight dimensions: fairness, explainability, privacy and security, safety, controllability, veracity and robustness, governance and transparency.

Question 3 · choose 1

A model scores 99% accuracy on its training data but only 71% on new data it has never seen. Which description of the model is correct?

  1. AIt is overfit, a high-variance model that does not generalize
  2. BIt is underfit, a high-bias model that is too simple for the data
  3. CIt is well fitted, because training accuracy is what matters
  4. DIt is hallucinating, because only generative models make this error
Show the answer and why
  • AIt is overfit, a high-variance model that does not generalize

    Correct

    Overfit models perform well on the training set but not on the test set because they fit the training data too closely; AWS describes this as high variance.

  • BIt is underfit, a high-bias model that is too simple for the data

    Incorrect

    Underfit models produce inaccurate results on both the training data and the test set. This model is very accurate on its training data.

  • CIt is well fitted, because training accuracy is what matters

    Incorrect

    A well-fitted model captures the dominant trend for both seen and unseen data. Accuracy on new data is what shows whether a model generalizes.

  • DIt is hallucinating, because only generative models make this error

    Incorrect

    Hallucination describes incorrect output from a generative model. The gap between training and test accuracy is the classic sign of overfitting in any ML model.

Great on training data and poor on new data means high variance (overfitting); poor on both means high bias (underfitting).

Question 4 · choose 1

A lender's credit model is 92% accurate across all applicants. A reviewer worries that it may serve some groups of applicants worse than others. What should the team do first to find out?

  1. ARetrain the model on more data so that the overall accuracy rises
  2. BMeasure the model's performance separately for each applicant group
  3. CRaise the approval threshold for every applicant by the same amount
  4. DRemove the reviewer's concern by publishing the overall accuracy
Show the answer and why
  • ARetrain the model on more data so that the overall accuracy rises

    Incorrect

    A higher overall number can still hide a group the model serves badly; the question is about differences between groups.

  • BMeasure the model's performance separately for each applicant group

    Correct

    Bias includes disparities in a model's performance across different groups, and cohort-level analysis examines how the model behaves for a particular group.

  • CRaise the approval threshold for every applicant by the same amount

    Incorrect

    Changing every decision equally does not reveal whether any group is treated differently, and it is not a measurement.

  • DRemove the reviewer's concern by publishing the overall accuracy

    Incorrect

    Publishing an aggregate figure does not test for bias; disparities across groups have to be measured.

Fairness problems hide inside averages. Subgroup analysis, comparing metrics group by group, is how they are found.

Question 5 · choose 2

A retailer plans to publish product descriptions and images created with a generative AI model. Which legal or reputational risks should its review cover? (Choose TWO.)

  1. AClaims that generated content infringes someone's copyright
  2. BFalse statements about products that mislead customers and erode trust
  3. CThe model's context window being too small for long prompts
  4. DHigher latency during peak shopping hours
  5. EPaying per token instead of per hour for inference
Show the answer and why
  • AClaims that generated content infringes someone's copyright

    Correct

    Copyright claims arising from generative output are a recognized risk; AWS even offers an IP indemnity for copyright claims on output from certain of its generative AI services.

  • BFalse statements about products that mislead customers and erode trust

    Correct

    Hallucinations are incorrect or misleading outputs, and they reduce user trust; published as product claims, they create risk for the business.

  • CThe model's context window being too small for long prompts

    Incorrect

    A context window limit is a technical design constraint, not a legal or reputational risk.

  • DHigher latency during peak shopping hours

    Incorrect

    Latency is a performance concern for model selection, not a legal risk.

  • EPaying per token instead of per hour for inference

    Incorrect

    Token-based pricing is a cost model. It does not create legal exposure.

Generative output carries intellectual property, accuracy and bias risks that a human review and clear policies must address before publication.

Question 6 · choose 1

A model performs poorly on its training data and equally poorly on new test data. The team suspects the model is too simple for the relationships in the data. Which description fits?

  1. AThe model is overfit and has high variance
  2. BThe model is well fitted and ready to deploy
  3. CThe model is underfit and has high bias
  4. DThe model has data drift from production traffic
Show the answer and why
  • AThe model is overfit and has high variance

    Incorrect

    Overfit models perform well on the training set but not on the test set; this model is poor on both.

  • BThe model is well fitted and ready to deploy

    Incorrect

    A well-fitted model captures the dominant trend for seen and unseen data, which this model does not.

  • CThe model is underfit and has high bias

    Correct

    Underfit models produce inaccurate results on both training and test data, for example because the model is too simple or was not trained long enough; AWS describes this as high bias.

  • DThe model has data drift from production traffic

    Incorrect

    The model already fails on its own training data, so the problem is the fit, not a change in production data.

Poor on both sets means high bias (underfitting); good on training but poor on new data means high variance (overfitting).

Question 7 · choose 1

A gaming community app uses a generative AI model. It must block hateful, insulting and violent content in both user messages and model replies, and the team wants to tune how strict each category is. Which Amazon Bedrock Guardrails safeguard fits?

  1. AContent filters
  2. BSensitive information filters
  3. CContextual grounding check
  4. DDenied topics
Show the answer and why
  • AContent filters

    Correct

    Content filters detect and filter harmful content in prompts and responses across categories such as hate, insults, sexual, violence and misconduct, with a configurable strength per category.

  • BSensitive information filters

    Incorrect

    Sensitive information filters detect and block or mask PII. They do not rate hate or violence.

  • CContextual grounding check

    Incorrect

    The contextual grounding check detects hallucinations against a reference source, not harmful content.

  • DDenied topics

    Incorrect

    Denied topics block specific subjects that are undesirable for the application, such as investment advice. Category-based harm filtering with adjustable strength is the content filter.

Guardrails safeguards each target a different risk: content filters for harm categories, denied topics for subjects, sensitive information filters for PII, grounding checks for hallucinations.

Question 8 · choose 2

A team training a hiring-support model finds that almost all of its training data comes from one country and one age group. Which changes to the dataset address the resulting bias risk? (Choose TWO.)

  1. ADuplicate the existing records to make the dataset larger
  2. BKeep only the most recent month of data to stay current
  3. CAdd data so that all relevant applicant groups are represented
  4. DTrain longer on the same data until accuracy stops improving
  5. EBalance the dataset so that no group dominates training
Show the answer and why
  • ADuplicate the existing records to make the dataset larger

    Incorrect

    Copying the same records adds volume but not diversity; the groups that are missing stay missing.

  • BKeep only the most recent month of data to stay current

    Incorrect

    Cutting the dataset to one month shrinks it further and does nothing to fix which groups it represents.

  • CAdd data so that all relevant applicant groups are represented

    Correct

    Pretraining bias occurs when training data does not fairly represent the real-world distribution, for example with too little data from specific age groups.

  • DTrain longer on the same data until accuracy stops improving

    Incorrect

    More training on unrepresentative data can fit it more closely, but it cannot represent groups the data does not contain.

  • EBalance the dataset so that no group dominates training

    Correct

    Imbalanced data can bias training, so the model performs well on the majority and poorly on the rest; a broad, diverse and unbiased dataset makes accurate responses more likely.

Inclusive, diverse and balanced datasets are the first defense against biased models; detect remaining bias by measuring results per group.

Question 9 · choose 1

A company wants its generative AI workload to use less energy and have a smaller carbon footprint while still meeting its performance goals. Which practice does the AWS Well-Architected Generative AI Lens recommend?

  1. AAlways choose the largest available model to reduce retries
  2. BKeep a fixed fleet of large GPU instances running around the clock, even when idle
  3. CFine-tune a new copy of the model for every customer request
  4. DUse smaller models with techniques such as quantization and distillation
Show the answer and why
  • AAlways choose the largest available model to reduce retries

    Incorrect

    The lens recommends the opposite direction: smaller, optimized models that meet the performance goal.

  • BKeep a fixed fleet of large GPU instances running around the clock, even when idle

    Incorrect

    The lens recommends auto scaling and serverless architectures to optimize resource utilization, not idle fixed capacity.

  • CFine-tune a new copy of the model for every customer request

    Incorrect

    Repeated training adds compute and energy use; the lens recommends efficient customization and smaller models.

  • DUse smaller models with techniques such as quantization and distillation

    Correct

    The lens recommends smaller models and optimization techniques such as quantization, pruning and distillation to reduce resource use and promote sustainability while meeting performance goals.

Responsible model selection includes environmental cost: pick the smallest model that meets the requirement, and optimize how it runs.

Question 10 · choose 1

A team evaluates its RAG assistant with Amazon Bedrock. It wants a built-in metric that flags generalized statements about individuals or groups of people in the responses. Which metric fits?

  1. AContext coverage
  2. BStereotyping
  3. CCitation precision
  4. DCompleteness
Show the answer and why
  • AContext coverage

    Incorrect

    Context coverage measures how much of the ground truth the retrieved texts cover.

  • BStereotyping

    Correct

    Stereotyping measures generalized statements about individuals or groups of people in responses.

  • CCitation precision

    Incorrect

    Citation precision measures how many cited passages were correctly cited.

  • DCompleteness

    Incorrect

    Completeness measures how well responses resolve all aspects of the questions.

Built-in responsible AI metrics such as stereotyping and harmfulness let teams monitor bias alongside quality metrics.

Practise domain 4 →Practise all domains →