Question 1 · choose 1
A banking assistant calls InvokeModel with a guardrail whose prompt attack filter is set to HIGH. The prompt sent to the model concatenates the developer's system instructions, retrieved account FAQs and the customer's message. Red-team tests show that injected instructions in the customer's message are not caught, and the team does not want its own system instructions evaluated as attacks. What should the developer change?
- ACreate a second guardrail with the prompt attack filter and apply it to the whole prompt, including the system instructions
- BSwitch the model to a reasoning model so that it recognizes injected instructions on its own
- CWrap only the customer's message in the guardrail input tags so that the prompt attack filter evaluates that section
- DAdd a denied topic named prompt injection with sample phrases such as "ignore all previous instructions"
Show the answer and why
ACreate a second guardrail with the prompt attack filter and apply it to the whole prompt, including the system instructions
Incorrect
Evaluating the whole prompt would also flag the developer's own instructions, which can resemble attacks, and untagged InvokeModel input is still not filtered for prompt attacks.
BSwitch the model to a reasoning model so that it recognizes injected instructions on its own
Incorrect
Model behavior is not a deterministic control. The guardrail must be configured to see the user input.
CWrap only the customer's message in the guardrail input tags so that the prompt attack filter evaluates that section
Correct
With InvokeModel, prompt attacks are filtered only inside the guardrail input tags. Tagging the user input makes the filter evaluate it while the developer's instructions stay out of the evaluation.
DAdd a denied topic named prompt injection with sample phrases such as "ignore all previous instructions"
Incorrect
Denied topics block subject areas in conversations. They are not a substitute for the prompt attack filter, which is not applied to untagged input in this setup.
Prompt attacks often look like system instructions, so the guardrail needs to know which part of the prompt came from the user. With InvokeModel and InvokeModelWithResponseStream you mark that part with input tags; without tags, prompt attacks are not filtered. With Converse you use guardContent blocks.
AWS documentation