Question 1 · choose 1
A team hosts a large language model on a SageMaker AI real-time endpoint with GPU instances for internal testing. The endpoint is idle every night and weekend, and the team wants it to drop to zero instances then while keeping the same real-time invocation interface during working hours. Cold starts of several minutes are acceptable. How should the ML engineer configure this?
- AMove the model to serverless inference so it scales to zero automatically
- BSet the variant minimum capacity to 0 in a target tracking policy on InvocationsPerInstance
- CDeploy it as an inference component with min 0 instances and step scaling
- DConvert the endpoint to asynchronous inference
Show the answer and why
AMove the model to serverless inference so it scales to zero automatically
Incorrect
Serverless inference does scale to zero, but it does not support GPUs, which this model needs.
BSet the variant minimum capacity to 0 in a target tracking policy on InvocationsPerInstance
Incorrect
Scaling to zero is tied to inference components; a variant without them cannot scale in to zero instances. Once at zero, a scale-out also needs a step scaling policy and alarm.
CDeploy it as an inference component with min 0 instances and step scaling
Correct
An endpoint can scale in to and out from zero instances only if it hosts inference components, with MinInstanceCount set to 0. A step scaling policy tied to an alarm provisions an instance again when requests arrive.
DConvert the endpoint to asynchronous inference
Incorrect
Asynchronous inference can scale to zero, but callers must place payloads in S3 and use asynchronous invocation, which changes the real-time interface the team wants to keep.
Real-time endpoints can scale to zero when models are deployed as inference components. Expect errors during the several minutes it takes to provision the first instance again.
AWS documentation