Question 1 · choose 1
A company serves a deep learning image model on SageMaker AI with high, steady request volume, and inference cost is now its largest ML bill. The model is built with PyTorch, and the team can compile it with the AWS Neuron SDK. Which instance family should the ML engineer evaluate first to lower the cost per inference?
- ALarger general purpose instances with more vCPUs
- BBurstable instances, to pay less for idle time
- CInf2 instances, based on AWS Inferentia2
- DThe same instance type with SageMaker AI managed warm pools
Show the answer and why
ALarger general purpose instances with more vCPUs
Incorrect
More CPU capacity can raise throughput, but general purpose CPUs are not built for deep learning inference, so cost per inference rarely drops this way.
BBurstable instances, to pay less for idle time
Incorrect
Burstable instances provide a baseline CPU level with the ability to burst, for workloads with low-to-moderate CPU use. A high, steady inference volume needs sustained performance.
CInf2 instances, based on AWS Inferentia2
Correct
Inferentia chips are designed for high performance at the lowest cost in Amazon EC2 for deep learning and generative AI inference, and Inf2 instances are built on Inferentia2. The Neuron SDK integrates with PyTorch to deploy models on them.
DThe same instance type with SageMaker AI managed warm pools
Incorrect
Warm pools keep training infrastructure ready between training jobs. They do not apply to endpoint inference cost.
For steady, high-volume deep learning inference, purpose-built accelerators such as Inferentia are a main cost lever. Use load tests, for example with Inference Recommender, to confirm the choice.
AWS documentation