Question 1 · choose 1
A media company runs a SageMaker AI model that analyzes uploaded video clips of up to 400 MB. Each analysis takes 10 to 20 minutes. Clips arrive at random times during the day, results are needed within the hour, and there are long periods with no uploads, during which the company does not want to pay for instances. Which inference option should the ML engineer use?
- AA real-time endpoint on a large GPU instance
- BAsynchronous inference that scales to zero when idle
- CServerless inference with the maximum memory setting
- DA nightly batch transform job over all clips uploaded that day
Show the answer and why
AA real-time endpoint on a large GPU instance
Incorrect
Real-time inference supports payloads up to 25 MB and 60-second processing for regular responses, and the instances keep running during idle periods.
BAsynchronous inference that scales to zero when idle
Correct
Asynchronous inference queues requests, accepts payloads up to 1 GB with processing times up to one hour, and can scale the instance count to zero when there are no requests to process.
CServerless inference with the maximum memory setting
Incorrect
Serverless inference supports payloads up to 4 MB and processing times up to 60 seconds, far below a 400 MB clip that takes 20 minutes.
DA nightly batch transform job over all clips uploaded that day
Incorrect
Batch transform suits offline processing of data that is available up front. Waiting for a nightly run would miss the one-hour target for clips uploaded in the morning.
Choose the inference option by payload size, processing time, latency target and traffic pattern. Large payloads with long processing and near real-time needs point to asynchronous inference, which can also scale to zero.
AWS documentation