Question 1 · choose 1
A media company runs a trained model that analyzes uploaded video files. Each file is about 500 MB and takes up to 20 minutes to process. Every upload must be queued for processing as soon as it arrives, uploads come in irregular bursts, and the company does not want to pay for idle instances between bursts. Which Amazon SageMaker AI inference option fits best?
- AAsynchronous inference
- BReal-time inference on a persistent endpoint
- CServerless inference
- DBatch transform jobs over the full dataset
Show the answer and why
AAsynchronous inference
Correct
Asynchronous inference queues requests, accepts payloads up to 1 GB with processing times up to one hour, and can scale the endpoint down to zero when there is nothing to process.
BReal-time inference on a persistent endpoint
Incorrect
Real-time inference is for low-latency online requests and supports payloads up to 25 MB with 60-second processing for regular responses, far below a 500 MB file that runs for 20 minutes.
CServerless inference
Incorrect
Serverless inference suits intermittent traffic, but it supports payloads only up to 4 MB and processing times up to 60 seconds.
DBatch transform jobs over the full dataset
Incorrect
Batch transform is for offline processing when a large dataset is available upfront. Here each file arrives on its own and must be queued the moment it is uploaded.
Large payloads, long processing times and bursty traffic with a wish to scale to zero are the profile SageMaker AI documents for asynchronous inference.
AWS documentation