Question 1 · choose 1
Every night a single 40 GB gzip-compressed CSV file is loaded into an Amazon Redshift table with one COPY command. The load takes hours, and most slices sit idle during it. How can a data engineer make the load faster?
- ASplit the export into many files and start one concurrent COPY command for each file
- BSplit the export into similar-sized gzip files, a multiple of the slice count, and run one COPY
- CReplace the COPY command with batched INSERT statements sent from a client script
- DConvert the export to one gzip-compressed JSON file and keep a single COPY command for the load
Show the answer and why
ASplit the export into many files and start one concurrent COPY command for each file
Incorrect
Several concurrent COPY commands into one table force a serialized load, which is much slower and may need a VACUUM afterwards.
BSplit the export into similar-sized gzip files, a multiple of the slice count, and run one COPY
Correct
Gzip-compressed CSV cannot be split automatically, so one file loads serially. AWS recommends files of similar size, 1 MB to 1 GB after compression, in a multiple of the number of slices, loaded by one COPY.
CReplace the COPY command with batched INSERT statements sent from a client script
Incorrect
COPY loads large amounts of data much more efficiently than INSERT statements.
DConvert the export to one gzip-compressed JSON file and keep a single COPY command for the load
Incorrect
Compressed JSON is not split automatically either, so one file is still loaded serially.
Redshift loads in parallel across slices only when there is more than one piece to load. A splittable format or several similar-sized files, loaded with one COPY, keeps every slice busy.
AWS documentation