Valid Associate-Developer-Apache-Spark Dumps shared by EduDump.com for Helping Passing Associate-Developer-Apache-Spark Exam! EduDump.com now offer the newest Associate-Developer-Apache-Spark exam dumps, the EduDump.com Associate-Developer-Apache-Spark exam questions have been updated and answers have been corrected get the newest EduDump.com Associate-Developer-Apache-Spark dumps with Test Engine here:
Which of the following code blocks reduces a DataFrame from 12 to 6 partitions and performs a full shuffle?
Correct Answer: E
Explanation DataFrame.repartition(6) Correct. repartition() always triggers a full shuffle (different from coalesce()). DataFrame.repartition(12) No, this would just leave the DataFrame with 12 partitions and not 6. DataFrame.coalesce(6) coalesce does not perform a full shuffle of the data. Whenever you see "full shuffle", you know that you are not dealing with coalesce(). While coalesce() can perform a partial shuffle when required, it will try to minimize shuffle operations, so the amount of data that is sent between executors. Here, 12 partitions can easily be repartitioned to be 6 partitions simply by stitching every two partitions into one. DataFrame.coalesce(6, shuffle=True) and DataFrame.coalesce(6).shuffle() These statements are not valid Spark API syntax. More info: Spark Repartition & Coalesce - Explained and Repartition vs Coalesce in Apache Spark - Rock the JVM Blog