如何在Spark中避免Hive分区Parquet文件拆分以减少输出文件(无Shuffle)
Hey there, let's tackle your problem step by step—you want to join your partitioned Parquet table A with a small table C (via MapJoin), write the result to table B, and keep each partition of B as a single file without any shuffle. Let's break down what's going on and how to fix it.
First: Why mapreduce.input.fileinputformat.split.minsize Didn't Work
That Hadoop configuration is technically meant to control split sizes, but Spark SQL doesn't prioritize it by default. Spark has its own parameter that overrides this for file-based data sources: spark.sql.files.maxPartitionBytes (default value is 128MB). This is the key setting that determines whether Spark splits a single input file into multiple RDD/DataFrame partitions. So even if you set the Hadoop split minsize larger than your files, Spark's own parameter was still splitting them if your files exceeded 128MB.
The No-Shuffle Fix: Keep Each Input File as a Single Spark Partition
To make sure each Parquet file in table A maps to exactly one Spark partition (no splitting), you need to adjust Spark's file partitioning settings:
1. Set spark.sql.files.maxPartitionBytes to Exceed Your Largest File Size
Configure this parameter to a value bigger than the largest Parquet file in table A. For example, if your biggest file is 2GB, set it to 2g or 2147483648 (bytes). This tells Spark not to split any of your input files.
You can set this when initializing your SparkSession:
val spark = SparkSession.builder() .appName("SingleFilePerPartition") .config("spark.sql.files.maxPartitionBytes", "2g") // Adjust to your max file size .enableHiveSupport() .getOrCreate()
2. Ensure MapJoin (Broadcast Join) is Used for Table C
Since table C is small, Spark should automatically broadcast it to all executors (no shuffle needed). The default threshold for auto-broadcast is 10MB, but if your table C is larger, adjust spark.sql.autoBroadcastJoinThreshold to a value bigger than table C's size:
// Add this to your SparkSession config if needed .config("spark.sql.autoBroadcastJoinThreshold", "50m") // Allow broadcasting up to 50MB
You can verify this works by checking the Spark UI's SQL tab—look for a BroadcastHashJoin in the execution plan, not a SortMergeJoin (which triggers shuffle).
3. Write to Table B Without Shuffle
When writing, since each Spark partition corresponds to one partition from table A (and thus one partition in table B), Spark will write each partition's data to a single file—no shuffle required. Just make sure you use the same partition column as table A/B:
val dfA = spark.table("A") val dfC = spark.table("C") // Perform the join (MapJoin will be used automatically) val joinedDF = dfA.join(dfC, Seq("your_join_key"), "inner") // Write to table B joinedDF.write .mode("append") // Use "overwrite" if you want to replace the empty table .partitionBy("your_partition_column") // Match A and B's partition column .saveAsTable("B")
Bonus: Prevent Accidental File Splitting During Write
If you still see multiple files per partition in B (rare if you followed the above), you can set spark.sql.files.maxRecordsPerFile to a very large number (like 1 billion) to ensure one file per Spark partition:
.config("spark.sql.files.maxRecordsPerFile", "1000000000")
Key Checks to Confirm No Shuffle
- Spark UI's Stages tab: No stages marked as "Shuffle Write"
- Execution plan: Only
BroadcastHashJoin(noSortMergeJoin) - Table B's partitions: Each has exactly one Parquet file matching the corresponding file in table A
内容的提问来源于stack exchange,提问作者Joha

