Spark 2.3动态分区在AWS EMR 5.13.0写入S3时失效求助
Yep, I’ve encountered this exact issue with EMR 5.13.0 and Spark 2.3’s dynamic partition overwrites to S3—temp directories get created but vanish without writing data to the final partition structure. Here are the solutions that resolved it for me and other community members:
1. EMR Version Compatibility is the Root Cause
EMR 5.13.0 has a known bug that breaks Spark 2.3’s dynamic partition handling for S3. AWS patched this in EMR 5.14.0 and later versions. If possible, upgrading your EMR cluster to 5.14.0+ is the most reliable fix.
2. Add Critical S3-Specific Spark Configurations
If upgrading isn’t an option right away, add these configurations to your SparkSession setup to work around the bug:
var spark = SparkSession .builder .appName(MyClass.getClass.getSimpleName) .config("spark.sql.sources.partitionOverwriteMode", "DYNAMIC") .config("spark.sql.parquet.output.committer.class", "org.apache.spark.sql.execution.datasources.parquet.ParquetOutputCommitter") .config("spark.hadoop.fs.s3a.impl", "org.apache.hadoop.fs.s3a.S3AFileSystem") .config("spark.hadoop.fs.s3a.fast.upload", "true") .getOrCreate()
- The
ParquetOutputCommitterreplaces EMR’s default committer which had issues with dynamic partition renames on S3. s3a.fast.uploadoptimizes S3 writes and reduces the chance of commit failures.
3. Adjust Your Write Logic
- Make sure you set
partitionOverwriteModebefore creating your Dataset (as shown above), not after getting the SparkSession. Some versions of Spark don’t pick up the config if set post-initialization. - Avoid using
coalesce(10)for large datasets—it can create bottlenecks during the commit phase. Userepartition(10)instead for more balanced parallelism.
4. Verify Permissions and S3 Settings
- Ensure your EMR cluster’s IAM role has full permissions for the target S3 bucket:
s3:PutObject,s3:DeleteObject,s3:ListBucket, ands3:GetObject. - Check that your S3 bucket’s object ownership is set correctly (e.g., bucket owner preferred) to avoid permission errors when moving files from temp to final partitions.
Debugging Tip
Enable DEBUG-level logging for the output committer to see exactly where the process fails:
spark.sparkContext.setLogLevel("DEBUG")
Look for log lines containing ParquetOutputCommitter or S3AFileSystem—they’ll show if temp files were created but couldn’t be renamed to the final partition paths.
内容的提问来源于stack exchange,提问作者David Costa Faidella

