You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark 2.3动态分区在AWS EMR 5.13.0写入S3时失效求助

Fix for Spark 2.3 Dynamic Partitioning Failure on EMR 5.13.0 Writing to S3

Yep, I’ve encountered this exact issue with EMR 5.13.0 and Spark 2.3’s dynamic partition overwrites to S3—temp directories get created but vanish without writing data to the final partition structure. Here are the solutions that resolved it for me and other community members:

1. EMR Version Compatibility is the Root Cause

EMR 5.13.0 has a known bug that breaks Spark 2.3’s dynamic partition handling for S3. AWS patched this in EMR 5.14.0 and later versions. If possible, upgrading your EMR cluster to 5.14.0+ is the most reliable fix.

2. Add Critical S3-Specific Spark Configurations

If upgrading isn’t an option right away, add these configurations to your SparkSession setup to work around the bug:

var spark = SparkSession
  .builder
  .appName(MyClass.getClass.getSimpleName)
  .config("spark.sql.sources.partitionOverwriteMode", "DYNAMIC")
  .config("spark.sql.parquet.output.committer.class", "org.apache.spark.sql.execution.datasources.parquet.ParquetOutputCommitter")
  .config("spark.hadoop.fs.s3a.impl", "org.apache.hadoop.fs.s3a.S3AFileSystem")
  .config("spark.hadoop.fs.s3a.fast.upload", "true")
  .getOrCreate()
  • The ParquetOutputCommitter replaces EMR’s default committer which had issues with dynamic partition renames on S3.
  • s3a.fast.upload optimizes S3 writes and reduces the chance of commit failures.

3. Adjust Your Write Logic

  • Make sure you set partitionOverwriteMode before creating your Dataset (as shown above), not after getting the SparkSession. Some versions of Spark don’t pick up the config if set post-initialization.
  • Avoid using coalesce(10) for large datasets—it can create bottlenecks during the commit phase. Use repartition(10) instead for more balanced parallelism.

4. Verify Permissions and S3 Settings

  • Ensure your EMR cluster’s IAM role has full permissions for the target S3 bucket: s3:PutObject, s3:DeleteObject, s3:ListBucket, and s3:GetObject.
  • Check that your S3 bucket’s object ownership is set correctly (e.g., bucket owner preferred) to avoid permission errors when moving files from temp to final partitions.

Debugging Tip

Enable DEBUG-level logging for the output committer to see exactly where the process fails:

spark.sparkContext.setLogLevel("DEBUG")

Look for log lines containing ParquetOutputCommitter or S3AFileSystem—they’ll show if temp files were created but couldn’t be renamed to the final partition paths.

内容的提问来源于stack exchange,提问作者David Costa Faidella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:04:36