EMR on EKS使用IBM S3 Shuffle Plugin时动态资源分配缩容异常求助
EMR on EKS + IBM S3 Shuffle Plugin v0.9.6:动态资源分配(DRA)无法正常缩容
已在EMR on EKS(Spark 3.5.0)上成功部署IBM S3 Shuffle Plugin v0.9.6,S3 shuffle功能运行正常,但动态资源分配(DRA)无法正常缩容:闲置Executor未按配置超时时间关闭,对应的EC2实例也未被Karpenter终止,造成不必要成本开销。已知该场景官方确认支持缩容。
当前配置
{ "spark.driver.container.image": "{{ ti.xcom_pull(task_ids='fetch_configs', key='custom_emr_image_uri') }}", "spark.executor.container.image": "{{ ti.xcom_pull(task_ids='fetch_configs', key='custom_emr_image_uri') }}", "spark.shuffle.s3.rootDir": "s3://{{ ti.xcom_pull(task_ids='fetch_configs', key='emr_data_bucket_name') }}/shuffle-data/", "spark.cleaner.periodicGC.interval": "1min", "spark.decommission.enabled": "true", "spark.dynamicAllocation.enabled": "true", "spark.dynamicAllocation.cachedExecutorIdleTimeout": "90s", "spark.dynamicAllocation.executorIdleTimeout": "60s", "spark.dynamicAllocation.schedulerBacklogTimeout": "1s", "spark.dynamicAllocation.shuffleTracking.timeout": "0", "spark.dynamicAllocation.initialExecutors": "10", "spark.dynamicAllocation.maxExecutors": "200", "spark.dynamicAllocation.minExecutors": "1", "spark.storage.decommission.enabled": "true", "spark.storage.decommission.rddBlocks.enabled": "true", "spark.storage.decommission.shuffleBlocks.enabled": "true", "spark.eventLog.enabled": "true", "spark.hadoop.fs.file.block.size": "134217728", "spark.shuffle.checksum.enabled": "true", "spark.shuffle.file.buffer.size": "1m", "spark.shuffle.manager": "org.apache.spark.shuffle.sort.S3ShuffleManager", "spark.shuffle.sort.io.plugin.class": "org.apache.spark.shuffle.S3ShuffleDataIO", "spark.sql.adaptive.enabled": "true", "spark.sql.adaptive.localShuffleReader.enabled": "false", "spark.serializer": "org.apache.spark.serializer.KryoSerializer" }
具体问题
- Executor能正常扩容,但闲置远超配置超时时间后仍不缩容
- 搭载闲置Executor的EC2实例未被Karpenter终止
疑问
- 使用S3 shuffle存储时,需额外配置或调整以实现Executor正常退役吗?
- 使用IBM S3 shuffle plugin时,
spark.dynamicAllocation.shuffleTracking.enabled应设为true还是false? - 期望实现的缩容功能是否可行?
环境信息
- EMR on EKS搭配Karpenter
- Spark 3.5.0(EMR 7.5.0,因插件兼容性使用旧版本)
- IBM S3 Shuffle Plugin v0.9.6
- 预安装插件JAR的自定义Docker镜像
问题分析与解决方案
核心配置修正
你的配置中spark.dynamicAllocation.shuffleTracking.timeout设为0,这会导致Shuffle跟踪永不超时,Spark判定Executor仍持有必要的Shuffle数据,不会触发缩容。结合IBM S3 Shuffle Plugin特性,需调整以下关键配置:
spark.dynamicAllocation.shuffleTracking.enabled:必须设为true。由于S3 Shuffle Plugin将Shuffle数据存放在S3而非本地磁盘,Spark需要通过Shuffle跟踪确认Executor是否还需为后续阶段保留Shuffle数据。禁用该配置会导致Spark默认认为所有Executor持有不可替代的Shuffle数据,拒绝缩容。spark.dynamicAllocation.shuffleTracking.timeout:设置为与spark.dynamicAllocation.executorIdleTimeout一致的值(如60s),确保闲置Executor超时后被标记为可回收。spark.shuffle.s3.deleteShuffleFiles:设为true,确保Executor退役后自动清理S3上的Shuffle文件,避免存储冗余。spark.shuffle.s3.decommission.enabled:确认该值为true(IBM插件专属配置),保障Executor退役时正确处理S3上的Shuffle资源。
补充缩容优化配置
新增或调整以下配置,确保Executor退役流程顺畅:
"spark.executor.decommission.awaitCompletionTimeout": "30s", "spark.dynamicAllocation.executorAllocationRatio": "1.0"
spark.executor.decommission.awaitCompletionTimeout:给Executor足够时间完成Shuffle数据清理与状态上报spark.dynamicAllocation.executorAllocationRatio:让Spark更精准地根据负载调整Executor数量
Karpenter配置验证
确认Karpenter Provisioner中ttlSecondsAfterEmpty设置合理(如30s),当节点上所有Pod(含Executor)被删除后,Karpenter会在超时后终止EC2实例。同时检查Karpenter日志,排除实例终止失败的错误(如节点被标记为不可调度但未删除)。
疑问解答
- 需要额外配置:除基础DRA配置外,必须启用Shuffle跟踪并设置合理超时,同时配置IBM插件专属的退役参数,让Spark识别到Executor已无Shuffle数据依赖,允许其退役。
- 必须设为
true:S3 Shuffle Plugin使用共享存储(S3)存放Shuffle数据,Spark依赖Shuffle跟踪机制判断Executor是否需要保留。设为false会导致闲置Executor永远无法被缩容。 - 完全可行:根据官方确认,IBM S3 Shuffle Plugin支持DRA缩容功能,配置正确后,闲置Executor会在超时后被回收,对应的EC2实例也会被Karpenter终止。
内容的提问来源于stack exchange,提问作者metersk
相关产品推荐
相关产品推荐

