Spark 3.3 on Kubernetes:每日120个ETL作业的Pod复用需求
解决方案:Spark on Kubernetes + Argo Workflows Pod复用优化
核心思路
针对Spark 3.3在K8s结合Argo Workflows的场景,实现Pod复用需依托Spark的动态资源分配与K8s端Pod复用机制,同时配合Argo的配置调整,减少重复创建Pod的资源开销。
具体实现步骤
1. 配置Spark开启Pod复用
在spark-submit命令或Spark配置文件中添加以下关键参数:
spark-submit \ --master kubernetes \ --conf spark.dynamicAllocation.enabled=true \ --conf spark.dynamicAllocation.shuffleTracking.enabled=true \ --conf spark.dynamicAllocation.executorIdleTimeout=300s \ --conf spark.kubernetes.executor.reuse.enabled=true \ my-etl.py
spark.kubernetes.executor.reuse.enabled=true:开启K8s executor Pod复用,空闲executor不会立即销毁spark.dynamicAllocation.enabled=true:允许Spark根据任务负载动态调整executor数量spark.executorIdleTimeout:设置executor空闲超时时间(示例为5分钟),可根据任务平均运行时长调整spark.dynamicAllocation.shuffleTracking.enabled=true:跟踪shuffle数据,避免复用executor时丢失关键数据
2. 调整Argo Workflows配置
- 复用
WorkflowTemplate:定义通用的Spark执行模板,让所有ETL任务共享该模板,避免重复创建Workflow资源 - 开启Pod复用策略:在Workflow级别配置
podReusePolicy,允许Argo复用已完成任务的Pod资源:
apiVersion: argoproj.io/v1alpha1 kind: Workflow metadata: generateName: spark-etl- spec: podReusePolicy: "OnCompletion" templates: - name: spark-etl-task container: image: your-spark-python-image:3.3 command: ["/bin/sh", "-c"] args: ["spark-submit --master kubernetes --conf ... my-etl.py"]
3. 优化任务调度与资源管控
- 任务分组:将120个每日任务按数据依赖或资源需求分组,用Argo的
DAG结构设置依赖,避免瞬间全量并行导致资源抢占 - 资源配额:通过K8s的
ResourceQuota和LimitRange,限制Spark所在Namespace的总资源使用,防止资源过载
注意事项
- 确保使用的Spark 3.3镜像包含完整的
spark-kubernetes模块,该版本正式支持executor Pod复用功能 - 空闲executor超时时间不宜过长,避免占用过多闲置资源
- 有状态ETL任务需确保shuffle tracking功能正常开启,防止数据丢失
内容的提问来源于stack exchange,提问作者Thomas Decaux
相关产品推荐
相关产品推荐

