Dataproc集群Spark作业资源问题及配置优化咨询
Dataproc集群Spark作业问题与配置优化咨询
集群初始配置
我们拥有一个Dataproc集群,初始配置如下:
- 主节点机型:
n1-highmem-16(16vCPU、104GB内存) - 工作节点:2台,初始机型为
n1-highmem-08(8vCPU、52GB内存)
机型参数详情:
Machine type vcpu memory n1-highmem-16 16 104 n1-highmem-08 08 52
作业运行属性
运行Spark数据导入作业时使用的属性:
Properties spark.submit.deployMode client spark.executor.cores 5 spark.dynamicAllocation.maxExecutors 2 spark.executor.memory 4g spark.driver.memory 4g spark.sql.autoBroadcastJoinThreshold 104857600 spark.dynamicAllocation.enabled true spark.executor.memoryOverhead 1g spark.executor.instances 1 spark.sql.shuffle.partitions 200
问题现象与疑问
近期作业突然耗时变长,等待资源约3小时,获取资源后仅需5分钟即可完成,同时存在以下疑问:
- Spark UI显示运行的容器数较少,按理解容器数应等于总vCore数,原因是什么?
- 基于节点vCPU和内存配置,Spark属性的最优组合是什么?
- 工作节点实际为8vCPU、52GB内存,但Spark UI显示单节点内存约185GB,且
Maximum Allocation <memory:40960, vCores:8>,该值如何计算并映射到Spark UI? - Spark UI中
Memory Total & Mem Avail的计算方式是什么? - 当前主节点
spark-defaults.conf配置如下,是否需要修正?
# Configure Spark on YARN spark.master=yarn spark.submit.deployMode=client spark.yarn.jars=local:/usr/lib/spark/jars/* # Dynamic allocation on YARN spark.dynamicAllocation.enabled=true spark.dynamicAllocation.minExecutors=1 spark.executor.instances=10000 spark.dynamicAllocation.maxExecutors=10000 spark.shuffle.service.enabled=true spark.scheduler.minRegisteredResourcesRatio=0.0 # This undoes setting hive.execution.engine to tez in hive-site.xml # It is not used by Spark spark.hadoop.hive.execution.engine=mr spark.rpc.message.maxSize=512 # Adding namespace to extract app_name and app_id for spark metrics spark.metrics.namespace=app_name:${spark.app.name}.app_id:${spark.app.id} spark.eventLog.enabled=true spark.eventLog.dir=hdfs://mbnl-pipe-emob-transform-v1-m/user/spark/eventlog spark.history.fs.logDirectory=hdfs://mbnl-pipe-emob-transform-v1-m/user/spark/eventlog spark.yarn.historyServer.address=mbnl-pipe-emob-transform-v1-m:18080 # User-supplied properties. #Wed Sep 20 12:11:20 UTC 2023 spark.port.maxRetries=1024 spark.sql.cbo.enabled=true spark.scheduler.mode=FAIR spark.executorEnv.OPENBLAS_NUM_THREADS=1 spark.executor.instances=2 spark.executor.memory=18619m spark.extraListeners=com.google.cloud.spark.performance.DataprocMetricsListener spark.history.fs.cleaner.interval=1d spark.driver.memory=26624m spark.history.fs.cleaner.maxAge=7d spark.history.fs.cleaner.enabled=true spark.executor.cores=2 spark.yarn.am.memory=640m spark.driver.maxResultSize=13312m
更新后的集群配置
2个工作节点机型统一为n1-highmem-08:
Machine type vcpu memory n1-highmem-08 08 52
当前主节点yarn-site.xml配置需调整:
yarn.nodemanager.resource.cpu-vcores设为8,应改为7(预留1vCPU给系统守护进程)yarn.nodemanager.resource.memory-mb设为190000,应改为49152(即48GB*1024,预留4GB给系统守护进程)
调整后每个工作节点可用资源为:7vCPU、48GB内存
计划配置合理性咨询
计划设置Spark属性如下,是否合理?
spark.executor.instances 2 spark.executor.cores=3 spark.executor.memory=14g spark.executor.memoryOverhead=2g
主节点配置建议
主节点机型为n1-highmem-16(16vCPU、104GB内存),请给出spark.driver.core、spark.driver.memory、spark.driver.memoryOverhead的理想配置。
内容的提问来源于stack exchange,提问作者Vikrant Singh Rana
相关产品推荐
相关产品推荐

