You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何理解Spark SQL执行计划中的前缀字符?

Spark SQL物理执行计划中+-、:-、: 前缀的含义解释

问题内容

以下面这条统计各部门员工数量的Spark SQL查询为例:

select 
   t2.dept_name, 
   coalesce(count(t1.emp_id),0) as emp_num  
from t2 
left join t1 on t2.dept_id = t1.dept_id
group by t2.dept_name 
order by t2.dept_name

对应的物理执行计划如下:

== Physical Plan ==
*(3) Sort [dept_name#13 ASC NULLS FIRST], true, 0
+- Exchange rangepartitioning(dept_name#13 ASC NULLS FIRST, 200), true, [id=#51]
   +- *(2) HashAggregate(keys=[dept_name#13], functions=[count(emp_id#4)], output=[dept_name#13, emp_num#16L])
      +- Exchange hashpartitioning(dept_name#13, 200), true, [id=#47]
         +- *(1) HashAggregate(keys=[dept_name#13], functions=[partial_count(emp_id#4)], output=[dept_name#13, count#21L])
            +- *(1) Project [dept_name#13, emp_id#4]
               +- *(1) BroadcastHashJoin [dept_id#12], [dept_id#5], LeftOuter, BuildRight
                  :- *(1) LocalTableScan [dept_id#12, dept_name#13]
                  +- BroadcastExchange HashedRelationBroadcastMode(List(cast(input[1, int, false] as bigint))), [id=#32]
                     +- LocalTableScan [emp_id#4, dept_id#5]

咨询:如何理解上述执行计划树中的+-、:-前缀?另外有时还会出现: (冒号后带空格)的前缀,其含义是什么?

含义解释

这些前缀是Spark用来可视化执行计划树结构的符号,核心作用是区分节点的层级和关系:

  • +-:表示当前节点是父节点的直接子节点,属于常规的树状层级延伸。执行计划里大部分串行的算子节点都会用这个符号,比如Sort的子节点Exchange、HashAggregate的子节点Exchange等,体现数据从上到下的流转关系。

  • :-:专门用于Join算子的左输入节点。例子中的BroadcastHashJoin是Join算子,:-对应的就是左表(t2)的扫描节点LocalTableScan,对应SQL里left join的左表数据源。

  • : :这是:-的变体,当Join算子的左输入节点还有更深层级的子节点时,左分支后续的节点会用: 标识,用来保持树结构对齐,明确属于左分支的后续算子。

简单来说,这三个符号都是为了让执行计划的树状结构更清晰,一眼就能看出节点间的父子关系,以及Join场景下左右分支的区分。

内容的提问来源于stack exchange,提问作者Tom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 07:10:10