You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark单行代码实现两DataFrame匹配替换字段值?

PySpark单行代码实现字段匹配替换

当然可以用单行代码实现需求,以下是两种常见场景的写法:

场景1:仅替换匹配到的行,未匹配行保留原ManagerJobcode值

df1.join(df2, df1.ManagerJobcode == df2.Managerjobcode, "left").withColumn("ManagerJobcode", coalesce(df2.ManagerID, df1.ManagerJobcode)).drop(df2.Managerjobcode)

代码说明:

  • 使用left左连接确保DF1的所有行都被保留
  • coalesce函数优先取匹配到的DF2.ManagerID,未匹配时保留原DF1.ManagerJobcode
  • 最后移除DF2中多余的Managerjobcode列

场景2:未匹配行的ManagerJobcode设为null

df1.join(df2, df1.ManagerJobcode == df2.Managerjobcode, "left").withColumn("ManagerJobcode", df2.ManagerID).drop(df2.Managerjobcode)

代码说明:

  • 同样用左连接保留DF1所有行
  • 直接将ManagerJobcode替换为DF2.ManagerID,未匹配时该字段会被设为null
  • 移除冗余列

内容的提问来源于stack exchange,提问作者Dhruv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 02:05:17