You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中基于多列执行Full Join(Outer Join)时如何避免生成带后缀的重复连接列

Polars中基于多列执行Full Join(Outer Join)时如何避免生成带后缀的重复连接列

我懂你碰到的这个麻烦——用Polars做全连接的时候,默认会把用来连接的列(比如这里的type和origin)分别保留左右表的版本,还自动加上_right后缀,看起来既冗余又不直观。其实解决这个问题只需要加一个参数就行,非常简单。

先回顾下你原来的代码和问题:

import polars as pl

df1 = pl.DataFrame({
    'type':   ['A', 'O', 'B', 'O'],
    'origin': ['EU', 'US', 'US', 'EU'],
    'qty1':   [343,11,22,-5]
})

df2 = pl.DataFrame({
    'type':   ['A', 'O', 'B', 'S'],
    'origin': ['EU', 'US', 'US', 'AS'],
    'qty2':   [-200,-12,-25,8]
})

# 原写法会生成重复的连接列
df1.join(df2, on=['type', 'origin'], how='full')

这个执行后会得到带type_right和origin_right的结果,而我们想要的是保留一套type和origin列,只把非连接的qty1和qty2合并进来。

解决方案:使用coalesce=True参数

Polars的join方法提供了coalesce参数,当设置为True时,它会自动把连接列的左右版本合并成单一列——优先取左表的值,如果左表对应位置是null(比如来自右表的独有序列),就用右表的值填充,完美适配全连接的场景。

修改后的代码如下:

# 添加coalesce=True,合并连接列
result = df1.join(df2, on=['type', 'origin'], how='full', coalesce=True)
print(result)

执行后会得到你想要的结果:

┌──────┬────────┬──────┬──────┐
│ type ┆ origin ┆ qty1 ┆ qty2 │
│ ---  ┆ ---    ┆ ---  ┆ ---  │
│ str  ┆ str    ┆ i64  ┆ i64  │
╞══════╪════════╪══════╪══════╡
│ A    ┆ EU     ┆ 343  ┆ -200 │
│ O    ┆ US     ┆ 11   ┆ -12  │
│ B    ┆ US     ┆ 22   ┆ -25  │
│ S    ┆ AS     ┆ null ┆ 8    │
│ O    ┆ EU     ┆ -5   ┆ null │
└──────┴────────┴──────┴──────┘

这样就不会再出现带后缀的重复连接列了,结果看起来清爽多了!

备注:内容来源于stack exchange,提问作者Phil-ZXX

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.16 07:38:02