You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

合并Polars DataFrame时忽略on列中的重复值

问题解决:Polars 避免重复键连接产生笛卡尔积

默认的join操作会匹配所有具有相同连接键的行,因此当A列存在重复值时,会产生笛卡尔积(2×2=4行)。要实现按行对应合并的需求,可以通过以下两种方法解决:

方法1:添加临时行索引列进行连接

给两个DataFrame添加唯一的行索引列,将其与A列共同作为连接键,确保每行仅匹配对应位置的行:

import polars as pl

df1 = pl.DataFrame({"A": [1, 1], "B": [3, 4]}).with_row_index("_temp_idx")
df2 = pl.DataFrame({"A": [1, 1], "C": [5, 6]}).with_row_index("_temp_idx")

result = df1.join(df2, on=["A", "_temp_idx"]).drop("_temp_idx")

执行后result的输出为:

shape: (2, 3)
┌─────┬─────┬─────┐
│ A   ┆ B   ┆ C   │
│ --- ┆ --- ┆ --- │
│ i64 ┆ i64 ┆ i64 │
╞═════╪═════╪═════╡
│ 1   ┆ 3   ┆ 5   │
│ 1   ┆ 4   ┆ 6   │
└─────┴─────┴─────┘

方法2:直接横向拼接(hstack)

如果可以确保两个DataFrame的行数完全相同且行顺序严格对应,可以直接使用hstack将两表横向拼接,无需执行连接操作:

result = df1.hstack(df2.select("C"))

该方法更简洁,适合行序完全匹配的场景。


内容的提问来源于stack exchange,提问作者vinnewbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 22:25:24