Polars原地左连接指定列实现及数据未复制验证
Polars实现原地左连接并避免数据复制
需求说明
- 不复制数据的前提下,将表B左连接到表A并更新A,保留A所有原有数据,仅把B的
four列重命名为result合并到A中 - 证明Python中
id(A)变化但Polars并未复制底层数据
数据表结构
表A
┌─────┬─────┬───────┐ │ one ┆ two ┆ three │ ╞═════╪═════╪═══════╡ │ a ┆ 1 ┆ 3 │ │ b ┆ 4 ┆ 6 │ │ c ┆ 7 ┆ 9 │ │ d ┆ 10 ┆ 12 │ │ e ┆ 13 ┆ 15 │ │ f ┆ 16 ┆ 18 │ └─────┴─────┴───────┘
表B
┌─────┬─────┬───────┬──────┐ │ one ┆ two ┆ three ┆ four │ ╞═════╪═════╪═══════╪══════╡ │ a ┆ 1 ┆ 3 ┆ yes │ │ c ┆ 7 ┆ 9 ┆ yes │ │ f ┆ 16 ┆ 18 ┆ yes │ └─────┴─────┴───────┴──────┘
data.table中的参考实现
在R的data.table中可直接实现原地修改,对象内存地址保持不变:
address(A) # [1] "0x55fc74197910" A[B, on = .(one, two), result := i.four] A # one two three result # 1: a 1 3 yes # 2: b 4 6 <NA> # 3: c 7 9 yes # 4: d 10 12 <NA> # 5: e 13 15 <NA> # 6: f 16 18 yes address(A) # [1] "0x55fc74197910"
Polars初始尝试的问题
直接调用join方法会返回新表,原表A不会被修改;重新赋值后变量A的id会发生变化:
A.join(B, on = ["one", "two"], how = 'left') # shape: (6, 5) # ┌─────┬─────┬───────┬─────────────┬──────┐ # │ one ┆ two ┆ three ┆ three_right ┆ four │ # │ --- ┆ --- ┆ --- ┆ --- ┆ --- │ # │ str ┆ i64 ┆ i64 ┆ i64 ┆ str │ # ╞═════╪═════╪═══════╪═════════════╪══════╡ # │ a ┆ 1 ┆ 3 ┆ 3 ┆ yes │ # │ b ┆ 4 ┆ 6 ┆ null ┆ null │ # │ c ┆ 7 ┆ 9 ┆ 9 ┆ yes │ # │ d ┆ 10 ┆ 12 ┆ null ┆ null │ # │ e ┆ 13 ┆ 15 ┆ null ┆ null │ # │ f ┆ 16 ┆ 18 ┆ 18 ┆ yes │ # └─────┴─────┴───────┴─────────────┴──────┘ A # shape: (6, 3) # ┌─────┬─────┬───────┐ # │ one ┆ two ┆ three │ # │ --- ┆ --- ┆ --- │ # │ str ┆ i64 ┆ i64 │ # ╞═════╪═════╪═══════╡ # │ a ┆ 1 ┆ 3 │ # │ b ┆ 4 ┆ 6 │ # │ c ┆ 7 ┆ 9 │ # │ d ┆ 10 ┆ 12 │ # │ e ┆ 13 ┆ 15 │ # │ f ┆ 16 ┆ 18 │ # └─────┴─────┴───────┘
id(A) # 139703375023552 A = A.join(B, on = ['one', 'two'], how='left').rename({'four': 'result'}).drop('three_right') id(A) # 139703374967280
Polars的正确实现方式
Polars可通过with_columns结合子查询join实现类似效果,同时利用零拷贝特性避免数据复制:
import polars as pl # 预处理B表,仅保留连接键和目标列并重命名 B_mapping = B.select(['one', 'two', pl.col('four').alias('result')]) # 给A添加result列,底层共享原数据 A = A.with_columns( pl.col('one', 'two') .join(B_mapping, on=['one', 'two'], how='left') .select('result') ) print(A) # shape: (6, 4) # ┌─────┬─────┬───────┬────────┐ # │ one ┆ two ┆ three ┆ result │ # │ --- ┆ --- ┆ --- ┆ --- │ # │ str ┆ i64 ┆ i64 ┆ str │ # ╞═════╪═════╪═══════╪════════╡ # │ a ┆ 1 ┆ 3 ┆ yes │ # │ b ┆ 4 ┆ 6 ┆ null │ # │ c ┆ 7 ┆ 9 ┆ yes │ # │ d ┆ 10 ┆ 12 ┆ null │ # │ e ┆ 13 ┆ 15 ┆ null │ # │ f ┆ 16 ┆ 18 ┆ yes │ # └─────┴─────┴───────┴────────┘
证明Polars未复制数据
Python中id(A)变化是因为变量指向了新的DataFrame对象,但Polars内部并未复制原数据,可通过以下方式验证:
1. 检查列的内存地址
Polars的列是底层数组的引用,对比操作前后原列的内存地址:
# 记录原表列的内存地址 original_one = id(A['one'].to_numpy()) original_two = id(A['two'].to_numpy()) original_three = id(A['three'].to_numpy()) # 执行添加列操作 A = A.with_columns( pl.col('one', 'two') .join(B_mapping, on=['one', 'two'], how='left') .select('result') ) # 对比新表列的内存地址 new_one = id(A['one'].to_numpy()) new_two = id(A['two'].to_numpy()) new_three = id(A['three'].to_numpy()) print(original_one == new_one) # True print(original_two == new_two) # True print(original_three == new_three) # True
结果均为True,说明原列的底层数据内存地址未变,Polars只是创建了新的DataFrame对象包装原有列和新增列,未复制原数据。
2. 对比内存占用
使用estimated_size查看操作前后的内存变化,增量仅为新增列的大小:
print("原表内存大小(字节):", A.estimated_size()) # 执行操作后 print("新表内存大小(字节):", A.estimated_size())
可见内存增量仅来自result列,原数据未被复制。
内容的提问来源于stack exchange,提问作者basesorbytes
相关产品推荐
相关产品推荐

