如何使用Polars Expressions修改DataFrame指定行列的值
Polars表达式如何修改指定行列的值
我了解Polars Expressions针对DataFrame的数据读取、计算和修改采用列并行处理,速度非常快,但找不到能像df[1,'a']=12那样修改指定行列值的Polars Expressions语法,该如何实现?
以下是期望的实现场景:
import polars as pl df = pl.DataFrame( { "a": [1,2,3,4,5,6,7,8], "b": [8,9,0,1,2,3,4,5], "c": [0,0,0,0,0,0,0,0], "d": [0,0,0,0,0,0,0,0], } ) df = df.with_columns(pl.all().cast(pl.Float64)) df[7,'c']=12 df[7,'d']=15 # 我希望用Polars Expressions实现上述操作,类似: # => df.select(?)( # pl.col('c').row(7) = 12 # pl.col('d').row(7) = 15 # ) print(df)
期望输出:
shape: (8, 4) ┌─────┬─────┬──────┬──────┐ │ a ┆ b ┆ c ┆ d │ │ --- ┆ --- ┆ --- ┆ --- │ │ f64 ┆ f64 ┆ f64 ┆ f64 │ ╞═════╪═════╪══════╪══════╡ │ 1.0 ┆ 8.0 ┆ 0.0 ┆ 0.0 │ │ 2.0 ┆ 9.0 ┆ 0.0 ┆ 0.0 │ │ 3.0 ┆ 0.0 ┆ 0.0 ┆ 0.0 │ │ 4.0 ┆ 1.0 ┆ 0.0 ┆ 0.0 │ │ 5.0 ┆ 2.0 ┆ 0.0 ┆ 0.0 │ │ 6.0 ┆ 3.0 ┆ 0.0 ┆ 0.0 │ │ 7.0 ┆ 4.0 ┆ 0.0 ┆ 0.0 │ │ 8.0 ┆ 5.0 ┆ 12.0 ┆ 15.0 │ └─────┴─────┴──────┴──────┘
解决方案
Polars是面向列的并行计算框架,不推荐直接修改单个元素(会破坏列并行优势),但可以通过条件判断+表达式的方式实现指定行列的修改,完全符合Polars的范式:
方法1:使用row_index()(Polars 0.17.0+支持)
pl.row_index()可以直接获取每行的索引,配合when/then/otherwise实现精准替换:
import polars as pl df = pl.DataFrame( { "a": [1,2,3,4,5,6,7,8], "b": [8,9,0,1,2,3,4,5], "c": [0,0,0,0,0,0,0,0], "d": [0,0,0,0,0,0,0,0], } ) df = df.with_columns(pl.all().cast(pl.Float64)) # 用Polars表达式修改指定行列 df = df.with_columns( pl.col("c").when(pl.row_index() == 7).then(12.0).otherwise(pl.col("c")), pl.col("d").when(pl.row_index() == 7).then(15.0).otherwise(pl.col("d")) ) print(df)
方法2:兼容旧版本的int_range写法
如果你的Polars版本低于0.17.0,可以用pl.int_range(0, pl.count())生成行索引序列:
df = df.with_columns( pl.col("c").when(pl.int_range(0, pl.count()) == 7).then(12.0).otherwise(pl.col("c")), pl.col("d").when(pl.int_range(0, pl.count()) == 7).then(15.0).otherwise(pl.col("d")) )
说明
这种写法保留了Polars列并行的性能优势,同时实现了指定行列的修改需求。如果需要修改多行或更复杂的条件,只需调整when中的判断逻辑即可。
内容的提问来源于stack exchange,提问作者young hwan Song
相关产品推荐
相关产品推荐

