如何简化Polars中指定多列的常量填充操作?
Polars批量将指定列设为常量的简洁方法
原始代码及效果
我用以下Polars代码处理DataFrame:
import polars as pl df = pl.DataFrame({ 'region': ['GB', 'FR', 'US'], 'qty': [3, 6, -8], 'price': [100, 102, 95], 'tenor': ['1Y', '6M', '2Y'], }) cols_to_set = ['price', 'tenor'] fill_val = '-' df.with_columns([pl.lit(fill_val).alias(c) for c in cols_to_set])
执行后输出:
shape: (3, 4) ┌────────┬─────┬───────┬───────┐ │ region ┆ qty ┆ price ┆ tenor │ │ --- ┆ --- ┆ --- ┆ --- │ │ str ┆ i64 ┆ str ┆ str │ ╞════════╪═════╪═══════╪═══════╡ │ GB ┆ 3 ┆ - ┆ - │ │ FR ┆ 6 ┆ - ┆ - │ │ US ┆ -8 ┆ - ┆ - │ └────────┴─────┴───────┴───────┘
尝试的错误写法
我不想用列表推导式,尝试用pl.lit(fill_val).alias(cols_to_set)实现,触发如下报错:
TypeError: argument 'name': 'list' object cannot be converted to 'PyString'
简洁解决方案
可以直接利用Polars的with_columns方法支持字典参数的特性,用字典推导式生成列名与常量的映射,代码更简洁:
df.with_columns({c: fill_val for c in cols_to_set})
Polars会自动将常量转换为pl.lit表达式,效果和原始代码完全一致。
如果想进一步简化,也可以借助pl.select配合with_columns批量生成表达式:
df.with_columns(pl.select([pl.lit(fill_val).alias(c) for c in cols_to_set]))
不过第一种字典推导的写法可读性更高,也更简洁。
内容的提问来源于stack exchange,提问作者Phil-ZXX
相关产品推荐
相关产品推荐

