在Polars中选择后缀匹配列并添加去除后缀的新列
Polars批量添加去后缀列的实现
问题背景
现有如下Polars DataFrame:
import polars as pl import numpy as np df = pl.DataFrame({ "nrs": [1, 2, 3, None, 5], "names_A0": ["foo", "ham", "spam", "egg", None], "random_A0": np.random.rand(5), "A_A2": [True, True, False, False, False], }) digit = 0
需求:对所有列名以suf = f'_A{digit}'(即_A0)结尾的列,添加一列内容完全相同但列名去除该后缀的新列。比如示例中要添加names和random列,分别与names_A0、random_A0内容一致,最终得到6列的DataFrame。
解决方案
通过Polars的with_columns方法结合列表推导式,可高效实现批量添加列的需求:
suf = f'_A{digit}' # 筛选出符合后缀条件的列 target_cols = [col for col in df.columns if col.endswith(suf)] # 批量添加去后缀的新列 df = df.with_columns( pl.col(col).alias(col.rstrip(suf)) for col in target_cols )
执行结果
运行上述代码后,得到的DataFrame如下:
shape: (5, 6) ┌──────┬──────────┬───────────┬───────┬───────┬──────────┐ │ nrs ┆ names_A0 ┆ random_A0 ┆ A_A2 ┆ names ┆ random │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ i64 ┆ str ┆ f64 ┆ bool ┆ str ┆ f64 │ ╞══════╪══════════╪═══════════╪═══════╪═══════╪══════════╡ │ 1 ┆ foo ┆ 0.274748 ┆ true ┆ foo ┆ 0.274748 │ │ 2 ┆ ham ┆ 0.26136 ┆ true ┆ ham ┆ 0.26136 │ │ 3 ┆ spam ┆ 0.718004 ┆ false ┆ spam ┆ 0.718004 │ │ null ┆ egg ┆ 0.303287 ┆ false ┆ egg ┆ 0.303287 │ │ 5 ┆ null ┆ 0.719041 ┆ false ┆ null ┆ 0.719041 │ └──────┴──────────┴───────────┴───────┴───────┴──────────┘
代码说明
- 筛选目标列:用列表推导式遍历所有列名,通过
endswith(suf)判断列是否符合后缀要求,得到需要处理的列集合。 - 生成新列:在
with_columns中,对每个目标列调用pl.col(col).alias(col.rstrip(suf)),保留原列数据的同时,将列名末尾的指定后缀去除作为新列名。
内容的提问来源于stack exchange,提问作者DeltaIV
相关产品推荐
相关产品推荐

