如何基于Label列批量修改Polars DataFrame的Feature列?
给定如下Polars DataFrame:
import polars as pl import numpy as np df = pl.DataFrame({ "id": [1, 2, 3, 4, 5], "feature_a": np.random.randint(0, 3, 5), "feature_b": np.random.randint(0, 3, 5), "label": [1, 0, 0, 1, 1], })
对应结构:
┌─────┬───────────┬───────────┬───────┐
│ id ┆ feature_a ┆ feature_b ┆ label │
│ --- ┆ --- ┆ --- ┆ --- │
│ i64 ┆ i64 ┆ i64 ┆ i64 │
╞═════╪═══════════╪═══════════╪═══════╡
│ 1 ┆ 2 ┆ 0 ┆ 1 │
│ 2 ┆ 1 ┆ 1 ┆ 0 │
│ 3 ┆ 2 ┆ 2 ┆ 0 │
│ 4 ┆ 1 ┆ 0 ┆ 1 │
│ 5 ┆ 0 ┆ 0 ┆ 1 │
└─────┴───────────┴───────────┴───────┘
需要实现:根据label列的值批量修改所有以feature_开头的列(label为1时设为1,否则设为0),最终移除label列,得到目标结构:
┌─────┬───────────┬───────────┐
│ id ┆ feature_a ┆ feature_b │
│ --- ┆ --- ┆ --- │
│ i64 ┆ i64 ┆ i64 │
╞═════╪═══════════╪═══════════╡
│ 1 ┆ 1 ┆ 1 │
│ 2 ┆ 0 ┆ 0 │
│ 3 ┆ 0 ┆ 0 │
│ 4 ┆ 1 ┆ 1 │
│ 5 ┆ 1 ┆ 1 │
└─────┴───────────┴───────────┘
实现代码
提供两种高效写法,按需选择:
写法1:性能优先(直接基于布尔值转整数批量赋值)
result_df = df.with_columns( # 将label的布尔判断结果转为整数,批量覆盖所有feature列 (pl.col("label") == 1).cast(pl.Int64).alias(pl.col(r"^feature_.*$")) ).drop("label")
写法2:逻辑直观(使用when/then/otherwise语法)
result_df = df.with_columns( # 对每个feature列应用条件逻辑 pl.col(r"^feature_.*$").map( lambda _: pl.when(pl.col("label") == 1).then(1).otherwise(0) ) ).drop("label")
关键说明
pl.col(r"^feature_.*$"):通过正则精准匹配所有以feature_开头的目标列(pl.col("label") == 1).cast(pl.Int64):把label的布尔判断结果转为整数1/0,比when/then语法性能更高.alias(pl.col(r"^feature_.*$")):将生成的结果批量命名为原feature列名,实现覆盖修改.drop("label"):移除最终不需要的label列
内容的提问来源于stack exchange,提问作者bkw1491

