Polars中能否用单个pl.when批量拆分数据至多列?
在Polars中高效拆分全名到多列(避免重复编写
pl.when) 测试数据
首先定义测试用的DataFrame:
import polars as pl data = [ {"name_full": "McCartney, Paul"}, {"name_last": "Lennon", "name_first": "John"}, {"name_full": "Starr, Ringo"}, {"name_last": "Harrison", "name_first": "George"} ] df = pl.DataFrame(data)
初始DataFrame结构:
shape: (4, 3) ┌─────────────────┬───────────┬────────────┐ │ name_full ┆ name_last ┆ name_first │ │ --- ┆ --- ┆ --- │ │ str ┆ str ┆ str │ ╞═════════════════╪═══════════╪════════════╡ │ McCartney, Paul ┆ null ┆ null │ │ null ┆ Lennon ┆ John │ │ Starr, Ringo ┆ null ┆ null │ │ null ┆ Harrison ┆ George │ └─────────────────┴───────────┴────────────┘
现有解法(扩展性差)
目前通过为每个目标列单独编写pl.when实现需求,但字段增多时重复代码会大幅增加:
( df.with_columns( pl.col("name_full").str.split(",").list.eval(pl.element().str.strip_chars()).alias("name_parts") ).with_columns( pl.when(pl.col("name_last").is_null()) .then(pl.col("name_parts").list.get(0, null_on_oob=True)) .otherwise(pl.col("name_last")).alias("name_last"), pl.when(pl.col("name_first").is_null()) .then(pl.col("name_parts").list.get(1, null_on_oob=True)) .otherwise(pl.col("name_first")).alias("name_first") ).select(pl.all().exclude("name_full", "name_parts")) )
输出结果:
shape: (4, 2) ┌───────────┬────────────┐ │ name_last ┆ name_first │ │ --- ┆ --- │ │ str ┆ str │ ╞═══════════╪════════════╡ │ McCartney ┆ Paul │ │ Lennon ┆ John │ │ Starr ┆ Ringo │ │ Harrison ┆ George │ └───────────┴────────────┘
推荐优化方法(利用Struct简化逻辑)
可以通过Struct类型结合pl.coalesce实现批量字段合并,无需为每个字段单独编写判断逻辑,扩展性更强:
( df.with_columns( # 将name_full拆分后转换为与目标字段匹配的Struct split_struct=pl.col("name_full") .str.split(",") .list.eval(pl.element().str.strip()) # 替代strip_chars(),精准去除空格 .struct.rename_fields(["name_last", "name_first"]) ) # 合并原有字段Struct和拆分后的Struct,优先保留原有非空值 .with_columns( pl.coalesce( pl.struct("name_last", "name_first"), pl.col("split_struct") ).alias("combined_names") ) # 展开Struct为独立列,并清理临时字段 .unnest("combined_names") .drop("name_full", "split_struct") )
逻辑说明:
- 拆分并转Struct:把
name_full按逗号拆分、清理空格后,直接转为包含name_last和name_first的Struct,与原有字段结构对齐; - 批量合并:用
pl.coalesce对比原有字段的Struct和拆分后的Struct,自动优先取非空的对应字段; - 展开Struct:将合并后的Struct拆分为独立列,完成最终数据整理。
这种方式的优势在于:如果后续需要处理更多字段(比如中间名),只需要调整拆分后的Struct字段名和原有Struct的字段列表,无需重复编写判断逻辑。
内容的提问来源于stack exchange,提问作者dewser_the_board
相关产品推荐
相关产品推荐

