You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中能否用单个pl.when批量拆分数据至多列?

在Polars中高效拆分全名到多列(避免重复编写pl.when)

测试数据

首先定义测试用的DataFrame:

import polars as pl
data = [
    {"name_full": "McCartney, Paul"},
    {"name_last": "Lennon", "name_first": "John"},
    {"name_full": "Starr, Ringo"},
    {"name_last": "Harrison", "name_first": "George"}
]
df = pl.DataFrame(data)

初始DataFrame结构:

shape: (4, 3)
┌─────────────────┬───────────┬────────────┐
│ name_full       ┆ name_last ┆ name_first │
│ ---             ┆ ---       ┆ ---        │
│ str             ┆ str       ┆ str        │
╞═════════════════╪═══════════╪════════════╡
│ McCartney, Paul ┆ null      ┆ null       │
│ null            ┆ Lennon    ┆ John       │
│ Starr, Ringo    ┆ null      ┆ null       │
│ null            ┆ Harrison  ┆ George     │
└─────────────────┴───────────┴────────────┘

现有解法(扩展性差)

目前通过为每个目标列单独编写pl.when实现需求,但字段增多时重复代码会大幅增加:

(
    df.with_columns(
        pl.col("name_full").str.split(",").list.eval(pl.element().str.strip_chars()).alias("name_parts")
    ).with_columns(
        pl.when(pl.col("name_last").is_null())
        .then(pl.col("name_parts").list.get(0, null_on_oob=True))
        .otherwise(pl.col("name_last")).alias("name_last"),
        pl.when(pl.col("name_first").is_null())
        .then(pl.col("name_parts").list.get(1, null_on_oob=True))
        .otherwise(pl.col("name_first")).alias("name_first")
    ).select(pl.all().exclude("name_full", "name_parts"))
)

输出结果:

shape: (4, 2)
┌───────────┬────────────┐
│ name_last ┆ name_first │
│ ---       ┆ ---        │
│ str       ┆ str        │
╞═══════════╪════════════╡
│ McCartney ┆ Paul       │
│ Lennon    ┆ John       │
│ Starr     ┆ Ringo      │
│ Harrison  ┆ George     │
└───────────┴────────────┘

推荐优化方法(利用Struct简化逻辑)

可以通过Struct类型结合pl.coalesce实现批量字段合并,无需为每个字段单独编写判断逻辑,扩展性更强:

(
    df.with_columns(
        # 将name_full拆分后转换为与目标字段匹配的Struct
        split_struct=pl.col("name_full")
            .str.split(",")
            .list.eval(pl.element().str.strip())  # 替代strip_chars(),精准去除空格
            .struct.rename_fields(["name_last", "name_first"])
    )
    # 合并原有字段Struct和拆分后的Struct,优先保留原有非空值
    .with_columns(
        pl.coalesce(
            pl.struct("name_last", "name_first"),
            pl.col("split_struct")
        ).alias("combined_names")
    )
    # 展开Struct为独立列,并清理临时字段
    .unnest("combined_names")
    .drop("name_full", "split_struct")
)

逻辑说明:

  1. 拆分并转Struct:把name_full按逗号拆分、清理空格后,直接转为包含name_last和name_first的Struct,与原有字段结构对齐;
  2. 批量合并:用pl.coalesce对比原有字段的Struct和拆分后的Struct,自动优先取非空的对应字段;
  3. 展开Struct:将合并后的Struct拆分为独立列,完成最终数据整理。

这种方式的优势在于:如果后续需要处理更多字段(比如中间名),只需要调整拆分后的Struct字段名和原有Struct的字段列表,无需重复编写判断逻辑。

内容的提问来源于stack exchange,提问作者dewser_the_board

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 16:20:55