如何在Polars中为每行创建含非空值唯一列表的新列?
在Polars中提取每行非空唯一货币代码集合
针对你的需求,Polars可以通过向量化操作高效实现,无需转成Pandas再处理,代码如下:
import polars as pl df = pl.DataFrame( { "col1cur": ["EUR", "EUR", "EUR"], "col2cur": [None, "EUR", None], "col3cur": ["EUR", None, None], "col4cur": ["EUR", "GBP", None], } ) # 生成目标列 result = df.with_columns( target=pl.concat_list(pl.col(r'^.*cur$')) .list.drop_nulls() .list.unique() ) print(result)
代码逻辑说明
pl.col(r'^.*cur$'):用正则筛选所有以cur结尾的货币列,匹配你提到的150个目标列pl.concat_list(...):将每行的所有货币列值合并为一个列表.list.drop_nulls():移除列表中的空值(None).list.unique():对列表中的元素去重,得到该行的唯一货币代码集合
输出结果
┌─────────┬─────────┬─────────┬─────────┬────────────────┐ │ col1cur ┆ col2cur ┆ col3cur ┆ col4cur ┆ target │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ str ┆ str ┆ str ┆ str ┆ list[str] │ ╞═════════╪═════════╪═════════╪═════════╪════════════════╡ │ EUR ┆ null ┆ EUR ┆ EUR ┆ ["EUR"] │ │ EUR ┆ EUR ┆ null ┆ GBP ┆ ["EUR", "GBP"] │ │ EUR ┆ null ┆ null ┆ null ┆ ["EUR"] │ └─────────┴─────────┴─────────┴─────────┴────────────────┘
优势说明
相比Pandas的apply逐行循环,Polars的向量化操作在处理大量列(比如你提到的150列)时,性能会显著提升,且代码更简洁易读。
内容的提问来源于stack exchange,提问作者Andrew
相关产品推荐
相关产品推荐

