如何用Polars获取DataFrame每行最右侧非空值并生成新列?
Polars实现取每行最右侧非空值的方案
可以用Polars的pl.coalesce函数轻松实现这个需求,核心思路是从右到左依次检查列的非空值,取第一个非null的结果:
import polars as pl data = { "product_id": ["1", "2", "3", "4", "5", "6", "7", "8", "9"], "col1": ["a", "a", "a", "a", "a", "a", "a", "a", "a"], "col2": ["b", None, "b", None, "b", None, "b", None, "b"], "col3": ["c", None, "c", None, None, None, None, None, None], "col4": [None, None, None, None, None, "d", None, None, "d"] } df = pl.DataFrame(data) # 添加desired_column,取col4到col1中第一个非null值 df = df.with_columns( desired_column=pl.coalesce(["col4", "col3", "col2", "col1"]) ) print(df)
原理说明
pl.coalesce的作用是按传入的顺序,返回每行第一个不为null的值。我们将列从最右侧的col4开始,依次传入col3、col2、col1,这样就会优先获取最右侧的非空值,完全匹配需求。
输出结果
执行后得到的DataFrame结构如下:
shape: (9, 6) ┌────────────┬──────┬──────┬──────┬──────┬────────────────┐ │ product_id ┆ col1 ┆ col2 ┆ col3 ┆ col4 ┆ desired_column │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ str ┆ str ┆ str ┆ str ┆ str ┆ str │ ╞════════════╪══════╪══════╪══════╪══════╪════════════════╡ │ 1 ┆ a ┆ b ┆ c ┆ null ┆ c │ │ 2 ┆ a ┆ null ┆ null ┆ null ┆ a │ │ 3 ┆ a ┆ b ┆ c ┆ null ┆ c │ │ 4 ┆ a ┆ null ┆ null ┆ null ┆ a │ │ 5 ┆ a ┆ b ┆ null ┆ null ┆ b │ │ 6 ┆ a ┆ null ┆ null ┆ d ┆ d │ │ 7 ┆ a ┆ b ┆ null ┆ null ┆ b │ │ 8 ┆ a ┆ null ┆ null ┆ null ┆ a │ │ 9 ┆ a ┆ b ┆ null ┆ d ┆ d │ └────────────┴──────┴──────┴──────┴──────┴────────────────┘
内容的提问来源于stack exchange,提问作者Okroshiashvili
相关产品推荐
相关产品推荐

