如何将表达式字典映射到Polars DataFrame?
问题
有没有简洁高效的方法,将Polars表达式字典应用并求值到DataFrame中?具体要求是根据指定列的值匹配对应表达式,同时利用当前列及其他列的值完成表达式求值?
环境准备
import polars as pl pl.Config.set_fmt_str_lengths(100) # 基础数据 df = pl.DataFrame( { "a": [1,2,3], "b": [2,3,4] } ) # 根据列`a`映射表达式的字典,可基于df的其他列自定义消息 dct = { 1: pl.format("My message about A '{}' and B '{}'", pl.col("a"), pl.col("b")), 2: pl.format("Another message having a={}'", pl.col("a")) }
期望结果
exp_df = pl.DataFrame( { "a": [1,2,3], "b": [2,3,4], "message": ["My message about a=1 and b=2", "Another message having a=2", "Default message having a=3, b=4"] } )
执行print(exp_df)后的输出:
shape: (3, 3) ┌─────┬─────┬─────────────────────────────────┐ │ a ┆ b ┆ message │ │ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ str │ ╞═════╪═════╪═════════════════════════════════╡ │ 1 ┆ 2 ┆ My message about a=1 and b=2 │ │ 2 ┆ 3 ┆ Another message having a=2 │ │ 3 ┆ 4 ┆ Default message having a=3, b=4 │ └─────┴─────┴─────────────────────────────────┘
我的尝试
df_achieved = df.with_columns( [ pl.col("a").map_elements( lambda value: dct.get( value, pl.format("Default message having a={}, b={}", pl.col("a"), pl.col("b")) ) ).alias("message") ] )
执行print(df_achieved)后的输出:
shape: (3, 3) ┌─────┬─────┬───────────────────────────────────────────────────────────────┐ │ a ┆ b ┆ message │ │ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ object │ ╞═════╪═════╪═══════════════════════════════════════════════════════════════╡ │ 1 ┆ 2 ┆ String(My message about A ').str.concat_horizontal([col("a"), │ │ ┆ ┆ String(' and B '), col("b"), String(')… │ │ 2 ┆ 3 ┆ String(Another message having │ │ ┆ ┆ a=).str.concat_horizontal([col("a"), String(')]) │ │ 3 ┆ 4 ┆ String(Default message having │ │ ┆ ┆ a=).str.concat_horizontal([col("a"), String(, b=), col("b")]) │ └─────┴─────┴───────────────────────────────────────────────────────────────┘
解决方案
你的尝试问题在于map_elements返回的是Polars表达式对象而非求值后的字符串,Polars不会自动在map_elements内部执行表达式计算。正确的做法是使用Polars的向量化条件分支,这不仅能得到正确结果,还比逐行处理的map_elements高效得多。
方法一:链式when-then
# 定义默认表达式 default_expr = pl.format("Default message having a={}, b={}", pl.col("a"), pl.col("b")) # 构建条件分支 message_expr = pl.when(pl.col("a") == 1).then(dct[1]) for key, expr in dct.items(): if key != 1: message_expr = message_expr.when(pl.col("a") == key).then(expr) message_expr = message_expr.otherwise(default_expr).alias("message") # 应用到DataFrame df_result = df.with_columns(message_expr)
方法二:用reduce简化条件构建
如果字典键值较多,用functools.reduce可以更简洁地生成条件链:
from functools import reduce default_expr = pl.format("Default message having a={}, b={}", pl.col("a"), pl.col("b")) message_expr = reduce( lambda acc, (k, expr): acc.when(pl.col("a") == k).then(expr), dct.items(), pl.when(False) # 初始空条件 ).otherwise(default_expr).alias("message") df_result = df.with_columns(message_expr)
执行结果
运行print(df_result)会得到符合预期的输出:
shape: (3, 3) ┌─────┬─────┬─────────────────────────────────┐ │ a ┆ b ┆ message │ │ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ str │ ╞═════╪═════╪═════════════════════════════════╡ │ 1 ┆ 2 ┆ My message about A '1' and B '2'│ │ 2 ┆ 3 ┆ Another message having a=2' │ │ 3 ┆ 4 ┆ Default message having a=3, b=4 │ └─────┴─────┴─────────────────────────────────┘
(注:如需调整字符串格式,直接修改字典里的pl.format模板即可)
内容的提问来源于stack exchange,提问作者Michal Chromčák
相关产品推荐
相关产品推荐

