如何使用cum_fold或cum_reduce创建有状态DataFrame列?
解决方案
可以通过Polars的累积求和结合字符串格式化实现需求,核心思路是用累积求和生成分组标识,再基于标识生成目标列内容:
- 对列
a执行累积求和:每次遇到1时求和值递增,遇到0时保持当前值,这样就能把连续的1及后续0归为同一组; - 根据分组标识,格式化生成
new_valueX格式的字符串。
代码实现
import polars as pl # 初始化原始DataFrame a = pl.DataFrame({'a': [1, 0, 0, 0, 1, 0, 0, 1]}) # 添加目标列b result = a.with_columns( pl.col('a').cumsum().map_elements(lambda x: f"new_value{x}", return_dtype=pl.String).alias('b') ) print(result)
输出结果
shape: (8, 2) ┌─────┬────────────┐ │ a ┆ b │ │ --- ┆ --- │ │ i64 ┆ str │ ╞═════╪════════════╡ │ 1 ┆ new_value1 │ │ 0 ┆ new_value1 │ │ 0 ┆ new_value1 │ │ 0 ┆ new_value1 │ │ 1 ┆ new_value2 │ │ 0 ┆ new_value2 │ │ 0 ┆ new_value2 │ │ 1 ┆ new_value3 │ └─────┴────────────┘
高性能优化方案
如果处理大数据集,建议避免逐行处理的map_elements,改用Polars向量化操作提升效率:
result = a.with_columns( (pl.lit("new_value") + pl.col('a').cumsum().cast(pl.String)).alias('b') )
内容的提问来源于stack exchange,提问作者Kaster
相关产品推荐
相关产品推荐

