You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用cum_fold或cum_reduce创建有状态DataFrame列?

解决方案

可以通过Polars的累积求和结合字符串格式化实现需求,核心思路是用累积求和生成分组标识,再基于标识生成目标列内容:

  1. 对列a执行累积求和:每次遇到1时求和值递增,遇到0时保持当前值,这样就能把连续的1及后续0归为同一组;
  2. 根据分组标识,格式化生成new_valueX格式的字符串。

代码实现

import polars as pl

# 初始化原始DataFrame
a = pl.DataFrame({'a': [1, 0, 0, 0, 1, 0, 0, 1]})

# 添加目标列b
result = a.with_columns(
    pl.col('a').cumsum().map_elements(lambda x: f"new_value{x}", return_dtype=pl.String).alias('b')
)

print(result)

输出结果

shape: (8, 2)
┌─────┬────────────┐
│ a   ┆ b          │
│ --- ┆ ---        │
│ i64 ┆ str        │
╞═════╪════════════╡
│ 1   ┆ new_value1 │
│ 0   ┆ new_value1 │
│ 0   ┆ new_value1 │
│ 0   ┆ new_value1 │
│ 1   ┆ new_value2 │
│ 0   ┆ new_value2 │
│ 0   ┆ new_value2 │
│ 1   ┆ new_value3 │
└─────┴────────────┘

高性能优化方案

如果处理大数据集,建议避免逐行处理的map_elements,改用Polars向量化操作提升效率:

result = a.with_columns(
    (pl.lit("new_value") + pl.col('a').cumsum().cast(pl.String)).alias('b')
)

内容的提问来源于stack exchange,提问作者Kaster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 01:20:38