You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Polars为DataFrame添加满足条件的首个后续Time列

Polars:为每行查找后续首个满足Value小于当前值的Time(带前瞻限制)

我有一个按Time列排序的Polars DataFrame:

import polars as pl

df = pl.from_repr("""
┌──────┬───────┐
│ Time ┆ Value │
│ ---  ┆ ---   │
│ i64  ┆ i64   │
╞══════╪═══════╡
│ 1    ┆ 100   │
│ 2    ┆ 75    │
│ 3    ┆ 70    │
│ 4    ┆ 105   │
│ 5    ┆ 140   │
│ 6    ┆ 220   │
│ 7    ┆ 65    │
│ 8    ┆ 180   │
│ 9    ┆ 150   │
└──────┴───────┘
""")

需要新增名为NewColumn的列,存储首个满足Value小于当前行Value的后续Time值。

补充说明

原始数据集包含超过1000万行,部分操作易导致内存溢出,因此可设置前瞻限制(例如最多往后查找100个Time值),返回范围内首个满足条件的Time。

预期结果示例

|  Time  | Value  | NewColumn |
|   1    |   100  |       2   | >> next Time with Value lower than 100
|   2    |    75  |       3   | >> next Time with Value lower than 75
|   3    |    70  |       7   | >> next Time with Value lower than 70
|   4    |   105  |       7   | >> next Time with Value lower than 105
|   5    |   140  |       7   | >> next Time with Value lower than 140
|   6    |   220  |       7   | >> next Time with Value lower than 220
|   7    |    65  |    null   | >> next Time with Value lower than 65
|   8    |   180  |       9   | >> next Time with Value lower than 180
|   9    |   150  |    null   | >> next Time with Value lower than 150

尝试的思路及遇到的问题

我尝试先创建临时列Slice_of_Value,存储从下一行到最后一行的Value切片,再推断出满足“小于当前Value”条件的首个匹配位置,但遇到了两个问题:

  • 步骤1:创建Slice_of_Value列时,最初尝试的代码无法在slice()中使用pl.col(),只能改用移位拼接的方式实现;
  • 步骤2:计算Position列时,触发“list.eval中不允许使用命名列”的报错。

基准测试结果

我在真实数据集上测试了3种方案:

  • 方案1(基于shift和coalesce):在1000万行数据上耗时0.83s,性能最优且代码简洁;
  • 方案2:耗时3.41s;
  • 方案3:因内存溢出无法运行。

不同前瞻大小下的性能对比也验证了方案1的优势。


内容的提问来源于stack exchange,提问作者Jona Rodrigues

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 21:52:00