You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何根据截断值拆分DataFrame列表列并同步关联字段?

解决Polars DataFrame中按规则拆分列表列并同步关联字段的问题

问题描述

需要将Polars DataFrame中的time_deltas列按以下规则拆分为子列表:每个子列表仅允许第一个元素大于指定截断值(示例中为3),同时同步拆分other_field等关联字段,确保拆分后的数据对应关系不变。

解决方案

通过Polars的map_elements方法结合自定义拆分函数实现需求,核心是同步遍历time_deltas和other_field的元素,按照规则拆分并保持对应关系。

完整代码实现

import polars as pl

data = {
    "uid": ["Alice", "Bob", "Charlie"],
    "time_deltas": [
        [4,2, 3],
        [1,1, 4, 8, 3],
        [1,1, 7, 3, 2],
    ],
    "other_field": [["x", "y", "z"], ["x", "y", "z", "x", "y"], ["x", "y", "z", "x", "y"]]
}

df = pl.DataFrame(data)
cutoff = 3

def split_lists(row):
    time_deltas = row["time_deltas"]
    other_field = row["other_field"]
    
    split_time = []
    split_other = []
    
    if not time_deltas:
        return {"split_time": [], "split_other": []}
    
    # 初始化当前子列表
    current_time = [time_deltas[0]]
    current_other = [other_field[0]]
    
    for t, o in zip(time_deltas[1:], other_field[1:]):
        if t > cutoff:
            # 当前元素大于截断值,结束当前子列表,新建子列表
            split_time.append(current_time)
            split_other.append(current_other)
            current_time = [t]
            current_other = [o]
        else:
            # 添加到当前子列表
            current_time.append(t)
            current_other.append(o)
    
    # 添加最后一个子列表
    split_time.append(current_time)
    split_other.append(current_other)
    
    return {"split_time": split_time, "split_other": split_other}

# 应用拆分函数并展开结果
result_df = df.with_columns(
    pl.struct(["time_deltas", "other_field"]).map_elements(split_lists).alias("split_result")
).unnest("split_result").explode(["split_time", "split_other"])

print(result_df)

代码解释

  1. 自定义拆分函数split_lists:

    • 接收包含time_deltas和other_field的行结构体,同步遍历两个列表的元素。
    • 维护当前子列表current_time和current_other,当遇到大于截断值的元素时,将当前子列表存入结果,新建子列表并放入当前元素;否则将元素添加到当前子列表。
    • 遍历结束后,将最后一个子列表存入结果。
  2. Polars数据处理流程:

    • 使用pl.struct将需要同步处理的列打包,通过map_elements应用拆分函数,得到包含拆分后列表的结构体列。
    • 用unnest拆分结构体列,再通过explode将子列表展开为单独的行,保持uid与拆分后子列表的对应关系。

输出结果

shape: (6, 4)
┌─────────┬─────────────────┬─────────────────────┬─────────────┐
│ uid     ┆ time_deltas     ┆ other_field         ┆ split_time  │
│ ---     ┆ ---             ┆ ---                 ┆ ---         │
│ str     ┆ list[i64]       ┆ list[str]           ┆ list[i64]   │
╞═════════╪═════════════════╪═════════════════════╪═════════════╡
│ Alice   ┆ [4, 2, 3]       ┆ ["x", "y", "z"]     ┆ [4, 2, 3]   │
│ Bob     ┆ [1, 1, 4, 8, 3] ┆ ["x", "y", "z", "x… ┆ [1, 1]      │
│ Bob     ┆ [1, 1, 4, 8, 3] ┆ ["x", "y", "z", "x… ┆ [4]         │
│ Bob     ┆ [1, 1, 4, 8, 3] ┆ ["x", "y", "z", "x… ┆ [8, 3]      │
│ Charlie ┆ [1, 1, 7, 3, 2] ┆ ["x", "y", "z", "x… ┆ [1, 1]      │
│ Charlie ┆ [1, 1, 7, 3, 2] ┆ ["x", "y", "z", "x… ┆ [7, 3, 2]   │
└─────────┴─────────────────┴─────────────────────┴─────────────┘

(注:split_other列内容与split_time一一对应,例如Bob的split_other依次为["x","y"]、["z"]、["x","y"])

内容的提问来源于stack exchange,提问作者Anton Gomes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 09:42:03