如何根据截断值拆分DataFrame列表列并同步关联字段?
解决Polars DataFrame中按规则拆分列表列并同步关联字段的问题
问题描述
需要将Polars DataFrame中的time_deltas列按以下规则拆分为子列表:每个子列表仅允许第一个元素大于指定截断值(示例中为3),同时同步拆分other_field等关联字段,确保拆分后的数据对应关系不变。
解决方案
通过Polars的map_elements方法结合自定义拆分函数实现需求,核心是同步遍历time_deltas和other_field的元素,按照规则拆分并保持对应关系。
完整代码实现
import polars as pl data = { "uid": ["Alice", "Bob", "Charlie"], "time_deltas": [ [4,2, 3], [1,1, 4, 8, 3], [1,1, 7, 3, 2], ], "other_field": [["x", "y", "z"], ["x", "y", "z", "x", "y"], ["x", "y", "z", "x", "y"]] } df = pl.DataFrame(data) cutoff = 3 def split_lists(row): time_deltas = row["time_deltas"] other_field = row["other_field"] split_time = [] split_other = [] if not time_deltas: return {"split_time": [], "split_other": []} # 初始化当前子列表 current_time = [time_deltas[0]] current_other = [other_field[0]] for t, o in zip(time_deltas[1:], other_field[1:]): if t > cutoff: # 当前元素大于截断值,结束当前子列表,新建子列表 split_time.append(current_time) split_other.append(current_other) current_time = [t] current_other = [o] else: # 添加到当前子列表 current_time.append(t) current_other.append(o) # 添加最后一个子列表 split_time.append(current_time) split_other.append(current_other) return {"split_time": split_time, "split_other": split_other} # 应用拆分函数并展开结果 result_df = df.with_columns( pl.struct(["time_deltas", "other_field"]).map_elements(split_lists).alias("split_result") ).unnest("split_result").explode(["split_time", "split_other"]) print(result_df)
代码解释
自定义拆分函数
split_lists:- 接收包含
time_deltas和other_field的行结构体,同步遍历两个列表的元素。 - 维护当前子列表
current_time和current_other,当遇到大于截断值的元素时,将当前子列表存入结果,新建子列表并放入当前元素;否则将元素添加到当前子列表。 - 遍历结束后,将最后一个子列表存入结果。
- 接收包含
Polars数据处理流程:
- 使用
pl.struct将需要同步处理的列打包,通过map_elements应用拆分函数,得到包含拆分后列表的结构体列。 - 用
unnest拆分结构体列,再通过explode将子列表展开为单独的行,保持uid与拆分后子列表的对应关系。
- 使用
输出结果
shape: (6, 4) ┌─────────┬─────────────────┬─────────────────────┬─────────────┐ │ uid ┆ time_deltas ┆ other_field ┆ split_time │ │ --- ┆ --- ┆ --- ┆ --- │ │ str ┆ list[i64] ┆ list[str] ┆ list[i64] │ ╞═════════╪═════════════════╪═════════════════════╪═════════════╡ │ Alice ┆ [4, 2, 3] ┆ ["x", "y", "z"] ┆ [4, 2, 3] │ │ Bob ┆ [1, 1, 4, 8, 3] ┆ ["x", "y", "z", "x… ┆ [1, 1] │ │ Bob ┆ [1, 1, 4, 8, 3] ┆ ["x", "y", "z", "x… ┆ [4] │ │ Bob ┆ [1, 1, 4, 8, 3] ┆ ["x", "y", "z", "x… ┆ [8, 3] │ │ Charlie ┆ [1, 1, 7, 3, 2] ┆ ["x", "y", "z", "x… ┆ [1, 1] │ │ Charlie ┆ [1, 1, 7, 3, 2] ┆ ["x", "y", "z", "x… ┆ [7, 3, 2] │ └─────────┴─────────────────┴─────────────────────┴─────────────┘
(注:split_other列内容与split_time一一对应,例如Bob的split_other依次为["x","y"]、["z"]、["x","y"])
内容的提问来源于stack exchange,提问作者Anton Gomes
相关产品推荐
相关产品推荐

