You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中高效取消分组Polars DataFrame并保留新增列?

高效展开Polars分组后带列表字段的DataFrame

问题背景

我有一个Polars DataFrame,file列存在重复值。先按file列分组聚合所有列,之后新增了folder列,得到含列表类型字段的分组DataFrame。现在需要取消分组,恢复原始行结构,同时保留新增的folder列。实际数据量很大,迭代处理效率极低,需要能作用于整个DataFrame的高效方法。

原始DataFrame结构

filecol1col2
Acell 1cell 2
Bcell 3cell 4
Acell 5cell 6
Bcell 7cell 8

分组并添加folder列后的DataFrame

filecol1col2folder
A[cell 1, cell 5][cell 2, cell 6][file1, file2]
B[cell 3, cell 7][cell 4, cell 8][file1, file2]

期望最终结果

fileheader 1header 2folder
Acell 1cell 2file1
Bcell 3cell 4file1
Acell 5cell 6file2
Bcell 7cell 8file2

我已执行的代码:

dfg = df.groupby('FILE').agg(pl.all())             # 第一次分组聚合
newdf = dfg.with_columns(pl.repeat([file1,file2,file3], dfg.height))  # 添加目标列

高效解决方案

直接使用Polars内置的explode方法,它是矢量化操作,能一次性展开所有列表类型的列,效率远高于逐行/逐列迭代。

优化后的完整代码:

import polars as pl

# 示例原始DataFrame
df = pl.DataFrame({
    "file": ["A", "B", "A", "B"],
    "col1": ["cell 1", "cell 3", "cell 5", "cell 7"],
    "col2": ["cell 2", "cell 4", "cell 6", "cell 8"]
})

# 分组聚合 + 添加folder列(确保folder列表长度和其他聚合列一致)
dfg = df.groupby("file").agg(pl.all())
newdf = dfg.with_columns(pl.lit(["file1", "file2"]).alias("folder"))

# 一次性展开所有列表列,保留file列不变
result = newdf.explode(pl.all().exclude("file"))

# 按需重命名列
result = result.rename({"col1": "header 1", "col2": "header 2"})

print(result)

核心说明

  1. explode(pl.all().exclude("file")):指定展开除file外的所有列表列,确保file列维持分组后的值,同时将col1、col2、folder的列表逐一拆分为单行。
  2. 矢量化操作完全基于Polars底层优化,避免Python层面的循环,处理大数据量时性能优势显著。
  3. 需保证folder列的列表长度与其他聚合后的列表列长度一致,否则explode会抛出长度不匹配的错误。

内容的提问来源于stack exchange,提问作者megha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 01:38:31