如何为Polars DataFrame添加值出现次数的序号列Occurrence_No?
Polars实现按列表重复行并添加出现序号列
问题背景
给定如下Polars DataFrame:
import polars as pl df = pl.from_repr(""" ┌───────┬───────┬───────┐ │ Col_1 ┆ Col_2 ┆ Col_3 │ │ --- ┆ --- ┆ --- │ │ str ┆ str ┆ i64 │ ╞═══════╪═══════╪═══════╡ │ A ┆ a ┆ 1 │ │ B ┆ b ┆ 2 │ │ C ┆ c ┆ 3 │ │ D ┆ d ┆ 4 │ └───────┴───────┴───────┘ """)
以及列表:
display_list = ['A','B','B','B','C','D','D','A']
期望得到包含Occurrence_No列的输出,该列统计每个Col_1值在结果中的出现序号,同时根据display_list中Col_1的出现次数重复对应行:
shape: (8, 4) ┌───────┬───────┬───────┬───────────────┐ │ Col_1 ┆ Col_2 ┆ Col_3 ┆ Occurrence_No │ │ --- ┆ --- ┆ --- ┆ --- │ │ str ┆ str ┆ i64 ┆ i64 │ ╞═══════╪═══════╪═══════╪═══════════════╡ │ A ┆ a ┆ 1 ┆ 1 │ │ B ┆ b ┆ 2 ┆ 1 │ │ B ┆ b ┆ 2 ┆ 2 │ │ B ┆ b ┆ 2 ┆ 3 │ │ C ┆ c ┆ 3 ┆ 1 │ │ D ┆ d ┆ 4 ┆ 1 │ │ D ┆ d ┆ 4 ┆ 2 │ │ A ┆ a ┆ 1 ┆ 1 │ └───────┴───────┴───────┴───────────────┘
现有代码能实现行重复,但无法生成Occurrence_No列:
df = df.with_columns(pl.col('Col_1').map_elements(lambda x: display_list.count(x)).alias('occur')) df = df.select(pl.exclude('occur').repeat_by('occur').explode())
解决方案
方法一:基于原DataFrame生成序号序列
先统计每个Col_1在列表中的出现次数,再生成对应范围的序号序列,最后一起展开:
import polars as pl # 统计每个Col_1的出现次数 count_df = pl.DataFrame({'Col_1': display_list}).group_by('Col_1').len().rename({'len': 'occur'}) df = df.join(count_df, on='Col_1') # 生成Occurrence_No并展开所有列 result = df.with_columns( Occurrence_No=pl.int_range(1, pl.col('occur') + 1, dtype=pl.Int64) ).explode(pl.all().exclude('occur')).drop('occur') print(result)
方法二:直接基于列表构建结果(更直观)
从display_list直接构建DataFrame,生成分组序号后关联原DataFrame,这种方式还能保留列表的原始顺序:
import polars as pl # 从列表构建DataFrame并生成Occurrence_No list_df = pl.DataFrame({'Col_1': display_list}).with_columns( Occurrence_No=pl.int_range(1, pl.count() + 1).over('Col_1') ) # 关联原数据得到完整结果 result = list_df.join(df, on='Col_1') print(result)
两种方法都能得到符合要求的输出,方法二更贴合需求逻辑,且不需要额外处理重复和展开步骤。
内容的提问来源于stack exchange,提问作者Shankze
相关产品推荐
相关产品推荐

