You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Polars DataFrame中另一列的数值实现explode并重复行

如何基于Polars DataFrame中另一列的数值实现explode并重复行

嘿,我来帮你搞定Polars里根据列值重复行的需求!先看你给出的DataFrame,NumberInst列混着字符串类型的数字和None,咱们得先把它转成数值类型,再根据里面的数值来重复对应行数,对吧?

首先,先把原始DataFrame里的NumberInst转成整数类型,同时把None处理成Polars能识别的缺失值:

import polars as pl

df = pl.from_repr("""
┌────────────┬────────────┬───────┐
│ numberType ┆ NumberInst ┆ Type  │
│ ---        ┆ ---        ┆ ---   │
│ i64        ┆ str        ┆ str   │
╞════════════╪════════════╪═══════╡
│ 1          ┆ None       ┆ Car   │
│ 2          ┆ 1          ┆ Bus   │
│ 3          ┆ 1          ┆ Plane │
└────────────┴────────────┴───────┘
""")

# 转换NumberInst为整数类型,自动将None转为null
df = df.with_columns(
    pl.col("NumberInst").cast(pl.Int64, strict=False)
)

接下来分两种常见场景给你方案:

场景1:保留NumberInst为null的行(默认重复1次)

如果希望None对应的行保留1行,其他行按NumberInst的数值重复对应次数(比如数值是2就把原行变成2行),可以这么写:

result = df.with_columns(
    # 把null填充为1,确保这类行至少保留1行
    pl.col("NumberInst").fill_null(1)
).explode(
    # 生成对应长度的占位符数组,explode后自动重复行
    pl.int_range(0, pl.col("NumberInst")).alias("temp")
).drop("temp")  # 删掉临时生成的占位符列

print(result)

运行后得到的结果:

┌────────────┬────────────┬───────┐
│ numberType ┆ NumberInst ┆ Type  │
│ ---        ┆ ---        ┆ ---   │
│ i64        ┆ i64        ┆ str   │
╞════════════╪════════════╪═══════╡
│ 1          ┆ null       ┆ Car   │
│ 2          ┆ 1          ┆ Bus   │
│ 3          ┆ 1          ┆ Plane │
└────────────┴────────────┴───────┘

场景2:丢弃NumberInst为null的行

如果不想保留None对应的行,就把null填充为0,再过滤掉重复次数为0的行:

result = df.with_columns(
    pl.col("NumberInst").fill_null(0)
).filter(
    pl.col("NumberInst") > 0  # 去掉不需要保留的行
).explode(
    pl.int_range(0, pl.col("NumberInst")).alias("temp")
).drop("temp")

print(result)

输出就只剩下有有效重复次数的行:

┌────────────┬────────────┬───────┐
│ numberType ┆ NumberInst ┆ Type  │
│ ---        ┆ ---        ┆ ---   │
│ i64        ┆ i64        ┆ str   │
╞════════════╪════════════╪═══════╡
│ 2          ┆ 1          ┆ Bus   │
│ 3          ┆ 1          ┆ Plane │
└────────────┴────────────┴───────┘

核心逻辑说明

这里的关键是用pl.int_range(0, pl.col("NumberInst"))生成一个长度等于NumberInst数值的整数数组——比如数值是1就生成[0],数值是3就生成[0,1,2]。然后explode方法会把这个数组“炸开”,数组里有几个元素,原行就会被重复几次,最后删掉临时的占位符列就搞定啦。

这种方法用的是Polars的矢量化运算,比循环或者apply高效太多,处理大数据量也不会卡顿~

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 12:18:08