You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars DataFrame unpivot过程中移除空值?

在Polars中Unpivot时直接过滤空值(避免内存过载)

针对含大量空值的大型Polars DataFrame(尤其是方阵结构),直接先执行unpivot再过滤空值会导致内存爆炸,你可以通过逐行提取非空条目再展开的方式,在unpivot过程中直接过滤空值,大幅降低内存占用。

解决方案代码

import polars as pl

# 示例数据(实际为16万行×16万列、100GB的数据集)
df = pl.DataFrame({
    "A": [None, 2, 3],
    "B": [None, None, 2],
    "C": [None, None, None], 
    "names": ["A", "B", "C"]
})

# 核心逻辑:逐行保留非空的(列名, 值)结构,再展开
result = (
    df
    .with_columns(
        # 对所有非index列,生成(列名, 距离值)的结构,过滤掉空值后收集为数组
        pl.struct(
            pl.col("*").exclude("names").name.map(lambda col_name: pl.lit(col_name).alias("names_2")),
            pl.col("*").exclude("names").alias("distance")
        )
        .filter(pl.col("distance").is_not_null())
        .alias("temp_entries")
    )
    .explode("temp_entries")  # 展开数组为行
    .select(
        "names",
        pl.col("temp_entries").struct.field("names_2"),
        pl.col("temp_entries").struct.field("distance")
    )
)

print(result)

输出结果

shape: (3, 3)
┌───────┬─────────┬──────────┐
│ names ┆ names_2 ┆ distance │
│ ---   ┆ ---     ┆ ---      │
│ str   ┆ str     ┆ i64      │
╞═══════╪═════════╪══════════╡
│ A     ┆ B       ┆ 2        │
│ A     ┆ C       ┆ 3        │
│ B     ┆ C       ┆ 2        │
└───────┴─────────┴──────────┘

效率说明

  • 传统unpivot+drop_nulls会先生成所有行(包括空值),对于16万×16万的方阵,会产生2.56×10¹⁰行的中间数据,直接耗尽内存。
  • 本方法是逐行处理:只保留每行中的非空条目,再将这些条目展开为行,内存占用仅与最终非空结果的大小成正比,完全避免了空值行的生成。

内容的提问来源于stack exchange,提问作者Nils R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 20:37:05