You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Polars+Python处理大LazyFrames时的内存优化方案

问题

我的应用部署在配置12GB内存的Kubernetes容器中,需要处理一份包含2亿条记录、磁盘大小约9GB的CSV文件,格式如下:

id,aligned_value,date_reference,model,predict,segment
00000000001,564,20240520,XPTO,,3
00000000002,741,20240520,XPTO,,6
00000000003,503,20240520,XPTO,,5
00000000004,200,20240520,XPTO,,0

我尝试按aligned_value(取值0-1000)和segment(取值0-6)分组聚合统计数量,但执行以下代码时内存持续增长直至容器被终止:

def get_summarized_base_df(self, filepath: str) -> pl.DataFrame:
    """
    Summarizes the base on a dataframe grouped by all fields included
    on this report's setup
    """

    # 此场景下返回的字段列表仅为 ["aligned_value", "segment"]
    required_fields = self.list_all_required_fields()

    base_lf = pl.scan_csv(filepath)

    summarized_base_df = base_lf.group_by(required_fields).agg(pl.count()).collect()

    return summarized_base_df

我曾尝试设置环境变量POLARS_MAX_MEMORY_MIB限制Polars内存使用,但无明显效果。请问是否有可降低内存占用的参数?或是我使用框架的方式有误?

补充信息:Python版本3.10.11,Polars版本0.20.18

解决方案
  • 仅加载必需字段:当前代码会加载CSV所有列,但实际仅用到aligned_value和segment,在scan_csv中指定columns参数可大幅削减内存占用:

    base_lf = pl.scan_csv(filepath, columns=["aligned_value", "segment"])
    

    避免加载id、date_reference等无关数据,直接降低内存开销。

  • 显式指定紧凑数据类型:aligned_value取值0-1000,可用UInt16类型;segment取值0-6,可用UInt8类型,比默认整数类型占用更少内存:

    schema = {
        "aligned_value": pl.UInt16,
        "segment": pl.UInt8
    }
    base_lf = pl.scan_csv(filepath, columns=["aligned_value", "segment"], schema=schema)
    
  • 调整Polars内存配置:旧版本Polars中POLARS_MAX_MEMORY_MIB对懒加载模式的控制有限,可通过代码直接设置内存限制并启用磁盘缓存:

    import polars as pl
    # 设置内存限制为10GB(单位:MiB)
    pl.Config.set_memory_limit(10240)
    # 指定临时缓存路径,内存不足时将中间数据写入磁盘
    pl.Config.set_temporary_cache_path("/tmp/polars_cache")
    pl.Config.set_allow_cache(True)
    
  • 利用取值范围小的特性优化聚合:由于aligned_value和segment的组合仅7007种(1001*7),可预先初始化计数字典,通过逐行读取统计替代分组聚合,进一步降低内存波动:

    from collections import defaultdict
    
    def count_groups(filepath: str):
        counter = defaultdict(int)
        with open(filepath, "r") as f:
            next(f)  # 跳过表头
            for line in f:
                parts = line.strip().split(",")
                aligned = int(parts[1])
                segment = int(parts[5])
                counter[(aligned, segment)] += 1
        # 转换为Polars DataFrame
        return pl.DataFrame(
            [(k[0], k[1], v) for k, v in counter.items()],
            schema=["aligned_value", "segment", "count"]
        )
    

    这种方式内存占用极低,仅需存储计数字典即可。

内容的提问来源于stack exchange,提问作者João Luiz dos Reis Santos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 11:08:11