You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Dash+Polars应用内存持续增长问题求助

Polars迁移后Dash应用内存持续膨胀问题排查与解决建议

问题概述

将本地Dash应用从pandas+duckdb迁移到Polars后,出现严重内存泄漏问题:pythonw.exe进程内存随每次回调持续增长,初始约100MB,每次回调增加5MB左右,内存达到1500MB时仍无停止趋势。对比测试显示:

  • Polars模式(polars_check=True):初始内存98MB,100次迭代后增至261MB
  • pandas模式(polars_check=False):内存始终稳定在98MB

复现示例代码

import pathlib, os, shutil
import polars as pl, pandas as pd, numpy as np, datetime as dt

from dash import Dash, dcc, html, Input, Output
import plotly.graph_objects as go


#Check-input
polars_check = True ### Whether the example returns with polars or with pandas.

if polars_check: #To accomdate the slower data retrieval with pandas.
    interval_time = 3E3
else:
    interval_time = 3E3

#Constants
folder = pathlib.Path(r'C:\PerovskiteCell example')

n_files = 100 #Number of files in folder
n_lines = 500000 #Number of total lines in folder
n_cols = 25


#Generating sample data in example folder (Only once).
if not folder.exists():

    size = int(n_lines / n_files)
    col = np.linspace(-1E3, 1E3, num=size)

    df = pl.DataFrame({f'col{n}': col for n in range(n_cols)})

    # Creating folder & files
    os.makedirs(folder)

    f_path0 = folder.joinpath('0.csv')
    df.write_csv(f_path0)

    for n in range(1, n_files):
        shutil.copy2(f_path0, folder.joinpath(f'{n}.csv'))


#Functions
def pl_data():
    """Retrieves data via the polars route"""
    
    lf = (pl.scan_csv(folder.joinpath(f'{n}.csv'),
                                schema={f'col{n}': pl.Float64 for n in range(n_cols)})
                            
                            .select(pl.all().get(n)) for n in range(n_files))
    
    lf = pl.concat(lf)
    lf = lf.select('col0', 'col1')

    return lf.collect()


def pd_data():
    """Retrieves data via the pandas route"""

    dfs = (pd.read_csv(folder.joinpath(f'{n}.csv'), usecols=['col0', 'col1']).iloc[n:n+1]
                                         for n in range(n_files))
    
    return pd.concat(dfs, ignore_index=True)



#App (initialization)
app = Dash()
app.layout = html.Div([dcc.Graph(id='graph'),
                        dcc.Interval(id = 'check', 
                                        interval = interval_time,
                                        max_intervals = 100)])


@app.callback(
    Output('graph', 'figure'),
    Input('check', 'n_intervals'))

def plot(_):

    #Data retrieval
    if polars_check:
        df = pl_data()
    else:
        df = pd_data()

    #Plotting
    fig = go.Figure()
    trace = go.Scattergl(x = list(df['col0']), y=list(df['col1']), mode='lines+markers')

    fig.add_trace(trace)
    fig.update_xaxes(title = str(dt.datetime.now()))

    return fig


if __name__ == '__main__':
    app.run(debug=False, port = 8050)

问题分析

Polars内存膨胀的核心原因通常与以下几点相关:

  1. 对象引用未被及时回收:Polars的LazyFrame/DataFrame底层基于Rust实现,若Python层的引用未被正确释放,会导致Rust侧内存无法回收
  2. LazyFrame操作的内存残留:循环生成LazyFrame再concat的过程中,可能产生未被清理的中间对象
  3. 数据转换的额外开销:将Polars Series转换为Python list时,会生成额外内存对象,未及时回收会累积占用内存

解决方案建议

1. 显式触发垃圾回收

在回调函数末尾手动触发Python垃圾回收,强制释放未被引用的Polars对象:

import gc

@app.callback(
    Output('graph', 'figure'),
    Input('check', 'n_intervals'))
def plot(_):
    df = pl_data() if polars_check else pd_data()
    
    fig = go.Figure()
    trace = go.Scattergl(x = list(df['col0']), y=list(df['col1']), mode='lines+markers')
    fig.add_trace(trace)
    fig.update_xaxes(title = str(dt.datetime.now()))
    
    # 显式清理对象并触发GC
    del df
    gc.collect()
    
    return fig

2. 优化Polars数据加载逻辑

避免循环生成LazyFrame再concat的低效方式,改用批量读取+筛选的逻辑,减少中间对象生成:

def pl_data():
    # 批量读取所有CSV文件并添加文件索引
    lf = pl.scan_csv(
        folder.glob("*.csv"),
        schema={f'col{n}': pl.Float64 for n in range(n_cols)},
        with_filename=True
    ).with_columns(
        file_idx=pl.col("filename").str.extract(r'(\d+)\.csv').cast(pl.Int32)
    )
    # 按文件索引筛选对应行,匹配pandas逻辑
    lf = lf.with_row_index().filter(pl.col("index") == pl.col("file_idx"))
    return lf.select('col0', 'col1').collect()

3. 优化Plotly数据传递方式

避免将Polars Series转换为Python list,改用to_numpy()直接传递numpy数组,减少内存开销:

trace = go.Scattergl(x=df['col0'].to_numpy(), y=df['col1'].to_numpy(), mode='lines+markers')

4. 禁用Polars字符串缓存

Polars默认启用字符串缓存,可能导致内存累积,可在初始化时禁用:

import polars as pl
pl.enable_string_cache(False)

5. 升级/回退Polars版本

部分版本的Polars存在已知内存泄漏问题,建议升级到最新稳定版(执行pip install --upgrade polars),或回退到之前验证过的稳定版本。

6. 限制对象作用域

确保Polars对象仅在回调函数内部作用域存在,避免全局引用或闭包持有对象导致无法回收。


内容的提问来源于stack exchange,提问作者PerovskiteCell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 05:30:17