Dash+Polars应用内存持续增长问题求助
Polars迁移后Dash应用内存持续膨胀问题排查与解决建议
问题概述
将本地Dash应用从pandas+duckdb迁移到Polars后,出现严重内存泄漏问题:pythonw.exe进程内存随每次回调持续增长,初始约100MB,每次回调增加5MB左右,内存达到1500MB时仍无停止趋势。对比测试显示:
- Polars模式(
polars_check=True):初始内存98MB,100次迭代后增至261MB - pandas模式(
polars_check=False):内存始终稳定在98MB
复现示例代码
import pathlib, os, shutil import polars as pl, pandas as pd, numpy as np, datetime as dt from dash import Dash, dcc, html, Input, Output import plotly.graph_objects as go #Check-input polars_check = True ### Whether the example returns with polars or with pandas. if polars_check: #To accomdate the slower data retrieval with pandas. interval_time = 3E3 else: interval_time = 3E3 #Constants folder = pathlib.Path(r'C:\PerovskiteCell example') n_files = 100 #Number of files in folder n_lines = 500000 #Number of total lines in folder n_cols = 25 #Generating sample data in example folder (Only once). if not folder.exists(): size = int(n_lines / n_files) col = np.linspace(-1E3, 1E3, num=size) df = pl.DataFrame({f'col{n}': col for n in range(n_cols)}) # Creating folder & files os.makedirs(folder) f_path0 = folder.joinpath('0.csv') df.write_csv(f_path0) for n in range(1, n_files): shutil.copy2(f_path0, folder.joinpath(f'{n}.csv')) #Functions def pl_data(): """Retrieves data via the polars route""" lf = (pl.scan_csv(folder.joinpath(f'{n}.csv'), schema={f'col{n}': pl.Float64 for n in range(n_cols)}) .select(pl.all().get(n)) for n in range(n_files)) lf = pl.concat(lf) lf = lf.select('col0', 'col1') return lf.collect() def pd_data(): """Retrieves data via the pandas route""" dfs = (pd.read_csv(folder.joinpath(f'{n}.csv'), usecols=['col0', 'col1']).iloc[n:n+1] for n in range(n_files)) return pd.concat(dfs, ignore_index=True) #App (initialization) app = Dash() app.layout = html.Div([dcc.Graph(id='graph'), dcc.Interval(id = 'check', interval = interval_time, max_intervals = 100)]) @app.callback( Output('graph', 'figure'), Input('check', 'n_intervals')) def plot(_): #Data retrieval if polars_check: df = pl_data() else: df = pd_data() #Plotting fig = go.Figure() trace = go.Scattergl(x = list(df['col0']), y=list(df['col1']), mode='lines+markers') fig.add_trace(trace) fig.update_xaxes(title = str(dt.datetime.now())) return fig if __name__ == '__main__': app.run(debug=False, port = 8050)
问题分析
Polars内存膨胀的核心原因通常与以下几点相关:
- 对象引用未被及时回收:Polars的LazyFrame/DataFrame底层基于Rust实现,若Python层的引用未被正确释放,会导致Rust侧内存无法回收
- LazyFrame操作的内存残留:循环生成LazyFrame再concat的过程中,可能产生未被清理的中间对象
- 数据转换的额外开销:将Polars Series转换为Python list时,会生成额外内存对象,未及时回收会累积占用内存
解决方案建议
1. 显式触发垃圾回收
在回调函数末尾手动触发Python垃圾回收,强制释放未被引用的Polars对象:
import gc @app.callback( Output('graph', 'figure'), Input('check', 'n_intervals')) def plot(_): df = pl_data() if polars_check else pd_data() fig = go.Figure() trace = go.Scattergl(x = list(df['col0']), y=list(df['col1']), mode='lines+markers') fig.add_trace(trace) fig.update_xaxes(title = str(dt.datetime.now())) # 显式清理对象并触发GC del df gc.collect() return fig
2. 优化Polars数据加载逻辑
避免循环生成LazyFrame再concat的低效方式,改用批量读取+筛选的逻辑,减少中间对象生成:
def pl_data(): # 批量读取所有CSV文件并添加文件索引 lf = pl.scan_csv( folder.glob("*.csv"), schema={f'col{n}': pl.Float64 for n in range(n_cols)}, with_filename=True ).with_columns( file_idx=pl.col("filename").str.extract(r'(\d+)\.csv').cast(pl.Int32) ) # 按文件索引筛选对应行,匹配pandas逻辑 lf = lf.with_row_index().filter(pl.col("index") == pl.col("file_idx")) return lf.select('col0', 'col1').collect()
3. 优化Plotly数据传递方式
避免将Polars Series转换为Python list,改用to_numpy()直接传递numpy数组,减少内存开销:
trace = go.Scattergl(x=df['col0'].to_numpy(), y=df['col1'].to_numpy(), mode='lines+markers')
4. 禁用Polars字符串缓存
Polars默认启用字符串缓存,可能导致内存累积,可在初始化时禁用:
import polars as pl pl.enable_string_cache(False)
5. 升级/回退Polars版本
部分版本的Polars存在已知内存泄漏问题,建议升级到最新稳定版(执行pip install --upgrade polars),或回退到之前验证过的稳定版本。
6. 限制对象作用域
确保Polars对象仅在回调函数内部作用域存在,避免全局引用或闭包持有对象导致无法回收。
内容的提问来源于stack exchange,提问作者PerovskiteCell
相关产品推荐
相关产品推荐

