使用mplfinance/matplotlib疑似内存泄漏,如何解决?
解决K线图批量生成中的内存泄漏问题
首先得说,你遇到的这种内存持续累积的问题,在批量生成matplotlib/mplfinance图表时真的很常见,尤其是在Jupyter Notebook这种交互式环境里。结合你的代码和运行环境,我整理了几个针对性的解决方案,按优先级排序:
1. 先脱离Jupyter,改用普通Python脚本运行
Jupyter的内核会默认保留所有变量的引用——哪怕你手动del了,gc.collect()的效果在交互式环境里也会大打折扣。把三个单元格的代码合并成一个.py脚本直接运行,是解决内存累积最快速的第一步。
你只需要把代码里的路径调整好,然后用python your_script.py命令启动,完全避开Jupyter的变量缓存问题。
2. 复用Figure/Axes对象,避免每次创建新实例
现在你的plot_candle函数每次调用mpf.plot都会生成全新的Figure和Axes对象,这是内存泄漏的核心原因之一。我们可以提前创建好一套可复用的绘图容器,每次只更新数据:
# 在循环外提前创建可复用的fig和axes fig, axes = mpf.plot(data1[symbols1[0]].head(tf), type='candle', style=s, volume=True, axisoff=True, figratio=(1,1), returnfig=True) plt.close(fig) # 先关闭,避免自动显示 def plot_candle(i,j,data,symbols,s,direct,img_size, tf, fig, axes): # 用loc切片避免链式索引产生不必要的DataFrame副本 data_temp = data.loc[i-tf:i, symbols[j]] # 复用已有的fig和axes更新图表 mpf.plot(data_temp, type='candle', style=s, volume=True, axisoff=True, figratio=(1,1), fig=fig, axes=axes) buf = io.BytesIO() # 直接用matplotlib的savefig,避免mpf封装带来的额外开销 fig.savefig(buf, pad_inches=0, bbox_inches='tight') buf.seek(0) # 用with上下文管理器自动关闭Image对象 with Image.open(buf) as im: im_resized = im.resize((img_size, img_size)) im_resized.save(f"{direct}/{symbols[j]}/{i-tf+1}.png", "PNG") # 清空axes内容,为下一次绘图准备 for ax in axes: ax.clear() # 手动删除临时对象,触发垃圾回收 del data_temp, buf, im_resized gc.collect()
然后在循环调用时传入提前创建好的fig和axes:
for j in range(len(symbols1)): symbol_dir = f"{direct}/{symbols1[j]}" if not os.path.exists(symbol_dir): os.mkdir(symbol_dir) for i in range(tf, len(data1)): img_path = f"{symbol_dir}/{i-tf+1}.png" if not os.path.exists(img_path): plot_candle(i, j, data1, symbols1, s, direct, img_size, tf, fig, axes)
3. 用进程池实现并行,彻底隔离内存空间
既然你想同时运行多个实例,不如直接用multiprocessing.Pool把每个交易标的的绘图任务分给独立的进程——每个进程的内存是完全隔离的,即使单个进程有内存泄漏,进程结束后系统也会自动回收所有内存。
修改循环部分为进程池模式:
from multiprocessing import Pool, cpu_count def process_single_symbol(symbol, data, s, mc, direct, img_size, tf): symbol_dir = f"{direct}/{symbol}" if not os.path.exists(symbol_dir): os.mkdir(symbol_dir) # 每个进程单独创建自己的fig和axes,避免跨进程共享资源 fig, axes = mpf.plot(data[symbol].head(tf), type='candle', style=s, volume=True, axisoff=True, figratio=(1,1), returnfig=True) plt.close(fig) for i in range(tf, len(data)): img_path = f"{symbol_dir}/{i-tf+1}.png" if not os.path.exists(img_path): data_temp = data.loc[i-tf:i, symbol] mpf.plot(data_temp, type='candle', style=s, volume=True, axisoff=True, figratio=(1,1), fig=fig, axes=axes) buf = io.BytesIO() fig.savefig(buf, pad_inches=0, bbox_inches='tight') buf.seek(0) with Image.open(buf) as im: im.resize((img_size, img_size)).save(img_path, "PNG") for ax in axes: ax.clear() del data_temp, buf gc.collect() # 进程结束前销毁绘图资源 plt.close(fig) del fig, axes gc.collect() if __name__ == "__main__": # 用cpu_count()-1避免占满所有核心,给系统留余量 with Pool(cpu_count()-1) as pool: pool.starmap(process_single_symbol, [(sym, data1, s, mc, direct, img_size, tf) for sym in symbols1])
这样每个进程只负责一个交易标的的绘图任务,完成后自动释放内存,主进程的内存不会出现持续累积的情况。
4. 其他细节优化
- 升级mplfinance版本:旧版本的mplfinance存在已知的内存泄漏bug,用
pip install --upgrade mplfinance升级到最新版能解决不少问题。 - 关闭matplotlib的GUI后端:在脚本开头添加
mpl.use('Agg'),避免加载不必要的交互式GUI组件,减少内存占用。 - 避免全局变量依赖:尽量把所有参数都传入函数,不要依赖全局变量,确保变量在函数结束后能被垃圾回收机制正确处理。
最后验证
你可以先拿小批量数据(比如1000张图)测试,观察内存变化。如果改用脚本+进程池的方案后内存不再持续上涨,就可以放心跑全量的170万张图了。
内容的提问来源于stack exchange,提问作者StatsNerd
相关产品推荐
相关产品推荐

