You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用mplfinance/matplotlib疑似内存泄漏,如何解决?

解决K线图批量生成中的内存泄漏问题

首先得说,你遇到的这种内存持续累积的问题,在批量生成matplotlib/mplfinance图表时真的很常见,尤其是在Jupyter Notebook这种交互式环境里。结合你的代码和运行环境,我整理了几个针对性的解决方案,按优先级排序:

1. 先脱离Jupyter,改用普通Python脚本运行

Jupyter的内核会默认保留所有变量的引用——哪怕你手动del了,gc.collect()的效果在交互式环境里也会大打折扣。把三个单元格的代码合并成一个.py脚本直接运行,是解决内存累积最快速的第一步。

你只需要把代码里的路径调整好,然后用python your_script.py命令启动,完全避开Jupyter的变量缓存问题。

2. 复用Figure/Axes对象,避免每次创建新实例

现在你的plot_candle函数每次调用mpf.plot都会生成全新的Figure和Axes对象,这是内存泄漏的核心原因之一。我们可以提前创建好一套可复用的绘图容器,每次只更新数据:

# 在循环外提前创建可复用的fig和axes
fig, axes = mpf.plot(data1[symbols1[0]].head(tf), type='candle', style=s, volume=True, axisoff=True, figratio=(1,1), returnfig=True)
plt.close(fig)  # 先关闭,避免自动显示

def plot_candle(i,j,data,symbols,s,direct,img_size, tf, fig, axes):
    # 用loc切片避免链式索引产生不必要的DataFrame副本
    data_temp = data.loc[i-tf:i, symbols[j]]
    
    # 复用已有的fig和axes更新图表
    mpf.plot(data_temp, type='candle', style=s, volume=True, axisoff=True, figratio=(1,1), fig=fig, axes=axes)
    
    buf = io.BytesIO()
    # 直接用matplotlib的savefig,避免mpf封装带来的额外开销
    fig.savefig(buf, pad_inches=0, bbox_inches='tight')
    buf.seek(0)
    
    # 用with上下文管理器自动关闭Image对象
    with Image.open(buf) as im:
        im_resized = im.resize((img_size, img_size))
        im_resized.save(f"{direct}/{symbols[j]}/{i-tf+1}.png", "PNG")
    
    # 清空axes内容,为下一次绘图准备
    for ax in axes:
        ax.clear()
    
    # 手动删除临时对象,触发垃圾回收
    del data_temp, buf, im_resized
    gc.collect()

然后在循环调用时传入提前创建好的fig和axes:

for j in range(len(symbols1)):
    symbol_dir = f"{direct}/{symbols1[j]}"
    if not os.path.exists(symbol_dir):
        os.mkdir(symbol_dir)
    for i in range(tf, len(data1)):
        img_path = f"{symbol_dir}/{i-tf+1}.png"
        if not os.path.exists(img_path):
            plot_candle(i, j, data1, symbols1, s, direct, img_size, tf, fig, axes)

3. 用进程池实现并行,彻底隔离内存空间

既然你想同时运行多个实例,不如直接用multiprocessing.Pool把每个交易标的的绘图任务分给独立的进程——每个进程的内存是完全隔离的,即使单个进程有内存泄漏,进程结束后系统也会自动回收所有内存。

修改循环部分为进程池模式:

from multiprocessing import Pool, cpu_count

def process_single_symbol(symbol, data, s, mc, direct, img_size, tf):
    symbol_dir = f"{direct}/{symbol}"
    if not os.path.exists(symbol_dir):
        os.mkdir(symbol_dir)
    
    # 每个进程单独创建自己的fig和axes,避免跨进程共享资源
    fig, axes = mpf.plot(data[symbol].head(tf), type='candle', style=s, volume=True, axisoff=True, figratio=(1,1), returnfig=True)
    plt.close(fig)
    
    for i in range(tf, len(data)):
        img_path = f"{symbol_dir}/{i-tf+1}.png"
        if not os.path.exists(img_path):
            data_temp = data.loc[i-tf:i, symbol]
            mpf.plot(data_temp, type='candle', style=s, volume=True, axisoff=True, figratio=(1,1), fig=fig, axes=axes)
            
            buf = io.BytesIO()
            fig.savefig(buf, pad_inches=0, bbox_inches='tight')
            buf.seek(0)
            
            with Image.open(buf) as im:
                im.resize((img_size, img_size)).save(img_path, "PNG")
            
            for ax in axes:
                ax.clear()
            del data_temp, buf
            gc.collect()
    
    # 进程结束前销毁绘图资源
    plt.close(fig)
    del fig, axes
    gc.collect()

if __name__ == "__main__":
    # 用cpu_count()-1避免占满所有核心,给系统留余量
    with Pool(cpu_count()-1) as pool:
        pool.starmap(process_single_symbol, [(sym, data1, s, mc, direct, img_size, tf) for sym in symbols1])

这样每个进程只负责一个交易标的的绘图任务,完成后自动释放内存,主进程的内存不会出现持续累积的情况。

4. 其他细节优化

  • 升级mplfinance版本:旧版本的mplfinance存在已知的内存泄漏bug,用pip install --upgrade mplfinance升级到最新版能解决不少问题。
  • 关闭matplotlib的GUI后端:在脚本开头添加mpl.use('Agg'),避免加载不必要的交互式GUI组件,减少内存占用。
  • 避免全局变量依赖:尽量把所有参数都传入函数,不要依赖全局变量,确保变量在函数结束后能被垃圾回收机制正确处理。

最后验证

你可以先拿小批量数据(比如1000张图)测试,观察内存变化。如果改用脚本+进程池的方案后内存不再持续上涨,就可以放心跑全量的170万张图了。

内容的提问来源于stack exchange,提问作者StatsNerd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:38:01