You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BaggingRegressor多CPU训练内存占用持续增长问题求助

BaggingRegressor多CPU训练内存泄漏问题解决方案

使用BaggingRegressor时,设置n_jobs=1可正常完成训练,但将n_jobs设为6启用多CPU后,内存占用随迭代持续上升且无法自动释放。相关代码如下:

test1= np.empty((788,))
test2= np.empty((788,))
pred_test= np.empty((788,))

for y in np.arange(0, 71, 1):
    for x in np.arange(0, 73, 1):

        dataY = rawdata[y, x, :]/10
        dataX= (fx + fxa)/2

        dataX = np.reshape(dataX , (4017, 1))
        dataY = np.reshape(dataY , (4017, 1))

        u = np.argwhere(np.isnan(dataY))
        dataY = np.delete(dataY, u)
        dataX= np.delete(dataX, u)
        l = len(dataY)
        dataY= np.reshape(dataY, (l, 1))
        dataX= np.reshape(dataX, (l, 1))

        X_train, X_test, y_train, y_test = train_test_split(
            dataX, dataY, test_size=0.20)

        xgb_reg = xgb.XGBRegressor()
        model = BaggingRegressor(
            base_estimator=xgb_reg, n_jobs=1, n_estimators=100)

        model.fit(X_train, y_train.ravel())
        pred_final = model.predict(X_test)

        # assign data into list
        test1[:] = np.squeeze(y_test)
        test2[:] = np.squeeze(X_test)
        pred_test[:] = np.squeeze(pred_final)

        del xgb_reg, model, X_train, X_test, y_train, y_test, dataY, dataX, u, l, pred_test
        gc.collect()

核心原因

  • 多进程模式下,BaggingRegressor依赖joblib创建子进程,若子进程未被正确回收,会导致内存累积。
  • XGBoost默认启用多线程(nthread=-1),与Bagging的多进程叠加会加剧内存消耗,甚至引发泄漏。
  • 循环内重复创建模型、频繁的数组操作(如np.delete)产生的内存碎片未被及时回收。

解决措施

1. 限制XGBoost线程数,避免资源叠加

Bagging已通过n_jobs启用多进程,每个子进程中的XGBoost模型应禁用多线程,同时添加内存优化参数:

xgb_reg = xgb.XGBRegressor(
    nthread=1,          # 每个子模型仅用1个线程
    tree_method='hist', # 直方图算法大幅降低内存占用
    max_depth=6,        # 限制树深减少内存开销
    subsample=0.8       # 采样减少计算量
)

2. 优化数据处理,减少内存碎片

用掩码替代np.delete,避免频繁创建新数组:

# 替换原有的NaN处理逻辑
mask = ~np.isnan(dataY).ravel()
dataX = dataX[mask].reshape(-1, 1)
dataY = dataY[mask].reshape(-1, 1)

3. 强制回收子进程资源

使用joblib上下文管理器明确控制进程池,避免子进程残留:

from joblib import parallel_backend

# 在循环内的fit阶段添加上下文管理
with parallel_backend('loky', n_jobs=6):
    model.fit(X_train, y_train.ravel())

# 清理后强制回收内存
del model, xgb_reg
gc.collect()

4. 避免循环内重复创建模型实例

将模型初始化移到循环外,确保每次fit重置状态:

# 循环外初始化模型
xgb_reg = xgb.XGBRegressor(nthread=1)
model = BaggingRegressor(
    base_estimator=xgb_reg, 
    n_jobs=6, 
    n_estimators=100,
    warm_start=False  # 强制每次fit重新训练
)

for y in np.arange(0, 71, 1):
    for x in np.arange(0, 73, 1):
        # ... 数据处理与训练 ...
        model.fit(X_train, y_train.ravel())
        # ... 预测与赋值 ...
        # 清理临时变量
        del X_train, X_test, y_train, y_test, dataY, dataX, pred_final
        gc.collect()

5. 定位泄漏点(可选)

用memory_profiler工具精准定位内存泄漏位置:

pip install memory-profiler

添加装饰器后运行:

from memory_profiler import profile

@profile
def train_loop():
    # 你的完整循环训练代码

内容的提问来源于stack exchange,提问作者Chun Yen Huang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 03:43:13