You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SMOTE-NC的fit_resample时内核崩溃问题求助

解决方案:从内存优化到替代方案

你的内核崩溃本质是SMOTE-NC的内存开销远高于普通SMOTE,结合2.2M行的数据集+8GB内存的M1 Mac,内存直接被打满了。以下是具体解决办法:

1. 先给数据集“瘦个身”——内存预处理

SMOTE-NC处理时会加载整个数据集到内存,先通过类型压缩减少内存占用:

  • 数值特征降精度:把64位类型转成32位甚至8位(确保数值范围允许)
  • 分类特征转category类型:比object类型节省大量内存
import numpy as np
import pandas as pd

# 压缩数值特征类型
for col in X.select_dtypes(include=['float64']).columns:
    X[col] = X[col].astype('float32')  # 精度损失可忽略,内存减半

for col in X.select_dtypes(include=['int64']).columns:
    # 检查数值范围是否适合转成int32
    if X[col].min() >= np.iinfo(np.int32).min and X[col].max() <= np.iinfo(np.int32).max:
        X[col] = X[col].astype('int32')

# 分类特征转category
for col in cat_cols:
    X[col] = X[col].astype('category')

2. 分批处理——避免一次性加载全量数据

SMOTE-NC不需要全量数据做合成,你可以拆分少数类样本,分批合成后再合并:

from sklearn.model_selection import train_test_split

# 分离多数类和少数类
X_majority = X[y == 0].reset_index(drop=True)
X_minority = X[y == 1].reset_index(drop=True)
y_majority = y[y == 0].reset_index(drop=True)
y_minority = y[y == 1].reset_index(drop=True)

# 定义批次大小,根据内存调整,比如500-2000
batch_size = 1000
resampled_minority_list = []

smote_cat = SMOTENC(random_state=42, categorical_features=cat_cols, k_neighbors=3)

for start_idx in range(0, len(X_minority), batch_size):
    # 取当前少数类批次
    minority_batch = X_minority[start_idx:start_idx+batch_size]
    # 匹配等量的多数类样本(保证批次内类别平衡)
    majority_batch, _ = train_test_split(X_majority, train_size=len(minority_batch), random_state=42)
    
    # 合并批次数据并做SMOTE-NC
    temp_X = pd.concat([majority_batch, minority_batch])
    temp_y = pd.concat([pd.Series([0]*len(majority_batch)), pd.Series([1]*len(minority_batch))])
    res_X, res_y = smote_cat.fit_resample(temp_X, temp_y)
    
    # 提取合成后的少数类样本,加入列表
    resampled_minority_list.append(res_X[res_y == 1])

# 合并所有合成样本和原多数类
X_res = pd.concat([X_majority] + resampled_minority_list)
y_res = pd.concat([y_majority] + [pd.Series([1]*len(b)) for b in resampled_minority_list])

3. 降低SMOTE-NC的计算开销

调整SMOTE-NC的参数减少内存占用:

  • 减小k_neighbors参数:默认是5,改成3,减少邻居计算时的内存消耗
  • 限制合成样本的数量:通过sampling_strategy指定少数类的目标数量,不要直接合成到和多数类等量(如果不需要的话)
# 比如只把少数类样本增加到多数类的20%
smote_cat = SMOTENC(
    random_state=42,
    categorical_features=cat_cols,
    k_neighbors=3,
    sampling_strategy=0.2
)
X_res, y_res = smote_cat.fit_resample(X, y)

4. 替代方案:用目标编码+普通SMOTE

如果SMOTE-NC实在跑不动,先对分类特征做目标编码,再用普通SMOTE(内存开销小很多):

from category_encoders import TargetEncoder
from imblearn.over_sampling import SMOTE

# 目标编码分类特征(用目标变量y做编码)
te = TargetEncoder(cols=cat_cols)
X_encoded = te.fit_transform(X, y)

# 用普通SMOTE做过采样
smote = SMOTE(random_state=42)
X_res, y_res = smote.fit_resample(X_encoded, y)

注意:目标编码容易过拟合,建议后续建模时配合交叉验证,或者在编码时加入平滑参数(smoothing)

5. 临时硬件优化(Mac M1专属)

  • 关闭所有后台无关应用,用Activity Monitor杀掉内存占用高的进程
  • 临时调整系统资源限制:
sudo launchctl limit maxfiles 65536 200000
sudo launchctl limit maxproc 2048 4096

内容的提问来源于stack exchange,提问作者dsk4ch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 22:27:10