使用SMOTE-NC的fit_resample时内核崩溃问题求助
解决方案:从内存优化到替代方案
你的内核崩溃本质是SMOTE-NC的内存开销远高于普通SMOTE,结合2.2M行的数据集+8GB内存的M1 Mac,内存直接被打满了。以下是具体解决办法:
1. 先给数据集“瘦个身”——内存预处理
SMOTE-NC处理时会加载整个数据集到内存,先通过类型压缩减少内存占用:
- 数值特征降精度:把64位类型转成32位甚至8位(确保数值范围允许)
- 分类特征转
category类型:比object类型节省大量内存
import numpy as np import pandas as pd # 压缩数值特征类型 for col in X.select_dtypes(include=['float64']).columns: X[col] = X[col].astype('float32') # 精度损失可忽略,内存减半 for col in X.select_dtypes(include=['int64']).columns: # 检查数值范围是否适合转成int32 if X[col].min() >= np.iinfo(np.int32).min and X[col].max() <= np.iinfo(np.int32).max: X[col] = X[col].astype('int32') # 分类特征转category for col in cat_cols: X[col] = X[col].astype('category')
2. 分批处理——避免一次性加载全量数据
SMOTE-NC不需要全量数据做合成,你可以拆分少数类样本,分批合成后再合并:
from sklearn.model_selection import train_test_split # 分离多数类和少数类 X_majority = X[y == 0].reset_index(drop=True) X_minority = X[y == 1].reset_index(drop=True) y_majority = y[y == 0].reset_index(drop=True) y_minority = y[y == 1].reset_index(drop=True) # 定义批次大小,根据内存调整,比如500-2000 batch_size = 1000 resampled_minority_list = [] smote_cat = SMOTENC(random_state=42, categorical_features=cat_cols, k_neighbors=3) for start_idx in range(0, len(X_minority), batch_size): # 取当前少数类批次 minority_batch = X_minority[start_idx:start_idx+batch_size] # 匹配等量的多数类样本(保证批次内类别平衡) majority_batch, _ = train_test_split(X_majority, train_size=len(minority_batch), random_state=42) # 合并批次数据并做SMOTE-NC temp_X = pd.concat([majority_batch, minority_batch]) temp_y = pd.concat([pd.Series([0]*len(majority_batch)), pd.Series([1]*len(minority_batch))]) res_X, res_y = smote_cat.fit_resample(temp_X, temp_y) # 提取合成后的少数类样本,加入列表 resampled_minority_list.append(res_X[res_y == 1]) # 合并所有合成样本和原多数类 X_res = pd.concat([X_majority] + resampled_minority_list) y_res = pd.concat([y_majority] + [pd.Series([1]*len(b)) for b in resampled_minority_list])
3. 降低SMOTE-NC的计算开销
调整SMOTE-NC的参数减少内存占用:
- 减小
k_neighbors参数:默认是5,改成3,减少邻居计算时的内存消耗 - 限制合成样本的数量:通过
sampling_strategy指定少数类的目标数量,不要直接合成到和多数类等量(如果不需要的话)
# 比如只把少数类样本增加到多数类的20% smote_cat = SMOTENC( random_state=42, categorical_features=cat_cols, k_neighbors=3, sampling_strategy=0.2 ) X_res, y_res = smote_cat.fit_resample(X, y)
4. 替代方案:用目标编码+普通SMOTE
如果SMOTE-NC实在跑不动,先对分类特征做目标编码,再用普通SMOTE(内存开销小很多):
from category_encoders import TargetEncoder from imblearn.over_sampling import SMOTE # 目标编码分类特征(用目标变量y做编码) te = TargetEncoder(cols=cat_cols) X_encoded = te.fit_transform(X, y) # 用普通SMOTE做过采样 smote = SMOTE(random_state=42) X_res, y_res = smote.fit_resample(X_encoded, y)
注意:目标编码容易过拟合,建议后续建模时配合交叉验证,或者在编码时加入平滑参数(smoothing)
5. 临时硬件优化(Mac M1专属)
- 关闭所有后台无关应用,用Activity Monitor杀掉内存占用高的进程
- 临时调整系统资源限制:
sudo launchctl limit maxfiles 65536 200000 sudo launchctl limit maxproc 2048 4096
内容的提问来源于stack exchange,提问作者dsk4ch
相关产品推荐
相关产品推荐

