如何优化代码避免Google Colab免费版内存耗尽崩溃?
Colab免费版内存耗尽与代码运行缓慢的优化方案
针对你在Colab免费版运行机器学习代码时遇到的内存崩溃、运行缓慢问题,结合提供的代码,给出以下具体优化建议:
一、数据读取与内存精简
- 指定数据类型读取:CSV默认的int64、float64类型会占用大量内存,给类别列指定
category类型,数值列用更小精度(如float32),能大幅降低内存占用:dtype_config = { 'Property Type': 'category', 'Old/New': 'category', 'Record Status - monthly file only': 'category', 'PPDCategory Type': 'category', 'County': 'category', 'District': 'category', 'Town/City': 'category', 'Duration': 'category', 'Price': 'float32' } df = pd.read_csv('path/beforeNeural.csv', dtype=dtype_config) - 删除无用列:
Transaction unique identifier是唯一标识,对建模无任何帮助,直接删除避免浪费内存:df = df.drop(columns=['Transaction unique identifier'])
二、特征处理纠错与优化
- 日期列错误修复:用
LabelEncoder处理日期完全不合理,会把时间顺序变成无序数值,正确做法是转成日期类型后提取有效时间特征:df['Date of Transfer'] = pd.to_datetime(df['Date of Transfer']) # 提取年、月特征替代原日期列 df['Transfer_Year'] = df['Date of Transfer'].dt.year df['Transfer_Month'] = df['Date of Transfer'].dt.month df = df.drop(columns=['Date of Transfer']) - 类别编码简化:无需重复创建
LabelEncoder,循环处理所有类别列即可:encoder = LabelEncoder() cat_columns = ['Property Type', 'Old/New', 'Record Status - monthly file only', 'PPDCategory Type', 'County', 'District', 'Town/City', 'Duration'] for col in cat_columns: df[col] = encoder.fit_transform(df[col])
三、XGBoost模型参数优化
XGBoost默认参数对内存消耗较大,调整以下参数可在保证效果的前提下降低内存占用、提升速度:
boostenc = XGBRegressor( n_estimators=100, # 减少树的数量(默认1000),按需调整 max_depth=6, # 限制树深度,避免过拟合+减少内存 subsample=0.8, # 采样训练数据,降低内存压力 colsample_bytree=0.8, # 采样特征,减少计算量 tree_method='hist', # 直方图算法,比精确算法更快更省内存 random_state=2 )
四、Colab内存管理技巧
- 主动释放内存:删除不用的变量并调用垃圾回收,避免内存泄漏:
import gc del encoder # 删除不再使用的对象 gc.collect() # 触发垃圾回收 - 查看内存占用:定位内存消耗大户,针对性优化:
# 查看DataFrame总内存(MB) print(f"DataFrame内存占用:{df.memory_usage(deep=True).sum() / 1024**2:.2f} MB") - 重启会话:如果内存累积导致崩溃,重启Colab会话后重新运行代码,能清空残留内存。
内容的提问来源于stack exchange,提问作者Samir Almeida
相关产品推荐
相关产品推荐

