MacBook Air M2运行随机森林模型时RStudio崩溃求助
内存耗尽导致RStudio崩溃的随机森林训练解决方案
问题背景
在MacBook Air M2(基础款)上处理3109403行航班数据,运行随机森林模型时先触发vector memory exhausted (limit reached?)错误,调整内存设置后RStudio直接崩溃,需每次重启R环境。
用户运行代码:
# TEST library(dplyr) # for data manipulation library(caret) # for machine learning # Select relevant columns for analysis selected_cols <- c("Year", "Month", "DayofMonth", "DayOfWeek", "DepTime", "CRSDepTime", "ArrTime", "CRSArrTime", "UniqueCarrier", "FlightNum", "Origin", "Dest", "Distance", "Cancelled", "CancellationCode") flight_data <- Flight_Details_2003 %>% select(selected_cols) # Create a delay variable by subtracting the scheduled departure time from the actual departure time. We will also filter out cancelled flights and negative delays flight_data <- flight_data %>% mutate(DepDelay = DepTime - CRSDepTime) %>% filter(Cancelled == 0, DepDelay >= 0) # create a new fata frame with the delay information by airport and date: airport_data <- flight_data %>% group_by(Year, Month, DayofMonth, Origin) %>% summarize(TotalDelay = sum(DepDelay)) # Merge the delay data with itself to create a delay-by-airport-by-date matrix: delay_matrix <- merge(airport_data, airport_data, by = c("Year", "Month", "DayofMonth")) # Create a new column that indicates if a delay occured in the destination airport due to a delay in the origin airport delay_matrix <- delay_matrix %>% mutate(CascadingDelay = ifelse(TotalDelay.x > 0 & TotalDelay.y == 0, 1, 0)) # Prepare data for machine learning # Drop unnecessary columns delay_matrix <- delay_matrix %>% select(-c("TotalDelay.x", "TotalDelay.y")) # Convert data to binary format delay_matrix$CascadingDelay <- factor(delay_matrix$CascadingDelay, levels = c(0,1), labels = c("No", "Yes")) # Split data into training and testing sets set.seed(123) train_index <- createDataPartition(delay_matrix$CascadingDelay, p = 0.7, list = FALSE) train_data <- delay_matrix[train_index,] test_data <- delay_matrix[-train_index,] # Randomly sampling a smaller portion of the training data because of this error: vector memory exhausted (limit reached?) set.seed(123) sample_size <- 10000 train_data_sample <- train_data[sample(seq_len(nrow(train_data)), size = sample_size), ] # train the machine learning model using a random forest algorithm: # Train the model model <- train(CascadingDelay ~ ., data = train_data, method = "rf", trControl = trainControl(method = "cv", number = 10)) # Check model accuracy confusionMatrix(model, test_data$CascadingDelay)
数据结构:
> str(flight_data) 'data.frame': 3109403 obs. of 16 variables: $ Year : int 2003 2003 2003 2003 2003 2003 2003 2003 2003 2003 ... $ Month : int 1 1 1 1 1 1 1 1 1 1 ... $ DayofMonth : int 31 2 5 1 4 5 6 7 13 16 ... $ DayOfWeek : int 5 4 7 3 6 7 1 2 1 4 ... $ DepTime : int 1724 1053 1035 1713 1710 1832 1710 1712 1714 1711 ... $ CRSDepTime : int 1655 1035 1035 1710 1710 1710 1710 1710 1710 1710 ... $ ArrTime : int 1936 1726 1636 1851 1835 1951 1843 1839 1835 1837 ... $ CRSArrTime : int 1913 1634 1634 1847 1847 1847 1847 1847 1847 1847 ... $ UniqueCarrier : chr "UA" "UA" "UA" "UA" ... $ FlightNum : int 1017 1018 1018 1020 1020 1020 1020 1020 1020 1020 ... $ Origin : chr "ORD" "OAK" "OAK" "IAD" ... $ Dest : chr "MSY" "ORD" "ORD" "BOS" ... $ Distance : int 837 1835 1835 413 413 413 413 413 413 413 ... $ Cancelled : int 0 0 0 0 0 0 0 0 0 0 ... $ CancellationCode: chr NA NA NA NA ... $ DepDelay : int 69 18 0 3 0 122 0 2 4 1 ...
核心问题分析
原代码中merge(airport_data, airport_data, by = c("Year", "Month", "DayofMonth"))的自合并操作会生成所有机场对的笛卡尔积,数据行数直接膨胀为N²(N为每日机场数量),再加上caret包10折交叉验证的内存开销,远超M2基础款的内存承载上限,这是崩溃的根本原因。
解决方案
1. 优化数据预处理,避免无意义的数据膨胀
放弃自合并逻辑,直接在航班数据层面关联出发/到达机场的延误状态,保持数据量与原数据集一致:
# 提取每个机场每日的延误状态(是否存在延误) airport_delay_status <- flight_data %>% group_by(Year, Month, DayofMonth, Origin) %>% summarize(HasDelay = as.integer(sum(DepDelay) > 0), .groups = "drop") # 关联航班与出发、到达机场的延误状态,定义级联延误 flight_cascade <- flight_data %>% left_join(airport_delay_status, by = c("Year", "Month", "DayofMonth", "Origin")) %>% left_join(airport_delay_status, by = c("Year", "Month", "DayofMonth", "Dest" = "Origin"), suffix = c("_origin", "_dest")) %>% mutate(CascadingDelay = factor(ifelse(HasDelay_origin == 1 & HasDelay_dest == 0, "Yes", "No"), levels = c("No", "Yes"))) %>% # 保留建模必需字段,减少内存占用 select(Year, Month, DayofMonth, DayOfWeek, UniqueCarrier, Origin, Dest, Distance, CascadingDelay)
2. 使用内存效率更高的模型实现
替换caret的train函数,改用ranger包(专为大内存优化,支持多核心并行):
install.packages("ranger") library(ranger) # 拆分数据集 set.seed(123) train_index <- createDataPartition(flight_cascade$CascadingDelay, p = 0.7, list = FALSE) train_data <- flight_cascade[train_index,] test_data <- flight_cascade[-train_index,] # 训练模型,控制参数减少内存消耗 model <- ranger( CascadingDelay ~ ., data = train_data, num.trees = 100, # 减少树的数量(默认500),可根据精度调整 min.node.size = 10, # 增大节点最小样本数,降低内存开销 num.threads = 4, # 利用M2的多核心加速 seed = 123 ) # 模型评估 pred <- predict(model, test_data) confusionMatrix(pred$predictions, test_data$CascadingDelay)
3. 合理调整R的内存限制(针对M系列Mac)
在RStudio启动前修改全局配置,避免动态调整内存导致崩溃:
- 打开终端,执行
open ~/.Rprofile - 添加以下配置(根据你的M2内存调整,8GB机型建议设为6GB):
options(mem.max_size = 16 * 1024^3) # 16GB内存限制
- 保存后重启RStudio。
4. 分层采样进一步压缩数据规模
如果内存仍有压力,对训练集进行分层采样(保持类别平衡):
set.seed(123) train_sample <- train_data %>% group_by(CascadingDelay) %>% sample_n(size = 50000) %>% # 每个类别采样5万行,可按需调整 ungroup()
内容的提问来源于stack exchange,提问作者Joseph Ng
相关产品推荐
相关产品推荐

