You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MacBook Air M2运行随机森林模型时RStudio崩溃求助

内存耗尽导致RStudio崩溃的随机森林训练解决方案

问题背景

在MacBook Air M2(基础款)上处理3109403行航班数据,运行随机森林模型时先触发vector memory exhausted (limit reached?)错误,调整内存设置后RStudio直接崩溃,需每次重启R环境。

用户运行代码:

# TEST
library(dplyr) # for data manipulation
library(caret) # for machine learning 

# Select relevant columns for analysis
selected_cols <- c("Year", "Month", "DayofMonth", "DayOfWeek", "DepTime", "CRSDepTime", "ArrTime", "CRSArrTime", "UniqueCarrier", "FlightNum", "Origin", "Dest", "Distance", "Cancelled", "CancellationCode")

flight_data <- Flight_Details_2003 %>% select(selected_cols)

# Create a delay variable by subtracting the scheduled departure time from the actual departure time. We will also filter out cancelled flights and negative delays
flight_data <- flight_data %>% mutate(DepDelay = DepTime - CRSDepTime) %>% 
  filter(Cancelled == 0, DepDelay >= 0)

# create a new fata frame with the delay information by airport and date:
airport_data <- flight_data %>% group_by(Year, Month, DayofMonth, Origin) %>% 
  summarize(TotalDelay = sum(DepDelay))

# Merge the delay data with itself to create a delay-by-airport-by-date matrix:
delay_matrix <- merge(airport_data, airport_data, by = c("Year", "Month", "DayofMonth"))

# Create a new column that indicates if a delay occured in the destination airport due to a delay in the origin airport
delay_matrix <- delay_matrix %>% mutate(CascadingDelay = ifelse(TotalDelay.x > 0 & TotalDelay.y == 0, 1, 0))

# Prepare data for machine learning
# Drop unnecessary columns
delay_matrix <- delay_matrix %>% select(-c("TotalDelay.x", "TotalDelay.y"))

# Convert data to binary format
delay_matrix$CascadingDelay <- factor(delay_matrix$CascadingDelay, levels = c(0,1), labels = c("No", "Yes"))

# Split data into training and testing sets
set.seed(123)
train_index <- createDataPartition(delay_matrix$CascadingDelay, p = 0.7, list = FALSE)
train_data <- delay_matrix[train_index,]
test_data <- delay_matrix[-train_index,]

# Randomly sampling a smaller portion of the training data because of this error: vector memory exhausted (limit reached?)
set.seed(123)
sample_size <- 10000
train_data_sample <- train_data[sample(seq_len(nrow(train_data)), size = sample_size), ]


# train the machine learning model using a random forest algorithm:

# Train the model
model <- train(CascadingDelay ~ ., data = train_data, method = "rf", trControl = trainControl(method = "cv", number = 10))

# Check model accuracy
confusionMatrix(model, test_data$CascadingDelay)

数据结构:

> str(flight_data)
'data.frame':   3109403 obs. of  16 variables:
 $ Year            : int  2003 2003 2003 2003 2003 2003 2003 2003 2003 2003 ...
 $ Month           : int  1 1 1 1 1 1 1 1 1 1 ...
 $ DayofMonth      : int  31 2 5 1 4 5 6 7 13 16 ...
 $ DayOfWeek       : int  5 4 7 3 6 7 1 2 1 4 ...
 $ DepTime         : int  1724 1053 1035 1713 1710 1832 1710 1712 1714 1711 ...
 $ CRSDepTime      : int  1655 1035 1035 1710 1710 1710 1710 1710 1710 1710 ...
 $ ArrTime         : int  1936 1726 1636 1851 1835 1951 1843 1839 1835 1837 ...
 $ CRSArrTime      : int  1913 1634 1634 1847 1847 1847 1847 1847 1847 1847 ...
 $ UniqueCarrier   : chr  "UA" "UA" "UA" "UA" ...
 $ FlightNum       : int  1017 1018 1018 1020 1020 1020 1020 1020 1020 1020 ...
 $ Origin          : chr  "ORD" "OAK" "OAK" "IAD" ...
 $ Dest            : chr  "MSY" "ORD" "ORD" "BOS" ...
 $ Distance        : int  837 1835 1835 413 413 413 413 413 413 413 ...
 $ Cancelled       : int  0 0 0 0 0 0 0 0 0 0 ...
 $ CancellationCode: chr  NA NA NA NA ...
 $ DepDelay        : int  69 18 0 3 0 122 0 2 4 1 ...

核心问题分析

原代码中merge(airport_data, airport_data, by = c("Year", "Month", "DayofMonth"))的自合并操作会生成所有机场对的笛卡尔积,数据行数直接膨胀为N²(N为每日机场数量),再加上caret包10折交叉验证的内存开销,远超M2基础款的内存承载上限,这是崩溃的根本原因。


解决方案

1. 优化数据预处理,避免无意义的数据膨胀

放弃自合并逻辑,直接在航班数据层面关联出发/到达机场的延误状态,保持数据量与原数据集一致:

# 提取每个机场每日的延误状态(是否存在延误)
airport_delay_status <- flight_data %>%
  group_by(Year, Month, DayofMonth, Origin) %>%
  summarize(HasDelay = as.integer(sum(DepDelay) > 0), .groups = "drop")

# 关联航班与出发、到达机场的延误状态,定义级联延误
flight_cascade <- flight_data %>%
  left_join(airport_delay_status, by = c("Year", "Month", "DayofMonth", "Origin")) %>%
  left_join(airport_delay_status, by = c("Year", "Month", "DayofMonth", "Dest" = "Origin"), suffix = c("_origin", "_dest")) %>%
  mutate(CascadingDelay = factor(ifelse(HasDelay_origin == 1 & HasDelay_dest == 0, "Yes", "No"), levels = c("No", "Yes"))) %>%
  # 保留建模必需字段,减少内存占用
  select(Year, Month, DayofMonth, DayOfWeek, UniqueCarrier, Origin, Dest, Distance, CascadingDelay)

2. 使用内存效率更高的模型实现

替换caret的train函数,改用ranger包(专为大内存优化,支持多核心并行):

install.packages("ranger")
library(ranger)

# 拆分数据集
set.seed(123)
train_index <- createDataPartition(flight_cascade$CascadingDelay, p = 0.7, list = FALSE)
train_data <- flight_cascade[train_index,]
test_data <- flight_cascade[-train_index,]

# 训练模型,控制参数减少内存消耗
model <- ranger(
  CascadingDelay ~ .,
  data = train_data,
  num.trees = 100,  # 减少树的数量(默认500),可根据精度调整
  min.node.size = 10,  # 增大节点最小样本数,降低内存开销
  num.threads = 4,  # 利用M2的多核心加速
  seed = 123
)

# 模型评估
pred <- predict(model, test_data)
confusionMatrix(pred$predictions, test_data$CascadingDelay)

3. 合理调整R的内存限制(针对M系列Mac)

在RStudio启动前修改全局配置,避免动态调整内存导致崩溃:

  • 打开终端,执行open ~/.Rprofile
  • 添加以下配置(根据你的M2内存调整,8GB机型建议设为6GB):
options(mem.max_size = 16 * 1024^3) # 16GB内存限制
  • 保存后重启RStudio。

4. 分层采样进一步压缩数据规模

如果内存仍有压力,对训练集进行分层采样(保持类别平衡):

set.seed(123)
train_sample <- train_data %>%
  group_by(CascadingDelay) %>%
  sample_n(size = 50000) %>% # 每个类别采样5万行,可按需调整
  ungroup()

内容的提问来源于stack exchange,提问作者Joseph Ng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 13:21:28