You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RStudio 4.2.2中memory.size()不再支持,求替代函数

替代memory.size()的内存统计方案

背景说明

在R 4.2+版本中,原Windows专属的memory.size()函数已被弃用,你可以用以下几种跨平台、更可靠的方法来统计机器学习过程中的内存使用情况。


方法1:利用内置gc()函数

gc()是R的内置垃圾回收函数,调用它会返回当前内存的使用详情,其中used列就是已使用的内存(单位:MB)。

修改你的代码示例:

result <- vector("list", 6L)
for (n in 1:5) {
  Train <- Training[sample(1:180568, 0.4*nrow(Training)),]
  # 初始内存统计
  initial_mem <- sum(gc()[,"used"])
  Time1 = Sys.time()
  fit_rand <- randomForest(Label~., data= Train)
  # 拟合后内存统计
  post_fit_mem <- sum(gc()[,"used"])
  Time2 = Sys.time()
  Train_time = (Time2 - Time1)
  print(Train_time)
  
  start_memory <- sum(gc()[,"used"])
  pred_rand <- predict(fit_rand, Testing)
  result[[n]] <- confusionMatrix(pred_rand, Testing$Label)
  print(result)
  Time3 = Sys.time()
  Validation_time <- (Time3 - Time2)
  print(Validation_time)
  
  stop_memory <- sum(gc()[,"used"])
  Validation_memory = (stop_memory - start_memory)
}

注意:每次调用gc()会触发垃圾回收,若不想主动回收,可改用gc(FALSE)获取当前内存状态而不执行回收。


方法2:使用pryr包的mem_used()

pryr包的mem_used()函数会直接返回R进程当前使用的总内存(单位:字节,可自行转换为MB),用法更简洁。

首先安装并加载包:

install.packages("pryr")
library(pryr)

修改代码:

result <- vector("list", 6L)
for (n in 1:5) {
  Train <- Training[sample(1:180568, 0.4*nrow(Training)),]
  # 初始内存(转换为MB)
  initial_mem <- mem_used() / 1024^2
  Time1 = Sys.time()
  fit_rand <- randomForest(Label~., data= Train)
  post_fit_mem <- mem_used() / 1024^2
  Time2 = Sys.time()
  Train_time = (Time2 - Time1)
  print(Train_time)
  
  start_memory <- mem_used() / 1024^2
  pred_rand <- predict(fit_rand, Testing)
  result[[n]] <- confusionMatrix(pred_rand, Testing$Label)
  print(result)
  Time3 = Sys.time()
  Validation_time <- (Time3 - Time2)
  print(Validation_time)
  
  stop_memory <- mem_used() / 1024^2
  Validation_memory = (stop_memory - start_memory)
}

方法3:系统级内存统计(processx包)

如果需要更精准的系统级内存占用(包括R进程的所有内存开销),可以用processx包获取当前R进程的内存使用情况。

安装加载包:

install.packages("processx")
library(processx)

修改代码:

result <- vector("list", 6L)
# 获取当前R进程ID
pid <- Sys.getpid()
for (n in 1:5) {
  Train <- Training[sample(1:180568, 0.4*nrow(Training)),]
  # 获取内存(转换为MB)
  initial_mem <- processx::ps_memory_info(pid)$rss / 1024^2
  Time1 = Sys.time()
  fit_rand <- randomForest(Label~., data= Train)
  post_fit_mem <- processx::ps_memory_info(pid)$rss / 1024^2
  Time2 = Sys.time()
  Train_time = (Time2 - Time1)
  print(Train_time)
  
  start_memory <- processx::ps_memory_info(pid)$rss / 1024^2
  pred_rand <- predict(fit_rand, Testing)
  result[[n]] <- confusionMatrix(pred_rand, Testing$Label)
  print(result)
  Time3 = Sys.time()
  Validation_time <- (Time3 - Time2)
  print(Validation_time)
  
  stop_memory <- processx::ps_memory_info(pid)$rss / 1024^2
  Validation_memory = (stop_memory - start_memory)
}

rss表示驻留集大小,即进程实际占用的物理内存,统计结果更贴近系统层面的真实内存使用。


内容的提问来源于stack exchange,提问作者Frenzy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 11:15:44