求助:在Ubuntu RStudio Workbench中实现R函数时间与CPU使用率基准测试
解决R函数基准测试中CPU使用率监控的方案
bench::mark确实只专注于执行时间、内存分配等指标,无法直接获取CPU使用率。在Ubuntu的RStudio Workbench环境下,你可以通过以下两种方案实现同时测试执行时间和最大CPU使用率:
方案一:使用processx包启动子进程监控
通过processx创建独立子进程执行目标函数,实时监控子进程的CPU使用情况,同时记录执行时间。这种方式能避免监控到RStudio主进程的CPU干扰,结果更准确。
代码示例
library(processx) library(bench) # 定义你要测试的目标函数 my_test_func <- function() { # 示例计算:模拟CPU密集型任务 x <- matrix(rnorm(1e7), ncol = 100) svd(x) } # 封装基准测试函数:返回执行时间和最大CPU使用率 benchmark_with_cpu <- function(target_func) { # 生成临时R脚本,用于子进程执行 temp_script <- tempfile(fileext = ".R") result_file <- tempfile(fileext = ".rds") writeLines( sprintf( 'result <- %s(); saveRDS(result, "%s")', deparse(substitute(target_func)), result_file ), temp_script ) # 启动子进程执行脚本 proc <- process$new( command = R.home("bin/R"), args = c("--slave", "-f", temp_script), stdout = "|", stderr = "|" ) max_cpu <- 0 # 循环监控子进程的CPU使用率 while (proc$is_alive()) { proc$wait(0.1) # 每0.1秒采样一次 cpu_times <- proc$get_cpu_times() elapsed <- proc$get_elapsed_time() if (elapsed > 0) { # 计算当前CPU使用率((用户态+内核态时间)/已运行时间 * 100) current_cpu <- (cpu_times$user + cpu_times$system) / elapsed * 100 if (current_cpu > max_cpu) max_cpu <- current_cpu } } # 获取总执行时间 total_time <- proc$get_elapsed_time() # 清理临时文件 unlink(c(temp_script, result_file)) list( elapsed_time_sec = round(total_time, 3), max_cpu_percent = round(max_cpu, 2), exit_status = proc$get_exit_status() ) } # 运行测试 test_result <- benchmark_with_cpu(my_test_func) print(test_result) # 如需更精准的时间统计,可结合bench::mark的结果 bench_time_result <- bench::mark(my_test_func(), iterations = 3) print(bench_time_result)
方案二:使用ps包监控当前进程CPU
ps包可以直接获取当前R进程的CPU使用率,通过后台采样的方式记录执行期间的最大值。这种方式无需创建子进程,但要注意RStudio Workbench环境中,主进程可能包含RStudio本身的CPU开销,建议在测试时尽量关闭其他无关任务。
代码示例
library(ps) library(bench) library(parallel) # 目标测试函数 my_test_func <- function() { x <- 1:1e6 y <- lapply(x, function(i) sum(rnorm(1e3))) mean(unlist(y)) } # 封装带CPU监控的基准测试 bench_with_cpu_ps <- function(target_func, sample_interval = 0.05) { current_pid <- Sys.getpid() proc_handle <- ps_handle(current_pid) max_cpu <- 0 start_time <- Sys.time() # 启动后台线程监控CPU monitor_thread <- mcparallel({ while (ps_is_running(proc_handle)) { current_cpu <- ps_cpu_percent(proc_handle) if (current_cpu > max_cpu) max_cpu <<- current_cpu Sys.sleep(sample_interval) } }) # 执行目标函数 func_output <- target_func() # 等待监控线程结束 mccollect(monitor_thread) total_time <- as.numeric(difftime(Sys.time(), start_time, units = "secs")) list( elapsed_time_sec = round(total_time, 3), max_cpu_percent = round(max_cpu, 2), function_result = func_output ) } # 运行测试 ps_test_result <- bench_with_cpu_ps(my_test_func) print(ps_test_result)
注意事项
- 采样间隔越小,CPU使用率的监控越精准,但会增加少量额外开销,可根据任务类型调整(CPU密集型任务可适当增大间隔)。
- 在RStudio Workbench中,使用子进程方案(方案一)能有效隔离RStudio主进程的CPU干扰,结果更可靠。
- 若需要多次迭代取平均值,可在封装函数中添加循环逻辑,结合
bench::mark的迭代机制。
内容的提问来源于stack exchange,提问作者Ai4l2s
相关产品推荐
相关产品推荐

