在远程集群使用R语言future_map函数时出现报错求助
问题与解决方案
问题描述
我正在使用学校集群执行计算任务,运行的代码如下:
cl <- makeCluster(detectCores()) plan(cluster, workers = cl) selected_a <- future_map(dat, ~abess(as.matrix(.x[, names(X)]), .x[, "errs"]), support.size = 10) stopCluster(cl)
执行后返回如下错误:
terminate called after throwing an instance of 'std::bad_alloc'
what(): std::bad_alloc
Erreur dans unserialize(node$con) :
ClusterFuture () failed to receive results from cluster SOCKnode #1 (on ‘localhost’). The reason reported was ‘erreur de lecture de la connexion’. Post-mortem diagnostic: The total size of the 8 globals exported is 1.18 MiB. The three largest globals are ‘...furrr_chunk_args’ (605.80 KiB of class ‘list’), ‘X’ (581.48 KiB of class ‘numeric’) and ‘as.matrix’ (11.45 KiB of class ‘function’)
该代码在本地机器运行正常,将future_map替换为map后,在本地和集群环境下均能正常执行,请问该如何解决此问题?
解决方案
限制集群核心数:集群节点的内存可能无法支撑全核心并行,手动减少核心数避免内存竞争:
cl <- makeCluster(floor(detectCores() / 2))优化全局变量传递:
- 显式指定
globals参数,只传递必要变量,减少冗余复制:selected_a <- future_map(dat, function(x) { abess(as.matrix(x[, names(X)]), x[, "errs"], support.size = 10) }, globals = list(X = X, abess = abess)) - 提前在集群节点加载
abess包,避免重复传递函数:cl <- makeCluster(floor(detectCores() / 2)) clusterEvalQ(cl, library(abess)) plan(cluster, workers = cl)
- 显式指定
降低
abess内存负载:- 尝试减小
support.size参数,降低模型计算的内存需求; - 检查
dat的每个子集规模,必要时精简特征或分块处理。
- 尝试减小
切换并行计划类型:改用
multisession计划,它的资源管理更适配集群环境:plan(multisession, workers = floor(detectCores() / 2)) selected_a <- future_map(dat, ~abess(as.matrix(.x[, names(X)]), .x[, "errs"]), support.size = 10) plan(sequential) # 执行完成后恢复默认顺序计划
内容的提问来源于stack exchange,提问作者entropy

