Ubuntu 20.04下Rscript运行含doParallel的并行FASTQ处理脚本时PSOCK集群初始化失败(ssh连接超时)的解决方法
Ubuntu 20.04下Rscript运行含doParallel的并行FASTQ处理脚本时PSOCK集群初始化失败(ssh连接超时)的解决方法
看起来你遇到的问题是非交互式运行Rscript时PSOCK集群无法通过ssh连接本地节点,而交互式R里能正常运行,核心原因有两个:一是makePSOCKcluster的参数配置错误,二是Rscript的运行环境和交互式R存在差异。下面是具体的解决步骤和优化后的脚本:
一、核心问题定位
你当前的代码里,makePSOCKcluster的names参数设置成了n.cores(数字),这会导致R尝试连接类似0.0.0.20这样的无效主机(当n.cores=20时),自然会触发ssh超时错误。因为PSOCK集群在本地单机器并行时,应该连接localhost的多个进程,而不是把核心数当成主机名。
二、分步解决方法
1. 修正PSOCK集群初始化代码
把原来的集群创建代码:
my.cluster <- makePSOCKcluster(nnodes = n.cores, names = n.cores, outfile="")
修改为以下两种方式之一:
方式A:去掉names参数(推荐)
让makePSOCKcluster自动使用localhost创建本地集群:
my.cluster <- makePSOCKcluster(nnodes = n.cores, outfile = "")
方式B:显式指定localhost作为所有节点的主机名
如果需要明确指定,重复localhost核心数次:
my.cluster <- makePSOCKcluster(nnodes = n.cores, names = rep("localhost", n.cores), outfile = "")
2. 确保本地ssh无密码登录(可选但保险)
虽然交互式R能运行,但Rscript的环境可能无法自动通过ssh认证localhost,建议配置本地无密码ssh登录:
- 生成ssh密钥(如果还没有):
按回车跳过所有密码设置,生成ssh-keygen -t ed25519~/.ssh/id_ed25519和~/.ssh/id_ed25519.pub。 - 将公钥添加到授权列表:
cat ~/.ssh/id_ed25519.pub >> ~/.ssh/authorized_keys - 修正权限(必须,否则ssh会拒绝):
chmod 700 ~/.ssh chmod 600 ~/.ssh/authorized_keys
3. 替代方案:改用doMC(更适合Unix本地并行)
如果你不需要跨机器并行,只是单机器多核心处理,doMC比doParallel更简单,不需要PSOCK集群和ssh配置:
- 替换并行相关的代码:
这样就彻底避免了ssh相关的问题,Unix系统下稳定性更高。# 替换原来的library(doParallel)和集群初始化代码 library(doMC) registerDoMC(cores = n.cores) # 去掉后面的stopCluster(my.cluster)
三、优化后的完整脚本
这里给出修正了PSOCK集群参数的完整脚本,标注了修改的部分:
#! /usr/bin/Rscript ###### ## Counting and excluding low read fastqs # This script counts the number of reads in a fastq and moves those under a certain # threshold to a separate file ## Collect arguments args <- commandArgs(TRUE) ## Parse arguments (we expect the form --arg=value) parseArgs <- function(x) strsplit(sub("^--", "", x), "=") argsL <- as.list(as.character(as.data.frame(do.call("rbind", parseArgs(args)))$V2)) names(argsL) <- as.data.frame(do.call("rbind", parseArgs(args)))$V1 args <- argsL rm(argsL) if("--help" %in% args | is.null(args$files) | is.null(args$min_reads)) { cat(" Counts the number of reads in the gzipped FASTQs in a folder, then moves files below a certain number of reads to a folder called 'less-than-[min_reads]' --files=string - Full path to folder with raw files --min_reads=numeric - Minimum number of reads for a file to be included --n_cores=numeric - Number of cores to use Example: Rscript count_and_move_raw_fastqs.R --files=raw/ \\ \n --min_reads=2e6 \\ \n --n_cores=10 \n\n") q(save="no") } ## Import arguments raw.path <- args$files min.reads <- args$min_reads n.cores <- args$n_cores # Import libraries library(foreach) library(doParallel) # 如果用doMC就换成library(doMC) library(tidyverse) # set wd to raw.path setwd(raw.path) # -------------------------- 修改的部分 -------------------------- # 修正PSOCK集群初始化,去掉错误的names参数 my.cluster <- makePSOCKcluster(nnodes = n.cores, outfile="") registerDoParallel(cl = my.cluster) # -------------------------- 修改结束 -------------------------- # Create folder for low read files to go output_dir <- paste("files-with-less-than", min.reads,"-reads", sep = "") dir.create(output_dir, showWarnings = FALSE) # 添加showWarnings避免重复创建时报错 # Create function to count the lines in the file, and move files < min.reads to a different directory process_file <- function(file){ # Get the base name (without R1) for each file file.name <- stringr::str_split_fixed(file, "_", n = 2)[,1] # Count the reads message(paste("Counting reads in ", file, " as of ", Sys.time(), sep = "")) reads <- as.numeric(system(paste("gunzip -c ", file, " | wc -l | awk '{print $1/4}'", sep = ""), intern = TRUE)) # 替换sed -n '$='为wc -l,更高效稳定 # Move files if <= min.reads if(reads <= min.reads){ # Find all files with the name files.to.move <- list.files(pattern = paste0("^", file.name)) # 添加^确保匹配开头,避免误匹配 # Move them purrr::map(files.to.move, function(x) file.rename(x, file.path(output_dir, x))) # 用file.path代替手动拼接路径,更跨平台 } # Create final df of reads and names output_txt <- file.path(getwd(), "reads_and_names_for_raw_files.txt") cat(paste(file.name, reads, "\n", sep = "\t"), file = output_txt, append = TRUE) } # Find the files to count files.to.count <- list.files(getwd(), pattern = "_R1_001.fastq.gz$") # Loop over each file in parallel foreach(file=files.to.count, .combine = c) %dopar%{ process_file(file) } stopCluster(my.cluster)
备注:内容来源于stack exchange,提问作者Alex Krohn
相关产品推荐
相关产品推荐

