You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu 20.04下Rscript运行含doParallel的并行FASTQ处理脚本时PSOCK集群初始化失败(ssh连接超时)的解决方法

Ubuntu 20.04下Rscript运行含doParallel的并行FASTQ处理脚本时PSOCK集群初始化失败(ssh连接超时)的解决方法

看起来你遇到的问题是非交互式运行Rscript时PSOCK集群无法通过ssh连接本地节点,而交互式R里能正常运行,核心原因有两个:一是makePSOCKcluster的参数配置错误,二是Rscript的运行环境和交互式R存在差异。下面是具体的解决步骤和优化后的脚本:

一、核心问题定位

你当前的代码里,makePSOCKcluster的names参数设置成了n.cores(数字),这会导致R尝试连接类似0.0.0.20这样的无效主机(当n.cores=20时),自然会触发ssh超时错误。因为PSOCK集群在本地单机器并行时,应该连接localhost的多个进程,而不是把核心数当成主机名。

二、分步解决方法

1. 修正PSOCK集群初始化代码

把原来的集群创建代码:

my.cluster <- makePSOCKcluster(nnodes = n.cores, names = n.cores, outfile="")

修改为以下两种方式之一:

方式A:去掉names参数(推荐)

让makePSOCKcluster自动使用localhost创建本地集群:

my.cluster <- makePSOCKcluster(nnodes = n.cores, outfile = "")

方式B:显式指定localhost作为所有节点的主机名

如果需要明确指定,重复localhost核心数次:

my.cluster <- makePSOCKcluster(nnodes = n.cores, names = rep("localhost", n.cores), outfile = "")

2. 确保本地ssh无密码登录(可选但保险)

虽然交互式R能运行,但Rscript的环境可能无法自动通过ssh认证localhost,建议配置本地无密码ssh登录:

  • 生成ssh密钥(如果还没有):
    ssh-keygen -t ed25519
    
    按回车跳过所有密码设置,生成~/.ssh/id_ed25519和~/.ssh/id_ed25519.pub。
  • 将公钥添加到授权列表:
    cat ~/.ssh/id_ed25519.pub >> ~/.ssh/authorized_keys
    
  • 修正权限(必须,否则ssh会拒绝):
    chmod 700 ~/.ssh
    chmod 600 ~/.ssh/authorized_keys
    

3. 替代方案:改用doMC(更适合Unix本地并行)

如果你不需要跨机器并行,只是单机器多核心处理,doMC比doParallel更简单,不需要PSOCK集群和ssh配置:

  • 替换并行相关的代码:
    # 替换原来的library(doParallel)和集群初始化代码
    library(doMC)
    registerDoMC(cores = n.cores)
    
    # 去掉后面的stopCluster(my.cluster)
    
    这样就彻底避免了ssh相关的问题,Unix系统下稳定性更高。

三、优化后的完整脚本

这里给出修正了PSOCK集群参数的完整脚本,标注了修改的部分:

#! /usr/bin/Rscript

######

## Counting and excluding low read fastqs

# This script counts the number of reads in a fastq and moves those under a certain
# threshold to a separate file

## Collect arguments
args <- commandArgs(TRUE)

## Parse arguments (we expect the form --arg=value)
parseArgs <- function(x) strsplit(sub("^--", "", x), "=")
argsL <- as.list(as.character(as.data.frame(do.call("rbind", parseArgs(args)))$V2))
names(argsL) <- as.data.frame(do.call("rbind", parseArgs(args)))$V1
args <- argsL
rm(argsL)

if("--help" %in% args | is.null(args$files) | is.null(args$min_reads)) {
  cat("
Counts the number of reads in the gzipped FASTQs in a folder, then moves
files below a certain number of reads to a folder called 'less-than-[min_reads]'

--files=string        - Full path to folder with raw files
--min_reads=numeric   - Minimum number of reads for a file to be included
--n_cores=numeric     - Number of cores to use

Example:
Rscript count_and_move_raw_fastqs.R --files=raw/ \\   \n --min_reads=2e6 \\ \n --n_cores=10 \n\n")
  q(save="no")
}

## Import arguments
raw.path <- args$files
min.reads <- args$min_reads
n.cores <- args$n_cores

# Import libraries
library(foreach)
library(doParallel)  # 如果用doMC就换成library(doMC)
library(tidyverse)

# set wd to raw.path
setwd(raw.path)

# -------------------------- 修改的部分 --------------------------
# 修正PSOCK集群初始化,去掉错误的names参数
my.cluster <- makePSOCKcluster(nnodes = n.cores, outfile="")
registerDoParallel(cl = my.cluster)
# -------------------------- 修改结束 --------------------------

# Create folder for low read files to go
output_dir <- paste("files-with-less-than", min.reads,"-reads", sep = "")
dir.create(output_dir, showWarnings = FALSE)  # 添加showWarnings避免重复创建时报错

# Create function to count the lines in the file, and move files < min.reads to a different directory
process_file <- function(file){
  # Get the base name (without R1) for each file
  file.name <- stringr::str_split_fixed(file, "_", n = 2)[,1]
  
  # Count the reads
  message(paste("Counting reads in ", file, " as of ", Sys.time(), sep = ""))
  reads <- as.numeric(system(paste("gunzip -c ", file, " | wc -l | awk '{print $1/4}'", sep = ""), intern = TRUE))
  # 替换sed -n '$='为wc -l,更高效稳定
  
  # Move files if <= min.reads
  if(reads <= min.reads){
    # Find all files with the name
    files.to.move <- list.files(pattern = paste0("^", file.name))  # 添加^确保匹配开头,避免误匹配
    # Move them
    purrr::map(files.to.move, function(x) file.rename(x, file.path(output_dir, x)))
    # 用file.path代替手动拼接路径,更跨平台
  }
  
  # Create final df of reads and names
  output_txt <- file.path(getwd(), "reads_and_names_for_raw_files.txt")
  cat(paste(file.name, reads, "\n", sep = "\t"), file = output_txt, append = TRUE)
}

# Find the files to count
files.to.count <-  list.files(getwd(), pattern = "_R1_001.fastq.gz$")

# Loop over each file in parallel
foreach(file=files.to.count, .combine = c) %dopar%{
  process_file(file)
}

stopCluster(my.cluster)

备注:内容来源于stack exchange,提问作者Alex Krohn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.20 08:39:29