You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量从不同目录的16份fastq文件生成readlength.tsv?

批量处理fastq文件生成readlength.tsv脚本方案

假设你已经将所有fastq文件的绝对路径整理到了一个文本文件中(比如命名为fastq_list.txt,每行一个文件路径),可以用以下两种方式实现批量处理:

1. 基础串行处理脚本

适合对资源占用要求不高的场景,逻辑简单易调试:

#!/bin/bash

# 读取文件列表中的每个fastq路径
while IFS= read -r fastq_path; do
    # 跳过空行
    [[ -z "$fastq_path" ]] && continue
    # 检查文件是否存在
    if [[ ! -f "$fastq_path" ]]; then
        echo "警告:文件 $fastq_path 不存在,跳过"
        continue
    fi
    # 生成对应输出文件路径(替换原文件后缀,保留原目录)
    output_path="${fastq_path%.fastq.gz}_readlength.tsv"
    # 优化原单文件脚本:用shell内置计算字符串长度替代外部wc命令,提升效率
    zcat "$fastq_path" | paste - - - - | cut -f1,2 | while read -r readID sequ; do
        len=${#sequ}
        echo -e "$readID\t$len"
    done > "$output_path"
    echo "处理完成:$fastq_path -> $output_path"
done < fastq_list.txt

使用步骤:

  1. 将上述代码保存为batch_fastq_process.sh
  2. 给脚本添加执行权限:chmod +x batch_fastq_process.sh
  3. 运行脚本:./batch_fastq_process.sh

2. 并行处理脚本(提速推荐)

如果你的服务器有多余CPU核心,可以用并行处理大幅缩短总耗时,以下示例用4个并行进程(可根据硬件调整-P后的数字):

#!/bin/bash

# 定义单个文件的处理函数
process_single_file() {
    fastq_path="$1"
    if [[ ! -f "$fastq_path" ]]; then
        echo "警告:文件 $fastq_path 不存在,跳过"
        return 1
    fi
    output_path="${fastq_path%.fastq.gz}_readlength.tsv"
    zcat "$fastq_path" | paste - - - - | cut -f1,2 | while read -r readID sequ; do
        len=${#sequ}
        echo -e "$readID\t$len"
    done > "$output_path"
    echo "处理完成:$fastq_path -> $output_path"
}

# 导出函数供子shell调用
export -f process_single_file

# 并行处理文件列表
xargs -P 4 -I {} bash -c 'process_single_file "$@"' _ {} < fastq_list.txt

注意事项:

  • 如果你的fastq文件是未压缩格式(后缀为.fastq),请将脚本中的zcat替换为cat,同时将${fastq_path%.fastq.gz}改为${fastq_path%.fastq}
  • 原脚本中wc -m会把换行符计入长度,而${#sequ}仅计算字符串本身长度;如果需要和原脚本结果完全一致,可将len=${#sequ}改为len=$(( ${#sequ} + 1 ))

内容的提问来源于stack exchange,提问作者pierogi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 17:36:11