如何批量从不同目录的16份fastq文件生成readlength.tsv?
批量处理fastq文件生成readlength.tsv脚本方案
假设你已经将所有fastq文件的绝对路径整理到了一个文本文件中(比如命名为fastq_list.txt,每行一个文件路径),可以用以下两种方式实现批量处理:
1. 基础串行处理脚本
适合对资源占用要求不高的场景,逻辑简单易调试:
#!/bin/bash # 读取文件列表中的每个fastq路径 while IFS= read -r fastq_path; do # 跳过空行 [[ -z "$fastq_path" ]] && continue # 检查文件是否存在 if [[ ! -f "$fastq_path" ]]; then echo "警告:文件 $fastq_path 不存在,跳过" continue fi # 生成对应输出文件路径(替换原文件后缀,保留原目录) output_path="${fastq_path%.fastq.gz}_readlength.tsv" # 优化原单文件脚本:用shell内置计算字符串长度替代外部wc命令,提升效率 zcat "$fastq_path" | paste - - - - | cut -f1,2 | while read -r readID sequ; do len=${#sequ} echo -e "$readID\t$len" done > "$output_path" echo "处理完成:$fastq_path -> $output_path" done < fastq_list.txt
使用步骤:
- 将上述代码保存为
batch_fastq_process.sh - 给脚本添加执行权限:
chmod +x batch_fastq_process.sh - 运行脚本:
./batch_fastq_process.sh
2. 并行处理脚本(提速推荐)
如果你的服务器有多余CPU核心,可以用并行处理大幅缩短总耗时,以下示例用4个并行进程(可根据硬件调整-P后的数字):
#!/bin/bash # 定义单个文件的处理函数 process_single_file() { fastq_path="$1" if [[ ! -f "$fastq_path" ]]; then echo "警告:文件 $fastq_path 不存在,跳过" return 1 fi output_path="${fastq_path%.fastq.gz}_readlength.tsv" zcat "$fastq_path" | paste - - - - | cut -f1,2 | while read -r readID sequ; do len=${#sequ} echo -e "$readID\t$len" done > "$output_path" echo "处理完成:$fastq_path -> $output_path" } # 导出函数供子shell调用 export -f process_single_file # 并行处理文件列表 xargs -P 4 -I {} bash -c 'process_single_file "$@"' _ {} < fastq_list.txt
注意事项:
- 如果你的fastq文件是未压缩格式(后缀为
.fastq),请将脚本中的zcat替换为cat,同时将${fastq_path%.fastq.gz}改为${fastq_path%.fastq} - 原脚本中
wc -m会把换行符计入长度,而${#sequ}仅计算字符串本身长度;如果需要和原脚本结果完全一致,可将len=${#sequ}改为len=$(( ${#sequ} + 1 ))
内容的提问来源于stack exchange,提问作者pierogi
相关产品推荐
相关产品推荐

