优化多.fasta文件合并Python脚本:提升GPD病毒库处理速度
高效合并FASTA文件的优化方案
一、优先使用系统原生命令(最快实现)
对于纯合并需求,无需解析FASTA序列的情况下,系统命令的速度远快于Python脚本——直接操作文本流,没有额外的对象解析开销。
适用场景:仅需合并所有FASTA文件,无需对序列做任何处理
使用find+cat组合命令(Linux/macOS):
find /path/to/main_folder -name "*.fasta" -type f -exec cat {} + > merged_output.fasta
- 说明:
find遍历所有子文件夹的.fasta文件,cat直接拼接文件内容到输出文件,全程无中间对象转换,速度最快。 - 若需跳过无法读取的文件,可添加错误重定向:
find /path/to/main_folder -name "*.fasta" -type f -exec cat {} 2>/dev/null + > merged_output.fasta
二、Python脚本优化方案(需保留Python处理逻辑时)
原脚本速度慢的核心原因是使用SeqIO解析/写入SeqRecord对象,该过程对纯合并需求来说完全冗余。以下是针对性优化:
优化点1:直接操作文本,跳过生物信息对象解析
直接按文本读取文件内容并写入输出,避免SeqRecord的序列化/反序列化开销:
from pathlib import Path import time def merge_all_fasta_from_subfolders(db_root: Path, output_fasta: Path): start_time = time.time() fasta_files = list(db_root.rglob("*.fasta")) if not fasta_files: print(f"⚠️ No FASTA files found in: {db_root}") return # 使用大缓冲区提升读写速度 with open(output_fasta, "w", buffering=1024*1024) as out_handle: total_files = 0 for file in fasta_files: try: # 直接读取文件全部内容写入 with open(file, "r") as in_handle: out_handle.write(in_handle.read()) total_files += 1 except Exception as e: print(f"❌ Error reading {file.name}: {e}") elapsed = time.time() - start_time print(f"✅ Merged {total_files} files in {elapsed:.2f} seconds.")
优化点2:可选并行读取(IO密集场景)
如果存储系统支持并行IO,可使用线程池加速文件读取(注意写入需串行,避免内容混乱):
from pathlib import Path import time from concurrent.futures import ThreadPoolExecutor def merge_all_fasta_from_subfolders(db_root: Path, output_fasta: Path, max_workers=4): start_time = time.time() fasta_files = list(db_root.rglob("*.fasta")) if not fasta_files: print(f"⚠️ No FASTA files found in: {db_root}") return # 线程池读取文件内容 def read_file(file_path): try: with open(file_path, "r") as f: return f.read() except Exception as e: print(f"❌ Error reading {file_path.name}: {e}") return "" # 串行写入输出文件 with open(output_fasta, "w", buffering=1024*1024) as out_handle: with ThreadPoolExecutor(max_workers=max_workers) as executor: for content in executor.map(read_file, fasta_files): out_handle.write(content) elapsed = time.time() - start_time print(f"✅ Merged {len(fasta_files)} files in {elapsed:.2f} seconds.")
优化点3:保留SeqIO时的小技巧(若需后续处理序列)
如果必须解析SeqRecord(比如去重、过滤),可通过以下方式提速:
- 增大
batch_size到更大值(如10000+),减少SeqIO.write的调用次数 - 使用
Bio.SeqIO.FastaIO.FastaIterator替代SeqIO.parse,减少内存开销 - 启用文件读写的大缓冲区(
buffering=1024*1024)
内容的提问来源于stack exchange,提问作者Andrea S.
相关产品推荐
相关产品推荐

