如何为Python TSV区间校验脚本添加多线程-p命令行参数
脚本修改实现方案
修改点说明
- 新增
-p命令行参数,支持自定义线程数,默认值为1 - 修复原脚本中误用路径变量
file_a替代DataFrame的bug - 提前提取
file_b的区间为元组列表,替代效率极低的iterrows遍历,先从算法层面提升校验速度 - 采用线程池方案实现多线程并行处理,将待校验的700万条数据拆分为对应线程数的分片,每个线程独立处理分片并返回统计结果,无共享资源竞争问题,性能更优
修改后完整代码
import argparse import pandas as pd from concurrent.futures import ThreadPoolExecutor def get_args(): ap = argparse.ArgumentParser() ap.add_argument("-p", "--thread_num", type=int, default=1, help="number of threads to use") ap.add_argument("-a", "--file_a", required=True, help="path to file_a") ap.add_argument("-b", "--file_b", required=True, help="path to file_b") return ap.parse_args() def main(): args = get_args() thread_num = args.thread_num # 读取输入文件 file_a_df = pd.read_table(args.file_a, header=None) file_b_df = pd.read_table(args.file_b, header=None) file_b_df.columns = ["seqname", "source", "feature", "start", "end", "score", "strand", "frame", "attribute"] # 提前提取校验区间,避免反复iterrows intervals = list(zip(file_b_df['start'], file_b_df['end'])) # 提取待校验的值列表 check_values = file_a_df.iloc[:, 1].tolist() # 拆分数据为对应线程数的分片 chunk_size = len(check_values) // thread_num + 1 value_chunks = [check_values[i:i+chunk_size] for i in range(0, len(check_values), chunk_size)] # 定义单分片处理逻辑 def process_chunk(chunk): contained = 0 not_contained = 0 for val in chunk: match_flag = False for start, end in intervals: if start <= val <= end: contained += 1 match_flag = True break if not match_flag: not_contained += 1 return contained, not_contained # 多线程并行处理 with ThreadPoolExecutor(max_workers=thread_num) as executor: results = list(executor.map(process_chunk, value_chunks)) # 汇总结果 total_contained = sum(res[0] for res in results) total_not_contained = sum(res[1] for res in results) print("Contained: ", total_contained) print("Not contained: ", total_not_contained) if __name__ == "__main__": main()
运行方式
直接使用你预期的命令即可:
python my_script.py -p 8 -a "path/to/file_a" -b "path/to/file_b"
额外提速建议
如果还需要进一步提升速度,可以提前对file_b的区间进行合并去重,减少无效的区间遍历,校验速度会更快。
内容的提问来源于stack exchange,提问作者Iacopo Passeri
相关产品推荐
相关产品推荐

