You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Python TSV区间校验脚本添加多线程-p命令行参数

脚本修改实现方案

修改点说明

  • 新增-p命令行参数,支持自定义线程数,默认值为1
  • 修复原脚本中误用路径变量file_a替代DataFrame的bug
  • 提前提取file_b的区间为元组列表,替代效率极低的iterrows遍历,先从算法层面提升校验速度
  • 采用线程池方案实现多线程并行处理,将待校验的700万条数据拆分为对应线程数的分片,每个线程独立处理分片并返回统计结果,无共享资源竞争问题,性能更优

修改后完整代码

import argparse
import pandas as pd
from concurrent.futures import ThreadPoolExecutor

def get_args():
    ap = argparse.ArgumentParser()
    ap.add_argument("-p", "--thread_num", type=int, default=1, help="number of threads to use")
    ap.add_argument("-a", "--file_a", required=True, help="path to file_a")
    ap.add_argument("-b", "--file_b", required=True, help="path to file_b")
    return ap.parse_args()

def main():
    args = get_args()
    thread_num = args.thread_num

    # 读取输入文件
    file_a_df = pd.read_table(args.file_a, header=None)
    file_b_df = pd.read_table(args.file_b, header=None)
    file_b_df.columns = ["seqname", "source", "feature", "start", "end", "score", "strand", "frame", "attribute"]
    
    # 提前提取校验区间,避免反复iterrows
    intervals = list(zip(file_b_df['start'], file_b_df['end']))
    # 提取待校验的值列表
    check_values = file_a_df.iloc[:, 1].tolist()

    # 拆分数据为对应线程数的分片
    chunk_size = len(check_values) // thread_num + 1
    value_chunks = [check_values[i:i+chunk_size] for i in range(0, len(check_values), chunk_size)]

    # 定义单分片处理逻辑
    def process_chunk(chunk):
        contained = 0
        not_contained = 0
        for val in chunk:
            match_flag = False
            for start, end in intervals:
                if start <= val <= end:
                    contained += 1
                    match_flag = True
                    break
            if not match_flag:
                not_contained += 1
        return contained, not_contained

    # 多线程并行处理
    with ThreadPoolExecutor(max_workers=thread_num) as executor:
        results = list(executor.map(process_chunk, value_chunks))
    
    # 汇总结果
    total_contained = sum(res[0] for res in results)
    total_not_contained = sum(res[1] for res in results)

    print("Contained: ", total_contained)
    print("Not contained: ", total_not_contained)

if __name__ == "__main__":
    main()

运行方式

直接使用你预期的命令即可:

python my_script.py -p 8 -a "path/to/file_a" -b "path/to/file_b"

额外提速建议

如果还需要进一步提升速度,可以提前对file_b的区间进行合并去重,减少无效的区间遍历,校验速度会更快。

内容的提问来源于stack exchange,提问作者Iacopo Passeri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 03:48:03