You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

优化运行缓慢的Python Genotype对比函数求助

Python基因型对比函数优化求助

我有一个运行速度极慢的Python函数,用于对比两个Genotype对象并返回样本基因型字典,其运行时长受Genotype对象大小及sample_list长度影响。以下是该函数的当前实现:

def compare_gt(gt1: Genotype, gt2: Genotype, sample_list: list = None):
    # is gt1 is None
    if not gt1 is None and not gt1.isValid(): 
        gt1=None
    if not gt2 is None and not gt2.isValid():
        gt2=None
    if gt1 is None and gt2 is None:
        return None
    if gt1 is None:
        if sample_list is None:
            sample_list = gt2.get_sample_list()
        return {sample: "6" + gt2.get_sample_genotypes(sample_list)[sample] for sample in sample_list}
    if gt2 is None: 
        if sample_list is None:
            sample_list = gt1.get_sample_list()
        return {sample: gt1.get_sample_genotypes(sample_list)[sample] + "6" for sample in sample_list}
    # compare two genotype
    if sample_list is None:
        sample_list = gt1.get_sample_list()
    return {sample: gt1.get_sample_genotypes(sample_list)[sample] + gt2.get_sample_genotypes(sample_list)[sample] for sample in sample_list}

我已尝试部分优化手段,比如避免重复计算sample_list、将gt1和gt2的结果存入变量以避免重复调用方法,但函数运行仍十分缓慢,希望能得到进一步优化建议。


更新内容

根据建议补充程序运行逻辑伪代码:

class gt_reader(): 
    def __init__(self, ...):
        # do something
    def update(self):
        # read next genotype line
    def get_sample_genotypes(self, sample_list) -> dict:
        # do something 
        return sample_gt
    # other methods

gt1 = gt_reader(...)
gt2 = gt_reader(...)

while gt1.update(): 
     # some code
    while gt2.update(): 
     # some code
        compare_gt(gt1, gt2)
        # some code
        if x: 
             break

除了待获取的优化建议外,我还对程序做了额外优化:确保get_sample_genotypes()的输出始终按查询排序,当gt1和gt2样本相同时无需持续通过索引获取返回值,这在大量对比操作中已提升了部分性能。


内容的提问来源于stack exchange,提问作者zhang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 14:42:54