You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多项式计算中abc.py __instancecheck__耗时过高的优化问询

多项式计算程序性能优化问题

我用Python编写了一个用于特定多项式计算的程序,整体运行速度尚可,但通过cProfile分析后发现了问题:某次运行总耗时296秒,其中abc.py的__instancecheck__函数累计耗时达43秒。我计划运行的一个计算按当前代码预计需要50天,这让我考虑改用其他语言,但希望能找到Python内的解决方案。

代码中的线程部分仅用于末尾的小模块,无性能问题,程序主体为单线程。我使用symengine或numpy进行多项式计算,可能是问题根源。

cProfile性能分析结果

ncalls  tottime  percall  cumtime  percall filename:lineno(function)
    201/1    0.001    0.000  296.104  296.104 {built-in method builtins.exec}
        1    0.000    0.000  296.104  296.104 double_samuel_schubmult7.py:1(<module>)
        1   85.132   85.132  295.717  295.717 double_samuel_schubmult7.py:205(schubmult)
    72618   34.608    0.000   94.825    0.001 double_samuel_schubmult7.py:247(<listcomp>)
 46620756   19.717    0.000   60.217    0.000 double_samuel_schubmult7.py:171(elem_sym_func)
        1    0.002    0.002   54.962   54.962 parallel.py:1000(__call__)
        1    0.584    0.584   54.886   54.886 parallel.py:960(retrieve)
    12039    0.013    0.000   54.181    0.005 pool.py:767(get)
    12039    0.008    0.000   54.160    0.004 pool.py:764(wait)
    12054    0.018    0.000   54.156    0.004 threading.py:589(wait)
      768    0.010    0.000   54.119    0.070 threading.py:288(wait)
     3126   54.105    0.017   54.105    0.017 {method 'acquire' of '_thread.lock' objects}
 58143528    7.986    0.000   43.820    0.000 abc.py:117(__instancecheck__)
 58143528   12.915    0.000   35.834    0.000 {built-in method _abc._abc_instancecheck}
 8188300/3054264   30.769    0.000   30.769    0.000 double_samuel_schubmult7.py:148(elem_sym_poly)
58143557/58143542    7.913    0.000   22.918    0.000 abc.py:121(__subclasscheck__)
58143557/58143542   15.006    0.000   15.006    0.000 {built-in method _abc._abc_subclasscheck}

主函数schubmult代码

from symengine import *
import numpy as np

# ...200 lines of omitted code...

def schubmult(perm_dict,v):
    vn1 = inverse(v)
    th = theta(vn1)
    if th[0]==0:
        return perm_dict        
    mu = permtrim(uncode(th))
    vmu = permtrim(mulperm(list(v),mu))
    inv_vmu = inv(vmu)
    inv_mu = inv(mu)
    ret_dict = {}
    vpaths = [([(vmu,0)],1)]
    while th[-1] == 0:
        th.pop()
    for i in range(len(th)):
        k = i+1
        vpaths2 = []
        for path,s in vpaths:
            last_perm = path[-1][0]
            newperms = kdown_perms(last_perm,th[i],k)
            for new_perm,s2,vdiff in newperms:
                new_perm2 = permtrim(new_perm)
                if i == len(th)-1 and (len(new_perm2) != 2 or new_perm2[0]!=1):
                    continue
                path2 = [*path,(new_perm2,vdiff)]
                vpaths2 += [(path2,s*s2)]
        vpaths = vpaths2
    arr0 = [0 for vpath in vpaths]
    for u,val in perm_dict.items():
        inv_u = inv(u)
        vpathsums = {u: val*np.array([vpath[1] for vpath in vpaths])}
        for index in range(len(th)):            
            newpathsums = {}
            for up, arr in vpathsums.items():
                inv_up = inv(up)
                newperms = elem_sym_perms(up,min(th[index],(inv_mu-(inv_up-inv_u))-inv_vmu),th[index])
                for up2, udiff in newperms:
                    newpathsums[up2] = newpathsums.get(up2,np.array(arr0))+arr*[elem_sym_func(th[index],index+1,up,up2,vpaths[i][0][index][0],vpaths[i][0][index+1][0],udiff,vpaths[i][0][index+1][1],var2,var3) for i in range(len(vpaths))]
            vpathsums = newpathsums
        if len(vpaths)<300:
            ret_dict = add_perm_dict({ep: np.sum(arr) for ep,arr in vpathsums.items()},ret_dict)
        else:
            ret_dict = add_perm_dict(dict(Parallel(n_jobs=-1,require='sharedmem')(delayed(pairsum)(ep,arr) for ep,arr in vpathsums.items())),ret_dict)            
    return ret_dict

核心疑问

  1. 这43秒的__instancecheck__耗时是否意味着改用C语言至少能节省43秒?
  2. 有没有办法在Python中抑制这部分耗时?

解决方案

关于改用C语言的效果

  • 不一定能直接节省43秒。__instancecheck__是Python类型检查机制带来的开销,若用C重写时完全规避这类类型检查逻辑,确实能省这部分时间,但C代码也可能需要做类型验证,只是开销不同。另外你的程序还有其他耗时模块(比如schubmult函数本身耗时85秒),整体优化效果还要看其他部分的重写效率。

Python内的优化方法

  • 减少类型检查次数:__instancecheck__通常由isinstance()触发,检查高频调用的函数(如elem_sym_func、elem_sym_poly)中是否存在大量不必要的isinstance()调用,能删除的直接删除,或提前做一次检查并缓存结果。
  • 替换symengine的类型检查逻辑:如果类型检查是symengine内部触发的,尝试直接使用symengine底层C接口(若有提供),或在不需要符号计算的阶段用numpy数组代替symengine符号对象,避免创建symengine实例。
  • 使用轻量类型替代:若多项式计算不需要完整符号运算能力,用numpy数值类型代替symengine符号类型,numpy的类型检查开销远低于Python的ABC实例检查。
  • 用Cython/Numba加速:将高频调用的函数(如elem_sym_func)用Cython重写,或用Numba的JIT编译,绕过Python的类型检查机制。Numba可在编译时确定类型,避免运行时的__instancecheck__开销。
  • 手动缓存类型检查结果:Python的ABC有缓存机制,但动态类型场景下缓存失效会导致重复检查。可手动用字典存储已检查过的类型,避免重复调用__instancecheck__。

内容的提问来源于stack exchange,提问作者Matt Samuel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 11:02:10