You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何NumPy比编译后的Mathematica慢?如何优化Python代码?

问题:优化Python多项式计算性能以超过编译后的Mathematica

这是上一个问题的延续,所用数据为维度(750000, 4)的floatMatrix。函数testFunction[x,y,w,z]是一个四变量多项式函数,返回4D向量,需应用于全部750000个向量。

Mathematica实现

使用带有Listable属性的Compile函数,设置Parallelization->True,具体实现如下:

Compile[{{f, _Real, 1}}, 
 {
   {0.011904761904761973` f[[2]]f[[1]]^3 + 
    0.002976190476190474` f[[1]]f[[2]]^3 - 0.020833333333333325` f[[3]] + 
    0.002976190476190474` f[[3]]^3 + 
    f[[2]]^2 (0.0029761904761904778` f[[3]] + ...)},
   {0.002976190476190483` f[[1]]^3 + 0.011904761904761906` f[[2]]^3 - 
    0.0875` f[[3]] + 0.0029761904761904765` f[[3]]^3 + 
    f[[1]]^2 (0.005952380952380952` f[[2]] + ...)}
 },
 CompilationTarget -> "C", 
 RuntimeAttributes -> {Listable}, 
 Parallelization -> True];

time = RepeatedTiming[testFunction[floatMatrix]];
Print["In Mathematica-C it takes an average of ", time[[1]], " secs."]

Python NumPy实现方案

方案1:函数内部转置

import numpy as np
import time

def testFunction(data):
    f1, f2, f3, f4 = data.T
    
    results = np.zeros((data.shape[0], 4))  # 初始化结果数组

    results[:, 0] = (0.011904761904761973*f2*f1**3 + 0.002976190476190474*f1*f2**3 - 
                    0.020833333333333325*f3 + 0.002976190476190474*f3**3 + f2**2* 
                    (0.0029761904761904778*f3 + ...))
    results[:, 1] = (0.002976190476190483*f1**3 + 0.011904761904761906*f2**3 - 0.0875*f3
                   + 0.0029761904761904765*f3**3 + f1**2*(0.005952380952380952*f2 + 
                     0.002976190476190469*f3 + 0.0029761904761904726*f4) + ...)
    return results

duration = 0
for i in range(10):
    start_time = time.time()
    testFunction(floatMatrix)
    end_time = time.time()
    duration += end_time - start_time

duration *= 0.1
print(f"NumPy(内部转置)平均耗时: {duration} 秒")

方案2:函数外部转置

import numpy as np
import time

def testFunction(f1, f2, f3, f4):
    results = np.zeros((f1.shape[0], 4))  # 初始化结果数组

    results[:, 0] = ...
    results[:, 1] = ...
    results[:, 2] = ...
    results[:, 3] = ...

    return results

# 在函数外部转置数据
f1, f2, f3, f4 = floatMatrix.T

duration = 0
for i in range(10):
    start_time = time.time()
    testFunction(f1, f2, f3, f4)
    end_time = time.time()
    duration += end_time - start_time

duration *= 0.1
print(f"NumPy(外部转置)平均耗时: {duration} 秒")

测试结果

  • Mathematica:0.119938秒
  • NumPy(内部转置):0.206754秒
  • NumPy(外部转置):0.20789377秒

我原本预期NumPy的速度会快很多,请问可以对Python代码做出哪些修改,使其速度超过编译后的Mathematica函数?


内容的提问来源于stack exchange,提问作者mmen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 11:54:51