You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

numpy使用frompyfunc时参数超32个无法构造ufunc怎么解决

问题描述

我知道Stack Overflow上存在类似问题,但我的应用场景和该问题不同。
我有一个包含32列的DataFrame,可通过以下代码生成:

import numpy as np
import pandas as pd
from io import StringIO
dfs = """
    M0  M1  M2  M3 M4  M5 M6 M7 M8 M9 M10 M11 M12 M13 M14 M15 M16 M17 M18 M19 M20 M21 M22 M23 M24 M25 M26 M27 M28 M29 M30  age 
1   1   2   3    4  5   6  1  2 3    4  5  6   1   2    3  4  5    6   7   8    9 1    2  3    4  5    6  1    2   3    4   3.2        
2   7   5   4    5  8   3  1  2 3    4  5  6   1   2    3  4  5    6   7   8    9 1    2  3    4  5    6  1    2   3    4   4.5
3   4   8   9    3  5   2  1  2 3    4  5  6   1   2    3  4  5    6   7   8    9 1    2  3    4  5    6  1    2   3    4   6.7
"""
df = pd.read_csv(StringIO(dfs.strip()), sep='\s+', )
df

基于业务逻辑我构建了向量化函数,当函数总入参数量小于32时运行正常:

M=["M0","M1","M2","M3","M4","M5","M6","M7","M8","M9","M10","M11","M12","M13","M14","M15","M16","M17","M18","M19",
       "M20","M21","M22","M23","M24","M25","M26","M27","M28","M29"]
    
def func2(df, M):
    return [df[i].values for i in M] 

def func(age,*Ms):
    newcol=np.prod(Ms[0:age])
    return newcol

vfunc = np.frompyfunc(func, len(M)+1, 1)

df['newcol']=vfunc(df['age'].values.astype(int), *func2(df,M))

为便于理解,func2仅用于简化代码,生成func的所有入参,不使用func2的等价代码如下:

def func(age,M0,M1,M2,...,M29):
    newcol=np.prod(Ms[0:age])
    return newcol

vfunc = np.frompyfunc(func, 31, 1)

df['newcol']=vfunc(df['age'].values.astype(int), df['M1'].values,...,df['M29'].values)

实际问题是当入参数量≥32时,比如以下代码(和上述代码的唯一差异是多了M30列):

M=["M0","M1","M2","M3","M4","M5","M6","M7","M8","M9","M10","M11","M12","M13","M14","M15","M16","M17","M18","M19",
           "M20","M21","M22","M23","M24","M25","M26","M27","M28","M29","M30"] # M30 is the only difference from the above function
        
def func2(df, M):
    return [df[i].values for i in M] 

def func(age,*Ms):
    newcol=np.prod(Ms[0:age])
    return newcol

vfunc = np.frompyfunc(func, len(M)+1, 1)

df['newcol']=vfunc(df['age'].values.astype(int), *func2(df,M))

就会抛出以下错误:

ValueError                                Traceback (most recent call last)
<ipython-input-66-9a042ad44f9b> in <module>()
     76     return newcol
     77 
---> 78 vfunc = np.frompyfunc(func, len(M)+1, 1)
     79 
     80 df['newcol']=vfunc(df['age'].values.astype(int), *func2(df,M))

ValueError: Cannot construct a ufunc with more than 32 operands (requested number were: inputs = 32 and outputs = 1)

我的实际业务场景中需要对超过100列数据使用np.prod计算,这个问题已经完全阻塞了开发进度,请问有什么可行的解决方案?

解决方案

核心思路是绕过np.frompyfunc的32个入参限制,不需要将每一列作为单独参数传入,直接把所有需要计算的M列打包为单个二维数组作为入参即可,修改后的代码如下:

# 包含任意数量M列,哪怕超过100列也可以
M=["M0","M1","M2","M3","M4","M5","M6","M7","M8","M9","M10","M11","M12","M13","M14","M15","M16","M17","M18","M19",
           "M20","M21","M22","M23","M24","M25","M26","M27","M28","M29","M30"]

# 直接把所有M列转成二维数组,shape为(行数, 列数)
m_values = df[M].values
ages = df['age'].values.astype(int)

# 逐行计算对应age长度的乘积,普通场景下性能足够
df['newcol'] = [np.prod(row[:age]) for row, age in zip(m_values, ages)]

如果需要更高的性能,也可以用numpy广播机制实现纯向量化计算,完全避免循环:

# 生成列索引掩码,小于对应age的位置为True,否则为False
mask = np.arange(len(M)) < ages[:, None]
# 用掩码把不需要参与计算的位置设为1,再按行求乘积
df['newcol'] = np.prod(np.where(mask, m_values, 1), axis=1)

两种方案都完全规避了np.frompyfunc的参数数量限制,支持任意数量的M列计算,性能也比原来的frompyfunc实现更优。

内容的提问来源于stack exchange,提问作者William

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 07:24:02