如何在Polars中使用输出类型与输入不匹配的Numba guvectorize函数?
解决Polars中Numba guvectorize函数输入输出类型不匹配的报错问题
问题背景
编写适配Polars的高性能Numba guvectorize函数时,遇到以下问题:
- 当输入输出类型一致(均为
float64)时,函数可正常在Polars中运行; - 当输入为
float64、输出为int64时,直接调用函数无问题,但在Polars中调用会报错,提示找不到匹配的循环; - 实际场景生成的整数位数超过15位,无法通过先输出
float64再转换为pl.Int64的方式解决。
环境信息:Ubuntu系统,numba 0.61.0,numpy 2.1.3,polars 1.27.1,python 3.12.8。
正常运行的代码示例(输入输出类型一致)
import numpy as np import polars as pl import numba as nb data = pl.DataFrame({'a': np.random.random(10), 'b': np.random.random(10)}) lazy = pl.LazyFrame(data) # 输入输出均为float64 @nb.guvectorize([(nb.float64[:], nb.float64[:], nb.float64[:])], '(n),(n)->(n)') def to_float(x,y, res): for i in range(x.shape[0]): res[i] = x[i] + y[i] lazy = lazy.with_columns(c=to_float(pl.col('a'), pl.col('b'))) print(lazy.collect())
报错代码示例(输入输出类型不一致)
import numpy as np import polars as pl import numba as nb data = pl.DataFrame({'a': np.random.random(10), 'b': np.random.random(10)}) lazy = pl.LazyFrame(data) # 输入为float64,输出为int64 @nb.guvectorize([(nb.float64[:], nb.float64[:], nb.int64[:])], '(n),(n)->(n)') def to_int(x,y, res): for i in range(x.shape[0]): res[i] = int(x[i] + y[i]) lazy = lazy.with_columns(c=to_int(pl.col('a'), pl.col('b'))) print(lazy.collect())
报错信息
Traceback (most recent call last): File "/tmp/fails.py", line 16, in <module> print(lazy.collect()) ^^^^^^^^^^^^^^ File "/home/ubuntu/miniforge3/lib/python3.12/site-packages/polars/_utils/deprecation.py", line 93, in wrapper return function(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/ubuntu/miniforge3/lib/python3.12/site-packages/polars/lazyframe/frame.py", line 2206, in collect return wrap_df(ldf.collect(engine, callback)) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ polars.exceptions.ComputeError: TypeError: No loop matching the specified signature and casting was found for ufunc to_int
解决方案
问题根源是Polars无法自动推断输入输出类型不一致的UDF返回类型,导致Numba找不到匹配的签名循环。解决方法是显式指定返回数据类型,通过map_batches方法配合return_dtype参数实现:
修改后的代码:
import numpy as np import polars as pl import numba as nb data = pl.DataFrame({'a': np.random.random(10), 'b': np.random.random(10)}) lazy = pl.LazyFrame(data) @nb.guvectorize([(nb.float64[:], nb.float64[:], nb.int64[:])], '(n),(n)->(n)') def to_int(x,y, res): for i in range(x.shape[0]): res[i] = int(x[i] + y[i]) # 使用map_batches并显式指定返回类型为pl.Int64 lazy = lazy.with_columns( c=pl.col('a', 'b').map_batches(to_int, return_dtype=pl.Int64) ) print(lazy.collect())
原理说明
通过return_dtype=pl.Int64告知Polars提前分配对应类型的输出数组,Numba就能找到匹配的float64输入转int64输出的签名循环,避免类型推断不匹配的问题。这种方式既保留了Numba guvectorize的高性能,又解决了Polars中的类型适配问题。
内容的提问来源于stack exchange,提问作者Stephen
相关产品推荐
相关产品推荐

