You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中使用输出类型与输入不匹配的Numba guvectorize函数?

解决Polars中Numba guvectorize函数输入输出类型不匹配的报错问题

问题背景

编写适配Polars的高性能Numba guvectorize函数时,遇到以下问题:

  • 当输入输出类型一致(均为float64)时,函数可正常在Polars中运行;
  • 当输入为float64、输出为int64时,直接调用函数无问题,但在Polars中调用会报错,提示找不到匹配的循环;
  • 实际场景生成的整数位数超过15位,无法通过先输出float64再转换为pl.Int64的方式解决。

环境信息:Ubuntu系统,numba 0.61.0,numpy 2.1.3,polars 1.27.1,python 3.12.8。

正常运行的代码示例(输入输出类型一致)

import numpy as np
import polars as pl
import numba as nb

data = pl.DataFrame({'a': np.random.random(10), 'b': np.random.random(10)})
lazy = pl.LazyFrame(data)

# 输入输出均为float64
@nb.guvectorize([(nb.float64[:], nb.float64[:], nb.float64[:])], '(n),(n)->(n)')
def to_float(x,y, res):
    for i in range(x.shape[0]):
        res[i] = x[i] + y[i]

lazy = lazy.with_columns(c=to_float(pl.col('a'), pl.col('b')))
print(lazy.collect())

报错代码示例(输入输出类型不一致)

import numpy as np
import polars as pl
import numba as nb

data = pl.DataFrame({'a': np.random.random(10), 'b': np.random.random(10)})
lazy = pl.LazyFrame(data)

# 输入为float64,输出为int64
@nb.guvectorize([(nb.float64[:], nb.float64[:], nb.int64[:])], '(n),(n)->(n)')
def to_int(x,y, res):
    for i in range(x.shape[0]):
        res[i] = int(x[i] + y[i])

lazy = lazy.with_columns(c=to_int(pl.col('a'), pl.col('b')))
print(lazy.collect())

报错信息

Traceback (most recent call last):
  File "/tmp/fails.py", line 16, in <module>
    print(lazy.collect())
          ^^^^^^^^^^^^^^
  File "/home/ubuntu/miniforge3/lib/python3.12/site-packages/polars/_utils/deprecation.py", line 93, in wrapper
    return function(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/ubuntu/miniforge3/lib/python3.12/site-packages/polars/lazyframe/frame.py", line 2206, in collect
    return wrap_df(ldf.collect(engine, callback))
                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
polars.exceptions.ComputeError: TypeError: No loop matching the specified signature and casting was found for ufunc to_int

解决方案

问题根源是Polars无法自动推断输入输出类型不一致的UDF返回类型,导致Numba找不到匹配的签名循环。解决方法是显式指定返回数据类型,通过map_batches方法配合return_dtype参数实现:

修改后的代码:

import numpy as np
import polars as pl
import numba as nb

data = pl.DataFrame({'a': np.random.random(10), 'b': np.random.random(10)})
lazy = pl.LazyFrame(data)

@nb.guvectorize([(nb.float64[:], nb.float64[:], nb.int64[:])], '(n),(n)->(n)')
def to_int(x,y, res):
    for i in range(x.shape[0]):
        res[i] = int(x[i] + y[i])

# 使用map_batches并显式指定返回类型为pl.Int64
lazy = lazy.with_columns(
    c=pl.col('a', 'b').map_batches(to_int, return_dtype=pl.Int64)
)
print(lazy.collect())

原理说明

通过return_dtype=pl.Int64告知Polars提前分配对应类型的输出数组,Numba就能找到匹配的float64输入转int64输出的签名循环,避免类型推断不匹配的问题。这种方式既保留了Numba guvectorize的高性能,又解决了Polars中的类型适配问题。

内容的提问来源于stack exchange,提问作者Stephen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 10:14:50