You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Polars中使用map_rows指定return_dtype仍报错“无法确定输出类型”

问题描述

以下代码可正常运行,创建包含struct类型列的DataFrame:

f = lambda t: {'x': 1, 'y': "abc"}
df = pl.DataFrame( {'c': [f(0)] } )

输出结果:

shape: (1, 1)
┌───────────┐
│ c         │
│ ---       │
│ struct[2] │
╞═══════════╡
│ {1,"abc"} │
└───────────┘

查看schema确认类型正确:

>>> df.schema
# Schema([('c', Struct({'x': Int64, 'y': String}))])

但使用map_rows传入相同函数f并指定返回类型时,触发报错:RuntimeError: BindingsError: "Could not determine output type"
报错代码:

(df.with_columns('c')
   .map_rows(f, return_dtype=
             pl.Struct([pl.Field('x', pl.Int64), pl.Field('y', pl.String)])
   )
)
问题原因

map_rows的函数参数接收的是整行数据(代表单行的对象),而非单个列值。原函数f设计为接受单个参数返回字典,但在map_rows调用时,实际传入的是整行,函数逻辑与输入不匹配,导致Polars无法正确解析返回值与指定类型的对应关系,从而抛出类型无法确定的错误。

解决方法

修改函数f,使其能正确处理整行输入,同时确保返回值结构与指定的return_dtype一致:

示例1:不依赖行数据生成结构

# 调整函数,接收整行对象,直接返回目标结构
f = lambda row: {'x': 1, 'y': "abc"}

df = pl.DataFrame( {'c': [{'x': 0, 'y': "test"}] } )

# 直接调用map_rows,无需冗余的with_columns操作
result = df.map_rows(f, return_dtype=pl.Struct([pl.Field('x', pl.Int64), pl.Field('y', pl.String)]))
print(result)

示例2:基于原列数据生成结构

如果需要用到原行的列值生成新结构:

# 基于原列c的内容生成新结构
f = lambda row: {'x': row['c']['x'] + 1, 'y': row['c']['y'].upper()}

df = pl.DataFrame( {'c': [{'x': 0, 'y': "test"}] } )

result = df.map_rows(f, return_dtype=pl.Struct([pl.Field('x', pl.Int64), pl.Field('y', pl.String)]))
print(result)

运行后会得到正确结果:

shape: (1, 2)
┌─────┬───────┐
│ x   ┆ y     │
│ --- ┆ ---   │
│ i64 ┆ str   │
╞═════╪═══════╡
│ 1   ┆ TEST  │
└─────┴───────┘

内容的提问来源于stack exchange,提问作者Des1303

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 19:22:43