在Polars中使用map_rows指定return_dtype仍报错“无法确定输出类型”
问题描述
以下代码可正常运行,创建包含struct类型列的DataFrame:
f = lambda t: {'x': 1, 'y': "abc"} df = pl.DataFrame( {'c': [f(0)] } )
输出结果:
shape: (1, 1) ┌───────────┐ │ c │ │ --- │ │ struct[2] │ ╞═══════════╡ │ {1,"abc"} │ └───────────┘
查看schema确认类型正确:
>>> df.schema # Schema([('c', Struct({'x': Int64, 'y': String}))])
但使用map_rows传入相同函数f并指定返回类型时,触发报错:RuntimeError: BindingsError: "Could not determine output type"
报错代码:
(df.with_columns('c') .map_rows(f, return_dtype= pl.Struct([pl.Field('x', pl.Int64), pl.Field('y', pl.String)]) ) )
问题原因
map_rows的函数参数接收的是整行数据(代表单行的对象),而非单个列值。原函数f设计为接受单个参数返回字典,但在map_rows调用时,实际传入的是整行,函数逻辑与输入不匹配,导致Polars无法正确解析返回值与指定类型的对应关系,从而抛出类型无法确定的错误。
解决方法
修改函数f,使其能正确处理整行输入,同时确保返回值结构与指定的return_dtype一致:
示例1:不依赖行数据生成结构
# 调整函数,接收整行对象,直接返回目标结构 f = lambda row: {'x': 1, 'y': "abc"} df = pl.DataFrame( {'c': [{'x': 0, 'y': "test"}] } ) # 直接调用map_rows,无需冗余的with_columns操作 result = df.map_rows(f, return_dtype=pl.Struct([pl.Field('x', pl.Int64), pl.Field('y', pl.String)])) print(result)
示例2:基于原列数据生成结构
如果需要用到原行的列值生成新结构:
# 基于原列c的内容生成新结构 f = lambda row: {'x': row['c']['x'] + 1, 'y': row['c']['y'].upper()} df = pl.DataFrame( {'c': [{'x': 0, 'y': "test"}] } ) result = df.map_rows(f, return_dtype=pl.Struct([pl.Field('x', pl.Int64), pl.Field('y', pl.String)])) print(result)
运行后会得到正确结果:
shape: (1, 2) ┌─────┬───────┐ │ x ┆ y │ │ --- ┆ --- │ │ i64 ┆ str │ ╞═════╪═══════╡ │ 1 ┆ TEST │ └─────┴───────┘
内容的提问来源于stack exchange,提问作者Des1303
相关产品推荐
相关产品推荐

