Polars中对列表类型调用map_batches触发InvalidOperationError求助
Polars列表列拼接报错:长度不匹配问题解决
当处理标量列时,以下代码可以正常运行并得到预期结果:
df = pl.DataFrame( { "int1": [1, 2, 3], "int2": [3, 2, 1] } ) df.with_columns( pl.struct('int1', 'int2') .map_batches(lambda x: x.struct.field('int1') + x.struct.field('int2')).alias('int3') )
输出:
shape: (3, 3) ┌──────┬──────┬──────┐ │ int1 ┆ int2 ┆ int3 │ │ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ i64 │ ╞══════╪══════╪══════╡ │ 1 ┆ 3 ┆ 4 │ │ 2 ┆ 2 ┆ 4 │ │ 3 ┆ 1 ┆ 4 │ └──────┴──────┴──────┘
但将列改为列表类型后,同样的写法会报错:
df = pl.DataFrame( { "int1": [[1], [2], [3]], "int2": [[3], [2], [1]] } ) df.with_columns( pl.struct('int1', 'int2') .map_batches(lambda x: x.struct.field('int1').to_list() + x.struct.field('int2').to_list()).alias('int3') )
报错信息:
InvalidOperationError: Series int3, length 1 doesn't match the DataFrame height of 3
期望输出:
┌───────────┬───────────┬───────────┐ │ int1 ┆ int2 ┆ int3 │ │ --- ┆ --- ┆ --- │ │ list[i64] ┆ list[i64] ┆ list[i64] │ ╞═══════════╪═══════════╪═══════════╡ │ [1] ┆ [3] ┆ [1, 3] │ │ [2] ┆ [2] ┆ [2, 2] │ │ [3] ┆ [1] ┆ [3, 1] │ └───────────┴───────────┴───────────┘
问题原因
map_batches是针对整个批次的Series进行操作,而非单个元素。在列表列场景中,x.struct.field('int1').to_list()会把整个Series转成Python列表(包含3个元素),和另一列的列表相加后得到一个包含6个元素的大列表,最终返回的是长度为1的Series(仅包含这个大列表),自然和原DataFrame的3行高度不匹配。
解决方案
方法1:直接用列表列相加(最简单)
Polars原生支持列表类型列的直接相加,底层会自动逐元素拼接列表,不需要额外的struct或map操作:
df.with_columns( (pl.col("int1") + pl.col("int2")).alias("int3") )
方法2:用map_elements处理单个struct元素
如果一定要保留struct的写法,换成map_elements——它会逐个处理每个struct元素,而非整个批次:
df.with_columns( pl.struct('int1', 'int2') .map_elements(lambda x: x['int1'] + x['int2']).alias('int3') )
方法3:用list.concat显式拼接
用Polars的list.concat函数,语义更清晰,适合需要拼接多个列表列的场景:
df.with_columns( pl.list.concat(pl.col("int1"), pl.col("int2")).alias("int3") )
以上三种方法都能得到你期望的输出结果。
内容的提问来源于stack exchange,提问作者erap129
相关产品推荐
相关产品推荐

