You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中extract_all提取列表最后元素报错的原因及解决

Polars中str.extract_all结合arr.last()报错的问题解决

问题背景

以下代码可在Polars中正常提取字符串里的首个年份:

path = ['some text 2020', '2021 text 2020', 'etxt 2022', '2023 text 2022']
names = ["Alice", "Bob", "Charlie", "David"]

df = pl.DataFrame({
    "path": path,
    "name": names
})

year = ['2019', '2020', '2021', '2022', '2023']
pattern = "((?i)" + '|'.join(year) + ")"
df.select(pl.col('path').str.extract(pattern))

输出结果:

shape: (4, 1)
path
str
"2020"
"2021"
"2022"
"2023"

但尝试用str.extract_all提取所有匹配年份,再通过.arr.last()获取列表最后一个元素时,代码报错:

year = ['2019', '2020', '2021', '2022', '2023']
pattern = "((?i)" + '|'.join(year) + ")"
df_new = df_new.with_columns(pl.col('path').str.extract_all(pattern).arr.last()).alias('year')

错误信息:

SchemaError: invalid series dtype: expected `FixedSizeList`, got `list[str]`

错误原因

.arr命名空间的方法仅适用于**FixedSizeList(固定长度列表)类型的列,而str.extract_all返回的是普通的List(可变长度列表)**类型列,类型不匹配导致报错。

正确实现

改用.list命名空间的.last()方法即可实现需求,正确代码如下:

path = ['some text 2020', '2021 text 2020', 'etxt 2022', '2023 text 2022']
names = ["Alice", "Bob", "Charlie", "David"]

df = pl.DataFrame({
    "path": path,
    "name": names
})

year = ['2019', '2020', '2021', '2022', '2023']
pattern = "((?i)" + '|'.join(year) + ")"
df.select(pl.col('path').str.extract_all(pattern).list.last().alias('year'))

输出结果:

shape: (4, 1)
year
str
"2020"
"2020"
"2022"
"2022"

内容的提问来源于stack exchange,提问作者Cam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 05:18:15