You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中从字符串列表选取最长字符串(含重叠匹配场景)

如何在Polars中从字符串列表选取最长字符串?

基础场景示例

假设我们有包含字符串列表的DataFrame:

import polars as pl

df = pl.DataFrame({
    "values": [
        ["the", "quickest", "brown", "fox"],
        ["jumps", "over", "the", "lazy", "dog"],
        []
    ]
})

要提取每个列表中的最长字符串,可使用以下代码:

result = df.with_columns(
    pl.col("values")
    .map_elements(
        lambda lst: max(lst, key=len) if lst else None,
        return_dtype=pl.String
    )
    .alias("longest_string")
)

print(result)

预期输出

┌──────────────────────────────┬────────────────┐
│ values                       ┆ longest_string │
│ ---                          ┆ ---            │
│ list[str]                    ┆ str            │
╞══════════════════════════════╪════════════════╡
│ ["the", "quickest", … "fox"] ┆ quickest       │
│ ["jumps", "over", … "dog"]   ┆ jumps          │
│ []                           ┆ null           │
└──────────────────────────────┴────────────────┘

实际业务场景:处理str.extract_many的重叠匹配结果

当使用Expr.str.extract_many开启重叠匹配后,会得到包含多个匹配项的列表,先看原始匹配结果:

df = pl.DataFrame({"values": ["discontent"]})

df_matches = df.with_columns(
    pl.col("values").str.extract_many(r"\w{4}").alias("matches"),
    pl.col("values").str.extract_many(r"\w{4}", overlapping=True).alias("matches_overlapping")
)

print(df_matches)

原始匹配输出

┌────────────┬───────────┬─────────────────────────────────┐
│ values     ┆ matches   ┆ matches_overlapping             │
│ ---        ┆ ---       ┆ ---                             │
│ str        ┆ list[str] ┆ list[str]                       │
╞════════════╪═══════════╪═════════════════════════════════╡
│ discontent ┆ ["disco"] ┆ ["disco", "onte", "discontent"] │
└────────────┴───────────┴─────────────────────────────────┘

要从matches_overlapping中提取最长匹配字符串,执行以下处理:

result = df_matches.with_columns(
    pl.col("matches_overlapping")
    .map_elements(lambda lst: max(lst, key=len) if lst else None, return_dtype=pl.String)
    .alias("longest_overlapping_match")
)

print(result)

最终输出

┌────────────┬───────────┬─────────────────────────────────┬────────────────────────┐
│ values     ┆ matches   ┆ matches_overlapping             ┆ longest_overlapping_match │
│ ---        ┆ ---       ┆ ---                             ┆ ---                    │
│ str        ┆ list[str] ┆ list[str]                       ┆ str                    │
╞════════════╪═══════════╪═════════════════════════════════╪════════════════════════╡
│ discontent ┆ ["disco"] ┆ ["disco", "onte", "discontent"] ┆ discontent             │
└────────────┴───────────┴─────────────────────────────────┴────────────────────────┘

核心说明

  • 通过map_elements遍历每个字符串列表,利用max(lst, key=len)筛选出最长字符串
  • 空列表会返回None,对应Polars中的null值,兼容空列表场景

内容的提问来源于stack exchange,提问作者conjuncts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 22:23:14