Polars中如何提取字符串中正则匹配的多个结果
Polars 原生提供 str.extract_all 方法,可提取字符串中所有符合正则规则的匹配项,每行的匹配结果会以列表类型返回。
针对你给出的示例,替换提取方法即可得到所有匹配的姓名:
import polars as pl df = pl.DataFrame( { "a": [ "name=John, name=Billy", "name=Jeff", "name=Taylor", ] } ) df.select( pl.col("a").str.extract_all(r"name=(\w+)", 1).alias("matched_names"), )
运行返回结果:
shape: (3, 1) ┌──────────────────┐ │ matched_names │ │ --- │ │ list[str] │ ╞══════════════════╡ │ ["John", "Billy"]│ │ ["Jeff"] │ │ ["Taylor"] │ └──────────────────┘
如果需要把列表内的多个匹配项拆分为独立行,只需要在提取后链式调用 explode 方法:
df.select( pl.col("a").str.extract_all(r"name=(\w+)", 1).alias("matched_name") ).explode("matched_name")
返回结果:
shape: (4, 1) ┌──────────────┐ │ matched_name │ │ --- │ │ str │ ╞══════════════╡ │ John │ │ Billy │ │ Jeff │ │ Taylor │ └──────────────┘
内容的提问来源于stack exchange,提问作者user6268172
相关产品推荐
相关产品推荐

