You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars DataFrame中截取到冒号前的子串

提取Polars DataFrame单元格中冒号左侧的子串

你可以通过以下两种简洁的方法实现需求:

方法1:拆分字符串后取首段

利用str.split按冒号分割字符串,再提取分割后的第一个元素,同时用str.strip清理可能存在的首尾空格:

import polars as pl

df = pl.DataFrame({
    "A": ["A|B:0d:cs", "A:1ds0", "QW|P:3dwsd"],
    "B": ["C:2ew2 ", "E|F:91we23", "A|Z:12w219"]
})

result_df = df.with_columns(
    pl.col("*").str.split(":").list.first().str.strip()
)

print(result_df)

方法2:正则表达式匹配

通过正则^[^:]+直接匹配字符串开头到第一个冒号前的所有内容,再清理空格:

result_df = df.with_columns(
    pl.col("*").str.extract(r"^[^:]+").str.strip()
)

print(result_df)

正则说明:

  • ^ 匹配字符串起始位置
  • [^:]+ 匹配任意非冒号字符,至少出现一次

两种方法都会得到你期望的结果:

shape: (3, 2)
┌───────┬───────┐
│ A     │ B     │
│ ---   │ ---   │
│ str   │ str   │
├───────┼───────┤
│ A|B   │ C     │
│ A     │ E|F   │
│ QW|P  │ A|Z   │
└───────┴───────┘

内容的提问来源于stack exchange,提问作者Míriam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 07:42:43