You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中从DataFrame文本列提取正则匹配项间的文本?

解决方法

可以通过Pandas结合正则表达式实现需求,核心思路是用正则提取每个Apple [字母][数字]开头到下一个匹配项(或文本结尾)的内容,再将提取结果展开为多行。

代码实现

import pandas as pd

# 构造示例DataFrame(替换为你的实际数据)
df = pd.DataFrame({
    'text': ['dd ee Apple A1 a b c d Apple A2 e f g Apple B1 hi g Apple C1 r 5 6 Apple D1...']
})

# 定义正则模式:匹配"Apple 字母数字"开头,直到下一个"Apple 字母数字"或文本结束的内容
pattern = r'(Apple [A-Za-z]\d.*?)(?=Apple [A-Za-z]\d|$)'

# 提取所有匹配片段并展开为多行
df['text_new'] = df['text'].str.findall(pattern)
df = df.explode('text_new').reset_index(drop=True)

# 可选:清理文本两端的多余空格
df['text_new'] = df['text_new'].str.strip()

# 查看结果
print(df)

代码说明

  • 正则解析:(Apple [A-Za-z]\d.*?)(?=Apple [A-Za-z]\d|$)中:
    • Apple [A-Za-z]\d 匹配目标起始标识(Apple+字母+数字)
    • .*? 非贪婪匹配后续内容,避免跨段抓取
    • (?=...) 正向预查,确保匹配到下一个起始标识或文本结尾时停止
  • 数据展开:str.findall()提取所有符合规则的片段为列表,explode()将列表元素拆分为单独行,实现每行对应一段目标文本。

内容的提问来源于stack exchange,提问作者FloriaT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 06:42:45