You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让str.findall()在指定位置终止匹配以拆分地址字段

解决正则提取第一个地址的问题

要让第一个地址的提取在“Line 2:”前终止,只需修改正则表达式、限制匹配边界即可,以下是两种常用方案:

方案一:使用正向前后瞻

把原正则替换为(?<=Line 1: ).*?(?= Line 2:),修改后的代码如下:

for column in maindf_clean:
    address1 = maindf_clean['Line 1: address Line 2: address2'].str.findall('(?<=Line 1: ).*?(?= Line 2:)')
    print(address1)
    print('-----')
  • (?<=Line 1: ):正向肯定后瞻,确保匹配内容紧跟在“Line 1: ”之后
  • .*?:非贪婪匹配任意字符,避免原正则中.*贪婪匹配到字符串末尾的问题
  • (?= Line 2:):正向肯定前瞻,确保匹配内容截止到“ Line 2:”之前

方案二:使用捕获组

更直观的写法是用捕获组直接提取“Line 1: ”和“ Line 2:”之间的内容,正则为Line 1: (.*?) Line 2:,代码如下:

for column in maindf_clean:
    address1 = maindf_clean['Line 1: address Line 2: address2'].str.findall('Line 1: (.*?) Line 2:')
    print(address1)
    print('-----')

这种方式下,findall会直接返回捕获组内的内容,也就是目标第一个地址。

另外注意:你的代码里循环for column in maindf_clean但未用到循环变量column,如果目标列固定,可直接指定列名,避免冗余代码。

内容的提问来源于stack exchange,提问作者IPadaJT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 08:42:06