You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则从Pandas列提取以3开头的6位固定长度标识符?

问题描述

需要从以下Pandas DataFrame的Incident_details列中提取符合规则的标识符:

df1=pd.DataFrame({'Incident_details':['324657_Sample text1 about the incident',
' 316678_sample text2 with details of incident',
'*DEPARTMENT LIST 316878-Sample text3 with information, ph: 01314522345',
'327787_34587621 (sample text4 with incident details)',
'Sample text5 with details',
'327997_1000587621 (sample text6 with incident info',
' 314489_incident text7 details',
'DEPARTMENT_LIST_325489_Text8 details',
'DEPARTMENT3_316489 text9 details',
'DEPARTMENT_LIST_326499',
'324512_1000257218',
'314656_text10(01345782345)',
'324757_03456789',
'DEPARTMENT_CDES_324903_35678910 (details text11)',
'326512_34500257218 - text12 details',
'Incident 325621_ 316512_ sample text 13']})

标识符规则

  • 始终以3开头,长度固定为6位;
  • 可出现在字符串开头、任意数量空格后或下划线后;
  • 单个字符串可能包含多个标识符,需全部提取。

当前使用的正则表达式无法得到正确结果,请求修正:

df1['Incident_id'] = df1['incident_details'].str \
   .findall(r'(?:^|\s|[^_])(\d{6})').str.join(", ")

修正方案

原正则存在三个核心问题:

  1. [^_]会匹配非下划线的任意字符,导致捕获的标识符前附带多余字符,同时无法匹配下划线后的目标标识符;
  2. 未限定标识符必须以3开头,会误匹配其他无关6位数字;
  3. 未处理后置边界,会从更长数字串中错误截取片段。

修正后的正则需要精准匹配规则,同时避免误捕:

完整修正代码

import pandas as pd

df1=pd.DataFrame({'Incident_details':['324657_Sample text1 about the incident',
' 316678_sample text2 with details of incident',
'*DEPARTMENT LIST 316878-Sample text3 with information, ph: 01314522345',
'327787_34587621 (sample text4 with incident details)',
'Sample text5 with details',
'327997_1000587621 (sample text6 with incident info',
' 314489_incident text7 details',
'DEPARTMENT_LIST_325489_Text8 details',
'DEPARTMENT3_316489 text9 details',
'DEPARTMENT_LIST_326499',
'324512_1000257218',
'314656_text10(01345782345)',
'324757_03456789',
'DEPARTMENT_CDES_324903_35678910 (details text11)',
'326512_34500257218 - text12 details',
'Incident 325621_ 316512_ sample text 13']})

# 修正后的正则表达式
df1['Incident_id'] = df1['Incident_details'].str.findall(r'(?:^|\s|_)(3\d{5})(?!\d)').str.join(", ")

# 输出结果查看
print(df1[['Incident_details', 'Incident_id']])

正则表达式说明

  • (?:^|\s|_):非捕获组,匹配字符串开头、任意空格或下划线,确保标识符的前置位置符合规则;
  • (3\d{5}):捕获组,精准匹配以3开头的6位数字(3 + 5位任意数字);
  • (?!\d):负向预查,确保标识符后无其他数字,避免从更长数字串中截取片段。

关键输出示例

Incident_detailsIncident_id
324657_Sample text1 about the incident324657
*DEPARTMENT LIST 316878-Sample text3...316878
Incident 325621_ 316512_ sample text 13325621, 316512
DEPARTMENT_LIST_325489_Text8 details325489

内容的提问来源于stack exchange,提问作者SriHari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 02:30:40