如何用正则从Pandas列提取以3开头的6位固定长度标识符?
问题描述
需要从以下Pandas DataFrame的Incident_details列中提取符合规则的标识符:
df1=pd.DataFrame({'Incident_details':['324657_Sample text1 about the incident', ' 316678_sample text2 with details of incident', '*DEPARTMENT LIST 316878-Sample text3 with information, ph: 01314522345', '327787_34587621 (sample text4 with incident details)', 'Sample text5 with details', '327997_1000587621 (sample text6 with incident info', ' 314489_incident text7 details', 'DEPARTMENT_LIST_325489_Text8 details', 'DEPARTMENT3_316489 text9 details', 'DEPARTMENT_LIST_326499', '324512_1000257218', '314656_text10(01345782345)', '324757_03456789', 'DEPARTMENT_CDES_324903_35678910 (details text11)', '326512_34500257218 - text12 details', 'Incident 325621_ 316512_ sample text 13']})
标识符规则
- 始终以
3开头,长度固定为6位; - 可出现在字符串开头、任意数量空格后或下划线后;
- 单个字符串可能包含多个标识符,需全部提取。
当前使用的正则表达式无法得到正确结果,请求修正:
df1['Incident_id'] = df1['incident_details'].str \ .findall(r'(?:^|\s|[^_])(\d{6})').str.join(", ")
修正方案
原正则存在三个核心问题:
[^_]会匹配非下划线的任意字符,导致捕获的标识符前附带多余字符,同时无法匹配下划线后的目标标识符;- 未限定标识符必须以
3开头,会误匹配其他无关6位数字; - 未处理后置边界,会从更长数字串中错误截取片段。
修正后的正则需要精准匹配规则,同时避免误捕:
完整修正代码
import pandas as pd df1=pd.DataFrame({'Incident_details':['324657_Sample text1 about the incident', ' 316678_sample text2 with details of incident', '*DEPARTMENT LIST 316878-Sample text3 with information, ph: 01314522345', '327787_34587621 (sample text4 with incident details)', 'Sample text5 with details', '327997_1000587621 (sample text6 with incident info', ' 314489_incident text7 details', 'DEPARTMENT_LIST_325489_Text8 details', 'DEPARTMENT3_316489 text9 details', 'DEPARTMENT_LIST_326499', '324512_1000257218', '314656_text10(01345782345)', '324757_03456789', 'DEPARTMENT_CDES_324903_35678910 (details text11)', '326512_34500257218 - text12 details', 'Incident 325621_ 316512_ sample text 13']}) # 修正后的正则表达式 df1['Incident_id'] = df1['Incident_details'].str.findall(r'(?:^|\s|_)(3\d{5})(?!\d)').str.join(", ") # 输出结果查看 print(df1[['Incident_details', 'Incident_id']])
正则表达式说明
(?:^|\s|_):非捕获组,匹配字符串开头、任意空格或下划线,确保标识符的前置位置符合规则;(3\d{5}):捕获组,精准匹配以3开头的6位数字(3+ 5位任意数字);(?!\d):负向预查,确保标识符后无其他数字,避免从更长数字串中截取片段。
关键输出示例
| Incident_details | Incident_id |
|---|---|
| 324657_Sample text1 about the incident | 324657 |
| *DEPARTMENT LIST 316878-Sample text3... | 316878 |
| Incident 325621_ 316512_ sample text 13 | 325621, 316512 |
| DEPARTMENT_LIST_325489_Text8 details | 325489 |
内容的提问来源于stack exchange,提问作者SriHari
相关产品推荐
相关产品推荐

