Python正则提取列表中三位数字组合元素失败,求解决方案
问题:从Python列表提取特定格式字符串
给定一个包含多种元素的Python字符串列表,需要提取其中由三个1-4位数字组成、以空格分隔的元素(如'101 2022 202')。但运行以下代码后始终得到空列表:
regex_id = re.compile(r"(\d{1,5})(\s)(\d{1,5})(\s)(\d{1,5})") id_list = [value for value in l if regex_id.search(str(value))==True] print(id_list)
问题原因
- 正则匹配结果判断错误:
re.search()成功匹配时返回的是Match对象,而非布尔值True。直接用==True比较永远不成立,导致过滤条件始终为假。 - 正则范围不符合需求:原正则用
\d{1,5}匹配1-5位数字,但需求是1-4位,虽然这不是空列表的直接原因,但需修正以符合预期。
正确实现方法
方法1:修正判断逻辑与正则范围
将判断条件改为检查search()结果是否非空,同时把数字位数调整为1-4位:
import re # 目标列表 l = ['tribunavtplus.client.db.PagingResultSet/2610158785', 'java.util.ArrayList/4159755760', 'tribunavtplus.client.db.ERGEBNISSTORE/3664619045', '605', 'I. Sozialversicherungsgerichtshof', '', 'a514394f9f18429995d4cd5c64fbc6de', 'Arrêt de la Ie Cour des assurances sociales du Tribunal cantonal', '605 2022 1', '3a4604ff7f9b404193346990b98f7179', '605 2021 234', 'Assurance-accidents, rente (capacité de travail, revenus retenus dans le calcul du taux d\\x27invalidité, exigibilité), atteinte à l\\x27intégrité.', '608', 'D:\\\\InetPubData\\\\PublicationDocuments\\\\80749472d88e4d45a0d3c340acbb670a.pdf', 'Polizeirichter Greyerz', 'D:\\\\InetPubData\\\\PublicationDocuments\\\\412d9325978b487f8b15a13380ba4436.pdf', 'Assurance-invalidité (refus de rente et de mesures d\\x27ordre professionnel).', '101', 'I. Zivilappellationshof', 'c321e55aa536438fb95a06eb798eff82', 'Arrêt de la Ie Cour d\\x27appel civil du Tribunal cantonal', '101 2022 202', '2022-09-16', 'D:\\\\InetPubData\\\\PublicationDocuments\\\\c321e55aa536438fb95a06eb798eff82.pdf', 'Eheschutzmassnahmen', 'Gericht Saane', 'Mesures protectrices de l\\x27union conjugale, droit de visite, contributions d\\x27entretien.', '603', 'III. Verwaltungsgerichtshof', '3a801995f04b48948c0d1d2224b8f4b9', 'Arrêt de la IIIe Cour administrative du Tribunal cantonal', '603 2022 110', 'D:\\\\InetPubData\\\\PublicationDocuments\\\\3a801995f04b48948c0d1d2224b8f4b9.pdf', 'Strassenverkehr und Transportwesen', 'Amt für Strassenverkehr und Schifffahrt (ASS)', '502 2022 24', '2022-09-22', 'D:\\\\InetPubData\\\\PublicationDocuments\\\\d2c0687313f24491870571dcfd64d6ad.pdf', '102', 'java.lang.Boolean/476441737'] # 修正后的正则与过滤逻辑 regex_id = re.compile(r"\d{1,4}\s\d{1,4}\s\d{1,4}") id_list = [value for value in l if regex_id.search(str(value))] print(id_list)
方法2:严格匹配整个字符串(更推荐)
使用re.fullmatch()确保元素完全符合格式,避免匹配到包含该格式的长字符串:
import re regex_id = re.compile(r"^\d{1,4}\s\d{1,4}\s\d{1,4}$") id_list = [value for value in l if regex_id.fullmatch(str(value))] print(id_list)
运行后会得到预期结果:
['605 2022 1', '605 2021 234', '101 2022 202', '603 2022 110', '502 2022 24']
内容的提问来源于stack exchange,提问作者taga
相关产品推荐
相关产品推荐

