You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则提取列表中三位数字组合元素失败,求解决方案

问题:从Python列表提取特定格式字符串

给定一个包含多种元素的Python字符串列表,需要提取其中由三个1-4位数字组成、以空格分隔的元素(如'101 2022 202')。但运行以下代码后始终得到空列表:

regex_id = re.compile(r"(\d{1,5})(\s)(\d{1,5})(\s)(\d{1,5})")
id_list = [value for value in l if regex_id.search(str(value))==True]
print(id_list)

问题原因

  1. 正则匹配结果判断错误:re.search()成功匹配时返回的是Match对象,而非布尔值True。直接用==True比较永远不成立,导致过滤条件始终为假。
  2. 正则范围不符合需求:原正则用\d{1,5}匹配1-5位数字,但需求是1-4位,虽然这不是空列表的直接原因,但需修正以符合预期。

正确实现方法

方法1:修正判断逻辑与正则范围

将判断条件改为检查search()结果是否非空,同时把数字位数调整为1-4位:

import re

# 目标列表
l = ['tribunavtplus.client.db.PagingResultSet/2610158785',
'java.util.ArrayList/4159755760',
'tribunavtplus.client.db.ERGEBNISSTORE/3664619045',
'605',
'I. Sozialversicherungsgerichtshof',
'',
'a514394f9f18429995d4cd5c64fbc6de',
'Arrêt de la Ie Cour des assurances sociales du Tribunal cantonal',
'605 2022 1',
'3a4604ff7f9b404193346990b98f7179',
'605 2021 234',
'Assurance-accidents, rente (capacité de travail, revenus retenus dans le calcul du taux d\\x27invalidité, exigibilité), atteinte à l\\x27intégrité.',
'608',
'D:\\\\InetPubData\\\\PublicationDocuments\\\\80749472d88e4d45a0d3c340acbb670a.pdf',
'Polizeirichter Greyerz',
'D:\\\\InetPubData\\\\PublicationDocuments\\\\412d9325978b487f8b15a13380ba4436.pdf',
'Assurance-invalidité (refus de rente et de mesures d\\x27ordre professionnel).',
'101',
'I. Zivilappellationshof',
'c321e55aa536438fb95a06eb798eff82',
'Arrêt de la Ie Cour d\\x27appel civil du Tribunal cantonal',
'101 2022 202',
'2022-09-16',
'D:\\\\InetPubData\\\\PublicationDocuments\\\\c321e55aa536438fb95a06eb798eff82.pdf',
'Eheschutzmassnahmen',
'Gericht Saane',
'Mesures protectrices de l\\x27union conjugale, droit de visite, contributions d\\x27entretien.',
'603',
'III. Verwaltungsgerichtshof',
'3a801995f04b48948c0d1d2224b8f4b9',
'Arrêt de la IIIe Cour administrative du Tribunal cantonal',
'603 2022 110',
'D:\\\\InetPubData\\\\PublicationDocuments\\\\3a801995f04b48948c0d1d2224b8f4b9.pdf',
'Strassenverkehr und Transportwesen',
'Amt für Strassenverkehr und Schifffahrt (ASS)',
'502 2022 24',
'2022-09-22',
'D:\\\\InetPubData\\\\PublicationDocuments\\\\d2c0687313f24491870571dcfd64d6ad.pdf',
'102',
'java.lang.Boolean/476441737']

# 修正后的正则与过滤逻辑
regex_id = re.compile(r"\d{1,4}\s\d{1,4}\s\d{1,4}")
id_list = [value for value in l if regex_id.search(str(value))]
print(id_list)

方法2:严格匹配整个字符串(更推荐)

使用re.fullmatch()确保元素完全符合格式,避免匹配到包含该格式的长字符串:

import re

regex_id = re.compile(r"^\d{1,4}\s\d{1,4}\s\d{1,4}$")
id_list = [value for value in l if regex_id.fullmatch(str(value))]
print(id_list)

运行后会得到预期结果:

['605 2022 1', '605 2021 234', '101 2022 202', '603 2022 110', '502 2022 24']

内容的提问来源于stack exchange,提问作者taga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 17:01:10