You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中使用re匹配任意顺序的多条目模式

问题:如何让正则支持任意顺序的键值对匹配?

需要捕获如下格式的输入内容:
name="Game Title" authors="John Doe" studios="Studio A,Studio B" licence=ABC123 url=https://example.com command="start game" type=action code=xyz78
其中name、authors、studios……code等条目可能以任意顺序出现,但当前正则模式要求条目严格按指定顺序匹配,请问如何修改才能支持任意顺序?

原代码如下:

import re

input_string = 'name="Game Title" authors="John Doe" studios="Studio A,Studio B" licence=ABC123 url=https://example.com command="start game" type=action code=xyz789'

ADD_GAME_PATERN = r'(?P<name>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))\s+' \
    r'licence=(?P<licence>[a-z0-9]*)\s+' \
    r'type=(?P<typeCode>[a-z0-9]*)\s+' \
    r'command=(?P<command>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))\s+' \
    r'url=(?P<url>\S+)\s+' \
    r'code=(?P<code>[a-z0-9]*)\s+' \
    r'studios=(?P<studios>.*)\s+' \
    r'authors=(?P<authors>.*)\s+'

match = re.match(ADD_GAME_PATERN, input_string)

if match:
    name = match.group('name')
    code = match.group('code')
    licence = match.group('licence')
    type_code = match.group('typeCode')
    command = match.group('command')
    url = match.group('url')
    studios = match.group('studios')
    authors = match.group('authors')

    print(f"Name: {name}")
    print(f"Code: {code}")
    print(f"Licence: {licence}")
    print(f"Type: {type_code}")
    print(f"Command: {command}")
    print(f"URL: {url}")
    print(f"Studios: {studios}")
    print(f"Authors: {authors}")
else:
    print("No correspondance founded.")

解决方案

要支持任意顺序的键值对匹配,需放弃固定顺序的正则结构,改为为每个字段编写独立匹配规则,同时通过前瞻断言确保必填字段存在,允许规则以任意顺序组合。

修改后的代码

import re

input_string = 'licence=ABC123 name="Game Title" command="start game" type=action url=https://example.com code=xyz789 studios="Studio A,Studio B" authors="John Doe"'

# 定义所有必填字段,用于前瞻断言确保存在
REQUIRED_FIELDS = ['name', 'licence', 'type', 'command', 'url', 'code', 'studios', 'authors']
# 生成前瞻断言:确保每个必填字段都在输入中
lookaheads = ''.join([f'(?=.*{field}=)' for field in REQUIRED_FIELDS])

# 每个字段的匹配规则,用|分隔表示任意顺序
field_patterns = r'(?:name=(?P<name>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))|' \
                 r'licence=(?P<licence>[a-z0-9]+)|' \
                 r'type=(?P<typeCode>[a-z0-9]+)|' \
                 r'command=(?P<command>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))|' \
                 r'url=(?P<url>\S+)|' \
                 r'code=(?P<code>[a-z0-9]+)|' \
                 r'studios=(?P<studios>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))|' \
                 r'authors=(?P<authors>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*)))'

# 最终正则:锚定整个字符串,先做前瞻检查,再匹配任意顺序的字段(字段间用空白分隔)
ADD_GAME_PATTERN = fr'^{lookaheads}\s*(?:{field_patterns}\s*)*$'

# 使用fullmatch匹配整个字符串,忽略大小写(可选)
match = re.fullmatch(ADD_GAME_PATTERN, input_string, re.IGNORECASE)

if match:
    # 提取并清理字段值(去除引号)
    name = match.group('name').strip('"\'')
    code = match.group('code')
    licence = match.group('licence')
    type_code = match.group('typeCode')
    command = match.group('command').strip('"\'')
    url = match.group('url')
    studios = match.group('studios').strip('"\'')
    authors = match.group('authors').strip('"\'')

    print(f"Name: {name}")
    print(f"Code: {code}")
    print(f"Licence: {licence}")
    print(f"Type: {type_code}")
    print(f"Command: {command}")
    print(f"URL: {url}")
    print(f"Studios: {studios}")
    print(f"Authors: {authors}")
else:
    print("No correspondance founded.")

关键修改说明

  1. 前瞻断言:(?=.*field=)确保每个必填字段都存在于输入中,避免遗漏核心字段。
  2. 任意顺序匹配:用|分隔所有字段的匹配规则,再通过(?: ... \s*)*允许规则以任意顺序出现,适配不同的输入排列。
  3. 全字符串匹配:使用re.fullmatch替代re.match,确保整个输入字符串都被校验,避免部分匹配导致的错误。
  4. 值清理:通过strip('"\'')去除字符串类型字段的引号,得到更干净的业务数据。

内容的提问来源于stack exchange,提问作者fauve

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 20:25:56