Python中使用re匹配任意顺序的多条目模式
问题:如何让正则支持任意顺序的键值对匹配?
需要捕获如下格式的输入内容:name="Game Title" authors="John Doe" studios="Studio A,Studio B" licence=ABC123 url=https://example.com command="start game" type=action code=xyz78
其中name、authors、studios……code等条目可能以任意顺序出现,但当前正则模式要求条目严格按指定顺序匹配,请问如何修改才能支持任意顺序?
原代码如下:
import re input_string = 'name="Game Title" authors="John Doe" studios="Studio A,Studio B" licence=ABC123 url=https://example.com command="start game" type=action code=xyz789' ADD_GAME_PATERN = r'(?P<name>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))\s+' \ r'licence=(?P<licence>[a-z0-9]*)\s+' \ r'type=(?P<typeCode>[a-z0-9]*)\s+' \ r'command=(?P<command>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))\s+' \ r'url=(?P<url>\S+)\s+' \ r'code=(?P<code>[a-z0-9]*)\s+' \ r'studios=(?P<studios>.*)\s+' \ r'authors=(?P<authors>.*)\s+' match = re.match(ADD_GAME_PATERN, input_string) if match: name = match.group('name') code = match.group('code') licence = match.group('licence') type_code = match.group('typeCode') command = match.group('command') url = match.group('url') studios = match.group('studios') authors = match.group('authors') print(f"Name: {name}") print(f"Code: {code}") print(f"Licence: {licence}") print(f"Type: {type_code}") print(f"Command: {command}") print(f"URL: {url}") print(f"Studios: {studios}") print(f"Authors: {authors}") else: print("No correspondance founded.")
解决方案
要支持任意顺序的键值对匹配,需放弃固定顺序的正则结构,改为为每个字段编写独立匹配规则,同时通过前瞻断言确保必填字段存在,允许规则以任意顺序组合。
修改后的代码
import re input_string = 'licence=ABC123 name="Game Title" command="start game" type=action url=https://example.com code=xyz789 studios="Studio A,Studio B" authors="John Doe"' # 定义所有必填字段,用于前瞻断言确保存在 REQUIRED_FIELDS = ['name', 'licence', 'type', 'command', 'url', 'code', 'studios', 'authors'] # 生成前瞻断言:确保每个必填字段都在输入中 lookaheads = ''.join([f'(?=.*{field}=)' for field in REQUIRED_FIELDS]) # 每个字段的匹配规则,用|分隔表示任意顺序 field_patterns = r'(?:name=(?P<name>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))|' \ r'licence=(?P<licence>[a-z0-9]+)|' \ r'type=(?P<typeCode>[a-z0-9]+)|' \ r'command=(?P<command>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))|' \ r'url=(?P<url>\S+)|' \ r'code=(?P<code>[a-z0-9]+)|' \ r'studios=(?P<studios>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*))|' \ r'authors=(?P<authors>(?:\"[^\"]*\"|\'[^\']*\'|[^\"\']*)))' # 最终正则:锚定整个字符串,先做前瞻检查,再匹配任意顺序的字段(字段间用空白分隔) ADD_GAME_PATTERN = fr'^{lookaheads}\s*(?:{field_patterns}\s*)*$' # 使用fullmatch匹配整个字符串,忽略大小写(可选) match = re.fullmatch(ADD_GAME_PATTERN, input_string, re.IGNORECASE) if match: # 提取并清理字段值(去除引号) name = match.group('name').strip('"\'') code = match.group('code') licence = match.group('licence') type_code = match.group('typeCode') command = match.group('command').strip('"\'') url = match.group('url') studios = match.group('studios').strip('"\'') authors = match.group('authors').strip('"\'') print(f"Name: {name}") print(f"Code: {code}") print(f"Licence: {licence}") print(f"Type: {type_code}") print(f"Command: {command}") print(f"URL: {url}") print(f"Studios: {studios}") print(f"Authors: {authors}") else: print("No correspondance founded.")
关键修改说明
- 前瞻断言:
(?=.*field=)确保每个必填字段都存在于输入中,避免遗漏核心字段。 - 任意顺序匹配:用
|分隔所有字段的匹配规则,再通过(?: ... \s*)*允许规则以任意顺序出现,适配不同的输入排列。 - 全字符串匹配:使用
re.fullmatch替代re.match,确保整个输入字符串都被校验,避免部分匹配导致的错误。 - 值清理:通过
strip('"\'')去除字符串类型字段的引号,得到更干净的业务数据。
内容的提问来源于stack exchange,提问作者fauve
相关产品推荐
相关产品推荐

