Python正则表达式无法匹配列表最后一项问题求助
Fixing Your Python Regex for Key-Value Pairs
Your current regex has a couple of straightforward issues that are preventing it from matching your target key-value pairs correctly. Let's break down the problems and fix them step by step.
Issues with the Original Regex
- The
^anchor: This forces matches to start at the very beginning of the string, so it won't pick up the second pair (some_key=[j7;k9]) in your input line. - Character class typo:
[a-zAZ]is incorrect—you need[a-zA-Z]to properly match both uppercase and lowercase letters. - Mandatory trailing semicolon: The
([a-zAZ]\d;)+pattern expects every element to end with a semicolon, but your last element doesn't have one, so the match fails entirely.
Corrected Regex
Here's the adjusted regex that fixes all these problems:
r'(?P<key_values_pair>(?P<text>\w+)\s*=\s*\[(?P<values>[a-zA-Z]\d(?:;[a-zA-Z]\d)*)\])'
Breakdown of the Corrected Regex
(?P<key_values_pair>...): Named group capturing the entire key-value pair.(?P<text>\w+): Captures the key (one or more word characters like letters, numbers, underscores).\s*=\s*: Matches the equals sign with optional whitespace on either side (handles cases likekey = [value]orkey=[value]).\[: Matches the opening square bracket.(?P<values>[a-zA-Z]\d(?:;[a-zA-Z]\d)*): Captures the list of values:[a-zA-Z]\d: Matches a single element (a letter followed by a digit).(?:;[a-zA-Z]\d)*: Non-capturing group that matches zero or more additional elements, each preceded by a semicolon. This lets the last element skip the trailing semicolon.
\]: Matches the closing square bracket.
Example Usage in Python
To extract all matching key-value pairs from your input string:
import re input_str = "title=[a3;d2;g5;a5] #comment # other comment some_key=[j7;k9]" pattern = r'(?P<key_values_pair>(?P<text>\w+)\s*=\s*\[(?P<values>[a-zA-Z]\d(?:;[a-zA-Z]\d)*)\])' matches = re.finditer(pattern, input_str) for match in matches: print(f"Key: {match.group('text')}") print(f"Values: {match.group('values')}") print(f"Full pair: {match.group('key_values_pair')}") print("---")
Output
Key: title Values: a3;d2;g5;a5 Full pair: title=[a3;d2;g5;a5] --- Key: some_key Values: j7;k9 Full pair: some_key=[j7;k9] ---
This will correctly match both key-value pairs in your input, even when the last element doesn't end with a semicolon.
内容的提问来源于stack exchange,提问作者tutizeri
相关产品推荐
相关产品推荐

