使用Regex解析含值内逗号的逗号分隔键值对
Nice problem! Parsing this structured comma-separated list requires a regex that can handle different value types (single quoted strings, unquoted booleans/numerics, and arrays of those) without splitting on commas that are inside arrays or quotes. Here's a complete solution tailored to your needs:
Step 1: Split the Input into Individual Key-Value Pairs
First, we need to split the input string into separate key-value pairs without breaking commas that are inside parentheses or single quotes. Use this regex to split on valid top-level commas:
,(?=(?:[^']*'[^']*')*[^']*$)(?=(?:[^()]*\([^()]*\))*[^()]*$)
How this works:
,: Matches a comma(?=(?:[^']*'[^']*')*[^']*$): Positive lookahead ensuring the comma is outside any single quotes (counts pairs of quotes to confirm we're not inside one)(?=(?:[^()]*\([^()]*\))*[^()]*$): Another positive lookahead ensuring the comma is outside any parentheses (counts pairs of parentheses to avoid splitting inside arrays)
Step 2: Parse Each Key-Value Pair
For each pair from the split, use this regex to extract the key and its corresponding value:
^(\w+)\s*=\s*((?:'[^']*')|(?:TRUE|FALSE|\d+)|(?:\((?:(?:'[^']*')|(?:TRUE|FALSE|\d+))(?:,\s*(?:'[^']*')|(?:TRUE|FALSE|\d+))*\)))$
Breakdown of the regex:
^(\w+): Captures the key (one or more word characters: letters, digits, underscores)\s*=\s*: Matches the equals sign with optional whitespace around it (remove\s*if your input has no whitespace)- The value group has three possible matches:
(?:'[^']*'): A single quoted string (e.g.,'Yoda')(?:TRUE|FALSE|\d+): Unquoted boolean (TRUE/FALSE) or numeric value (e.g.,3,72)(?:\((?:...)...\)): An array enclosed in parentheses, containing comma-separated elements (each element is either a quoted string, boolean, or numeric)
Step 3: Process Array Values
For values wrapped in parentheses, you'll want to extract and clean the inner elements:
- Strip the leading
(and trailing)from the value - Split the inner string using the same top-level comma regex from Step 1
- For each element:
- If it starts and ends with
', strip the quotes to get the raw string - If it's
TRUE/FALSE, convert to a boolean type - If it's digits, convert to a numeric type
- If it starts and ends with
Example Implementation (Python)
Here's how you might turn this into working code to get a clean dictionary of key-value pairs:
import re input_str = "TeacherName='Yoda',TeacherIsJedi=TRUE,TeachersAidNames=('Mace'),TeachersAidAlive=(FALSE),TeachersAidAges=(72),NumberOfStudents=3,StudentAges=(42,59,19)" # Split into key-value pairs pair_split_regex = r",(?=(?:[^']*'[^']*')*[^']*$)(?=(?:[^()]*\([^()]*\))*[^()]*$)" pairs = re.split(pair_split_regex, input_str) # Parse each pair into a dictionary key_value_regex = r"^(\w+)\s*=\s*((?:'[^']*')|(?:TRUE|FALSE|\d+)|(?:\((?:(?:'[^']*')|(?:TRUE|FALSE|\d+))(?:,\s*(?:'[^']*')|(?:TRUE|FALSE|\d+))*\)))$" result = {} for pair in pairs: match = re.match(key_value_regex, pair) if not match: continue # Skip invalid pairs if any key = match.group(1) value = match.group(2) # Handle array values if value.startswith('(') and value.endswith(')'): inner_content = value[1:-1].strip() elements = re.split(pair_split_regex, inner_content) processed_elements = [] for elem in elements: elem = elem.strip() if elem.startswith("'") and elem.endswith("'"): processed_elements.append(elem[1:-1]) elif elem == 'TRUE': processed_elements.append(True) elif elem == 'FALSE': processed_elements.append(False) elif elem.isdigit(): processed_elements.append(int(elem)) result[key] = processed_elements else: # Handle single values if value.startswith("'") and value.endswith("'"): result[key] = value[1:-1] elif value == 'TRUE': result[key] = True elif value == 'FALSE': result[key] = False elif value.isdigit(): result[key] = int(value) print(result)
Output:
{ 'TeacherName': 'Yoda', 'TeacherIsJedi': True, 'TeachersAidNames': ['Mace'], 'TeachersAidAlive': [False], 'TeachersAidAges': [72], 'NumberOfStudents': 3, 'StudentAges': [42, 59, 19] }
Notes:
- This regex assumes boolean values are exactly
TRUE/FALSE(uppercase). Adjust to(?:true|false|TRUE|FALSE)if you need to handle mixed/lowercase. - Numeric values are treated as integers here. To support floats, modify the numeric part of the regex to
\d+(?:\.\d+)?and update the processing code to usefloat()where appropriate. - The regex handles optional whitespace around equals signs and commas in arrays—remove
\s*if your input has no whitespace for better performance.
内容的提问来源于stack exchange,提问作者Kevin

