使用正则表达式去除冗余空格以获取预期输出的技术问题求助
Got it, let's work through this problem together. The core issue with your current approach is that splitting on spaces treats the square brackets and the number inside as separate elements—since you're replacing all excess spaces with single spaces, test3 [ 23 ] becomes test3 [ 23 ] (with single spaces around the brackets), then splitting on spaces breaks it into four parts instead of keeping the bracketed content as one unified value.
Here's a step-by-step solution that handles both regular key-value pairs and the bracketed case correctly:
Step 1: Split the Raw Data into Individual Entries
First, split your input using ; as the delimiter (the actual separator between your key-value pairs), and clean up any leading/trailing whitespace from each entry:
import re raw_input = "test1 21; test2 22; test3 [ 23 ];" # Split by semicolons, filter out empty strings, and strip whitespace from each entry entries = [entry.strip() for entry in raw_input.split(';') if entry.strip()]
Step 2: Parse Each Entry into Key-Value Pairs
For each entry, extract the key (the text before the first space) and handle the value differently depending on whether it's bracketed:
parsed_result = [] for entry in entries: # Split the entry only once on the first space—keeps the rest of the content intact key, value_part = entry.split(maxsplit=1) # Check if the value is enclosed in brackets if value_part.startswith('[') and value_part.endswith(']'): # Clean up the bracketed value: remove brackets, strip whitespace, wrap in a list cleaned_value = [value_part.strip('[]').strip()] else: # Regular value, keep it as is cleaned_value = value_part parsed_result.append([key, cleaned_value])
Step 3: Output the Result
If you print each item in parsed_result, you'll get exactly what you're expecting:
for item in parsed_result: print(item)
Output:
['test1', '21'] ['test2', '22'] ['test3', ['23']]
Why This Works Better Than Your Original Approach
Instead of globally replacing spaces and splitting blindly, we:
- First isolate each entry using the semicolon delimiter (the logical separator between your pairs).
- Split each entry only once on the first space—this ensures everything after the key stays together, even if it has internal spaces (like inside the brackets).
- Explicitly handle the bracketed value case to clean it up and format it as a list, matching your desired output.
If you prefer a regex-focused approach, you can use a pattern to match both key and value in one go:
parsed_result = [] pattern = re.compile(r'^(\w+)\s+(?:\[?\s*(\d+)\s*\]?)$') for entry in entries: match = pattern.match(entry) if match: key = match.group(1) num = match.group(2) # Wrap in list if brackets were present in the original entry value = [num] if '[' in entry else num parsed_result.append([key, value])
Either method will avoid splitting the bracketed value into separate elements and give you the formatted list you need.
内容的提问来源于stack exchange,提问作者Tony Montana

