编写正则表达式解析含引号子串的查询并返回嵌套列表
Hey there! I know regex can feel overwhelming when you're starting out, but let's break this problem down into manageable parts. Your goal is to split a string like "green lizards" like to sit "in the sun" into a nested list where quoted substrings become inner lists of words, and regular words stay as strings—totally doable with a targeted regex and a bit of post-processing.
Step 1: The Regex Pattern
First, we need a regex that can match two types of elements:
- Substrings wrapped in single or double quotes
- Regular words (sequences of non-whitespace characters)
Here's the pattern we'll use:
"([^"]+)"|'([^']+)'|(\S+)
Let's break down each part:
"([^"]+)": Matches text wrapped in double quotes. The[^"]+means "one or more characters that are NOT a double quote", and the parentheses capture the text inside the quotes (so we can extract it later).'([^']+)': Same as above, but for single quotes.(\S+): Matches any sequence of non-whitespace characters (your regular words), capturing them as a group.
The | acts like an OR—so the regex will try to match each part in order, stopping at the first match it finds.
Step 2: Processing the Matches
Once we have all the matches, we need to convert them into your desired nested list structure. Let's use Python for this example (it's readable and widely used):
import re def parse_quoted_text(input_string): # Define our regex pattern pattern = r'"([^"]+)"|\'([^\']+)\'|(\S+)' # Get all matches from the input string all_matches = re.findall(pattern, input_string) result = [] for match_group in all_matches: # Check which group has content (only one will per match) if match_group[0]: # Double-quoted content: split into words and add as a list result.append(match_group[0].split()) elif match_group[1]: # Single-quoted content: same as above result.append(match_group[1].split()) else: # Regular word: add as a string result.append(match_group[2]) return result # Test it out! test_input = '"green lizards" like to sit "in the sun"' print(parse_quoted_text(test_input)) # Output: [['green', 'lizards'], 'like', 'to', 'sit', ['in', 'the', 'sun']]
Handling Edge Cases (Optional)
If you need to handle escaped quotes (like "She said \"Hi!\""), we'll need to adjust the regex to ignore escaped quotes inside the quoted substring. Here's the updated pattern:
"((?:\\.|[^"])+)"|'((?:\\.|[^'])+)'|(\S+)
The (?:\\.|[^"])+ means:
\\.: Match any escaped character (like\"or\')[^"]: Match any character that's not a quote- The
(?:...)is a non-capturing group—we use it to group these two options without creating an extra capture group.
Final Notes
- The regex works because it prioritizes matching quoted substrings first (before regular words), so it won't split quoted content by spaces prematurely.
- The post-processing step is simple: we just check which capture group has content for each match, then convert quoted content into a list of words.
内容的提问来源于stack exchange,提问作者J. Taylor

