如何在Python中识别[[ ]]分隔符内单词单复数并生成对应列表
Got it, let's tackle this problem step by step. You already know how to extract the words inside [[ ]], but the missing piece is detecting that trailing s (for plural) right after the closing brackets. Here's a straightforward way to get both the word list and the plural marker list:
The key is updating our regex to capture two pieces of information: the word inside the brackets, and any optional s that immediately follows the closing ]].
1. Updated Regex Pattern
Instead of just grabbing the inner word, we'll use a regex that captures two groups:
import re # Your example text input_text = "I have a red [[pen]], two blue [[pen]]s, two black [[pencil]]s and a green [[pencil]]" # Capture (word inside [[ ]], optional trailing 's') matches = re.findall(r'\[\[(.*?)\]\](s?)', input_text)
Let's break down the regex:
\[\[(.*?)\]\]: Non-greedily captures the word inside the double brackets (so it stops at the first]]instead of matching everything to the last one)(s?): Captures an optionals—this group will be either's'(for plural) or an empty string (for singular)
2. Build the Two Lists
Now we just loop through the matches to populate our lists:
word_list = [] plural_markers = [] for word, plural_suffix in matches: word_list.append(word) # Mark 1 if there's an 's' after the bracket, else 0 plural_markers.append(1 if plural_suffix else 0) # Let's check the output print("Word List:", word_list) print("Plural Markers:", plural_markers)
3. Verify the Result
Running this code with your example text will produce exactly what you need:
Word List: ['pen', 'pen', 'pencil', 'pencil']
Plural Markers: [0, 1, 1, 0]
Quick Notes for Edge Cases
- If there's whitespace between
]]and thes(like[[pen]] s), adjust the regex tor'\[\[(.*?)\]\]\s*(s?)'to allow optional spaces. - For plural forms that don't use an
s(like[[child]]ren), you'd need to expand the regex to match those suffixes, but this solution works perfectly for your given example.
内容的提问来源于stack exchange,提问作者Sali

