Python正则表达式提取关键词前括号内的冠军与挑战者信息
Got it, let's figure out why your regex is only catching the underdog and not the champion. The key here is to craft a pattern that matches both the champion and underdog patterns consistently, capturing the name from the parentheses right before each keyword.
First, let's use a typical example of the text you're working with (mirroring the structure you described):
"(Mike Tyson) champion, the heavyweight title holder from Brooklyn, NY. (Kamil Kubaru) underdog, the challenger from Alexandria, Virginia."
The Correct Regex Pattern
We'll use named capture groups to make it easy to map results to roles:
\((?P<name>.*?)\)\s+(?P<role>champion|underdog)
Let's break down what each part does:
\(: Escapes the opening parenthesis (since parentheses are special in regex syntax)(?P<name>.*?): A named groupnamethat captures the content inside the parentheses non-greedily (this ensures it stops at the closing parenthesis instead of over-matching other text)\): Escapes the closing parenthesis\s+: Matches one or more spaces between the parentheses and the keyword (handles any extra spacing that might exist)(?P<role>champion|underdog): A named grouprolethat matches either "champion" or "underdog"
Example Implementation (Python)
Here's how you can use this regex to extract both the champion and challenger:
import re # Replace this with your actual input text input_text = "(Mike Tyson) champion, the heavyweight title holder from Brooklyn, NY. (Kamil Kubaru) underdog, the challenger from Alexandria, Virginia." # Our regex pattern pattern = r"\((?P<name>.*?)\)\s+(?P<role>champion|underdog)" # Find all matches in the text matches = re.finditer(pattern, input_text) # Organize results into a readable dictionary fight_roles = {} for match in matches: role = match.group("role") name = match.group("name") if role == "champion": fight_roles["current_champion"] = name elif role == "underdog": fight_roles["challenger"] = name print(fight_roles) # Output: {'current_champion': 'Mike Tyson', 'challenger': 'Kamil Kubaru'}
Why Your Previous Regex Failed
Chances are one of these issues was tripping you up:
- You only targeted the
underdogkeyword instead of including bothchampionandunderdogin your pattern - You used
re.search()instead ofre.finditer()/re.findall()—search()only finds the first match, so you'd miss one role if it came after the other - Your grouping logic was incorrect, so you weren't properly capturing the name associated with the champion keyword
This pattern should work for any text where the name is in parentheses immediately followed by either "champion" or "underdog" (with any amount of spacing in between).
内容的提问来源于stack exchange,提问作者i.n.n.m

