正则分组匹配词边界问题:无法提取赛事名称,求修正方案
Fixing the Regex for Post-Match Thread Extraction
Your original regex is failing for a few key reasons, and we can adjust it to correctly extract the home team, away team, and competition (EuroLeague) from the thread title. Let's break down the fixes and solution:
Why the Original Regex Failed
- Missing spaces around the hyphen: The title uses
Barcelona - Olimpia Milano(with spaces around the hyphen), but your regex uses-(.*?)which looks for a hyphen without surrounding spaces—so it couldn't find the team separator. - Missing closing bracket: The regex ends with
\binstead of matching the closing]at the end of the competition section, so it couldn't fully match the title structure. - Unanchored match: Without start (
^) and end ($) anchors, the regex might fail to match if there's any unexpected content, even if most parts align. - Overly broad competition capture: Your original competition group tried to capture everything inside the brackets, but you only need the first word (
EuroLeague), and the non-greedy match without a clear end point led to no valid capture.
Corrected Regex
Here's the adjusted regex that addresses all these issues:
title_group_regex = r"^Post-Match Thread: (?P<home_team>.*?) - (?P<away_team>.*?) \[(?P<competition>\w+).*?\]$"
Regex Breakdown
^Post-Match Thread:: Anchors the match to the start of the string, ensuring we only process valid post-match threads.(?P<home_team>.*?): Non-greedily captures the home team name until it hits the-separator.-: Explicitly matches the space-hyphen-space separator between home and away teams.(?P<away_team>.*?): Non-greedily captures the away team name until it hits the[(space + opening bracket) leading to the competition details.\[(?P<competition>\w+).*?\]$: Captures the first word inside the brackets (your desiredEuroLeague) using\w+, then ignores the rest of the text inside the brackets until the closing]and end of the string.
Usage Example
Here's how to implement this correctly, including strip() to handle any accidental leading/trailing whitespace:
import re thread_title = "Post-Match Thread: Barcelona - Olimpia Milano [EuroLeague Regular Season, Round 24]" title_group_regex = r"^Post-Match Thread: (?P<home_team>.*?) - (?P<away_team>.*?) \[(?P<competition>\w+).*?\]$" match = re.search(title_group_regex, thread_title) if match: home_team = match.group('home_team').strip() away_team = match.group('away_team').strip() competition = match.group('competition').strip() print(f"Home Team: {home_team}") print(f"Away Team: {away_team}") print(f"Competition: {competition}") else: print("No match found—check the title format.")
Output
Home Team: Barcelona Away Team: Olimpia Milano Competition: EuroLeague
This will avoid the NoneType error because the regex now fully matches the title structure, and correctly extracts all three pieces of data you need.
内容的提问来源于stack exchange,提问作者Joao Pereira
相关产品推荐
相关产品推荐

