Scrapy环境下正则表达式匹配失败求助:无法提取课程缩写与名称
First off, let's dive into why your regex works in testing tools but fails in Scrapy—this is a common pitfall with HTML entity handling!
Key Problem 1: Mishandling Non-Breaking Spaces
In your regex, you're targeting the string [ ], but Scrapy automatically resolves HTML entities like to their actual Unicode character (U+00A0, a non-breaking space) when using text(). So the text you're trying to match doesn't contain the literal   string—it contains actual whitespace characters (non-breaking spaces) instead.
Key Problem 2: Redundant Repeated Capturing Group
Your course code group (\w+\s\w+)+ uses a + outside the parentheses, which in Python's re module will only capture the last iteration of the group (not the full course code like "ECON 114"). This is unnecessary here since your target course code is exactly one instance of \w+\s\w+.
Corrected Regex
Here's the fixed regex that addresses both issues:
r'(?i)(\w+\s\w+)\s*-\s*\w+[\s\u00A0]+([\w\s]+)'
Let's break down the changes:
\s*-\s*: Allows optional spaces around the hyphen (more flexible for varying HTML formatting)[\s\u00A0]+: Matches any combination of regular spaces and non-breaking spaces (the actual characters present in Scrapy's extracted text)- Removed the redundant
+from the course code group to ensure full capture of "ECON 114"
Additional Checks to Validate
- Ensure Text Isn't Split Across Nodes: If the
<a>tag containing the course info has child elements (like<span>),text()might split the content into multiple strings. To fix this, usenormalize-space()to get the full concatenated text for each course:course_texts = response.xpath('//a[contains(@id, "class_id")]/normalize-space()').getall() classes = [] for text in course_texts: match = pythonRe.search(r'(?i)(\w+\s\w+)\s*-\s*\w+[\s\u00A0]+([\w\s]+)', text) if match: classes.extend(match.groups()) - Test with Actual Extracted Text: Add a quick debug line to see what text Scrapy is actually pulling:
This will show you exactly what strings your regex is trying to match, helping you spot any unexpected formatting.print(response.xpath('//a[contains(@id, "class_id")]/text()').getall())
Adjusted Parse Method Code
Here's how your parse method might look with these fixes:
def parse(self, response): def professor_filter(item): if (pythonRe.search(r'\w\.', item) or "Staff" in item): return True # Compile regex once for efficiency class_regex = pythonRe.compile(r'(?i)(\w+\s\w+)\s*-\s*\w+[\s\u00A0]+([\w\s]+)') page = response.url.split("/")[-2] classDict = {} # Get full normalized text for each course link course_texts = response.xpath('//a[contains(@id, "class_id")]/normalize-space()').getall() classes = [] for text in course_texts: match = class_regex.search(text) if match: classes.extend(match.groups()) professors = response.xpath('//div[contains(@class, "col-xs-6 col-sm-3")]/text()').getall() professors_filtered = list(filter(professor_filter, professors)) # Populate classDict (note: classes is a flat list of code/name pairs) for i in range(0, len(classes), 2): if i+1 < len(classes) and i < len(professors_filtered): classDict[classes[i]] = {'professor': professors_filtered[i]} print(classes) print(len(classes)) print(professors_filtered) print(len(professors_filtered)) print(classDict) filename = f'class-{page}.html' with open(filename, 'wb') as f: f.write(response.body) self.log(f'Saved file {filename}')
These changes should resolve the empty match array issue and correctly capture your course codes and names.
内容的提问来源于stack exchange,提问作者John Jacob

