You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy环境下正则表达式匹配失败求助:无法提取课程缩写与名称

Troubleshooting Your Scrapy Regex Issue

First off, let's dive into why your regex works in testing tools but fails in Scrapy—this is a common pitfall with HTML entity handling!

Key Problem 1: Mishandling Non-Breaking Spaces

In your regex, you're targeting the string [ ], but Scrapy automatically resolves HTML entities like   to their actual Unicode character (U+00A0, a non-breaking space) when using text(). So the text you're trying to match doesn't contain the literal   string—it contains actual whitespace characters (non-breaking spaces) instead.

Key Problem 2: Redundant Repeated Capturing Group

Your course code group (\w+\s\w+)+ uses a + outside the parentheses, which in Python's re module will only capture the last iteration of the group (not the full course code like "ECON 114"). This is unnecessary here since your target course code is exactly one instance of \w+\s\w+.

Corrected Regex

Here's the fixed regex that addresses both issues:

r'(?i)(\w+\s\w+)\s*-\s*\w+[\s\u00A0]+([\w\s]+)'

Let's break down the changes:

  • \s*-\s*: Allows optional spaces around the hyphen (more flexible for varying HTML formatting)
  • [\s\u00A0]+: Matches any combination of regular spaces and non-breaking spaces (the actual characters present in Scrapy's extracted text)
  • Removed the redundant + from the course code group to ensure full capture of "ECON 114"

Additional Checks to Validate

  1. Ensure Text Isn't Split Across Nodes: If the <a> tag containing the course info has child elements (like <span>), text() might split the content into multiple strings. To fix this, use normalize-space() to get the full concatenated text for each course:
    course_texts = response.xpath('//a[contains(@id, "class_id")]/normalize-space()').getall()
    classes = []
    for text in course_texts:
        match = pythonRe.search(r'(?i)(\w+\s\w+)\s*-\s*\w+[\s\u00A0]+([\w\s]+)', text)
        if match:
            classes.extend(match.groups())
    
  2. Test with Actual Extracted Text: Add a quick debug line to see what text Scrapy is actually pulling:
    print(response.xpath('//a[contains(@id, "class_id")]/text()').getall())
    
    This will show you exactly what strings your regex is trying to match, helping you spot any unexpected formatting.

Adjusted Parse Method Code

Here's how your parse method might look with these fixes:

def parse(self, response):
    def professor_filter(item):
        if (pythonRe.search(r'\w\.', item) or "Staff" in item):
            return True

    # Compile regex once for efficiency
    class_regex = pythonRe.compile(r'(?i)(\w+\s\w+)\s*-\s*\w+[\s\u00A0]+([\w\s]+)')
    
    page = response.url.split("/")[-2]
    classDict = {}
    
    # Get full normalized text for each course link
    course_texts = response.xpath('//a[contains(@id, "class_id")]/normalize-space()').getall()
    classes = []
    for text in course_texts:
        match = class_regex.search(text)
        if match:
            classes.extend(match.groups())
    
    professors = response.xpath('//div[contains(@class, "col-xs-6 col-sm-3")]/text()').getall()
    professors_filtered = list(filter(professor_filter, professors))
    
    # Populate classDict (note: classes is a flat list of code/name pairs)
    for i in range(0, len(classes), 2):
        if i+1 < len(classes) and i < len(professors_filtered):
            classDict[classes[i]] = {'professor': professors_filtered[i]}
    
    print(classes)
    print(len(classes))
    print(professors_filtered)
    print(len(professors_filtered))
    print(classDict)
    
    filename = f'class-{page}.html'
    with open(filename, 'wb') as f:
        f.write(response.body)
    self.log(f'Saved file {filename}')

These changes should resolve the empty match array issue and correctly capture your course codes and names.

内容的提问来源于stack exchange,提问作者John Jacob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 23:44:07