You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy自定义parse_check回调函数未触发问题求助

问题:Scrapy定时请求中自定义parse_check回调未触发

我需要定时请求页面判断内容是否更新,但自定义的parse_check回调函数始终没触发。相关配置和代码如下:

配置信息

allowed_domains = ['www1.hkexnews.hk']
start_urls = 'https://www1.hkexnews.hk/search/predefineddoc.xhtml?lang=zh&predefineddocuments=9'

parse方法代码

# Crawl all data first at each start
def parse(self, response):
    Total_records = int(re.findall("\d+",response.xpath("//div[@class='PD-TotalRecords']/text()").extract()[0])[0])
    dict = {}
    is_Latest = True
    global Latest_info
    global previous_hash

    for i in range(1, Total_records + 1):
        content = response.xpath("//table/tbody/tr[{}]//text()".format(i)).extract()

        # Use the group function to group the list by key
        result = list(group(content, self.keys))
        Time = dict['Time'] = result[0].get(self.keys[0])
        Code = dict['Code'] = result[1].get(self.keys[1])
        dict['Name'] = result[2].get(self.keys[2])
        if is_Latest:
            Latest_info = str(Time) + " | " + str(Code)
            is_Latest = False

        yield dict

    previous_hash = get_hash(Latest_info.encode('utf-8'))
    # Monitor data updates and crawl for new data
    while True:
        time.sleep(10)
        # Request website content and calculate hash values
        yield scrapy.Request(url=self.start_urls, callback=self.parse_check, dont_filter=True)

parse_check回调代码

def parse_check(self, response):
    global previous_hash
    global Latest_info
    dict = {}
    content = response.xpath("//table/tbody/tr[1]//text()").extract()
    # Use the group function to group the list by key
    result = list(group(content, self.keys))
    Time =  result[0].get(self.keys[0])
    Code = result[1].get(self.keys[1])

    current_info = str(Time) + " | " + str(Code)
    current_hash = get_hash(current_info.encode('utf-8'))

    # Compare hash values to determine if website content is updated
    if current_hash != previous_hash:

        dict['Time'] = Time
        dict['Code'] = Code
        dict['Name'] = result[2].get(self.keys[2])

        previous_hash = current_hash
        Latest_info = current_info
    yield dict

已尝试errback无输出,用requests.get能正常请求页面,但找不到回调未触发的原因。


排查与修复方案
  • 核心问题:死循环阻塞Scrapy引擎
    Scrapy是异步非阻塞框架,parse方法里的while True+time.sleep(10)会彻底卡住当前线程,导致后续生成的scrapy.Request根本无法被引擎调度执行。这种同步死循环的写法完全违背了Scrapy的异步设计逻辑。

  • 替换为Twisted定时器实现定时请求
    利用Scrapy底层依赖的Twisted框架的定时器来实现定时任务,避免阻塞:

    from twisted.internet import reactor
    
    def parse(self, response):
        # 原有爬取全量数据的逻辑...
        
        # 初始化状态到实例属性(替代全局变量)
        self.latest_info = str(Time) + " | " + str(Code)
        self.previous_hash = get_hash(self.latest_info.encode('utf-8'))
        # 启动定时检查
        self.schedule_next_check()
    
    def schedule_next_check(self):
        # 10秒后触发下一次检查请求
        reactor.callLater(10, self.send_check_request)
    
    def send_check_request(self):
        # 生成检查请求
        yield scrapy.Request(
            url=self.start_urls,
            callback=self.parse_check,
            dont_filter=True
        )
        # 递归调度下一次检查
        self.schedule_next_check()
    
  • 移除全局变量,改用实例属性
    全局变量在异步环境下容易出现竞态问题,把previous_hash和Latest_info改为Spider的实例属性(self.previous_hash、self.latest_info),避免多请求之间的状态干扰。

  • 添加异常捕获排查回调错误
    在parse_check中加入异常捕获,确认是否是回调内部报错导致无输出:

    def parse_check(self, response):
        try:
            content = response.xpath("//table/tbody/tr[1]//text()").extract()
            result = list(group(content, self.keys))
            Time = result[0].get(self.keys[0])
            Code = result[1].get(self.keys[1])
    
            current_info = str(Time) + " | " + str(Code)
            current_hash = get_hash(current_info.encode('utf-8'))
    
            if current_hash != self.previous_hash:
                yield {
                    'Time': Time,
                    'Code': Code,
                    'Name': result[2].get(self.keys[2])
                }
                self.previous_hash = current_hash
                self.latest_info = current_info
        except Exception as e:
            self.logger.error(f"parse_check执行出错: {str(e)}", exc_info=True)
    
  • 确认Request去重机制
    虽然已经添加dont_filter=True,但可以在Scrapy配置中临时关闭去重(DUPEFILTER_CLASS = 'scrapy.dupefilters.BaseDupeFilter'),确认是否是去重规则影响了请求调度。


内容的提问来源于stack exchange,提问作者yanis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 02:57:51