Scrapy自定义parse_check回调函数未触发问题求助
问题:Scrapy定时请求中自定义parse_check回调未触发
我需要定时请求页面判断内容是否更新,但自定义的parse_check回调函数始终没触发。相关配置和代码如下:
配置信息
allowed_domains = ['www1.hkexnews.hk'] start_urls = 'https://www1.hkexnews.hk/search/predefineddoc.xhtml?lang=zh&predefineddocuments=9'
parse方法代码
# Crawl all data first at each start def parse(self, response): Total_records = int(re.findall("\d+",response.xpath("//div[@class='PD-TotalRecords']/text()").extract()[0])[0]) dict = {} is_Latest = True global Latest_info global previous_hash for i in range(1, Total_records + 1): content = response.xpath("//table/tbody/tr[{}]//text()".format(i)).extract() # Use the group function to group the list by key result = list(group(content, self.keys)) Time = dict['Time'] = result[0].get(self.keys[0]) Code = dict['Code'] = result[1].get(self.keys[1]) dict['Name'] = result[2].get(self.keys[2]) if is_Latest: Latest_info = str(Time) + " | " + str(Code) is_Latest = False yield dict previous_hash = get_hash(Latest_info.encode('utf-8')) # Monitor data updates and crawl for new data while True: time.sleep(10) # Request website content and calculate hash values yield scrapy.Request(url=self.start_urls, callback=self.parse_check, dont_filter=True)
parse_check回调代码
def parse_check(self, response): global previous_hash global Latest_info dict = {} content = response.xpath("//table/tbody/tr[1]//text()").extract() # Use the group function to group the list by key result = list(group(content, self.keys)) Time = result[0].get(self.keys[0]) Code = result[1].get(self.keys[1]) current_info = str(Time) + " | " + str(Code) current_hash = get_hash(current_info.encode('utf-8')) # Compare hash values to determine if website content is updated if current_hash != previous_hash: dict['Time'] = Time dict['Code'] = Code dict['Name'] = result[2].get(self.keys[2]) previous_hash = current_hash Latest_info = current_info yield dict
已尝试errback无输出,用requests.get能正常请求页面,但找不到回调未触发的原因。
排查与修复方案
核心问题:死循环阻塞Scrapy引擎
Scrapy是异步非阻塞框架,parse方法里的while True+time.sleep(10)会彻底卡住当前线程,导致后续生成的scrapy.Request根本无法被引擎调度执行。这种同步死循环的写法完全违背了Scrapy的异步设计逻辑。替换为Twisted定时器实现定时请求
利用Scrapy底层依赖的Twisted框架的定时器来实现定时任务,避免阻塞:from twisted.internet import reactor def parse(self, response): # 原有爬取全量数据的逻辑... # 初始化状态到实例属性(替代全局变量) self.latest_info = str(Time) + " | " + str(Code) self.previous_hash = get_hash(self.latest_info.encode('utf-8')) # 启动定时检查 self.schedule_next_check() def schedule_next_check(self): # 10秒后触发下一次检查请求 reactor.callLater(10, self.send_check_request) def send_check_request(self): # 生成检查请求 yield scrapy.Request( url=self.start_urls, callback=self.parse_check, dont_filter=True ) # 递归调度下一次检查 self.schedule_next_check()移除全局变量,改用实例属性
全局变量在异步环境下容易出现竞态问题,把previous_hash和Latest_info改为Spider的实例属性(self.previous_hash、self.latest_info),避免多请求之间的状态干扰。添加异常捕获排查回调错误
在parse_check中加入异常捕获,确认是否是回调内部报错导致无输出:def parse_check(self, response): try: content = response.xpath("//table/tbody/tr[1]//text()").extract() result = list(group(content, self.keys)) Time = result[0].get(self.keys[0]) Code = result[1].get(self.keys[1]) current_info = str(Time) + " | " + str(Code) current_hash = get_hash(current_info.encode('utf-8')) if current_hash != self.previous_hash: yield { 'Time': Time, 'Code': Code, 'Name': result[2].get(self.keys[2]) } self.previous_hash = current_hash self.latest_info = current_info except Exception as e: self.logger.error(f"parse_check执行出错: {str(e)}", exc_info=True)确认Request去重机制
虽然已经添加dont_filter=True,但可以在Scrapy配置中临时关闭去重(DUPEFILTER_CLASS = 'scrapy.dupefilters.BaseDupeFilter'),确认是否是去重规则影响了请求调度。
内容的提问来源于stack exchange,提问作者yanis
相关产品推荐
相关产品推荐

