Scrapy+Firefox无法抓取页面URL:目标列表为空问题排查
问题描述
上月可正常运行的Scrapy爬虫(配合Firefox)本月失效,无法从目标允许爬取的网站获取URL列表,最终url_authority_list为空。爬虫日志显示已成功获取目标页面(状态码200),但XPath未匹配到任何元素。
运行日志
2024-06-29 14:30:25 [scrapy.extensions.telnet] INFO: Telnet console listening on 127.0.0.1:6023 2024-06-29 14:30:25 [filelock] DEBUG: Attempting to acquire lock 2630883719752 on C:\Users\andre\Anaconda3\lib\site-packages\tldextract\.suffix_cache/publicsuffix.org-tlds\de84b5ca2167d4c83e38fb162f2e8738.tldextract.json.lock 2024-06-29 14:30:25 [filelock] DEBUG: Lock 2630883719752 acquired on C:\Users\andre\Anaconda3\lib\site-packages\tldextract\.suffix_cache/publicsuffix.org-tlds\de84b5ca2167d4c83e38fb162f2e8738.tldextract.json.lock 2024-06-29 14:30:25 [filelock] DEBUG: Attempting to acquire lock 2630883914632 on C:\Users\andre\Anaconda3\lib\site-packages\tldextract\.suffix_cache/urls\62bf135d1c2f3d4db4228b9ecaf507a2.tldextract.json.lock 2024-06-29 14:30:25 [filelock] DEBUG: Lock 2630883914632 acquired on C:\Users\andre\Anaconda3\lib\site-packages\tldextract\.suffix_cache/urls\62bf135d1c2f3d4db4228b9ecaf507a2.tldextract.json.lock 2024-06-29 14:30:25 [filelock] DEBUG: Attempting to release lock 2630883914632 on C:\Users\andre\Anaconda3\lib\site-packages\tldextract\.suffix_cache/urls\62bf135d1c2f3d4db4228b9ecaf507a2.tldextract.json.lock 2024-06-29 14:30:25 [filelock] DEBUG: Lock 2630883914632 released on C:\Users\andre\Anaconda3\lib\site-packages\tldextract\.suffix_cache/urls\62bf135d1c2f3d4db4228b9ecaf507a2.tldextract.json.lock 2024-06-29 14:30:25 [filelock] DEBUG: Attempting to release lock 2630883719752 on C:\Users\andre\Anaconda3\lib\site-packages\tldextract\.suffix_cache/publicsuffix.org-tlds\de84b5ca2167d4c83e38fb162f2e8738.tldextract.json.lock 2024-06-29 14:30:25 [filelock] DEBUG: Lock 2630883719752 released on C:\Users\andre\Anaconda3\lib\site-packages\tldextract\.suffix_cache/publicsuffix.org-tlds\de84b5ca2167d4c83e38fb162f2e8738.tldextract.json.lock 2024-06-29 14:30:25 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://ratings.food.gov.uk/open-data> (referer: None)
原代码
import scrapy from urllib.parse import urljoin import os class foodstandardsagencySpider(scrapy.Spider): name = "foodstandardsagency" allowed_domains = ["ratings.food.gov.uk"] start_urls = ["https://ratings.food.gov.uk/open-data"] # will need a new token if using geckodriver os.environ['GH_TOKEN'] = "ghp_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx" def parse(self, response): #create a list of urls url_authority_list = [] for href in response.xpath('//tr/td/a[text()[contains(.,"English")]]/@href'): url = urljoin('http://ratings.food.gov.uk/',href.extract()) url_authority_list.append(url) print(url_authority_list)
问题排查与修复
核心原因
- 目标网站页面结构变更,原XPath表达式已无法匹配到包含"English"的链接元素;
- URL拼接使用的
http协议与目标网站的https协议不一致,可能引发路径错误。
修复步骤
调试XPath有效性:
执行Scrapy Shell命令进入调试环境:scrapy shell https://ratings.food.gov.uk/open-data在Shell中测试XPath,确认是否能匹配到元素:
response.xpath('//tr/td/a[text()[contains(.,"English")]]/@href').getall()若返回空列表,说明页面结构已变,可调整XPath为忽略大小写的鲁棒写法,适配文本格式变化:
//tbody/tr/td/a[contains(translate(text(), 'ABCDEFGHIJKLMNOPQRSTUVWXYZ', 'abcdefghijklmnopqrstuvwxyz'), 'english')]/@href修正URL协议:将
urljoin的基础URL改为https://ratings.food.gov.uk/,避免协议不一致问题。规范元素提取:使用
getall()批量获取所有匹配的href,替代循环单个节点的写法,同时用get()替代extract()(更符合Scrapy最佳实践)。
修正后的代码
import scrapy from urllib.parse import urljoin import os class foodstandardsagencySpider(scrapy.Spider): name = "foodstandardsagency" allowed_domains = ["ratings.food.gov.uk"] start_urls = ["https://ratings.food.gov.uk/open-data"] # will need a new token if using geckodriver os.environ['GH_TOKEN'] = "ghp_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx" def parse(self, response): url_authority_list = [] # 忽略大小写匹配包含English的链接,适配页面结构变化 hrefs = response.xpath('//tbody/tr/td/a[contains(translate(text(), "ABCDEFGHIJKLMNOPQRSTUVWXYZ", "abcdefghijklmnopqrstuvwxyz"), "english")]/@href').getall() for href in hrefs: url = urljoin('https://ratings.food.gov.uk/', href) url_authority_list.append(url) print(url_authority_list)
额外建议
- 静态爬虫失效优先排查页面结构变化,这是最常见的失效原因;
- 尽量结合元素的class、id属性编写XPath,减少对文本内容的依赖,提升鲁棒性;
- 开启Scrapy的
DEBUG日志等级,查看更多请求响应细节,便于定位问题。
内容的提问来源于stack exchange,提问作者nevster
相关产品推荐
相关产品推荐

