Scrapy脚本与Shell执行结果不一致问题排查
问题原因及解决办法
核心问题:请求URL格式错误
你的爬虫start_urls里的URL包含了HTML实体转义字符("和&),这会导致Scrapy请求的页面和你在Scrapy Shell中测试的页面完全不是同一个页面。
你在Shell里应该用了正确的URL(没有这些多余的转义字符),所以能正常提取到£250;但爬虫里的URL被转义后,实际请求的页面根本没有目标内容,自然返回None。
比如你当前的URL片段:
https://dvlaregistrations.dvla.gov.uk/search/results.html?search=CO11CTD"&"action=index"&"pricefrom=0"...
正确的URL需要把"全部删除,&替换成&,修正后如下:
https://dvlaregistrations.dvla.gov.uk/search/results.html?search=CO11CTD&action=index&pricefrom=0&priceto=&prefixmatches=¤tmatches=&limitprefix=&limitcurrent=&limitauction=&searched=true&openoption=&language=en&prefix2=Search&super=&super_pricefrom=&super_priceto=
次要问题:代码语法错误
你在循环里写的yield print(element.css('::text').get())是错误的:
print()函数返回None,yield None没有任何意义,应该返回实际提取到的数据,比如yield {'price': element.get()}- 另外
element已经是p::text的选择结果,直接用element.get()就能拿到文本,没必要再调用element.css('::text').get()
修正后的完整代码
import scrapy from scrapy.crawler import CrawlerProcess class TestSpiderSpider(scrapy.Spider): name = 'test_spider' allowed_domains = ['dvlaregistrations.direct.gov.uk'] # 替换为修正后的正确URL start_urls = [ 'https://dvlaregistrations.dvla.gov.uk/search/results.html?search=CO11CTD&action=index&pricefrom=0&priceto=&prefixmatches=¤tmatches=&limitprefix=&limitcurrent=&limitauction=&searched=true&openoption=&language=en&prefix2=Search&super=&super_pricefrom=&super_priceto=' ] def parse(self, response): price = response.css('div.resultsstrip p::text').get() print(price) print('---+---') all_prices = response.css('div.resultsstrip p::text') for element in all_prices: # 正确返回提取到的价格数据 yield {'price': element.get()} process = CrawlerProcess() process.crawl(TestSpiderSpider) process.start()
验证步骤
- 把修正后的URL复制到Scrapy Shell中测试,确认能正常获取
£250 - 运行修正后的爬虫,此时
print(price)会输出预期的价格值
内容的提问来源于stack exchange,提问作者xlmaster
相关产品推荐
相关产品推荐

