Scrapy爬虫执行无输出:未生成目标HTML文件问题求助
Scrapy爬虫未生成HTML文件的问题排查
你的代码里存在明显语法错误,导致爬虫无法正常请求目标页面,因此没有生成HTML文件:
- URL列表语法错误:
urls列表中两个URL字符串之间缺少逗号,Python会自动将相邻字符串拼接成一个无效地址'https://quotes.toscrape.com/page/1/https://quotes.toscrape.com/page/2/',Scrapy请求该无效地址失败,自然不会生成文件。
修正后的代码如下:
import scrapy class QuotesSpider(scrapy.Spider): name ="quotes" def start_requests(self): urls =[ 'https://quotes.toscrape.com/page/1/', 'https://quotes.toscrape.com/page/2/' ] for url in urls: yield scrapy.Request(url=url, callback=self.parse) def parse(self, response): page = response.url.split("/")[-2] filename = f'quotes-{page}.html' with open(filename, 'wb') as f: f.write(response.body) self.log(f'Saved file {filename}')
同时确认运行方式是否正确:
- 打开VS Code终端,进入包含
scrapy.cfg的Scrapy项目目录 - 执行命令
scrapy crawl quotes启动爬虫,不要直接右键运行该py文件
修正后重新运行,即可在当前目录下看到生成的quotes-1.html和quotes-2.html文件。
内容的提问来源于stack exchange,提问作者Nguyễn Vũ Trần Minh
相关产品推荐
相关产品推荐

