Python脚本运行Scrapy爬虫抓取成功但输出文件0KB无数据
问题描述
运行Python脚本启动Scrapy爬虫时,控制台显示数据可正常抓取,但最终生成的输出文件大小为0KB,文件中没有写入任何爬取到的数据。相关实现代码如下:
相关实现代码
Scrapy新闻爬虫主体
#Importing Scrapy library import scrapy #Defining spider's url,headers class DawnSpider(scrapy.Spider): name = 'dawn' allowed_domains = ['www.dawn.com'] #Channel link # start_urls = ['https://www.dawn.com/archive/2022-02-09'] # url = ['https://www.dawn.com'] # page = 1 # 定义设置请求头、指定爬取起始链接的函数 def start_requests(self): yield scrapy.Request(url='https://www.dawn.com/archive/2022-03-21', callback=self.parse, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'}) # 获取新闻标题及对应链接 def parse(self, response): titles = response.xpath("//h2[@class = 'story__title text-6 font-bold font-merriweather pt-1 pb-2 ']/a") for title in titles: headline = title.xpath(".//text()").get() headline_link = title.xpath(".//@href").get() # 遍历新闻标题链接 yield response.follow(url=headline_link, callback=self.parse_headline, meta={'heading': headline}, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'}) # 翻页访问历史页面代码 prev_page = response.xpath("//li[1]/a/@href").get() prev = 'https://www.dawn.com' + str(prev_page) yield scrapy.Request(url=prev, callback=self.parse, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'}) # 遍历标题链接,获取新闻详情与发布日期时间 def parse_headline(self, response): headline = response.request.meta['heading'] # logging.info(response.url) full_detail = response.xpath("//div[contains(@class , story__content)]/p[1]") date_and_time = response.xpath("//span[@class='timestamp--date']/text()").get() for detail in full_detail: data = detail.xpath(".//text()").get() yield { 'headline': headline, 'date_and_time': date_and_time, 'details': data }
独立Python运行脚本
from scrapy import cmdline cmdline.execute("scrapy crawl dawn -o data.csv".split(" "))
问题原因
共4个核心问题导致数据无法正常写入文件:
- XPath语法错误:
parse_headline方法中正文匹配的XPath//div[contains(@class , story__content)]/p[1]未给class属性值加引号,XPath解析直接失效,匹配不到任何内容,没有有效数据可yield。 - 翻页逻辑缩进错误:翻页代码写在了
for title in titles循环内部,每遍历一个新闻标题就会发起一次翻页请求,产生大量重复无效请求,极易触发站点反爬中断爬取流程。 - 翻页XPath匹配错误:当前翻页使用的
//li[1]/a/@href未限定匹配范围,会抓取页面所有li标签下第一个a标签的链接,大概率匹配不到正确的历史归档页地址,爬取会跳转到无关页面后中断。 - 输出配置缺失:直接导出csv默认编码不是utf-8,即使抓取到数据也可能因为编码兼容问题写入失败。
- 标题XPath容错性差:列表页标题匹配硬写了多段连续空格的class规则,只要页面模板调整空格数量就会完全匹配不到标题。
修复方案
1. 修正XPath语法与匹配规则
修正正文XPath的引号问题,同时将标题匹配改为contains模式降低匹配失败概率:
# 列表页标题匹配改为class包含匹配 titles = response.xpath("//h2[contains(@class, 'story__title')]/a") # 正文匹配补全class值的引号 full_detail = response.xpath("//div[contains(@class, 'story__content')]/p[1]")
2. 调整翻页逻辑位置与匹配规则
将翻页代码移到标题遍历循环外部,改用更精准的上一页按钮匹配规则,同时增加空值判断避免无效请求:
def parse(self, response): titles = response.xpath("//h2[contains(@class, 'story__title')]/a") for title in titles: headline = title.xpath(".//text()").get() headline_link = title.xpath(".//@href").get() yield response.follow(url=headline_link, callback=self.parse_headline, meta={'heading': headline}, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'}) # 翻页逻辑移到循环外,匹配rel属性为prev的上一页按钮 prev_page = response.xpath("//a[contains(@rel, 'prev')]/@href").get() if prev_page: yield response.follow(url=prev_page, callback=self.parse, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'})
3. 补全输出编码配置
修改启动脚本,指定csv导出编码为utf-8,避免写入失败:
from scrapy import cmdline cmdline.execute("scrapy crawl dawn -o data.csv -s FEED_EXPORT_ENCODING=utf-8".split(" "))
调试方法
如果修改后仍有问题,可以在parse_headline中增加打印语句,确认数据是否真的被抓取到:
def parse_headline(self, response): headline = response.request.meta['heading'] full_detail = response.xpath("//div[contains(@class, 'story__content')]/p[1]") date_and_time = response.xpath("//span[@class='timestamp--date']/text()").get() for detail in full_detail: data = detail.xpath(".//text()").get() # 打印确认抓取结果 print(f"抓取成功:标题[{headline}],发布时间[{date_and_time}],正文摘要[{data}]") yield { 'headline': headline, 'date_and_time': date_and_time, 'details': data }
内容的提问来源于stack exchange,提问作者SAAN
相关产品推荐
相关产品推荐

