You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本运行Scrapy爬虫抓取成功但输出文件0KB无数据

问题描述

运行Python脚本启动Scrapy爬虫时,控制台显示数据可正常抓取,但最终生成的输出文件大小为0KB,文件中没有写入任何爬取到的数据。相关实现代码如下:

相关实现代码

Scrapy新闻爬虫主体

#Importing Scrapy library
import scrapy


#Defining spider's url,headers
class DawnSpider(scrapy.Spider):
    name = 'dawn'
    allowed_domains = ['www.dawn.com']    #Channel link
    # start_urls = ['https://www.dawn.com/archive/2022-02-09']    
    # url = ['https://www.dawn.com']
    # page = 1

    # 定义设置请求头、指定爬取起始链接的函数
    def start_requests(self):
        yield scrapy.Request(url='https://www.dawn.com/archive/2022-03-21', callback=self.parse, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'})

    # 获取新闻标题及对应链接
    def parse(self, response):
        titles = response.xpath("//h2[@class = 'story__title      text-6  font-bold  font-merriweather      pt-1  pb-2  ']/a")    

        for title in titles:
            headline = title.xpath(".//text()").get()
            headline_link = title.xpath(".//@href").get()
# 遍历新闻标题链接

            yield response.follow(url=headline_link,  callback=self.parse_headline, meta={'heading': headline}, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'})

# 翻页访问历史页面代码
            prev_page = response.xpath("//li[1]/a/@href").get()
            prev = 'https://www.dawn.com' + str(prev_page)

            yield scrapy.Request(url=prev, callback=self.parse, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'})
            
    # 遍历标题链接,获取新闻详情与发布日期时间
    def parse_headline(self, response):
        headline = response.request.meta['heading']
        # logging.info(response.url)
        full_detail = response.xpath("//div[contains(@class , story__content)]/p[1]")
        date_and_time = response.xpath("//span[@class='timestamp--date']/text()").get()
        for detail in full_detail:
            data = detail.xpath(".//text()").get()
            yield {
                'headline': headline,
                'date_and_time': date_and_time,
                'details': data
            }

独立Python运行脚本

from scrapy import cmdline

cmdline.execute("scrapy crawl dawn -o data.csv".split(" "))
问题原因

共4个核心问题导致数据无法正常写入文件:

  • XPath语法错误:parse_headline方法中正文匹配的XPath//div[contains(@class , story__content)]/p[1]未给class属性值加引号,XPath解析直接失效,匹配不到任何内容,没有有效数据可yield。
  • 翻页逻辑缩进错误:翻页代码写在了for title in titles循环内部,每遍历一个新闻标题就会发起一次翻页请求,产生大量重复无效请求,极易触发站点反爬中断爬取流程。
  • 翻页XPath匹配错误:当前翻页使用的//li[1]/a/@href未限定匹配范围,会抓取页面所有li标签下第一个a标签的链接,大概率匹配不到正确的历史归档页地址,爬取会跳转到无关页面后中断。
  • 输出配置缺失:直接导出csv默认编码不是utf-8,即使抓取到数据也可能因为编码兼容问题写入失败。
  • 标题XPath容错性差:列表页标题匹配硬写了多段连续空格的class规则,只要页面模板调整空格数量就会完全匹配不到标题。
修复方案

1. 修正XPath语法与匹配规则

修正正文XPath的引号问题,同时将标题匹配改为contains模式降低匹配失败概率:

# 列表页标题匹配改为class包含匹配
titles = response.xpath("//h2[contains(@class, 'story__title')]/a")

# 正文匹配补全class值的引号
full_detail = response.xpath("//div[contains(@class, 'story__content')]/p[1]")

2. 调整翻页逻辑位置与匹配规则

将翻页代码移到标题遍历循环外部,改用更精准的上一页按钮匹配规则,同时增加空值判断避免无效请求:

def parse(self, response):
    titles = response.xpath("//h2[contains(@class, 'story__title')]/a")    
    for title in titles:
        headline = title.xpath(".//text()").get()
        headline_link = title.xpath(".//@href").get()
        yield response.follow(url=headline_link,  callback=self.parse_headline, meta={'heading': headline}, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'})

    # 翻页逻辑移到循环外,匹配rel属性为prev的上一页按钮
    prev_page = response.xpath("//a[contains(@rel, 'prev')]/@href").get()
    if prev_page:
        yield response.follow(url=prev_page, callback=self.parse, headers={'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'})

3. 补全输出编码配置

修改启动脚本,指定csv导出编码为utf-8,避免写入失败:

from scrapy import cmdline
cmdline.execute("scrapy crawl dawn -o data.csv -s FEED_EXPORT_ENCODING=utf-8".split(" "))

调试方法

如果修改后仍有问题,可以在parse_headline中增加打印语句,确认数据是否真的被抓取到:

def parse_headline(self, response):
    headline = response.request.meta['heading']
    full_detail = response.xpath("//div[contains(@class, 'story__content')]/p[1]")
    date_and_time = response.xpath("//span[@class='timestamp--date']/text()").get()
    for detail in full_detail:
        data = detail.xpath(".//text()").get()
        # 打印确认抓取结果
        print(f"抓取成功:标题[{headline}],发布时间[{date_and_time}],正文摘要[{data}]")
        yield {
            'headline': headline,
            'date_and_time': date_and_time,
            'details': data
        }

内容的提问来源于stack exchange,提问作者SAAN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 15:54:07