You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy在Jupyter Notebook中导出空CSV文件问题求助

解决Scrapy爬虫生成空CSV文件的问题

你的爬虫之所以生成空CSV,核心原因是没有提取并输出任何可被保存的数据:当前代码仅发起了详情页的请求,但既没有在首页提取数据,也没有定义详情页的处理逻辑来提取数据,Scrapy没有收集到任何可写入文件的内容。

修改后的代码示例(提取首页国家数据)

import scrapy

class WorldometersSpider(scrapy.Spider):
    name = 'worldometers'
    allowed_domains = ['www.worldometers.info']
    start_urls = ['https://www.worldometers.info/world-population/population-by-country/']
    
    def parse(self, response):
        # 定位首页表格的每一行数据
        rows = response.xpath('//table[@id="example2"]/tbody/tr')
        for row in rows:
            # 提取需要的字段并yield字典(这是Scrapy写入CSV的核心)
            yield {
                '国家名称': row.xpath('.//td[2]/a/text()').get(),
                '2023年人口': row.xpath('.//td[3]/text()').get(),
                '年增长率': row.xpath('.//td[4]/text()').get(),
                '年度净增长': row.xpath('.//td[5]/text()').get(),
                '详情页链接': response.urljoin(row.xpath('.//td[2]/a/@href').get())
            }

from scrapy.crawler import CrawlerProcess
process = CrawlerProcess(settings = {
    'FEEDS': {'output3.csv': {'format':'csv', 'encoding': 'utf-8'}},
    # 添加UA头避免被反爬识别
    'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
})
process.crawl(WorldometersSpider)
process.start()

关键修改说明

  • 必须通过yield输出包含字段的字典:Scrapy只有接收到这类字典,才会将数据写入CSV文件。
  • 添加USER_AGENT设置:避免网站将请求识别为爬虫而返回空内容。
  • 确保XPath准确性:如果网站结构变动,需要用浏览器开发者工具重新验证元素定位的XPath表达式。

可选:爬取详情页数据

如果需要获取每个国家的详细数据,可添加详情页回调函数:

# 在WorldometersSpider类中新增以下方法
def parse_country_detail(self, response):
    yield {
        '国家名称': response.xpath('//h1/text()').get(),
        '当前总人口': response.xpath('//div[@class="maincounter-number"]/span/text()').get(),
        '国土面积': response.xpath('//div[@class="row"]/div[2]/div/div[2]/text()').get().strip()
        # 更多字段可根据页面结构自行添加
    }

同时在parse方法中取消注释这段代码:

# link = row.xpath('.//td[2]/a/@href').get()
# yield response.follow(url=link, callback=self.parse_country_detail)

内容的提问来源于stack exchange,提问作者DevSmartz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 10:51:19