You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy多站点爬取数据导出Google Sheets:解决yield无变量问题

解决Scrapy中yield数据无法存储为列表的问题

方法1:通过Item Pipeline统一收集并处理数据

Scrapy的Pipeline是专门处理爬虫产出数据的组件,可用来收集所有yield的item,之后统一导出到Google Sheets。

  1. 在项目的pipelines.py中添加如下代码:
import gspread
from oauth2client.service_account import ServiceAccountCredentials

class GoogleSheetsPipeline:
    def __init__(self):
        # 初始化存储数据的列表
        self.data_list = []
        # 配置Google Sheets授权(替换为你的密钥文件路径)
        scope = ["https://spreadsheets.google.com/feeds", "https://www.googleapis.com/auth/drive"]
        creds = ServiceAccountCredentials.from_json_keyfile_name('your-service-account-key.json', scope)
        self.client = gspread.authorize(creds)

    def process_item(self, item, spider):
        # 将每个yield的item添加到列表
        self.data_list.append(dict(item))
        return item

    def close_spider(self, spider):
        # 爬虫结束后写入Google Sheets
        sheet = self.client.open("你的表格名称").sheet1
        # 写入表头
        headers = list(self.data_list[0].keys())
        sheet.append_row(headers)
        # 写入数据行
        rows = [list(item.values()) for item in self.data_list]
        sheet.append_rows(rows)
  1. 在settings.py中启用该Pipeline:
ITEM_PIPELINES = {
    '你的项目名.pipelines.GoogleSheetsPipeline': 300,
}

爬虫运行时,所有yield的item会自动存入列表,爬虫结束后直接写入Google Sheets。

方法2:在爬虫类中用实例变量存储数据

若不想使用Pipeline,可直接在Spider类内定义列表存储数据,单爬虫实例下无需担心线程安全问题:

修改你的爬虫代码:

import scrapy
import gspread
from oauth2client.service_account import ServiceAccountCredentials

class GainesvilleSpider(scrapy.Spider):
    name = 'gainesvillecoins'
    allowed_domains = ['gainesvillecoins.com']
    start_urls = [
        'https://www.gainesvillecoins.com/products/158985/1-oz-american-gold-eagle-coins',
        'https://www.gainesvillecoins.com/products/158550/1-2-oz-american-gold-eagle-coins',
        'https://www.gainesvillecoins.com/products/180843/2022-1-4-oz-american-gold-eagle',
        'https://www.gainesvillecoins.com/products/180331/2021-1-10-oz-gold-american-eagle-brilliant-uncirculated'
    ]

    def __init__(self):
        self.data_list = []
        # 初始化Google Sheets客户端
        scope = ["https://spreadsheets.google.com/feeds", "https://www.googleapis.com/auth/drive"]
        creds = ServiceAccountCredentials.from_json_keyfile_name('your-service-account-key.json', scope)
        self.client = gspread.authorize(creds)

    def parse(self, response):
        coin_name = response.xpath("//h1/text()").extract_first()
        coin_price = response.xpath("//td[2]/text()").extract_first()
        item = {"coin name": coin_name, "coin price": coin_price}
        self.data_list.append(item)
        yield item

    def closed(self, reason):
        # 爬虫结束后写入Google Sheets
        sheet = self.client.open("你的表格名称").sheet1
        headers = list(self.data_list[0].keys())
        sheet.append_row(headers)
        rows = [list(item.values()) for item in self.data_list]
        sheet.append_rows(rows)

方法3:先导出为CSV再导入Google Sheets

如果不想编写Google Sheets交互代码,可先用Scrapy的Feed Exports导出CSV,再导入到Google Sheets:

运行爬虫时执行命令:

scrapy crawl gainesvillecoins -o coins_data.csv

导出完成后,在Google Sheets中选择「文件」→「导入」→「上传」,选择生成的CSV文件即可完成数据导入。


内容的提问来源于stack exchange,提问作者brandon_abraham

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 14:45:33