Scrapy多站点爬取数据导出Google Sheets:解决yield无变量问题
解决Scrapy中yield数据无法存储为列表的问题
方法1:通过Item Pipeline统一收集并处理数据
Scrapy的Pipeline是专门处理爬虫产出数据的组件,可用来收集所有yield的item,之后统一导出到Google Sheets。
- 在项目的
pipelines.py中添加如下代码:
import gspread from oauth2client.service_account import ServiceAccountCredentials class GoogleSheetsPipeline: def __init__(self): # 初始化存储数据的列表 self.data_list = [] # 配置Google Sheets授权(替换为你的密钥文件路径) scope = ["https://spreadsheets.google.com/feeds", "https://www.googleapis.com/auth/drive"] creds = ServiceAccountCredentials.from_json_keyfile_name('your-service-account-key.json', scope) self.client = gspread.authorize(creds) def process_item(self, item, spider): # 将每个yield的item添加到列表 self.data_list.append(dict(item)) return item def close_spider(self, spider): # 爬虫结束后写入Google Sheets sheet = self.client.open("你的表格名称").sheet1 # 写入表头 headers = list(self.data_list[0].keys()) sheet.append_row(headers) # 写入数据行 rows = [list(item.values()) for item in self.data_list] sheet.append_rows(rows)
- 在
settings.py中启用该Pipeline:
ITEM_PIPELINES = { '你的项目名.pipelines.GoogleSheetsPipeline': 300, }
爬虫运行时,所有yield的item会自动存入列表,爬虫结束后直接写入Google Sheets。
方法2:在爬虫类中用实例变量存储数据
若不想使用Pipeline,可直接在Spider类内定义列表存储数据,单爬虫实例下无需担心线程安全问题:
修改你的爬虫代码:
import scrapy import gspread from oauth2client.service_account import ServiceAccountCredentials class GainesvilleSpider(scrapy.Spider): name = 'gainesvillecoins' allowed_domains = ['gainesvillecoins.com'] start_urls = [ 'https://www.gainesvillecoins.com/products/158985/1-oz-american-gold-eagle-coins', 'https://www.gainesvillecoins.com/products/158550/1-2-oz-american-gold-eagle-coins', 'https://www.gainesvillecoins.com/products/180843/2022-1-4-oz-american-gold-eagle', 'https://www.gainesvillecoins.com/products/180331/2021-1-10-oz-gold-american-eagle-brilliant-uncirculated' ] def __init__(self): self.data_list = [] # 初始化Google Sheets客户端 scope = ["https://spreadsheets.google.com/feeds", "https://www.googleapis.com/auth/drive"] creds = ServiceAccountCredentials.from_json_keyfile_name('your-service-account-key.json', scope) self.client = gspread.authorize(creds) def parse(self, response): coin_name = response.xpath("//h1/text()").extract_first() coin_price = response.xpath("//td[2]/text()").extract_first() item = {"coin name": coin_name, "coin price": coin_price} self.data_list.append(item) yield item def closed(self, reason): # 爬虫结束后写入Google Sheets sheet = self.client.open("你的表格名称").sheet1 headers = list(self.data_list[0].keys()) sheet.append_row(headers) rows = [list(item.values()) for item in self.data_list] sheet.append_rows(rows)
方法3:先导出为CSV再导入Google Sheets
如果不想编写Google Sheets交互代码,可先用Scrapy的Feed Exports导出CSV,再导入到Google Sheets:
运行爬虫时执行命令:
scrapy crawl gainesvillecoins -o coins_data.csv
导出完成后,在Google Sheets中选择「文件」→「导入」→「上传」,选择生成的CSV文件即可完成数据导入。
内容的提问来源于stack exchange,提问作者brandon_abraham
相关产品推荐
相关产品推荐

