将集成Google Sheets API的Scrapy爬虫部署到Heroku遇阻求助
排查Heroku上Scrapyd爬虫无法写入Google Sheets的问题
从你提供的信息来看,Scrapyd服务已成功启动,但无数据写入Google Sheets,以下是逐步排查方案:
1. 确认爬虫是否已部署并执行
- 部署验证:在本地执行
scrapyd-deploy heroku -p quotes(确保scrapyd-client已配置),部署完成后访问https://scrapy-test555.herokuapp.com/,查看页面Projects列表是否包含quotes项目。 - 手动触发爬虫:用curl命令触发任务:
执行后查看Heroku日志(curl -X POST https://scrapy-test555.herokuapp.com/schedule.json -d project=quotes -d spider=你的爬虫名称heroku logs --tail),若无爬虫启动日志,说明爬虫未成功部署或触发失败。
2. 检查Google Sheets认证配置
本地正常但云端失败,大概率是认证信息未正确配置:
- 环境变量替代本地文件:不要将Google Service Account的JSON文件上传到Heroku,而是将JSON内容转为环境变量:
- 在Heroku控制台
Settings→Config Vars中添加GOOGLE_SERVICE_ACCOUNT_KEY,值为JSON文件的完整内容。 - 修改代码中加载认证的逻辑:
import os import json from google.oauth2.service_account import Credentials def get_google_creds(): creds_json = json.loads(os.environ.get("GOOGLE_SERVICE_ACCOUNT_KEY")) return Credentials.from_service_account_info( creds_json, scopes=["https://www.googleapis.com/auth/spreadsheets"] )
- 在Heroku控制台
- 共享权限验证:将Service Account的邮箱(格式如
xxx@xxx.iam.gserviceaccount.com)添加到目标Google Sheet的共享列表,授予编辑权限。
3. 修正Scrapyd端口配置
Heroku使用动态端口,需让Scrapyd监听PORT环境变量:
- 修改
Procfile内容为:
重新部署后,Scrapyd会自动使用Heroku分配的端口,避免端口不匹配问题。web: scrapyd --port $PORT
4. 清理冗余依赖
你的requirements.txt包含大量本地开发依赖(如chromedriver-binary-auto、docker、selenium等),若爬虫不需要这些库,建议删除以减少部署冲突:
- 重新生成精简版requirements.txt:仅保留爬虫必需的依赖(如
Scrapy、scrapyd、gspread、google-auth等)。
5. 查看爬虫运行日志
触发爬虫后,通过以下方式获取详细日志:
- 查看Heroku实时日志:
heroku logs --tail,若爬虫启动,会输出执行日志,包括管道的错误信息。 - 通过Scrapyd API获取特定任务日志:
curl https://scrapy-test555.herokuapp.com/logs/quotes/你的爬虫名/[任务ID].log,替换[任务ID]为触发爬虫后返回的jobid。
6. 处理Dyno休眠问题
免费Heroku Dyno会在30分钟无请求后休眠,若需定时运行爬虫:
- 使用Heroku Scheduler添加定时任务,触发Scrapyd的调度API(如每天执行一次
curl -X POST https://scrapy-test555.herokuapp.com/schedule.json -d project=quotes -d spider=你的爬虫名)。
内容的提问来源于stack exchange,提问作者Reggie18
相关产品推荐
相关产品推荐

