GitHub Actions定时网页爬取至Google Sheets任务执行失败求助
GitHub Actions定时爬虫任务失败修复及脚本优化
一、GitHub Actions配置错误修复
报错every step must define a uses or run key的核心原因:配置文件中Set Google Sheets Credentials步骤仅定义了环境变量,未指定任何执行动作(uses或run),不符合GitHub Actions的步骤规则。
修改后的完整配置文件 .github/workflows/actions.yml
name: Scheduled Web Scraping on: schedule: - cron: '0 15 * * *' # UTC时间15点对应北京时间23点,实现每日23点执行 workflow_dispatch: # 支持手动触发测试 jobs: build: runs-on: ubuntu-latest steps: - name: 检出仓库代码 uses: actions/checkout@v4 - name: 设置Python环境 uses: actions/setup-python@v5 with: python-version: '3.11' - name: 安装依赖包 run: | python -m pip install --upgrade pip pip install requests beautifulsoup4 pandas gspread oauth2client numpy # 标准库无需安装 - name: 执行爬虫脚本 env: GOOGLE_SHEETS_CREDENTIALS: ${{ secrets.GOOGLE_SHEETS_CREDENTIALS }} run: python heat_stress_at_work_warning.py - name: 提交更新日志 run: | git config --local user.email "action@github.com" git config --local user.name "GitHub Action" git add -A git diff-index --quiet HEAD || (git commit -a -m "更新爬取日志" --allow-empty) - name: 推送更改到仓库 uses: ad-m/github-push-action@master with: github_token: ${{ secrets.GITHUB_TOKEN }} branch: main
配置修改要点
- 移除无效的单独环境变量定义步骤,将
GOOGLE_SHEETS_CREDENTIALS直接注入到执行脚本的步骤中,确保每个步骤都包含uses或run。 - 修正依赖包名:
BeautifulSoup的正确pip包名为beautifulsoup4,datetime、urllib、time等属于Python标准库,无需通过pip安装。 - 更新Action组件到最新版本,提升兼容性和稳定性。
- 调整Cron表达式为
0 15 * * *,适配GitHub Actions的UTC时区,确保北京时间每日23点执行任务。
二、Python脚本核心问题修复
1. Google Sheets凭证加载错误
原脚本硬编码了Windows本地路径的凭证文件,GitHub Actions运行在Ubuntu环境无法访问该路径,需改为从环境变量加载凭证:
原代码片段:
creds = ServiceAccountCredentials.from_json_keyfile_name('C:/Users/User/Documents/self-learn/heat-stress-index-15bbb5903054.json', scope)
修改为:
# 直接使用环境变量中的凭证字典创建凭证 creds = ServiceAccountCredentials.from_json_keyfile_dict(creds_dict, scope)
2. 数据类型一致性问题
原脚本中type字段混合了字符串('Warning'/'Cancellation')和数字(0),导致后续汇总时过滤逻辑失效,需统一为字符串类型:
修改默认记录的type值:
all_articles.append({ # ... 其他字段不变 'type': 'No Record', # 替换原0值为字符串 # ... 其他字段不变 })
同步更新汇总函数的过滤条件:
def generate_summary_data(detailed_df): filtered_df = detailed_df[(detailed_df['type'] == 'Cancellation') | (detailed_df['type'] == 'No Record')] summary_df = filtered_df.groupby('date', as_index=False).agg({'no_of_hours': 'sum'}) return summary_df
3. 数据处理顺序优化
原脚本先上传未计算时长的原始数据,后续计算时长后未重新上传,需调整执行顺序,确保所有数据处理完成后再上传:
在main函数中,注释掉提前上传的代码,待所有字段计算完成后再执行上传:
# 移除提前上传的代码:spreadsheet = upload_to_google_sheets(new_df, "Heat_Stress_Index") # ... 完成duration和no_of_hours的计算后,再执行上传 spreadsheet = upload_to_google_sheets(new_df, "Heat_Stress_Index")
三、验证步骤
- 确认GitHub仓库的Secrets中已添加
GOOGLE_SHEETS_CREDENTIALS,值为Google服务账户的JSON完整内容。 - 将Google Sheets共享给服务账户邮箱(JSON文件中的
client_email字段),并赋予编辑权限。 - 通过GitHub仓库的Actions页面手动触发Workflow,验证任务是否正常执行并写入数据到Google Sheets。
内容的提问来源于stack exchange,提问作者ronzenith
相关产品推荐
相关产品推荐

