如何通过Webhooks或其他方式自动捕获Yahoo Finance分析师报告更新?
解决方案:自动捕获Yahoo Finance公开分析师报告信息
一、Zapier 配置步骤(定时抓取+去重写入)
Yahoo Finance没有官方推送API,只能通过定时检查网页+抓取公开内容实现,结合Zapier完成自动化:
设置定时触发
- 选择
Schedule by Zapier作为触发应用,频率设为工作日每1-2小时(避免频繁请求被网站拦截),时区匹配Yahoo报告发布的时区。
- 选择
添加网页请求动作
- 选
Webhooks by Zapier,方法设为GET,URL填https://finance.yahoo.com/research。 - 在Headers里添加
User-Agent(比如Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36),模拟浏览器请求防止被屏蔽。
- 选
解析HTML提取公开字段
- 用
Code by Zapier(Python模式)写简单抓取代码,示例如下(注意:页面元素类名可能随时变化,需自己用浏览器F12工具确认实际结构):from bs4 import BeautifulSoup import requests headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} response = requests.get("https://finance.yahoo.com/research", headers=headers) soup = BeautifulSoup(response.text, 'html.parser') reports = [] # 假设公开报告容器的类名为research-report-item,需自行验证 for item in soup.select('.research-report-item'): report_date = item.select_one('.report-date').text.strip() if item.select_one('.report-date') else None ticker = item.select_one('.ticker').text.strip() if item.select_one('.ticker') else None company = item.select_one('.company-name').text.strip() if item.select_one('.company-name') else None title = item.select_one('.report-title').text.strip() if item.select_one('.report-title') else None url = "https://finance.yahoo.com" + item.select_one('.report-title a')['href'] if item.select_one('.report-title a') else None price = item.select_one('.current-price').text.strip() if item.select_one('.current-price') else None summary = item.select_one('.summary').text.strip() if item.select_one('.summary') else None if report_date and ticker and title: reports.append({ "发布日期": report_date, "股票代码": ticker, "公司名称": company, "报告标题": title, "报告URL": url, "当前股价": price, "摘要": summary }) return {"new_reports": reports}
- 用
去重过滤
- 添加
Filter by Zapier,对比目标表格中已有的报告URL,只让未出现过的条目进入下一步,避免重复写入。
- 添加
写入结构化表格
- 选择
Google Sheets或Excel Online作为动作应用,动作选Create Spreadsheet Row,把解析出的字段一一对应到表格列即可。
- 选择
二、替代方案(新手友好型)
如果Zapier的代码环节有难度,试试这些选项:
1. Make(原Integromat)
- 自带可视化网页抓取模块,无需写代码,直接用鼠标选择要抓取的页面元素;定时触发、去重、写入表格的逻辑都能通过拖拽配置完成。
2. Python脚本+免费云函数
- 写个简单的Python脚本(用
BeautifulSoup+pandas),部署在Google Cloud Functions或AWS Lambda免费层,设置定时运行:- 核心逻辑:抓取页面→解析字段→对比已存数据→写入Google Sheets/CSV
- 优势:完全自定义,成本低,适合后续扩展功能
3. 浏览器扩展(半自动)
- 用
Web Scraper(Chrome/Firefox扩展),可视化配置抓取规则,一键导出数据到CSV,再导入表格;适合不需要完全自动化,每天手动操作一次的场景。
注意事项
- 不要设置过于频繁的抓取频率,工作日每1小时一次足够,避免触发网站反爬机制。
- 页面结构更新时,要及时调整抓取规则或代码中的元素选择器。
- 仅抓取公开内容,不要尝试绕过付费墙,遵守网站服务条款。
内容的提问来源于stack exchange,提问作者stacksncodes
相关产品推荐
相关产品推荐

