如何用Beautiful Soup抓取Yahoo Finance利润表的可折叠/展开区域数据
解决Yahoo Finance利润表折叠区域数据抓取问题
你当前用requests+BeautifulSoup的方式只能获取服务器返回的初始HTML,但Yahoo Finance的折叠区域数据要么是通过JavaScript动态渲染加载(初始HTML中不存在,页面加载后通过AJAX请求补充),要么是初始HTML中存在但被CSS样式隐藏,而你的选择器未覆盖到这些隐藏元素。因此直接爬取静态HTML无法拿到隐藏数据。
以下是两种可行的解决方案:
方案一:用Selenium模拟浏览器展开区域后抓取
Selenium可以模拟浏览器的交互操作(比如点击展开按钮),让页面渲染出所有隐藏内容,再抓取完整的页面源代码。
步骤:
- 安装Selenium:
pip install selenium - 下载对应浏览器的驱动(比如ChromeDriver,需和浏览器版本匹配)
- 编写代码模拟点击展开并抓取:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup # 配置Chrome选项,可选无头模式(不显示浏览器窗口) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:89.0) Gecko/20100101 Firefox/89.0") # 初始化浏览器驱动 driver = webdriver.Chrome(options=chrome_options) target_url = "https://finance.yahoo.com/quote/AMZN/financials?p=AMZN" driver.get(target_url) # 点击所有折叠按钮,展开隐藏区域 # 定位所有带expandable类的折叠按钮 expand_buttons = driver.find_elements(By.CSS_SELECTOR, ".expandable") for btn in expand_buttons: try: # 用JS点击避免元素遮挡问题 driver.execute_script("arguments[0].click();", btn) except Exception: pass # 跳过已展开的按钮 # 获取渲染后的页面源代码 page_source = driver.page_source driver.quit() # 解析页面数据 soup = BeautifulSoup(page_source, "html.parser") table_rows = soup.find_all(class_="M(0) Whs(n) BdEnd Bdc($seperatorColor) D(itb)") for row in table_rows: # 格式化输出每行数据 print(row.get_text(strip=True, separator=" | "))
方案二:直接调用Yahoo Finance的API接口(推荐)
Yahoo Finance的财务数据实际是通过后台API加载的,直接请求API可以获取结构化的JSON数据,无需处理页面渲染,效率更高。
代码示例:
import requests import json # 利润表数据的API端点,modules参数指定获取利润表历史数据 api_url = "https://query1.finance.yahoo.com/v10/finance/quoteSummary/AMZN?modules=incomeStatementHistory" headers = { "User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:89.0) Gecko/20100101 Firefox/89.0", "Referer": "https://finance.yahoo.com/quote/AMZN/financials?p=AMZN" } # 发送请求获取JSON数据 response = requests.get(api_url, headers=headers) data = json.loads(response.text) # 解析并输出利润表数据 income_statements = data['quoteSummary']['result'][0]['incomeStatementHistory']['incomeStatementHistory'] for stmt in income_statements: print(f"报告日期: {stmt['endDate']['fmt']}") # 遍历每个财务指标,输出格式化后的值 for key, value in stmt.items(): if key != 'endDate' and 'fmt' in value: print(f" {key}: {value['fmt']}") print("-" * 50)
说明:
- API的
modules参数可以调整,比如incomeStatementHistoryQuarterly获取季度数据,balanceSheetHistory获取资产负债表等。 - 需保持
Referer和User-Agent头,避免被API拦截。
内容的提问来源于stack exchange,提问作者Mulak
相关产品推荐
相关产品推荐

