You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup抓取HTML表格数据?AAPL财报页面爬取遇阻

问题:无法抓取AlphaQuery股票收益历史页面的表格

我尝试抓取AlphaQuery上AAPL的收益历史页面表格,但始终无法成功,甚至找不到目标表格。以下是我编写的代码:

import requests
from bs4 import BeautifulSoup

def get_eps(ticker):
    url = f"https://www.alphaquery.com/stock/{ticker}/earnings-history"
    page = requests.get(url)
    soup = BeautifulSoup(page.content, 'html.parser')
    
    # 尝试通过检查特定表头来更可靠地查找表格
    table = None
    for table_candidate in soup.find_all("table"):
        headers = [th.get_text(strip=True) for th in table_candidate.find_all("th")]
        if "Estimated EPS" in headers:
            table = table_candidate
            break
    if table:
        rows = table.find_all('tr')[1:6]  
        for row in rows:
            cells = row.find_all('td')
            if len(cells) >= 4:  # 确保行中有足够的列
                try:
                    est_eps = cells[2].text.strip().replace('$', '').replace(',', '')
                except ValueError:
                    continue  # 跳过无法转换为浮点数的行
    else:
        print(f"Failed to find earnings table for {ticker}")

    return est_eps

# 示例用法
ticker = 'AAPL'
beats = get_eps(ticker)
print(f'{ticker}  estimates {est_eps}')

问题原因分析

  • 动态内容加载:目标页面的表格是通过JavaScript动态渲染的,requests.get()只能获取静态HTML源码,无法捕获JS加载后的表格内容,导致BeautifulSoup找不到目标表格。
  • 代码变量问题:
    • est_eps仅在循环内部赋值,若循环未执行到有效行,会返回未定义变量导致报错。
    • 示例调用中,打印时使用的est_eps是函数内部的局部变量,外部无法访问,会触发NameError。

解决方案

方案1:用Selenium处理动态渲染内容

Selenium可模拟浏览器加载页面,获取JS渲染后的完整HTML:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def get_eps(ticker):
    url = f"https://www.alphaquery.com/stock/{ticker}/earnings-history"
    # 初始化Chrome浏览器(需对应版本的ChromeDriver)
    driver = webdriver.Chrome()
    driver.get(url)
    
    try:
        # 等待表格加载完成,超时10秒
        table = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.XPATH, "//table[contains(., 'Estimated EPS')]"))
        )
        
        rows = table.find_elements(By.TAG_NAME, 'tr')[1:6]
        est_eps_list = []
        for row in rows:
            cells = row.find_elements(By.TAG_NAME, 'td')
            if len(cells) >= 4:
                est_eps = cells[2].text.strip().replace('$', '').replace(',', '')
                est_eps_list.append(est_eps)
        
        return est_eps_list if est_eps_list else None
    finally:
        driver.quit()

# 示例用法
ticker = 'AAPL'
est_eps_list = get_eps(ticker)
if est_eps_list:
    print(f'{ticker} 最近5期预估EPS:{est_eps_list}')
else:
    print(f'未找到{ticker}的收益表格')

方案2:直接调用页面API(更高效)

打开浏览器开发者工具查看网络请求,找到加载表格数据的API接口,直接请求接口获取JSON格式数据,无需渲染页面。例如该页面可能存在类似/api/earnings-history的接口,直接解析返回的JSON即可提取所需数据。

代码修正核心要点

  • 处理动态内容:优先选择API接口,其次用Selenium模拟浏览器
  • 变量作用域:给est_eps设置初始值,或收集所有有效结果返回列表
  • 错误处理:增加变量未定义、接口请求失败等场景的判断

内容的提问来源于stack exchange,提问作者Gloria Dalla Costa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 10:46:12