如何用BeautifulSoup抓取HTML表格数据?AAPL财报页面爬取遇阻
问题:无法抓取AlphaQuery股票收益历史页面的表格
我尝试抓取AlphaQuery上AAPL的收益历史页面表格,但始终无法成功,甚至找不到目标表格。以下是我编写的代码:
import requests from bs4 import BeautifulSoup def get_eps(ticker): url = f"https://www.alphaquery.com/stock/{ticker}/earnings-history" page = requests.get(url) soup = BeautifulSoup(page.content, 'html.parser') # 尝试通过检查特定表头来更可靠地查找表格 table = None for table_candidate in soup.find_all("table"): headers = [th.get_text(strip=True) for th in table_candidate.find_all("th")] if "Estimated EPS" in headers: table = table_candidate break if table: rows = table.find_all('tr')[1:6] for row in rows: cells = row.find_all('td') if len(cells) >= 4: # 确保行中有足够的列 try: est_eps = cells[2].text.strip().replace('$', '').replace(',', '') except ValueError: continue # 跳过无法转换为浮点数的行 else: print(f"Failed to find earnings table for {ticker}") return est_eps # 示例用法 ticker = 'AAPL' beats = get_eps(ticker) print(f'{ticker} estimates {est_eps}')
问题原因分析
- 动态内容加载:目标页面的表格是通过JavaScript动态渲染的,
requests.get()只能获取静态HTML源码,无法捕获JS加载后的表格内容,导致BeautifulSoup找不到目标表格。 - 代码变量问题:
est_eps仅在循环内部赋值,若循环未执行到有效行,会返回未定义变量导致报错。- 示例调用中,打印时使用的
est_eps是函数内部的局部变量,外部无法访问,会触发NameError。
解决方案
方案1:用Selenium处理动态渲染内容
Selenium可模拟浏览器加载页面,获取JS渲染后的完整HTML:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def get_eps(ticker): url = f"https://www.alphaquery.com/stock/{ticker}/earnings-history" # 初始化Chrome浏览器(需对应版本的ChromeDriver) driver = webdriver.Chrome() driver.get(url) try: # 等待表格加载完成,超时10秒 table = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//table[contains(., 'Estimated EPS')]")) ) rows = table.find_elements(By.TAG_NAME, 'tr')[1:6] est_eps_list = [] for row in rows: cells = row.find_elements(By.TAG_NAME, 'td') if len(cells) >= 4: est_eps = cells[2].text.strip().replace('$', '').replace(',', '') est_eps_list.append(est_eps) return est_eps_list if est_eps_list else None finally: driver.quit() # 示例用法 ticker = 'AAPL' est_eps_list = get_eps(ticker) if est_eps_list: print(f'{ticker} 最近5期预估EPS:{est_eps_list}') else: print(f'未找到{ticker}的收益表格')
方案2:直接调用页面API(更高效)
打开浏览器开发者工具查看网络请求,找到加载表格数据的API接口,直接请求接口获取JSON格式数据,无需渲染页面。例如该页面可能存在类似/api/earnings-history的接口,直接解析返回的JSON即可提取所需数据。
代码修正核心要点
- 处理动态内容:优先选择API接口,其次用Selenium模拟浏览器
- 变量作用域:给
est_eps设置初始值,或收集所有有效结果返回列表 - 错误处理:增加变量未定义、接口请求失败等场景的判断
内容的提问来源于stack exchange,提问作者Gloria Dalla Costa
相关产品推荐
相关产品推荐

