You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何BeautifulSoup无法抓取Yahoo Finance的Growth Estimates表格?

问题原因
  • 请求头未模拟浏览器:Yahoo Finance会校验请求的User-Agent,默认requests请求头会被识别为非浏览器请求,返回的页面不包含目标表格。
  • 类名依赖不可靠:你使用的表格类名包含CSS变量(如Bdc($seperatorColor)),这是Yahoo Finance的动态样式变量,实际返回的HTML中会被替换为具体值,导致类名定位失败。
  • 潜在动态加载:部分表格内容可能通过JavaScript异步渲染,requests只能获取初始HTML,无法加载JS生成的内容。
解决方案

方案1:添加浏览器请求头+优化表格定位

通过模拟浏览器请求头,改用更稳定的属性(如aria-label)定位表格:

import requests
from bs4 import BeautifulSoup

def get_growth_data(symbol):
    url = f"https://finance.yahoo.com/quote/{symbol}/analysis?p={symbol}"
    # 模拟Chrome浏览器请求头
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, "html.parser")

    # 通过aria-label定位Growth Estimates表格,比类名更稳定
    table = soup.find("table", {"aria-label": "Growth Estimates"})

    if table is None:
        print("Table not found.")
        return []

    # 提取表格数据(跳过表头行)
    growth_values = []
    rows = table.find_all("tr")
    for row in rows[1:]:
        columns = row.find_all("td")
        if len(columns) >= 2:
            growth_values.append(columns[1].text.strip())

    return growth_values

symbol = 'AAPL'
growth_data = get_growth_data(symbol)
print(growth_data)

方案2:处理动态加载(方案1失效时用)

如果表格是JS动态渲染的,用selenium模拟浏览器加载完整页面:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import time

def get_growth_data(symbol):
    url = f"https://finance.yahoo.com/quote/{symbol}/analysis?p={symbol}"
    # 配置无头浏览器
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    
    driver = webdriver.Chrome(options=chrome_options)
    driver.get(url)
    # 等待页面加载完成
    time.sleep(3)

    # 定位Growth Estimates表格
    try:
        table = driver.find_element(By.CSS_SELECTOR, 'table[aria-label="Growth Estimates"]')
    except:
        print("Table not found.")
        driver.quit()
        return []

    # 提取数据(跳过表头行)
    growth_values = []
    rows = table.find_elements(By.TAG_NAME, "tr")
    for row in rows[1:]:
        columns = row.find_elements(By.TAG_NAME, "td")
        if len(columns) >= 2:
            growth_values.append(columns[1].text.strip())
    
    driver.quit()
    return growth_values

symbol = 'AAPL'
growth_data = get_growth_data(symbol)
print(growth_data)

注意:使用selenium需提前安装对应浏览器驱动(如ChromeDriver),确保驱动版本与浏览器版本匹配。

内容的提问来源于stack exchange,提问作者Jack Jackie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 06:33:32