You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取Yahoo Finance利润表中的子类别财务数据?

问题:如何抓取Yahoo Finance利润表的子类别财务数据?

我正在练习数据爬取,尝试从Yahoo Finance抓取AAPL的财务数据(利润表页面),但目前的代码只能获取顶级分类数据(比如Total Revenue),抓不到子类别(比如Total Revenue下的Operating Revenue)——部分股票会有多个子类别及对应数值。以下是我的Python代码,求修改以同时抓取子类别数据:

import requests
from bs4 import BeautifulSoup
import pandas as pd

stock_abb = ['AAPL']

df = pd.DataFrame()

for s in stock_abb:
    url = 'https://finance.yahoo.com/quote/' + s + '/financials?p=' + s
    header = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.0.0 Safari/537.36'}
    r = requests.get(url, headers={"User-Agent": "Requests"}).text
    soup = BeautifulSoup(r, 'html.parser')


    #Scrape first part

    financials_number_1 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg D(tbc)'})
    financials_title = soup.find_all('div', {'class': 'D(ib) Va(m) Ell Mt(-3px) W(215px)--mv2 W(200px) undefined'})  
    financial_period_2 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg D(ib) Fw(b)'})

    df_first_part = pd.DataFrame ({'Company':[],'Period':[], 'Type':[],'Value':[]})

    t = 0
    i = 0
    for t in range(0,len(financials_number_1)): 
        try:
            f1 = financials_number_1[t].find_all('span')[0].get_text()

            if (t % 2):
                Type = financials_title[i].find_all('span')[0].get_text()
                Year = financial_period_2[1].find_all('span')[0].get_text()
                i = i + 1
            else:
                Type = financials_title[i].find_all('span')[0].get_text()
                Year = financial_period_2[0].find_all('span')[0].get_text()


            df2 = pd.DataFrame ({'Company':[s],'Period':[Year], 'Type':[Type],'Value':[f1]})
            data = [df_first_part, df2]
            df_first_part = pd.concat(data)

        except:
            if (t % 2):
                i = i  + 1



    #Scrape second part           

    financials_number_2 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg Bgc($lv1BgColor) fi-row:h_Bgc($hoverBgColor) D(tbc)'})
    financials_title = soup.find_all('div', {'class': 'D(ib) Va(m) Ell Mt(-3px) W(215px)--mv2 W(200px) undefined'})  
    financial_period_1 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg D(ib) Fw(b) Tt(u) Bgc($lv1BgColor)'})
    financial_period_3 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg D(ib) Fw(b) Bgc($lv1BgColor)'})

    df_second_part = pd.DataFrame ({'Period':[], 'Type':[],'Value':[]})

    t = 0
    i = 0
    u = 0

    for t in range(0,len(financials_number_2)):

        try:
            f2 = financials_number_2[t].find_all('span')[0].get_text()

            if t == 0 or (t % 3 == 0):
                Year = financial_period_1[0].find_all('span')[0].get_text()
            else:
                if u == 0:
                    Year =  financial_period_3[0].find_all('span')[0].get_text()
                    u = u + 1
                else:
                    Year =  financial_period_3[1].find_all('span')[0].get_text()
                    u = 0

            if t == 0:
                Type = financials_title[0].find_all('span')[0].get_text()  
            elif (t % 3 == 0):
                i = i + 1
                Type = financials_title[i].find_all('span')[0].get_text()

            else:
                Type = financials_title[i].find_all('span')[0].get_text()

            df2 = pd.DataFrame ({'Company':[s], 'Period':[Year], 'Type':[Type],'Value':[f2]})
            data = [df_second_part, df2]
            df_second_part = pd.concat(data)
        except:
            if (t % 3 != 0):
                if u == 0:
                    u = u + 1
            if (t % 3 == 0):
                    i = i + 1
            else: u = 0


    data = [df,df_second_part,df_first_part]

    df = pd.concat(data)

注:营收分为顶级分类和子类别,部分股票的子类别会有多个不同数值。


解决方案

原代码的问题在于没识别利润表中子类别的层级结构:Yahoo Finance的子类别行通常带fi-sub类标识,且标题宽度类和父行不同(父行是W(215px)--mv2 W(200px),子行是W(190px)--mv2 W(170px))。修改思路是直接遍历所有财务行,区分父行和子行,同时抓取每一行对应的所有周期数值,避免拆分处理的复杂逻辑。

修改后的代码如下:

import requests
from bs4 import BeautifulSoup
import pandas as pd

stock_abb = ['AAPL']
df = pd.DataFrame()

# 定义User-Agent,避免被反爬拦截
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.0.0 Safari/537.36'}

for s in stock_abb:
    url = f'https://finance.yahoo.com/quote/{s}/financials?p={s}'
    r = requests.get(url, headers=headers)
    soup = BeautifulSoup(r.text, 'html.parser')
    
    # 获取所有统计周期(比如2023、2022等)
    periods = [span.get_text() for span in soup.select('div.Ta(c).Py(6px).Bxz(bb).BdB.Bdc($seperatorColor).Miw(120px).Miw(100px)--pnclg.D(ib).Fw(b) span')]
    
    # 获取所有财务行(包含父行和子行)
    financial_rows = soup.select('div.fi-row')
    
    current_parent = ""
    for row in financial_rows:
        # 判断当前行是否为子类别行
        is_sub = 'fi-sub' in row.get('class', [])
        
        # 获取行标题文本
        title_elem = row.select_one('div.D(ib).Va(m).Ell.Mt(-3px)')
        if not title_elem:
            continue
        title = title_elem.get_text(strip=True)
        
        # 记录父行标题,子行继承关联
        if not is_sub:
            current_parent = title
        else:
            # 可自定义子标题格式,比如拼接父类别名称
            title = f"{current_parent} - {title}"
        
        # 获取当前行的所有周期数值
        values = [span.get_text(strip=True) for span in row.select('div.Ta(c).Py(6px).Bxz(bb).BdB.Bdc($seperatorColor).Miw(120px).Miw(100px)--pnclg.D(tbc) span')]
        
        # 将行数据批量加入DataFrame
        for period, value in zip(periods, values):
            df = pd.concat([df, pd.DataFrame({
                'Company': [s],
                'Parent_Type': [current_parent if is_sub else None],
                'Type': [title],
                'Period': [period],
                'Value': [value]
            })], ignore_index=True)

# 输出结果示例
print(df.head(20))

关键修改点:

  1. 用fi-row类直接定位所有财务行,包含父行和子行
  2. 通过fi-sub类判断子行,自动关联对应的父类别
  3. 一次性抓取所有周期和对应数值,避免原代码拆分处理的混乱逻辑
  4. 新增Parent_Type字段,清晰区分子类别所属的顶级分类
  5. 修复原代码中定义了User-Agent但未正确使用的问题

内容的提问来源于stack exchange,提问作者Basilicoq13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 00:21:41