如何抓取Yahoo Finance利润表中的子类别财务数据?
问题:如何抓取Yahoo Finance利润表的子类别财务数据?
我正在练习数据爬取,尝试从Yahoo Finance抓取AAPL的财务数据(利润表页面),但目前的代码只能获取顶级分类数据(比如Total Revenue),抓不到子类别(比如Total Revenue下的Operating Revenue)——部分股票会有多个子类别及对应数值。以下是我的Python代码,求修改以同时抓取子类别数据:
import requests from bs4 import BeautifulSoup import pandas as pd stock_abb = ['AAPL'] df = pd.DataFrame() for s in stock_abb: url = 'https://finance.yahoo.com/quote/' + s + '/financials?p=' + s header = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.0.0 Safari/537.36'} r = requests.get(url, headers={"User-Agent": "Requests"}).text soup = BeautifulSoup(r, 'html.parser') #Scrape first part financials_number_1 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg D(tbc)'}) financials_title = soup.find_all('div', {'class': 'D(ib) Va(m) Ell Mt(-3px) W(215px)--mv2 W(200px) undefined'}) financial_period_2 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg D(ib) Fw(b)'}) df_first_part = pd.DataFrame ({'Company':[],'Period':[], 'Type':[],'Value':[]}) t = 0 i = 0 for t in range(0,len(financials_number_1)): try: f1 = financials_number_1[t].find_all('span')[0].get_text() if (t % 2): Type = financials_title[i].find_all('span')[0].get_text() Year = financial_period_2[1].find_all('span')[0].get_text() i = i + 1 else: Type = financials_title[i].find_all('span')[0].get_text() Year = financial_period_2[0].find_all('span')[0].get_text() df2 = pd.DataFrame ({'Company':[s],'Period':[Year], 'Type':[Type],'Value':[f1]}) data = [df_first_part, df2] df_first_part = pd.concat(data) except: if (t % 2): i = i + 1 #Scrape second part financials_number_2 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg Bgc($lv1BgColor) fi-row:h_Bgc($hoverBgColor) D(tbc)'}) financials_title = soup.find_all('div', {'class': 'D(ib) Va(m) Ell Mt(-3px) W(215px)--mv2 W(200px) undefined'}) financial_period_1 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg D(ib) Fw(b) Tt(u) Bgc($lv1BgColor)'}) financial_period_3 = soup.find_all('div', {'class': 'Ta(c) Py(6px) Bxz(bb) BdB Bdc($seperatorColor) Miw(120px) Miw(100px)--pnclg D(ib) Fw(b) Bgc($lv1BgColor)'}) df_second_part = pd.DataFrame ({'Period':[], 'Type':[],'Value':[]}) t = 0 i = 0 u = 0 for t in range(0,len(financials_number_2)): try: f2 = financials_number_2[t].find_all('span')[0].get_text() if t == 0 or (t % 3 == 0): Year = financial_period_1[0].find_all('span')[0].get_text() else: if u == 0: Year = financial_period_3[0].find_all('span')[0].get_text() u = u + 1 else: Year = financial_period_3[1].find_all('span')[0].get_text() u = 0 if t == 0: Type = financials_title[0].find_all('span')[0].get_text() elif (t % 3 == 0): i = i + 1 Type = financials_title[i].find_all('span')[0].get_text() else: Type = financials_title[i].find_all('span')[0].get_text() df2 = pd.DataFrame ({'Company':[s], 'Period':[Year], 'Type':[Type],'Value':[f2]}) data = [df_second_part, df2] df_second_part = pd.concat(data) except: if (t % 3 != 0): if u == 0: u = u + 1 if (t % 3 == 0): i = i + 1 else: u = 0 data = [df,df_second_part,df_first_part] df = pd.concat(data)
注:营收分为顶级分类和子类别,部分股票的子类别会有多个不同数值。
解决方案
原代码的问题在于没识别利润表中子类别的层级结构:Yahoo Finance的子类别行通常带fi-sub类标识,且标题宽度类和父行不同(父行是W(215px)--mv2 W(200px),子行是W(190px)--mv2 W(170px))。修改思路是直接遍历所有财务行,区分父行和子行,同时抓取每一行对应的所有周期数值,避免拆分处理的复杂逻辑。
修改后的代码如下:
import requests from bs4 import BeautifulSoup import pandas as pd stock_abb = ['AAPL'] df = pd.DataFrame() # 定义User-Agent,避免被反爬拦截 headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.0.0 Safari/537.36'} for s in stock_abb: url = f'https://finance.yahoo.com/quote/{s}/financials?p={s}' r = requests.get(url, headers=headers) soup = BeautifulSoup(r.text, 'html.parser') # 获取所有统计周期(比如2023、2022等) periods = [span.get_text() for span in soup.select('div.Ta(c).Py(6px).Bxz(bb).BdB.Bdc($seperatorColor).Miw(120px).Miw(100px)--pnclg.D(ib).Fw(b) span')] # 获取所有财务行(包含父行和子行) financial_rows = soup.select('div.fi-row') current_parent = "" for row in financial_rows: # 判断当前行是否为子类别行 is_sub = 'fi-sub' in row.get('class', []) # 获取行标题文本 title_elem = row.select_one('div.D(ib).Va(m).Ell.Mt(-3px)') if not title_elem: continue title = title_elem.get_text(strip=True) # 记录父行标题,子行继承关联 if not is_sub: current_parent = title else: # 可自定义子标题格式,比如拼接父类别名称 title = f"{current_parent} - {title}" # 获取当前行的所有周期数值 values = [span.get_text(strip=True) for span in row.select('div.Ta(c).Py(6px).Bxz(bb).BdB.Bdc($seperatorColor).Miw(120px).Miw(100px)--pnclg.D(tbc) span')] # 将行数据批量加入DataFrame for period, value in zip(periods, values): df = pd.concat([df, pd.DataFrame({ 'Company': [s], 'Parent_Type': [current_parent if is_sub else None], 'Type': [title], 'Period': [period], 'Value': [value] })], ignore_index=True) # 输出结果示例 print(df.head(20))
关键修改点:
- 用
fi-row类直接定位所有财务行,包含父行和子行 - 通过
fi-sub类判断子行,自动关联对应的父类别 - 一次性抓取所有周期和对应数值,避免原代码拆分处理的混乱逻辑
- 新增
Parent_Type字段,清晰区分子类别所属的顶级分类 - 修复原代码中定义了User-Agent但未正确使用的问题
内容的提问来源于stack exchange,提问作者Basilicoq13
相关产品推荐
相关产品推荐

