Python开发通用Yahoo Finance财报爬虫:适配任意股票代码
通用Yahoo Finance财报爬虫实现(支持任意股票代码)
依赖库安装
先确保安装所需Python库:
pip install pandas requests beautifulsoup4
核心函数实现
针对三类报表分别编写独立函数,每个函数接收股票代码ticker作为参数,返回对应报表的DataFrame。
1. 资产负债表爬虫函数
import pandas as pd import requests from bs4 import BeautifulSoup def get_balance_sheet(ticker): # 构造资产负债表URL url = f"https://finance.yahoo.com/quote/{ticker}/balance-sheet?p={ticker}" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 请求页面 response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 提取报表数据 table = soup.find('div', class_='M(0) Whs(n) BdEnd Bdc($seperatorColor) D(itb)') if not table: return pd.DataFrame() # 无数据时返回空DataFrame # 解析表头和行数据 headers = [th.get_text() for th in table.find_all('th')] rows = [] for tr in table.find_all('tr')[1:]: row_data = [td.get_text() for td in tr.find_all('td')] if row_data: rows.append(row_data) # 转换为DataFrame并整理格式 df = pd.DataFrame(rows, columns=headers) df.set_index(headers[0], inplace=True) # 将数值列转换为数值类型(处理逗号和负数) for col in df.columns[1:]: df[col] = df[col].replace(',', '', regex=True).replace('-', 'NaN').astype(float) return df
2. 利润表爬虫函数
def get_income_statement(ticker): url = f"https://finance.yahoo.com/quote/{ticker}/financials?p={ticker}" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') table = soup.find('div', class_='M(0) Whs(n) BdEnd Bdc($seperatorColor) D(itb)') if not table: return pd.DataFrame() headers = [th.get_text() for th in table.find_all('th')] rows = [] for tr in table.find_all('tr')[1:]: row_data = [td.get_text() for td in tr.find_all('td')] if row_data: rows.append(row_data) df = pd.DataFrame(rows, columns=headers) df.set_index(headers[0], inplace=True) for col in df.columns[1:]: df[col] = df[col].replace(',', '', regex=True).replace('-', 'NaN').astype(float) return df
3. 现金流量表爬虫函数
def get_cash_flow(ticker): url = f"https://finance.yahoo.com/quote/{ticker}/cash-flow?p={ticker}" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') table = soup.find('div', class_='M(0) Whs(n) BdEnd Bdc($seperatorColor) D(itb)') if not table: return pd.DataFrame() headers = [th.get_text() for th in table.find_all('th')] rows = [] for tr in table.find_all('tr')[1:]: row_data = [td.get_text() for td in tr.find_all('td')] if row_data: rows.append(row_data) df = pd.DataFrame(rows, columns=headers) df.set_index(headers[0], inplace=True) for col in df.columns[1:]: df[col] = df[col].replace(',', '', regex=True).replace('-', 'NaN').astype(float) return df
使用示例
# 获取苹果公司财报 aapl_balance = get_balance_sheet("AAPL") aapl_income = get_income_statement("AAPL") aapl_cashflow = get_cash_flow("AAPL") # 获取微软公司财报 msft_balance = get_balance_sheet("MSFT") # 查看数据 print(aapl_income.head())
注意事项
- 反爬机制:必须添加
User-Agent请求头,否则Yahoo Finance会拒绝请求。请求频繁时建议添加随机延迟(如time.sleep(random.uniform(1,3)))。 - 页面结构变化:Yahoo Finance页面布局可能随时调整,若函数返回空DataFrame,需检查页面元素的class或标签是否变更,更新解析逻辑。
- 数据完整性:部分股票可能没有完整历史报表数据,函数会返回空DataFrame,使用前需判断。
内容的提问来源于stack exchange,提问作者Bash
相关产品推荐
相关产品推荐

