You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python开发通用Yahoo Finance财报爬虫:适配任意股票代码

通用Yahoo Finance财报爬虫实现(支持任意股票代码)

依赖库安装

先确保安装所需Python库:

pip install pandas requests beautifulsoup4

核心函数实现

针对三类报表分别编写独立函数,每个函数接收股票代码ticker作为参数,返回对应报表的DataFrame。

1. 资产负债表爬虫函数

import pandas as pd
import requests
from bs4 import BeautifulSoup

def get_balance_sheet(ticker):
    # 构造资产负债表URL
    url = f"https://finance.yahoo.com/quote/{ticker}/balance-sheet?p={ticker}"
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    # 请求页面
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # 提取报表数据
    table = soup.find('div', class_='M(0) Whs(n) BdEnd Bdc($seperatorColor) D(itb)')
    if not table:
        return pd.DataFrame()  # 无数据时返回空DataFrame
    
    # 解析表头和行数据
    headers = [th.get_text() for th in table.find_all('th')]
    rows = []
    for tr in table.find_all('tr')[1:]:
        row_data = [td.get_text() for td in tr.find_all('td')]
        if row_data:
            rows.append(row_data)
    
    # 转换为DataFrame并整理格式
    df = pd.DataFrame(rows, columns=headers)
    df.set_index(headers[0], inplace=True)
    # 将数值列转换为数值类型(处理逗号和负数)
    for col in df.columns[1:]:
        df[col] = df[col].replace(',', '', regex=True).replace('-', 'NaN').astype(float)
    
    return df

2. 利润表爬虫函数

def get_income_statement(ticker):
    url = f"https://finance.yahoo.com/quote/{ticker}/financials?p={ticker}"
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    table = soup.find('div', class_='M(0) Whs(n) BdEnd Bdc($seperatorColor) D(itb)')
    if not table:
        return pd.DataFrame()
    
    headers = [th.get_text() for th in table.find_all('th')]
    rows = []
    for tr in table.find_all('tr')[1:]:
        row_data = [td.get_text() for td in tr.find_all('td')]
        if row_data:
            rows.append(row_data)
    
    df = pd.DataFrame(rows, columns=headers)
    df.set_index(headers[0], inplace=True)
    for col in df.columns[1:]:
        df[col] = df[col].replace(',', '', regex=True).replace('-', 'NaN').astype(float)
    
    return df

3. 现金流量表爬虫函数

def get_cash_flow(ticker):
    url = f"https://finance.yahoo.com/quote/{ticker}/cash-flow?p={ticker}"
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    table = soup.find('div', class_='M(0) Whs(n) BdEnd Bdc($seperatorColor) D(itb)')
    if not table:
        return pd.DataFrame()
    
    headers = [th.get_text() for th in table.find_all('th')]
    rows = []
    for tr in table.find_all('tr')[1:]:
        row_data = [td.get_text() for td in tr.find_all('td')]
        if row_data:
            rows.append(row_data)
    
    df = pd.DataFrame(rows, columns=headers)
    df.set_index(headers[0], inplace=True)
    for col in df.columns[1:]:
        df[col] = df[col].replace(',', '', regex=True).replace('-', 'NaN').astype(float)
    
    return df

使用示例

# 获取苹果公司财报
aapl_balance = get_balance_sheet("AAPL")
aapl_income = get_income_statement("AAPL")
aapl_cashflow = get_cash_flow("AAPL")

# 获取微软公司财报
msft_balance = get_balance_sheet("MSFT")

# 查看数据
print(aapl_income.head())

注意事项

  • 反爬机制:必须添加User-Agent请求头,否则Yahoo Finance会拒绝请求。请求频繁时建议添加随机延迟(如time.sleep(random.uniform(1,3)))。
  • 页面结构变化:Yahoo Finance页面布局可能随时调整,若函数返回空DataFrame,需检查页面元素的class或标签是否变更,更新解析逻辑。
  • 数据完整性:部分股票可能没有完整历史报表数据,函数会返回空DataFrame,使用前需判断。

内容的提问来源于stack exchange,提问作者Bash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 18:33:11