You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从roic.ai爬取财务报表至DataFrame?页面源码嵌套复杂

从roic.ai爬取财务报表并存储至DataFrame的实现方案

针对roic.ai页面嵌套较深的结构,以下是直接从HTML表格元素提取数据并转换为DataFrame的实现代码,完全不依赖页面中的#__NEXT_DATA__数据源:

from gazpacho import get, Soup
import pandas as pd

def extract_financials(ticker, freq='annual'):
    # 构建目标URL
    url = f'https://roic.ai/financials/{ticker}?fs={freq}'
    html = get(url)
    soup = Soup(html)
    
    # 定位所有财务报表模块(每个报表对应一个带space-y-8的flex-col容器)
    report_containers = soup.find('div', {'class': 'flex-col space-y-8'}, mode='all')
    
    financial_data = {}
    
    for container in report_containers:
        # 提取报表名称
        report_name = container.find('h3').text.strip()
        
        # 提取表头(年份列)
        headers = [th.text.strip() for th in container.find('th', mode='all')]
        
        # 提取每一行的科目和对应数值
        rows = []
        for tr in container.find('tr', mode='all')[1:]:  # 跳过表头行
            cols = tr.find('td', mode='all')
            row_data = {headers[0]: cols[0].text.strip()}
            # 遍历数值列,去除千位分隔符
            for idx, col in enumerate(cols[1:]):
                row_data[headers[idx+1]] = col.text.strip().replace(',', '')
            rows.append(row_data)
        
        # 将当前报表转为DataFrame并存入字典
        financial_data[report_name] = pd.DataFrame(rows)
    
    return financial_data

# 示例:提取AAPL年度财务数据
ticker = 'aapl'
financials = extract_financials(ticker)

# 输出利润表示例
print(financials['Income Statement'])

代码说明

  • 定位报表容器:页面中每个财务报表(利润表、资产负债表等)都包裹在带有flex-col space-y-8类的div中,通过mode='all'获取所有报表模块。
  • 提取表头与行数据:每个报表的表头是<th>元素,数据行是<tr>下的<td>元素,逐行解析科目和对应年份的数值,同时去除千位分隔符便于后续数值处理。
  • 整理为DataFrame:将每个报表的行数据转换为DataFrame,最终以字典形式返回所有报表,键为报表名称,值为对应DataFrame。

内容的提问来源于stack exchange,提问作者Petr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 04:25:39