如何从roic.ai爬取财务报表至DataFrame?页面源码嵌套复杂
从roic.ai爬取财务报表并存储至DataFrame的实现方案
针对roic.ai页面嵌套较深的结构,以下是直接从HTML表格元素提取数据并转换为DataFrame的实现代码,完全不依赖页面中的#__NEXT_DATA__数据源:
from gazpacho import get, Soup import pandas as pd def extract_financials(ticker, freq='annual'): # 构建目标URL url = f'https://roic.ai/financials/{ticker}?fs={freq}' html = get(url) soup = Soup(html) # 定位所有财务报表模块(每个报表对应一个带space-y-8的flex-col容器) report_containers = soup.find('div', {'class': 'flex-col space-y-8'}, mode='all') financial_data = {} for container in report_containers: # 提取报表名称 report_name = container.find('h3').text.strip() # 提取表头(年份列) headers = [th.text.strip() for th in container.find('th', mode='all')] # 提取每一行的科目和对应数值 rows = [] for tr in container.find('tr', mode='all')[1:]: # 跳过表头行 cols = tr.find('td', mode='all') row_data = {headers[0]: cols[0].text.strip()} # 遍历数值列,去除千位分隔符 for idx, col in enumerate(cols[1:]): row_data[headers[idx+1]] = col.text.strip().replace(',', '') rows.append(row_data) # 将当前报表转为DataFrame并存入字典 financial_data[report_name] = pd.DataFrame(rows) return financial_data # 示例:提取AAPL年度财务数据 ticker = 'aapl' financials = extract_financials(ticker) # 输出利润表示例 print(financials['Income Statement'])
代码说明
- 定位报表容器:页面中每个财务报表(利润表、资产负债表等)都包裹在带有
flex-col space-y-8类的div中,通过mode='all'获取所有报表模块。 - 提取表头与行数据:每个报表的表头是
<th>元素,数据行是<tr>下的<td>元素,逐行解析科目和对应年份的数值,同时去除千位分隔符便于后续数值处理。 - 整理为DataFrame:将每个报表的行数据转换为DataFrame,最终以字典形式返回所有报表,键为报表名称,值为对应DataFrame。
内容的提问来源于stack exchange,提问作者Petr
相关产品推荐
相关产品推荐

