You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从SEC获取S-1表单及相关财务数据?

解决S-1表单财务数据获取问题

问题原因

你之前的代码失效是因为传入的是S-1的HTML页面链接,但xbrl_to_json需要的是**XBRL实例文件(.xml)**的URL。多数S-1的HTM页面仅用于展示,不包含可解析的XBRL结构化数据,必须获取对应的XBRL XML文件才能提取财务报表。

可行方案与脚本

方案1:使用Sec API获取XBRL文件并解析

先通过Sec API的Filings接口找到目标S-1对应的XBRL文件URL,再用xbrl_to_json解析:

import pandas as pd
from sec_api import XbrlApi, QueryApi

API_KEY = "你的API密钥"  # 替换为实际密钥

# 初始化API
xbrl_api = XbrlApi(API_KEY)
query_api = QueryApi(API_KEY)

# 1. 查询目标S-1的 filing 信息,获取XBRL文件URL
query = {
    "query": {
        "query_string": "formType:\"S-1\" AND cik:\"1802255\" AND filingDate:[2023-01-01 TO 2023-12-31]"
    },
    "from": "0",
    "size": "1"
}

filings = query_api.get_filings(query)
# 获取XBRL实例文件的URL(通常在filing['xbrlFiles']中)
xbrl_file_url = None
for filing in filings['filings']:
    for xbrl_file in filing['xbrlFiles']:
        if xbrl_file['type'] == "EX-101.INS":  # 核心XBRL实例文件
            xbrl_file_url = xbrl_file['url']
            break
    if xbrl_file_url:
        break

if not xbrl_file_url:
    raise ValueError("未找到对应的XBRL实例文件")

# 2. 解析XBRL文件为JSON
xbrl_json = xbrl_api.xbrl_to_json(xbrl_url=xbrl_file_url)

# 3. 提取利润表/资产负债表(兼容S-1的XBRL结构)
def extract_financial_statement(xbrl_json, statement_type):
    statement_store = {}
    if statement_type not in xbrl_json:
        return pd.DataFrame({"提示": [f"未找到{statement_type}数据"]})
    
    for us_gaap_item in xbrl_json[statement_type]:
        values = []
        indices = []
        for fact in xbrl_json[statement_type][us_gaap_item]:
            if 'segment' not in fact:
                # S-1的期间可能是年度或累计,处理日期格式
                if 'startDate' in fact['period']:
                    index = f"{fact['period']['startDate']}至{fact['period']['endDate']}"
                else:
                    index = fact['period']['instant']  # 资产负债表是时点数据
                if index not in indices:
                    values.append(fact['value'])
                    indices.append(index)
        statement_store[us_gaap_item] = pd.Series(values, index=indices)
    
    return pd.DataFrame(statement_store).T

# 提取利润表和资产负债表
income_statement = extract_financial_statement(xbrl_json, 'StatementsOfIncome')
balance_sheet = extract_financial_statement(xbrl_json, 'BalanceSheets')

print("利润表:")
print(income_statement)
print("\n资产负债表:")
print(balance_sheet)

方案2:使用EDGAR库直接抓取解析

如果不想依赖付费API,可以使用edgar库结合xbrl库解析(需先安装:pip install edgar xbrl pandas):

import pandas as pd
import edgar
from xbrl import XBRL

# 设置EDGAR邮箱(SEC要求)
edgar.set_agent("你的邮箱地址")

# 获取目标公司的S-1 filings
cik = "1802255"
filings = edgar.get_filings(cik, form_type="S-1")

# 获取最新的S-1对应的XBRL文件
latest_filing = filings[0]
xbrl_url = latest_filing.xbrl_url()

if not xbrl_url:
    raise ValueError("该S-1未包含XBRL文件")

# 解析XBRL文件
xbrl = XBRL(xbrl_url)

# 提取利润表数据
income_statement_data = {}
for concept in xbrl.concepts:
    if concept.namespace == "http://fasb.org/us-gaap/2023-01-31":
        # 筛选利润表相关项目(可根据US GAAP标签调整)
        if "IncomeStatement" in concept.label or "Revenue" in concept.label or "Expense" in concept.label:
            values = []
            indices = []
            for fact in xbrl.facts[concept]:
                if not fact.context.segment:
                    period = fact.context.period
                    if period.is_duration:
                        index = f"{period.start_date}至{period.end_date}"
                    else:
                        index = period.instant_date
                    values.append(fact.value)
                    indices.append(index)
            income_statement_data[concept.label] = pd.Series(values, index=indices)

income_statement = pd.DataFrame(income_statement_data).T
print("利润表:")
print(income_statement)

注意事项

  • 部分S-1可能未提交XBRL文件,这种情况只能通过OCR解析HTML页面,但难度较高。
  • 使用SEC相关工具时,需遵守SEC的访问规则(比如设置用户代理邮箱),避免被封禁IP。
  • US GAAP标签可能因年份或公司而异,需根据实际XBRL结构调整筛选逻辑。

内容的提问来源于stack exchange,提问作者william duran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 13:22:06