You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python解析银行Call Report的XBRL文件并转换为可分析格式

解析FFIEC银行Call Report XBRL文件并转换为可分析格式的解决方案

问题描述

从FFIEC平台下载的银行Call Report XBRL文件,使用xbrl库解析后,无法将数据转换为数组、DataFrame、字典等便于财务分析的格式。

用户原代码及输出

import xbrl
import bs4 as bs
from xbrl import XBRLParser, GAAP, GAAPSerializer

xbrl_file = "Call_Cert57944_063023.XBRL"

xbrl_parser = XBRLParser()
xbrl_document = xbrl_parser.parse(xbrl_file)
print(xbrl_document)

输出示例:

<cc:rcons550 contextref="CI_3357219_2023-06-30" decimals="0" unitref="USD">0</cc:rcons550>
<cc:rcon6835 contextref="CI_3357219_2023-06-30" decimals="0" unitref="USD">1577000</cc:rcon6835>
<cc:rcons556 contextref="CI_3357219_2023-06-30" decimals="0" unitref="USD">0</cc:rcons556>
<cc:rcons555 contextref="CI_3357219_2023-06-30" decimals="0" unitref="USD">0</cc:rcons555>
bank_call = xbrl_document.prettify()
print(bank_call)

输出示例:

0
   </cc:rcons490>
   <cc:rcons491 contextref="CI_3357219_2023-06-30" decimals="0" unitref="USD">
    0
   </cc:rcons491>
   <cc:rcons496 contextref="CI_3357219_2023-06-30" decimals="0" unitref="USD">
    0
   </cc:rcons496>
custom_obj = xbrl_parser.parseCustom(xbrl_document)
# sample output below
print(custom_obj())

输出示例:

dict_items([('rcona573', '321334000'), ('riad4356', '0'), ('rcona571', '52305000'), 
('rcona570', '476535000'), ('rcons498', '0'), ('rcona575', '0'), ('rcona574', '35287000'), 
('rcon5370', '20000'), ('rcons452', '0'), ('rcons453', '0'), ('rcons450', '0'),

解决方案

1. 快速转换为字典与DataFrame

parseCustom()返回的dict_items可直接转为Python字典,进而生成DataFrame:

import pandas as pd

# 转换为字典
call_report_dict = dict(custom_obj())

# 生成DataFrame
df = pd.DataFrame.from_dict(call_report_dict, orient='index', columns=['数值'])
df.index.name = '指标代码'
print(df.head())

2. 提取完整上下文信息(日期、单位)

若需要保留报告日期、单位等元数据,需从原始XBRL文件中解析上下文节点:

from bs4 import BeautifulSoup

# 用BeautifulSoup读取并解析原始XBRL文件
with open(xbrl_file, 'r', encoding='utf-8') as f:
    soup = BeautifulSoup(f, 'xml')

# 提取所有指标元素
call_items = soup.find_all(['cc:' + tag for tag in call_report_dict.keys()])

# 整理包含元数据的完整数据集
data = []
for item in call_items:
    tag_name = item.name.replace('cc:', '')
    value = item.get_text(strip=True)
    context_ref = item.get('contextref')
    
    # 从上下文节点提取报告日期
    context = soup.find('context', id=context_ref)
    date_node = context.find('endDate') or context.find('instant')
    report_date = date_node.get_text()
    
    # 提取单位
    unit = item.get('unitref')
    
    data.append({
        '指标代码': tag_name,
        '数值': int(value) if value.isdigit() else value,
        '报告日期': report_date,
        '单位': unit
    })

# 生成带完整元数据的DataFrame
full_df = pd.DataFrame(data)
print(full_df.head())

3. 指标代码映射为可读名称

可根据FFIEC提供的Call Report数据字典,构建指标代码与中文名称的映射表,提升数据可读性:

# 示例映射表,可根据需求扩展
code_mapping = {
    'rcona573': '总资产',
    'rcona571': '流动资产',
    'rcona570': '总负债',
    # 更多指标映射补充
}

full_df['指标名称'] = full_df['指标代码'].map(code_mapping)
print(full_df[['指标名称', '数值', '报告日期']].head())

内容的提问来源于stack exchange,提问作者johnsmith_228

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 17:03:27