You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现多XBRL文件合并并导出为单Excel/CSV文件

问题描述

需要将指定文件夹中数量不固定的多个XBRL文件合并为单个Excel/CSV文件。现有代码可实现单个文件导出为CSV,所有文件的代码/键均一致,仅文件日期和数值不同。目前卡在如何遍历file_list并整合解析数据以完成导出,次要目标是创建一个包含所有单个报告字典的总字典。

现有代码
import os, glob, pandas as pd, xbrl, csv
from xbrl import XBRLParser, GAAP, GAAPSerializer

path = 'P:/Bank Research/Bank Automation/Calls'

callCert = '57944'
dates = ['063023','123122','123121','123120','123119','123118']
xbrl_parser = XBRLParser()


file_list = glob.glob(path + "/*.XBRL")

文件列表示例:

['P:/Bank Research/Bank Automation/Calls\\Call_Cert57944_033123.XBRL',
'Call_Cert57944_063023.XBRL',  'Call_Cert57944_123118.XBRL',  'Call_Cert57944_123119.XBRL']
xbrl_file = "Call_Cert"+callCert+"_"+dates[0]+"\.XBRL"    
xbrl_document = xbrl_parser.parse(xbrl_file)
custom_obj = xbrl_parser.parseCustom(xbrl_document)  


list_bank_numbers = []
list_bank_keys = []
for i in custom_obj():
    list_bank_numbers.append(i[1])
    list_bank_keys.append(i[0])

bank_dict = {list_bank_keys[i]: list_bank_numbers[i] for i in range(len(list_bank_keys))}


def export_dict_to_csv(dictionary, output_file):

    keys = dictionary.keys()
    values = dictionary.values()

    with open(output_file, 'w', newline='') as csv_file:
        writer = csv.writer(csv_file)
        writer.writerow(keys)
        writer.writerow(values)
        

export_dict_to_csv(bank_dict, callCert+'.csv')
解决方案

以下是修改后的代码,实现遍历所有XBRL文件、整合数据并导出,同时生成包含所有报告的总字典:

import os, glob, pandas as pd, xbrl, csv
from xbrl import XBRLParser, GAAP, GAAPSerializer

path = 'P:/Bank Research/Bank Automation/Calls'
callCert = '57944'
xbrl_parser = XBRLParser()

# 获取文件夹下所有XBRL文件
file_list = glob.glob(os.path.join(path, "*.XBRL"))

# 总字典:键为文件日期,值为对应报告的数据字典
total_bank_dict = {}
# 存储所有带日期的记录,用于生成合并表格
all_records = []

for file_path in file_list:
    # 从文件名提取日期(例如从Call_Cert57944_063023.XBRL中提取063023)
    file_name = os.path.basename(file_path)
    date = file_name.split('_')[-1].replace('.XBRL', '')
    
    # 解析当前XBRL文件
    xbrl_doc = xbrl_parser.parse(file_path)
    custom_data = xbrl_parser.parseCustom(xbrl_doc)
    
    # 生成当前文件的数据字典
    current_dict = {item[0]: item[1] for item in custom_data()}
    # 将当前字典存入总字典
    total_bank_dict[date] = current_dict
    
    # 构造带日期的记录,方便后续合并成表格
    record = {'日期': date}
    record.update(current_dict)
    all_records.append(record)

# 使用pandas将所有记录合并为DataFrame并导出
merged_df = pd.DataFrame(all_records)
# 导出为CSV
merged_df.to_csv(f"{callCert}_合并数据.csv", index=False, encoding='utf-8-sig')
# 导出为Excel(需提前安装openpyxl:pip install openpyxl)
merged_df.to_excel(f"{callCert}_合并数据.xlsx", index=False)

# 打印总字典(按需使用)
print("所有报告的总字典:")
print(total_bank_dict)

关键说明

  • 自动提取文件名中的日期,无需依赖硬编码的dates列表,适配任意数量的XBRL文件
  • 遍历所有文件时,同步完成解析、数据字典生成及总字典的构建
  • 利用pandas的DataFrame自动对齐所有键(因所有文件键一致,无需担心列缺失),导出操作更简洁高效
  • 同时支持CSV和Excel两种导出格式,满足不同需求

内容的提问来源于stack exchange,提问作者johnsmith_228

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 08:02:12