You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将PDF提取的JSON格式银行账单数据转CSV/HTML?

解决方案:将Unstructured导出的银行账单JSON转为CSV/HTML

以下是直接可用的Python脚本,针对Unstructured.io导出的含Title、NarrativeText、Table字段的JSON文件,优先生成结构化CSV,也支持生成HTML。

核心思路

  1. 读取并解析JSON文件,提取标题、说明文本和表格数据
  2. 将标题和说明文本作为CSV的注释行(以#开头),与表格数据分隔
  3. 解析表格的行与单元格,转换为CSV可识别的结构化行
  4. 可选:用Pandas简化复杂表格的处理,或生成HTML格式

方法1:原生Python(无需额外依赖)

适合轻量场景,无需安装第三方库:

import json
import csv

# 读取JSON文件(替换为你的文件路径)
with open("bank_statement.json", "r", encoding="utf-8") as json_file:
    statement_data = json.load(json_file)

# 提取关键内容
bill_title = statement_data.get("Title", "未命名账单")
bill_note = statement_data.get("NarrativeText", "无说明文本")
bill_tables = statement_data.get("Table", [])

# 构建CSV内容行
csv_output = []
# 添加标题和说明作为注释
csv_output.append([f"# {bill_title}"])
csv_output.append([f"# {bill_note}"])
csv_output.append([])  # 空行分隔注释与表格

# 遍历所有表格并转换为CSV行
for table in bill_tables:
    table_rows = table.get("rows", [])
    if not table_rows:
        continue
    # 提取每一行的单元格内容
    for row in table_rows:
        csv_output.append(row.get("cells", []))

# 保存为CSV文件(替换为你的输出路径)
with open("bank_statement.csv", "w", newline="", encoding="utf-8") as csv_file:
    csv_writer = csv.writer(csv_file)
    csv_writer.writerows(csv_output)

print("CSV转换完成:bank_statement.csv")

方法2:使用Pandas(适合复杂表格)

如果JSON中的表格存在多表头、合并单元格等复杂结构,Pandas能更高效处理:

import json
import pandas as pd

# 读取JSON文件
with open("bank_statement.json", "r", encoding="utf-8") as json_file:
    statement_data = json.load(json_file)

bill_title = statement_data.get("Title", "未命名账单")
bill_note = statement_data.get("NarrativeText", "无说明文本")
bill_tables = statement_data.get("Table", [])

# 转换表格为DataFrame
table_dfs = []
for table in bill_tables:
    table_rows = table.get("rows", [])
    if not table_rows:
        continue
    # 提取表头和数据行
    header = table_rows[0]["cells"]
    data_rows = [row["cells"] for row in table_rows[1:]]
    table_df = pd.DataFrame(data_rows, columns=header)
    table_dfs.append(table_df)

# 合并所有表格并保存CSV
if table_dfs:
    combined_df = pd.concat(table_dfs, ignore_index=True)
    # 先写入标题和说明,再写入表格数据
    with open("bank_statement_pandas.csv", "w", encoding="utf-8") as csv_file:
        csv_file.write(f"# {bill_title}\n")
        csv_file.write(f"# {bill_note}\n\n")
        combined_df.to_csv(csv_file, index=False)

    print("CSV转换完成:bank_statement_pandas.csv")

生成HTML格式(可选)

基于Pandas快速生成带格式的HTML账单:

# 接上述Pandas代码,生成HTML
html_content = f"<h2>{bill_title}</h2>"
html_content += f"<p>{bill_note}</p>"
for df in table_dfs:
    html_content += df.to_html(index=False, border=1, classes="table table-striped")

with open("bank_statement.html", "w", encoding="utf-8") as html_file:
    html_file.write(html_content)

print("HTML转换完成:bank_statement.html")

注意事项

  • 如果你的JSON中cells字段不是直接字符串(比如包含text子字段),需修改单元格提取逻辑,例如:[cell.get("text", "") for cell in row.get("cells", [])]
  • 若存在多个表格,脚本会自动合并到同一文件中(CSV用空行分隔,HTML用连续表格展示)
  • 编码统一使用UTF-8,避免中文乱码

内容的提问来源于stack exchange,提问作者SAGE KHAN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 18:22:49