You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Adobe PDF Services Extract API将PDF数据整理为Excel/CSV结构化文件?

解决Adobe PDF Services Extract API提取数据结构化到Excel/CSV的问题

你的核心问题在于没有充分利用API输出的结构化JSON数据——官方示例生成的zip包中包含structuredData.json,里面存储了文本、表格的层级和布局信息,只需解析这个JSON就能将数据整理为可导入Excel/CSV的格式。以下是具体实现步骤:

1. 优化API提取配置

先修改你的提取选项,明确指定JSON输出并保留结构元素,让后续解析更精准:

# 需先导入ExtractOutputFormat
from adobe.pdfservices.operation.pdfops.options.extractpdf.extract_output_format import ExtractOutputFormat

# 替换原有的ExtractPDFOptions构建代码
extract_pdf_options: ExtractPDFOptions = ExtractPDFOptions.builder() \
    .with_element_to_extract(ExtractElementType.TEXT) \
    .with_element_to_extract(ExtractElementType.TABLES) \
    .with_output_format(ExtractOutputFormat.JSON)  # 强制输出结构化JSON
    .with_include_structure_elements(True)  # 保留布局结构信息
    .build()

2. 解析提取后的JSON文件

每个生成的zip包解压后会得到structuredData.json,以下代码用json和pandas解析表格与文本,并导出为CSV:

import json
import pandas as pd
import zipfile
import shutil
import os

def parse_extracted_zip(zip_file_path):
    # 临时解压zip
    temp_dir = "temp_pdf_extract"
    with zipfile.ZipFile(zip_file_path, 'r') as zip_ref:
        zip_ref.extractall(temp_dir)
    
    # 读取结构化JSON
    with open(f"{temp_dir}/structuredData.json", 'r', encoding='utf-8') as f:
        doc_data = json.load(f)
    
    # 提取表格数据(处理单元格合并、文本拼接)
    table_list = []
    for elem in doc_data["elements"]:
        if elem["type"] == "Table":
            table_rows = []
            for row in elem["rows"]:
                row_content = []
                for cell in row["cells"]:
                    # 拼接单元格内的所有文本片段
                    cell_text = " ".join([t["content"] for t in cell["elements"] if t["type"] == "Text"])
                    row_content.append(cell_text)
                table_rows.append(row_content)
            table_list.append(pd.DataFrame(table_rows))
    
    # 提取文本数据(按段落整理)
    text_paragraphs = []
    for elem in doc_data["elements"]:
        if elem["type"] == "Paragraph":
            para_text = " ".join([t["content"] for t in elem["elements"] if t["type"] == "Text"])
            text_paragraphs.append({"段落内容": para_text})
    text_df = pd.DataFrame(text_paragraphs)
    
    # 清理临时文件
    shutil.rmtree(temp_dir)
    return table_list, text_df

3. 批量处理你的100个PDF结果

在原代码的循环后,添加批量解析与导出逻辑:

# 原代码的循环结束后,添加这段
base_path = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
for x in range(0, 100):
    s = str(x)
    zip_path = f"{base_path}/output/ExtractTextTableInfoFromPDF{s}.zip"
    tables, text_df = parse_extracted_zip(zip_path)
    
    # 导出表格(多表格按序号区分)
    for idx, table in enumerate(tables):
        table.to_csv(f"{base_path}/output/table_{s}_{idx}.csv", index=False, encoding="utf-8-sig")
    
    # 导出文本
    text_df.to_csv(f"{base_path}/output/text_{s}.csv", index=False, encoding="utf-8-sig")

关键注意事项

  • 若PDF存在合并单元格,JSON中的cells会包含rowSpan和colSpan字段,如需精准还原表格结构,需额外处理这些字段。
  • 可利用JSON中的boundingBox字段,将文本与表格按原PDF的位置关联,实现更贴合原布局的结构化输出。
  • 若不想依赖pandas,也可使用Python内置的csv模块手动遍历数据写入文件。

内容的提问来源于stack exchange,提问作者Light143

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 12:07:45