You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Textract异步解析PDF表格部分CELL未识别问题求助

问题解决思路与脚本建议

关于Textract参数与控制台差异

  • 先确认异步调用StartDocumentAnalysis时,是否明确指定了FeatureTypes=["TABLES"]——虽然你POC验证可行,但部分特殊文档可能因参数细节遗漏导致识别差异(不过大概率不是核心原因)。
  • AWS Textract控制台确实会在基础API响应之上做额外后处理:通过文本块的坐标匹配,将未自动关联到CELL的WORD/LINE块,根据其BoundingBox与表格单元格的位置重叠度,手动映射到对应单元格中。这部分是控制台的增强逻辑,官方SDK/CLI返回的基础响应不会包含该步骤,所以你会看到CLI和控制台的JSON结果存在差异。
  • 可以尝试改用同步的AnalyzeDocument接口调用(3页文档完全在同步调用的限制范围内),对比异步响应的结果,看是否能让更多文本关联到CELL块。

高效解析JSON并补全缺失内容的脚本建议

核心思路是先提取表格结构,再通过坐标匹配补全未关联的文本块,以下是Python实现示例:

步骤1:解析Textract响应,提取表格与游离文本块

import json
from collections import defaultdict

def parse_textract_response(response):
    tables = []
    free_text_blocks = []
    cell_ids = set()

    # 提取表格结构及CELL块ID集合
    for block in response["Blocks"]:
        if block["BlockType"] == "TABLE":
            table_cells = []
            for rel in block.get("Relationships", []):
                if rel["Type"] == "CHILD":
                    for cell_id in rel["Ids"]:
                        cell = next(b for b in response["Blocks"] if b["Id"] == cell_id)
                        table_cells.append({
                            "row": cell["RowIndex"],
                            "col": cell["ColumnIndex"],
                            "bbox": cell["BoundingBox"],
                            "text": ""
                        })
                        cell_ids.add(cell_id)
            tables.append({"cells": table_cells})
        # 收集未关联到CELL的WORD块
        elif block["BlockType"] == "WORD":
            is_attached_to_cell = False
            for rel in block.get("Relationships", []):
                if rel["Type"] == "CHILD" and any(id in cell_ids for id in rel["Ids"]):
                    is_attached_to_cell = True
                    break
            if not is_attached_to_cell:
                free_text_blocks.append({
                    "text": block["Text"],
                    "bbox": block["BoundingBox"]
                })
    return tables, free_text_blocks

步骤2:通过坐标重叠匹配,补全单元格文本

def calculate_overlap(bbox1, bbox2):
    # 计算两个BoundingBox的重叠率(Textract坐标为页面相对值,范围0-1)
    x_left = max(bbox1["Left"], bbox2["Left"])
    y_top = max(bbox1["Top"], bbox2["Top"])
    x_right = min(bbox1["Left"] + bbox1["Width"], bbox2["Left"] + bbox2["Width"])
    y_bottom = min(bbox1["Top"] + bbox1["Height"], bbox2["Top"] + bbox2["Height"])
    
    if x_left >= x_right or y_top >= y_bottom:
        return 0.0
    overlap_area = (x_right - x_left) * (y_bottom - y_top)
    bbox_area = bbox1["Width"] * bbox1["Height"]
    return overlap_area / bbox_area

def fill_missing_cell_text(tables, free_text_blocks):
    for table in tables:
        for cell in table["cells"]:
            best_match = None
            max_overlap = 0.3  # 重叠率阈值可根据PDF格式调整
            for text_block in free_text_blocks:
                overlap_rate = calculate_overlap(cell["bbox"], text_block["bbox"])
                if overlap_rate > max_overlap:
                    max_overlap = overlap_rate
                    best_match = text_block
            if best_match:
                cell["text"] = best_match["text"]
                free_text_blocks.remove(best_match)
    return tables

步骤3:将表格结构转换为CSV

def tables_to_csv(tables):
    csv_content = ""
    for table in tables:
        # 按行号、列号排序单元格
        row_map = defaultdict(list)
        max_col = 0
        for cell in table["cells"]:
            row_map[cell["row"]].append(cell)
            if cell["col"] > max_col:
                max_col = cell["col"]
        # 生成每行CSV内容
        for row_num in sorted(row_map.keys()):
            row_cells = sorted(row_map[row_num], key=lambda c: c["col"])
            row_data = [""] * max_col
            for cell in row_cells:
                row_data[cell["col"] - 1] = cell["text"].replace(",", "")  # 避免CSV逗号冲突
            csv_content += ",".join(row_data) + "\n"
        csv_content += "\n"  # 不同表格间用空行分隔
    return csv_content

使用示例

# 加载Textract返回的JSON文件
with open("textract_output.json", "r") as f:
    response = json.load(f)

tables, free_text = parse_textract_response(response)
tables = fill_missing_cell_text(tables, free_text)
csv_result = tables_to_csv(tables)

# 保存为CSV文件
with open("final_output.csv", "w") as f:
    f.write(csv_result)

效率优化建议

  • 若处理大量文本块,可引入rtree空间索引库替代全量遍历,大幅提升坐标匹配速度。
  • 对于LINE块,可先合并同一行的WORD块再进行匹配,避免拆分的文本块分散匹配。
  • 针对特定PDF格式,可调整重叠率阈值,平衡匹配准确率和召回率。

内容的提问来源于stack exchange,提问作者Juloblairot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 05:35:28