You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用Google DocumentAI提取文档复选框值返回空的解决方法

图片描述

问题说明

需要通过Python调用Google DocumentAI接口提取文档内复选框的取值,现有代码运行时复选框对应提取结果为空字符串,无法获取复选框勾选状态。
现有实现代码如下:

import pandas as pd
from google.cloud import documentai_v1 as documentai
import re

def get_text(doc_element: dict, document: dict):
 
    response = ""
    # 原代码此处存在多余缩进,运行会报语法错
    for segment in doc_element.text_anchor.text_segments:
        start_index = (
            int(segment.start_index)
            if segment in doc_element.text_anchor.text_segments
            else 0
        )
        end_index = int(segment.end_index)
        response += document.text[start_index:end_index]
    return response

def getIndexes(dfObj, value):
    
    listOfPos = list()
    result = dfObj.isin([value])
    seriesObj = result.any()
    columnNames = list(seriesObj[seriesObj == True].index)
    
    for col in columnNames:
        rows = list(result[col][result[col] == True].index)
        for row in rows:
            listOfPos.append(row)
   
    return listOfPos

def online_process(
    project_id: str,
    location: str,
    processor_id: str,
    file_path: str,
    mime_type: str,
) -> documentai.Document:
   
    opts = {"api_endpoint": f"{location}-documentai.googleapis.com"}


    documentai_client = documentai.DocumentProcessorServiceClient(client_options=opts)

    resource_name = documentai_client.processor_path(project_id, location, processor_id)

   
    with open(file_path, "rb") as image:
        image_content = image.read()

        
        raw_document = documentai.RawDocument(
            content=image_content, mime_type=mime_type
        )

      
        request = documentai.ProcessRequest(
            name=resource_name, raw_document=raw_document
        )

        result = documentai_client.process_document(request=request)

        return result.document
问题原因

现有代码的get_text方法只适配了普通文本块的提取逻辑,而Google DocumentAI不会把复选框作为普通文本返回:复选框属于表单字段,单独存储在响应的pages[].form_fields结构中,走文本截取逻辑自然会拿到空值。

正确实现方案

前置要求

  • 必须使用表单解析器(Form Parser) 或支持表单字段识别的Document AI处理器,纯通用OCR处理器不会返回复选框状态字段
  • 不需要通过文本锚点拼接内容,直接读取表单字段的内置属性即可拿到勾选状态

核心逻辑

每个表单字段(form_field)包含两个核心部分:

  • field_name:复选框对应的旁侧标签文本,比如“已阅读协议”“选项1”这类说明文字
  • field_value:字段值,如果是复选框类型,会自带value_checkbox属性,直接读取该属性的checked布尔值就能得到勾选状态,True为已勾选,False为未勾选

可直接运行的提取代码

在原有代码基础上,添加复选框提取逻辑即可:

def extract_checkboxes(document: documentai.Document):
    """
    从DocumentAI返回结果中提取所有复选框的标签和勾选状态
    返回格式:列表,每个元素为(复选框标签文本, 是否勾选)
    """
    checkbox_result = []
    for page in document.pages:
        for form_field in page.form_fields:
            # 校验当前字段是否为复选框类型
            if hasattr(form_field.field_value, 'value_checkbox'):
                # 提取复选框对应的标签文本
                field_name_text = get_text(form_field.field_name, document).strip()
                # 直接读取内置的勾选状态属性
                is_checked = form_field.field_value.value_checkbox.checked
                checkbox_result.append( (field_name_text, is_checked) )
    return checkbox_result

# 调用示例
if __name__ == "__main__":
    # 替换为实际的配置参数
    doc = online_process(
        project_id="你的项目ID",
        location="处理器所在地域,比如us",
        processor_id="你的表单解析器ID",
        file_path="待识别文件路径",
        mime_type="文件MIME类型,比如application/pdf或者image/png"
    )
    checkboxes = extract_checkboxes(doc)
    # 打印输出结果
    for label, checked in checkboxes:
        print(f"选项:{label},勾选状态:{'已勾选' if checked else '未勾选'}")

额外说明

  • 记得先修正原get_text函数里for循环前的多余缩进,否则会直接触发语法错误
  • 如果需要把复选框和表格结构做关联,可以结合每个form_field的bounding_poly坐标信息,和表格单元格的坐标做位置匹配即可

内容的提问来源于stack exchange,提问作者kavin iyal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 00:36:27