You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python识别非结构化PDF中的输入字段?求解单选按钮提取问题

解析非结构化PDF中单选按钮等输入字段的解决方案

针对非结构化PDF的表单字段提取问题,以下是几种精准可行的方案:

1. OCR结合图像预处理(适用于扫描件/图像型PDF)

非结构化PDF如果是扫描生成的图像型文件,PyPDF2这类文本解析工具无法生效,需要用OCR技术识别内容,并通过坐标匹配关联单选按钮和对应标签:

from pdf2image import convert_from_path
import pytesseract
import cv2
import numpy as np

# 将PDF页面转为图像
pages = convert_from_path("target.pdf")
for page_idx, page in enumerate(pages):
    # 转换为OpenCV可处理的格式
    img = cv2.cvtColor(np.array(page), cv2.COLOR_RGB2BGR)
    # 图像预处理:灰度化+二值化,提升OCR识别率
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1]
    
    # 提取文本及坐标信息
    custom_config = r'--oem 3 --psm 6'
    text_data = pytesseract.image_to_data(thresh, output_type=pytesseract.Output.DICT, config=custom_config)
    
    # 匹配单选按钮与对应标签
    for i in range(len(text_data['text'])):
        option = text_data['text'][i].strip()
        if option in ['YES', 'NO']:
            opt_x, opt_y, opt_h = text_data['left'][i], text_data['top'][i], text_data['height'][i]
            # 查找垂直对齐且位于单选按钮左侧的标签
            for j in range(len(text_data['text'])):
                label = text_data['text'][j].strip()
                if label and j != i:
                    label_x, label_y = text_data['left'][j], text_data['top'][j]
                    if abs(label_y - opt_y) < opt_h and label_x + text_data['width'][j] < opt_x:
                        print(f"页面{page_idx+1}: {label} -> {option}")

2. 使用pdfplumber进行布局感知解析(适用于文本型非结构化PDF)

pdfplumber比PyPDF2更擅长处理复杂布局,能减少随机空格问题,同时保留文本的位置信息,方便关联字段:

import pdfplumber

with pdfplumber.open("target.pdf") as pdf:
    for page in pdf.pages:
        # 提取带位置信息的文本块,调整x/y容错值适配布局
        text_blocks = page.extract_words(x_tolerance=5, y_tolerance=5)
        radio_options = []
        labels = []
        
        # 分类存储标签和单选按钮选项
        for block in text_blocks:
            text = block['text'].strip()
            if text in ['YES', 'NO']:
                radio_options.append((text, block['x0'], block['top']))
            elif text:
                labels.append((text, block['x1'], block['top']))
        
        # 匹配最近的标签与单选按钮
        for opt_text, opt_x, opt_y in radio_options:
            matched_label = None
            min_distance = float('inf')
            for label_text, label_x, label_y in labels:
                if abs(label_y - opt_y) < 10 and label_x < opt_x:
                    distance = opt_x - label_x
                    if distance < min_distance:
                        min_distance = distance
                        matched_label = label_text
            if matched_label:
                print(f"{matched_label} -> {opt_text}")

3. 调用商业OCR/表单解析API(适用于高精度需求)

如果PDF布局复杂或对精度要求极高,可以使用AWS Textract、Google Cloud Vision这类商业API,它们内置了表单字段识别逻辑,能直接关联标签和单选按钮值:

import boto3

textract = boto3.client('textract')

with open("target.pdf", "rb") as f:
    response = textract.analyze_document(
        Document={'Bytes': f.read()},
        FeatureTypes=['FORMS']
    )

# 解析返回的表单字段
for block in response['Blocks']:
    if block['BlockType'] == 'KEY_VALUE_SET':
        key_text = ""
        value_text = ""
        # 获取标签文本
        if 'KEY' in block['EntityTypes']:
            key_id = block['Relationships'][0]['Ids'][0]
            key_text = next(b['Text'] for b in response['Blocks'] if b['Id'] == key_id)
        # 获取单选按钮值
        if 'VALUE' in block['EntityTypes']:
            value_id = block['Relationships'][0]['Ids'][0]
            value_text = next(b['Text'] for b in response['Blocks'] if b['Id'] == value_id)
        if key_text and value_text:
            print(f"{key_text} -> {value_text}")

注意事项

  • 非结构化PDF没有统一布局,需要根据实际文档调整位置匹配的判断条件(比如部分单选按钮可能在标签下方,需修改y轴比对逻辑)
  • 图像型PDF的OCR精度依赖预处理步骤,可根据实际情况添加降噪、倾斜校正等操作
  • 文本型PDF优先使用pdfplumber,能解决PyPDF2提取文本时的随机空格问题

内容的提问来源于stack exchange,提问作者Priyanshu Lahiri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 11:16:00