如何用Python识别非结构化PDF中的输入字段?求解单选按钮提取问题
解析非结构化PDF中单选按钮等输入字段的解决方案
针对非结构化PDF的表单字段提取问题,以下是几种精准可行的方案:
1. OCR结合图像预处理(适用于扫描件/图像型PDF)
非结构化PDF如果是扫描生成的图像型文件,PyPDF2这类文本解析工具无法生效,需要用OCR技术识别内容,并通过坐标匹配关联单选按钮和对应标签:
from pdf2image import convert_from_path import pytesseract import cv2 import numpy as np # 将PDF页面转为图像 pages = convert_from_path("target.pdf") for page_idx, page in enumerate(pages): # 转换为OpenCV可处理的格式 img = cv2.cvtColor(np.array(page), cv2.COLOR_RGB2BGR) # 图像预处理:灰度化+二值化,提升OCR识别率 gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # 提取文本及坐标信息 custom_config = r'--oem 3 --psm 6' text_data = pytesseract.image_to_data(thresh, output_type=pytesseract.Output.DICT, config=custom_config) # 匹配单选按钮与对应标签 for i in range(len(text_data['text'])): option = text_data['text'][i].strip() if option in ['YES', 'NO']: opt_x, opt_y, opt_h = text_data['left'][i], text_data['top'][i], text_data['height'][i] # 查找垂直对齐且位于单选按钮左侧的标签 for j in range(len(text_data['text'])): label = text_data['text'][j].strip() if label and j != i: label_x, label_y = text_data['left'][j], text_data['top'][j] if abs(label_y - opt_y) < opt_h and label_x + text_data['width'][j] < opt_x: print(f"页面{page_idx+1}: {label} -> {option}")
2. 使用pdfplumber进行布局感知解析(适用于文本型非结构化PDF)
pdfplumber比PyPDF2更擅长处理复杂布局,能减少随机空格问题,同时保留文本的位置信息,方便关联字段:
import pdfplumber with pdfplumber.open("target.pdf") as pdf: for page in pdf.pages: # 提取带位置信息的文本块,调整x/y容错值适配布局 text_blocks = page.extract_words(x_tolerance=5, y_tolerance=5) radio_options = [] labels = [] # 分类存储标签和单选按钮选项 for block in text_blocks: text = block['text'].strip() if text in ['YES', 'NO']: radio_options.append((text, block['x0'], block['top'])) elif text: labels.append((text, block['x1'], block['top'])) # 匹配最近的标签与单选按钮 for opt_text, opt_x, opt_y in radio_options: matched_label = None min_distance = float('inf') for label_text, label_x, label_y in labels: if abs(label_y - opt_y) < 10 and label_x < opt_x: distance = opt_x - label_x if distance < min_distance: min_distance = distance matched_label = label_text if matched_label: print(f"{matched_label} -> {opt_text}")
3. 调用商业OCR/表单解析API(适用于高精度需求)
如果PDF布局复杂或对精度要求极高,可以使用AWS Textract、Google Cloud Vision这类商业API,它们内置了表单字段识别逻辑,能直接关联标签和单选按钮值:
import boto3 textract = boto3.client('textract') with open("target.pdf", "rb") as f: response = textract.analyze_document( Document={'Bytes': f.read()}, FeatureTypes=['FORMS'] ) # 解析返回的表单字段 for block in response['Blocks']: if block['BlockType'] == 'KEY_VALUE_SET': key_text = "" value_text = "" # 获取标签文本 if 'KEY' in block['EntityTypes']: key_id = block['Relationships'][0]['Ids'][0] key_text = next(b['Text'] for b in response['Blocks'] if b['Id'] == key_id) # 获取单选按钮值 if 'VALUE' in block['EntityTypes']: value_id = block['Relationships'][0]['Ids'][0] value_text = next(b['Text'] for b in response['Blocks'] if b['Id'] == value_id) if key_text and value_text: print(f"{key_text} -> {value_text}")
注意事项
- 非结构化PDF没有统一布局,需要根据实际文档调整位置匹配的判断条件(比如部分单选按钮可能在标签下方,需修改y轴比对逻辑)
- 图像型PDF的OCR精度依赖预处理步骤,可根据实际情况添加降噪、倾斜校正等操作
- 文本型PDF优先使用pdfplumber,能解决PyPDF2提取文本时的随机空格问题
内容的提问来源于stack exchange,提问作者Priyanshu Lahiri
相关产品推荐
相关产品推荐

