如何从带有复选框的PDF文件中提取被勾选框对应的值
PDF复选框勾选状态及对应旁侧内容提取方案
针对交互式原生表单PDF
这类PDF的复选框是标准表单控件,可直接通过PDF解析库读取状态和关联内容:
- 使用
pdfplumber(推荐,表单解析能力更强)
代码示例:import pdfplumber with pdfplumber.open("待解析文件.pdf") as pdf: for page in pdf.pages: form_fields = page.form_fields for field in form_fields: # 筛选已勾选的复选框控件 if field.get("field_type") == "checkbox" and field.get("value") == "/Yes": # 优先取控件预设的标签文本 if field.get("label"): print(field["label"].strip()) continue # 无预设标签则提取复选框右侧同水平行的文本 x1, y1, x2, y2 = field["x"], field["y"]-3, page.width, field["y"]+3 print(page.within_bbox((x1+5, y1, x2, y2)).extract_text().strip()) - 使用
PyPDF2
代码示例:from PyPDF2 import PdfReader reader = PdfReader("待解析文件.pdf") for page in reader.pages: if "/Annots" not in page: continue for annot in page["/Annots"]: field = annot.get_object() # 筛选已勾选的按钮类控件(复选框属于按钮类表单) if field.get("/FT") == "/Btn" and field.get("/V") == "/Yes": print(field.get("/T", "无关联标签").strip()) # 如需提取旁侧文本可结合PyMuPDF等库的文本位置解析能力实现
针对扫描件/扁平化非交互式PDF
这类PDF的复选框是图片像素,没有表单控件属性,需要结合OCR和图像识别处理:
- 先通过
opencv或pillow定位所有复选框的坐标位置 - 用模板匹配或轻量二分类模型判断复选框是否被勾选
- 调用
pytesseract或PaddleOCR识别全页文本及对应坐标 - 将已勾选的复选框坐标和同水平行的文本做匹配,即可得到对应内容
提示:如果是固定版式的PDF,提前预设复选框坐标模板可以大幅提升识别效率和准确率。
内容的提问来源于stack exchange,提问作者kk97
相关产品推荐
相关产品推荐

