PyMuPDF提取非可编辑PDF键值对及复选框数据遇问题求助
问题:PyMuPDF提取非可编辑PDF字段与复选框数据异常
遇到两个核心问题:
- 提取的字典中字段与值对应错误,出现先列所有键再单独列对应值的情况
- 无法读取任何复选框数据
尝试的代码
import fitz import pandas as pd import re # Function to clean text def clean_text(text): return re.sub('\\s+', ' ', text).strip() def is_field_name(line): # A field name is likely to end with a colon followed by an optional space return bool(re.match(r'.*:\\s*$', line)) # Function to determine if the next line is a checkbox indicator def is_checkbox(line): # Looking for lines that have a checkbox indication, e.g., "\[X\] Yes" or "\[ \] No" return bool(re.match(r'\[(X| )\]\\s\*(Yes|No)', line)) pdf_document = fitz.open('PTO 2024.pdf') page = pdf_document[0] text = page.get_text("text") pdf_document.close() lines = text.split('\\n') extracted_data = {} current_field = None # Process each line, assuming that fields are followed by their values or checkboxes for i, line in enumerate(lines): line = clean_text(line) if is_field_name(line): # The current line is a field name current_field = line[:-1] # Remove the colon at the end elif current_field: if is_checkbox(line): # Line is a checkbox indicator, e.g., "\[X\] Yes" extracted_data[current_field] = 'Checked' if 'Yes' in line else 'Unchecked' current_field = None # Reset current field after capturing checkbox elif line: # Line has content and is not a checkbox, so it's a value for the current field extracted_data[current_field] = line current_field = None # Reset current field after capturing value df_extracted = pd.DataFrame(list(extracted_data.items()), columns=['Field', 'Value']) print("Extracted Lines:") for line in lines: print(line) print("\nExtracted Data:") print(df_extracted.head())
问题分析与解决方案
问题1:字段与值对应错误
原因
原代码假设字段名的下一行必然是对应值,但实际PDF文本提取时,字段与值之间可能存在空行,或者值跨多行,导致current_field未被及时赋值,后续新字段覆盖旧字段,最终出现键值错位。
修复方案
- 过滤空行,避免空行干扰字段匹配逻辑
- 允许字段值跨多行累积,直到遇到下一个字段名
- 优化字段名的清理逻辑,兼容冒号后带空格的情况
问题2:无法读取复选框数据
原因
- 正则表达式错误:原正则
r'\[(X| )\]\\s\*(Yes|No)'存在转义错误(\\s*应为\s*),且仅匹配Yes/No选项,兼容性差 - 非可编辑PDF的复选框可能是图形元素而非文本,
get_text("text")无法提取这类内容
修复方案
- 修正正则表达式,匹配更通用的复选框格式(如
[X]或[ ]开头的任意行) - 添加图形复选框提取逻辑:通过
get_drawings()获取页面图形,识别复选框形状(正方形),并匹配附近的字段文本
修改后的完整代码
import fitz import pandas as pd import re # 清理文本,合并多余空格但保留必要分隔 def clean_text(text): return re.sub(r'\s+', ' ', text).strip() def is_field_name(line): # 匹配以冒号结尾的字段名,兼容冒号后带空格的情况 return bool(re.match(r'.*:\s*$', line)) def is_checkbox(line): # 匹配[X]或[ ]开头的行,支持后续任意文本 return bool(re.match(r'\[(X| )\]\s*', line)) pdf_document = fitz.open('PTO 2024.pdf') page = pdf_document[0] # 提取并预处理文本行,过滤空行 text = page.get_text("text") lines = [clean_text(line) for line in text.split('\n') if clean_text(line)] # 提取页面图形,用于识别非文本型复选框 drawings = page.get_drawings() checkbox_shapes = [] for draw in drawings: # 筛选正方形图形(假设复选框为10-20px的正方形,可根据实际调整) rect = draw['rect'] if abs(rect.width - rect.height) < 2 and 10 < rect.width < 20: checkbox_shapes.append(rect) # 提取所有文本的位置信息,用于匹配图形复选框对应的字段 word_list = page.get_text("words") # 格式:(x0, y0, x1, y1, text, block_no, line_no, word_no) extracted_data = {} current_field = None # 处理文本型字段与复选框 for line in lines: if is_field_name(line): # 清理字段名,去掉末尾的冒号和空格 current_field = line.rstrip(': ').strip() elif current_field: if is_checkbox(line): # 判断复选框状态 extracted_data[current_field] = 'Checked' if '[X]' in line else 'Unchecked' # 提取复选框对应的选项文本(可选) option_text = line.split(']')[-1].strip() if option_text: extracted_data[f"{current_field} (选项)"] = option_text current_field = None else: # 处理多行字段值,累积内容 if current_field in extracted_data: extracted_data[current_field] += f" {line}" else: extracted_data[current_field] = line # 处理图形型复选框(文本提取不到的情况) for rect in checkbox_shapes: # 找到复选框附近的文本(交集判断) nearby_words = [word for word in word_list if fitz.Rect(word[:4]).intersects(rect)] if nearby_words: # 合并附近文本作为字段名 field_name = ' '.join([word[4] for word in nearby_words]) # 判断复选框是否被勾选(通过填充颜色判断,白色为未勾选) is_checked = any(draw['fill'] != (1.0, 1.0, 1.0) for draw in drawings if draw['rect'] == rect) extracted_data[field_name] = 'Checked' if is_checked else 'Unchecked' # 转换为DataFrame并输出 df_extracted = pd.DataFrame(list(extracted_data.items()), columns=['Field', 'Value']) print("提取的文本行:") for line in lines: print(line) print("\n提取的数据:") print(df_extracted) pdf_document.close()
内容的提问来源于stack exchange,提问作者Mar
相关产品推荐
相关产品推荐

