如何用Python识别PDF高亮文本并提取选择题答案
问题描述
- 核心需求:从多份含选择题的PDF中,分别生成仅包含题目的PDF/TXT文件,以及按
1-D、2-C格式记录正确答案的独立PDF/TXT文件 - 已知条件:
- 题目页以
Question开头,易定位;答案页必含B、C选项,可快速识别 - 题目页与对应答案页几乎连续,题目基本为单页显示,跨页情况极少
- 现存痛点:无法通过常规文本提取判断哪个选项是高亮的正确答案,备选方案(手动提取、转图片检测高亮)效率极低
- 题目页以
解决方案
由于常规文本提取无法直接获取高亮信息,需通过PDF的注释属性或专业PDF处理库来识别高亮选项,以下是两种可行实现方案:
方案一:基于pypdf提取高亮注释
PDF的高亮通常以Highlight类型注释存在,可通过遍历页面注释定位高亮区域,提取对应选项:
from pypdf import PdfReader, PdfWriter import glob # 初始化存储变量 question_text = "" answers = [] question_counter = 0 # 遍历目标PDF文件 for file_path in glob.glob('samplepath/*.pdf'): reader = PdfReader(file_path) print(f"处理文件: {file_path}") for page_num, page in enumerate(reader.pages): page_text = page.extract_text() # 处理题目页 if "Question" in page_text: question_counter += 1 question_text += f"Question {question_counter}\n{page_text}\n\n" # 处理答案页 elif "\nB" in page_text and "\nC" in page_text: highlighted_option = "" # 遍历页面所有注释 for annot in page.annotations: if annot.subtype == "/Highlight": # 优先读取注释自带的选中文本 if hasattr(annot, 'contents') and annot.contents.strip(): highlighted_option = annot.contents.strip().split('.')[0] else: # 若注释无文本,通过坐标匹配选项 lines = page_text.split('\n') for line in lines: if line.strip().startswith(('A.', 'B.', 'C.', 'D.')): # 获取文本坐标与高亮区域坐标对比(简化版) text_boxes = page.extract_text(extraction_mode="layout").split('\n') for box in text_boxes: if line.strip() in box: coords = box.split(']')[0][1:].split(',') x1, y1 = map(float, coords[:2]) quad_points = annot.get_object()['/QuadPoints'] if (quad_points[0]-10 <= x1 <= quad_points[2]+10) and (quad_points[1]-10 <= y1 <= quad_points[3]+10): highlighted_option = line.strip().split('.')[0] break if highlighted_option: break if highlighted_option: answers.append(f"{question_counter}-{highlighted_option}") # 保存纯题目TXT with open('questions.txt', 'w', encoding='utf-8') as f: f.write(question_text) # 保存答案TXT with open('answers.txt', 'w', encoding='utf-8') as f: f.write('\n'.join(answers)) # 保存原格式题目PDF question_writer = PdfWriter() for file_path in glob.glob('samplepath/*.pdf'): reader = PdfReader(file_path) for page in reader.pages: if "Question" in page.extract_text(): question_writer.add_page(page) with open('questions.pdf', 'wb') as f: question_writer.write(f)
方案二:使用pdfplumber简化高亮识别
pdfplumber对文本与注释的处理更精细,无需手动处理坐标,直接提取高亮区域文本:
import pdfplumber import glob question_text = "" answers = [] question_counter = 0 for file_path in glob.glob('samplepath/*.pdf'): with pdfplumber.open(file_path) as pdf: print(f"处理文件: {file_path}") for page in pdf.pages: page_text = page.extract_text() # 处理题目页 if "Question" in page_text: question_counter += 1 question_text += f"Question {question_counter}\n{page_text}\n\n" # 处理答案页 elif "\nB" in page_text and "\nC" in page_text: highlights = page.highlights if highlights: # 提取高亮区域内的选项字母 for highlight in highlights: highlight_content = page.within_bbox(highlight.bbox).extract_text().strip() option = highlight_content.split('.')[0] answers.append(f"{question_counter}-{option}") # 保存文件逻辑同方案一 with open('questions.txt', 'w', encoding='utf-8') as f: f.write(question_text) with open('answers.txt', 'w', encoding='utf-8') as f: f.write('\n'.join(answers))
注意事项
- 如果PDF的高亮是通过灰色背景文本框而非标准注释实现的,才需要考虑图像识别(如OpenCV)检测灰度区域,这类情况确实效率较低,建议优先尝试上述两种方案
- 先单文件测试逻辑,确认能正确识别题目与高亮选项后,再进行批量处理
内容的提问来源于stack exchange,提问作者Uğur Dinç
相关产品推荐
相关产品推荐

