You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python识别PDF高亮文本并提取选择题答案

问题描述
  • 核心需求:从多份含选择题的PDF中,分别生成仅包含题目的PDF/TXT文件,以及按1-D、2-C格式记录正确答案的独立PDF/TXT文件
  • 已知条件:
    • 题目页以Question开头,易定位;答案页必含B、C选项,可快速识别
    • 题目页与对应答案页几乎连续,题目基本为单页显示,跨页情况极少
    • 现存痛点:无法通过常规文本提取判断哪个选项是高亮的正确答案,备选方案(手动提取、转图片检测高亮)效率极低
解决方案

由于常规文本提取无法直接获取高亮信息,需通过PDF的注释属性或专业PDF处理库来识别高亮选项,以下是两种可行实现方案:

方案一:基于pypdf提取高亮注释

PDF的高亮通常以Highlight类型注释存在,可通过遍历页面注释定位高亮区域,提取对应选项:

from pypdf import PdfReader, PdfWriter
import glob

# 初始化存储变量
question_text = ""
answers = []
question_counter = 0

# 遍历目标PDF文件
for file_path in glob.glob('samplepath/*.pdf'):
    reader = PdfReader(file_path)
    print(f"处理文件: {file_path}")
    
    for page_num, page in enumerate(reader.pages):
        page_text = page.extract_text()
        
        # 处理题目页
        if "Question" in page_text:
            question_counter += 1
            question_text += f"Question {question_counter}\n{page_text}\n\n"
        
        # 处理答案页
        elif "\nB" in page_text and "\nC" in page_text:
            highlighted_option = ""
            # 遍历页面所有注释
            for annot in page.annotations:
                if annot.subtype == "/Highlight":
                    # 优先读取注释自带的选中文本
                    if hasattr(annot, 'contents') and annot.contents.strip():
                        highlighted_option = annot.contents.strip().split('.')[0]
                    else:
                        # 若注释无文本,通过坐标匹配选项
                        lines = page_text.split('\n')
                        for line in lines:
                            if line.strip().startswith(('A.', 'B.', 'C.', 'D.')):
                                # 获取文本坐标与高亮区域坐标对比(简化版)
                                text_boxes = page.extract_text(extraction_mode="layout").split('\n')
                                for box in text_boxes:
                                    if line.strip() in box:
                                        coords = box.split(']')[0][1:].split(',')
                                        x1, y1 = map(float, coords[:2])
                                        quad_points = annot.get_object()['/QuadPoints']
                                        if (quad_points[0]-10 <= x1 <= quad_points[2]+10) and (quad_points[1]-10 <= y1 <= quad_points[3]+10):
                                            highlighted_option = line.strip().split('.')[0]
                                            break
                            if highlighted_option:
                                break
            if highlighted_option:
                answers.append(f"{question_counter}-{highlighted_option}")

# 保存纯题目TXT
with open('questions.txt', 'w', encoding='utf-8') as f:
    f.write(question_text)

# 保存答案TXT
with open('answers.txt', 'w', encoding='utf-8') as f:
    f.write('\n'.join(answers))

# 保存原格式题目PDF
question_writer = PdfWriter()
for file_path in glob.glob('samplepath/*.pdf'):
    reader = PdfReader(file_path)
    for page in reader.pages:
        if "Question" in page.extract_text():
            question_writer.add_page(page)
with open('questions.pdf', 'wb') as f:
    question_writer.write(f)

方案二:使用pdfplumber简化高亮识别

pdfplumber对文本与注释的处理更精细,无需手动处理坐标,直接提取高亮区域文本:

import pdfplumber
import glob

question_text = ""
answers = []
question_counter = 0

for file_path in glob.glob('samplepath/*.pdf'):
    with pdfplumber.open(file_path) as pdf:
        print(f"处理文件: {file_path}")
        for page in pdf.pages:
            page_text = page.extract_text()
            
            # 处理题目页
            if "Question" in page_text:
                question_counter += 1
                question_text += f"Question {question_counter}\n{page_text}\n\n"
            
            # 处理答案页
            elif "\nB" in page_text and "\nC" in page_text:
                highlights = page.highlights
                if highlights:
                    # 提取高亮区域内的选项字母
                    for highlight in highlights:
                        highlight_content = page.within_bbox(highlight.bbox).extract_text().strip()
                        option = highlight_content.split('.')[0]
                        answers.append(f"{question_counter}-{option}")

# 保存文件逻辑同方案一
with open('questions.txt', 'w', encoding='utf-8') as f:
    f.write(question_text)
with open('answers.txt', 'w', encoding='utf-8') as f:
    f.write('\n'.join(answers))
注意事项
  • 如果PDF的高亮是通过灰色背景文本框而非标准注释实现的,才需要考虑图像识别(如OpenCV)检测灰度区域,这类情况确实效率较低,建议优先尝试上述两种方案
  • 先单文件测试逻辑,确认能正确识别题目与高亮选项后,再进行批量处理

内容的提问来源于stack exchange,提问作者Uğur Dinç

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 01:32:48