如何批量截取PDF数学教材中的完整Theorem定理模块?
问题需求
需要从数学PDF教材中截取所有带有红色标题、黄底色的「Theorem」定理框,保存至本地文件夹用于快速查阅重要定理。现有Python代码仅能截取到"Theorem"单词本身,无法完整捕获整个定理模块;尝试放大固定尺寸的bounding box,但因各定理尺寸不一效果不佳,寻求可适配不同尺寸定理框的完整截取方案。
解决方案思路
核心是通过识别黄底色连续区域定位完整定理框,而非仅依赖文本位置,具体步骤:
- 生成整页高分辨率图片,遍历像素识别所有黄底色的连续区块
- 结合"Theorem"文本的坐标,匹配包含该文本的黄底色区块,确认目标定理框
- 对匹配到的区块进行精准截取,避免固定尺寸的局限性
修改后的代码
import os import fitz # PyMuPDF import re from PIL import Image import numpy as np # 配置参数 pdf_file = r'C:\Users\Me\Desktop\Textbook\MathTextbookPDF.pdf' output_folder = r'C:\Users\Me\Desktop\Theorems' start_page = 100 end_page = 200 dpi = 500 # 黄底色的RGB范围(可根据实际PDF调整) YELLOW_LOW = np.array([250, 220, 180]) YELLOW_HIGH = np.array([255, 245, 220]) # 红色标题的RGB范围(可选,用于二次验证) RED_LOW = np.array([230, 0, 0]) RED_HIGH = np.array([255, 100, 100]) # 创建输出文件夹 if not os.path.exists(output_folder): os.makedirs(output_folder) # 打开PDF pdf_document = fitz.open(pdf_file) def find_yellow_regions(pil_image): """识别图片中所有黄底色的连续区域,返回每个区域的PDF页面坐标(x0,y0,x1,y1)""" img_np = np.array(pil_image) # 筛选黄底色像素 mask = np.all((img_np >= YELLOW_LOW) & (img_np <= YELLOW_HIGH), axis=-1) # 寻找连续区域 from skimage.measure import label, regionprops labeled_mask = label(mask) regions = regionprops(labeled_mask) # 转换为PDF页面的原始坐标(还原dpi缩放比例) scale = 72 / dpi # PDF默认72DPI,当前图片为dpi分辨率,需反向缩放 region_coords = [] page_height = pil_image.height * scale # PDF页面总高度 for region in regions: min_row, min_col, max_row, max_col = region.bbox # 图片坐标转PDF页面坐标,注意PDF y轴从下往上 y0_img = min_row * scale y1_img = max_row * scale x0 = min_col * scale x1 = max_col * scale y0_pdf = page_height - y1_img y1_pdf = page_height - y0_img region_coords.append((x0, y0_pdf, x1, y1_pdf)) return region_coords def has_red_title(page, region): """验证区域内是否包含红色标题,用于精准过滤非目标区块""" x0, y0, x1, y1 = region # 截取区块顶部的标题区域(高度可根据实际调整) title_region = (x0, y0, x1, y0 + 20) pix = page.get_pixmap(matrix=fitz.Matrix(dpi/72, dpi/72), clip=title_region) img = Image.frombytes("RGB", [pix.width, pix.height], pix.samples) img_np = np.array(img) # 检查是否存在红色像素 red_mask = np.all((img_np >= RED_LOW) & (img_np <= RED_HIGH), axis=-1) return np.any(red_mask) # 遍历目标页面 for page_idx in range(start_page - 1, end_page): page = pdf_document[page_idx] page_num = page_idx + 1 # 获取整页高分辨率图片 full_pix = page.get_pixmap(matrix=fitz.Matrix(dpi/72, dpi/72)) full_img = Image.frombytes("RGB", [full_pix.width, full_pix.height], full_pix.samples) # 识别页面中所有黄底色区域 yellow_regions = find_yellow_regions(full_img) if not yellow_regions: continue # 找到页面中所有"Theorem"文本的位置 theorem_instances = page.search_for(r'Theorem') if not theorem_instances: continue # 匹配文本对应的黄底色定理框 for idx, theorem_rect in enumerate(theorem_instances): tx0, ty0, tx1, ty1 = theorem_rect matched_region = None # 寻找包含当前Theorem文本的黄底色区域 for region in yellow_regions: rx0, ry0, rx1, ry1 = region if rx0 <= tx0 and ry0 <= ty0 and rx1 >= tx1 and ry1 >= ty1: matched_region = region break if matched_region and has_red_title(page, matched_region): # 截取并保存目标定理框 pix = page.get_pixmap(matrix=fitz.Matrix(dpi/72, dpi/72), clip=matched_region) save_path = os.path.join(output_folder, f'page_{page_num}_theorem_{idx+1}.png') pix.save(save_path) # 关闭PDF文档 pdf_document.close()
代码说明
- 黄底色区域识别:借助
numpy筛选黄底色像素,结合skimage找到连续区块,确保覆盖整个定理框范围 - 坐标转换:将图片像素坐标还原为PDF页面原始坐标,保证截取位置精准
- 文本-区块匹配:通过"Theorem"文本位置锁定对应的黄底色区块,避免误截取
- 二次验证:添加红色标题识别逻辑,进一步过滤非目标黄底内容
内容的提问来源于stack exchange,提问作者raysrule81
相关产品推荐
相关产品推荐

