You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量截取PDF数学教材中的完整Theorem定理模块?

问题需求

需要从数学PDF教材中截取所有带有红色标题、黄底色的「Theorem」定理框,保存至本地文件夹用于快速查阅重要定理。现有Python代码仅能截取到"Theorem"单词本身,无法完整捕获整个定理模块;尝试放大固定尺寸的bounding box,但因各定理尺寸不一效果不佳,寻求可适配不同尺寸定理框的完整截取方案。

解决方案思路

核心是通过识别黄底色连续区域定位完整定理框,而非仅依赖文本位置,具体步骤:

  • 生成整页高分辨率图片,遍历像素识别所有黄底色的连续区块
  • 结合"Theorem"文本的坐标,匹配包含该文本的黄底色区块,确认目标定理框
  • 对匹配到的区块进行精准截取,避免固定尺寸的局限性

修改后的代码

import os
import fitz  # PyMuPDF
import re
from PIL import Image
import numpy as np

# 配置参数
pdf_file = r'C:\Users\Me\Desktop\Textbook\MathTextbookPDF.pdf'
output_folder = r'C:\Users\Me\Desktop\Theorems'
start_page = 100
end_page = 200
dpi = 500
# 黄底色的RGB范围(可根据实际PDF调整)
YELLOW_LOW = np.array([250, 220, 180])
YELLOW_HIGH = np.array([255, 245, 220])
# 红色标题的RGB范围(可选,用于二次验证)
RED_LOW = np.array([230, 0, 0])
RED_HIGH = np.array([255, 100, 100])

# 创建输出文件夹
if not os.path.exists(output_folder):
    os.makedirs(output_folder)

# 打开PDF
pdf_document = fitz.open(pdf_file)

def find_yellow_regions(pil_image):
    """识别图片中所有黄底色的连续区域,返回每个区域的PDF页面坐标(x0,y0,x1,y1)"""
    img_np = np.array(pil_image)
    # 筛选黄底色像素
    mask = np.all((img_np >= YELLOW_LOW) & (img_np <= YELLOW_HIGH), axis=-1)
    # 寻找连续区域
    from skimage.measure import label, regionprops
    labeled_mask = label(mask)
    regions = regionprops(labeled_mask)
    
    # 转换为PDF页面的原始坐标(还原dpi缩放比例)
    scale = 72 / dpi  # PDF默认72DPI,当前图片为dpi分辨率,需反向缩放
    region_coords = []
    page_height = pil_image.height * scale  # PDF页面总高度
    
    for region in regions:
        min_row, min_col, max_row, max_col = region.bbox
        # 图片坐标转PDF页面坐标,注意PDF y轴从下往上
        y0_img = min_row * scale
        y1_img = max_row * scale
        x0 = min_col * scale
        x1 = max_col * scale
        
        y0_pdf = page_height - y1_img
        y1_pdf = page_height - y0_img
        region_coords.append((x0, y0_pdf, x1, y1_pdf))
    
    return region_coords

def has_red_title(page, region):
    """验证区域内是否包含红色标题,用于精准过滤非目标区块"""
    x0, y0, x1, y1 = region
    # 截取区块顶部的标题区域(高度可根据实际调整)
    title_region = (x0, y0, x1, y0 + 20)
    pix = page.get_pixmap(matrix=fitz.Matrix(dpi/72, dpi/72), clip=title_region)
    img = Image.frombytes("RGB", [pix.width, pix.height], pix.samples)
    img_np = np.array(img)
    
    # 检查是否存在红色像素
    red_mask = np.all((img_np >= RED_LOW) & (img_np <= RED_HIGH), axis=-1)
    return np.any(red_mask)

# 遍历目标页面
for page_idx in range(start_page - 1, end_page):
    page = pdf_document[page_idx]
    page_num = page_idx + 1
    
    # 获取整页高分辨率图片
    full_pix = page.get_pixmap(matrix=fitz.Matrix(dpi/72, dpi/72))
    full_img = Image.frombytes("RGB", [full_pix.width, full_pix.height], full_pix.samples)
    
    # 识别页面中所有黄底色区域
    yellow_regions = find_yellow_regions(full_img)
    if not yellow_regions:
        continue
    
    # 找到页面中所有"Theorem"文本的位置
    theorem_instances = page.search_for(r'Theorem')
    if not theorem_instances:
        continue
    
    # 匹配文本对应的黄底色定理框
    for idx, theorem_rect in enumerate(theorem_instances):
        tx0, ty0, tx1, ty1 = theorem_rect
        matched_region = None
        
        # 寻找包含当前Theorem文本的黄底色区域
        for region in yellow_regions:
            rx0, ry0, rx1, ry1 = region
            if rx0 <= tx0 and ry0 <= ty0 and rx1 >= tx1 and ry1 >= ty1:
                matched_region = region
                break
        
        if matched_region and has_red_title(page, matched_region):
            # 截取并保存目标定理框
            pix = page.get_pixmap(matrix=fitz.Matrix(dpi/72, dpi/72), clip=matched_region)
            save_path = os.path.join(output_folder, f'page_{page_num}_theorem_{idx+1}.png')
            pix.save(save_path)

# 关闭PDF文档
pdf_document.close()

代码说明

  1. 黄底色区域识别:借助numpy筛选黄底色像素,结合skimage找到连续区块,确保覆盖整个定理框范围
  2. 坐标转换:将图片像素坐标还原为PDF页面原始坐标,保证截取位置精准
  3. 文本-区块匹配:通过"Theorem"文本位置锁定对应的黄底色区块,避免误截取
  4. 二次验证:添加红色标题识别逻辑,进一步过滤非目标黄底内容

内容的提问来源于stack exchange,提问作者raysrule81

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 21:05:55