You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于坐标切割PDF图像:已获取题号坐标后的切割实现求助

解决方案:基于OCR坐标切割PDF页面为独立题目块

你的核心问题是当前用PyPDF2获取的坐标不准确——match.start()返回的是文本字符的索引,不是页面上的实际空间坐标,导致y坐标完全无法用来切割。下面是可落地的解决流程:

工具选择

放弃PyPDF2的文本位置提取,改用pdfplumber——它能精准获取每个文本块的边界框(bbox),包含真实的x/y坐标;配合Pillow处理图像切割。

先安装依赖:

pip install pdfplumber pillow

完整实现代码

import pdfplumber
from PIL import Image
import re
import os

# 定义题号匹配正则(匹配101.到200.的格式)
QUESTION_NUM_PATTERN = re.compile(r'^(1\d{2}|200)\.$')

def extract_question_coords(pdf_path):
    page_question_coords = {}
    with pdfplumber.open(pdf_path) as pdf:
        for page_num, page in enumerate(pdf.pages, start=1):
            page_width = page.width
            page_height = page.height
            coords = []
            
            # 遍历页面所有文本字符块,提取题号的真实坐标
            for char in page.chars:
                text = char['text'].strip()
                if QUESTION_NUM_PATTERN.match(text):
                    bbox = char['bbox']
                    coords.append({
                        'num': text,
                        'top_y': bbox[1],   # PDF坐标系下的顶部y值
                        'bottom_y': bbox[3] # PDF坐标系下的底部y值
                    })
            
            # 按顶部y坐标从上到下排序,确保题号顺序正确
            coords.sort(key=lambda x: x['top_y'])
            page_question_coords[page_num] = {
                'coords': coords,
                'page_size': (page_width, page_height)
            }
    return page_question_coords

def split_page_to_images(pdf_path, page_question_coords, output_dir='split_questions'):
    os.makedirs(output_dir, exist_ok=True)
    
    with pdfplumber.open(pdf_path) as pdf:
        for page_num, data in page_question_coords.items():
            page = pdf.pages[page_num-1]
            page_width, page_height = data['page_size']
            coords = data['coords']
            if not coords:
                continue
            
            # 将PDF页面转换为PIL图像(图像坐标系原点在左上角)
            img = page.to_image().original
            img_width, img_height = img.size
            
            # 生成切割区间(转换为图像坐标系)
            split_intervals = []
            # 1. 页面顶部到第一个题号的顶部
            first_top_img_y = page_height - coords[0]['top_y']
            split_intervals.append((0, first_top_img_y, f'page{page_num}_top_to_{coords[0]["num"]}'))
            
            # 2. 相邻题号之间的区间
            for i in range(len(coords)-1):
                current_bottom_img_y = page_height - coords[i]['bottom_y']
                next_top_img_y = page_height - coords[i+1]['top_y']
                split_intervals.append((current_bottom_img_y, next_top_img_y, f'page{page_num}_{coords[i]["num"]}_to_{coords[i+1]["num"]}'))
            
            # 3. 最后一个题号的底部到页面底部
            last_bottom_img_y = page_height - coords[-1]['bottom_y']
            split_intervals.append((last_bottom_img_y, img_height, f'page{page_num}_{coords[-1]["num"]}_to_bottom'))
            
            # 执行切割并保存
            for start_y, end_y, filename in split_intervals:
                crop_box = (0, start_y, img_width, end_y)
                cropped_img = img.crop(crop_box)
                cropped_img.save(os.path.join(output_dir, f'{filename}.png'))
                print(f"已保存:{filename}.png")

# 主执行流程
if __name__ == '__main__':
    pdf_path = 'document.pdf'
    coords_data = extract_question_coords(pdf_path)
    split_page_to_images(pdf_path, coords_data)

关键说明

  1. 坐标转换:PDF的原点在左下角,而PIL图像原点在左上角,必须用page_height - pdf_y转换y坐标,保证切割位置精准。
  2. 题号排序:提取的题号可能因OCR识别顺序混乱,必须按顶部y坐标排序,确保切割区间从上到下符合题目顺序。
  3. 边界处理:用题号的底部y作为上一块的结束,下一题号的顶部y作为下一块的开始,避免题号被切割到两个块中。

优化建议

  • 如果OCR识别有误差,可增加对bbox尺寸的过滤(比如只保留宽度在20-50之间的块),排除误识别的无效字符。
  • 若需要输出PDF而非图像,可改用pdfplumber的crop方法提取页面区域,再拼接保存为新PDF文件。

内容的提问来源于stack exchange,提问作者sibi kanagaraj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 01:20:46