基于坐标切割PDF图像:已获取题号坐标后的切割实现求助
解决方案:基于OCR坐标切割PDF页面为独立题目块
你的核心问题是当前用PyPDF2获取的坐标不准确——match.start()返回的是文本字符的索引,不是页面上的实际空间坐标,导致y坐标完全无法用来切割。下面是可落地的解决流程:
工具选择
放弃PyPDF2的文本位置提取,改用pdfplumber——它能精准获取每个文本块的边界框(bbox),包含真实的x/y坐标;配合Pillow处理图像切割。
先安装依赖:
pip install pdfplumber pillow
完整实现代码
import pdfplumber from PIL import Image import re import os # 定义题号匹配正则(匹配101.到200.的格式) QUESTION_NUM_PATTERN = re.compile(r'^(1\d{2}|200)\.$') def extract_question_coords(pdf_path): page_question_coords = {} with pdfplumber.open(pdf_path) as pdf: for page_num, page in enumerate(pdf.pages, start=1): page_width = page.width page_height = page.height coords = [] # 遍历页面所有文本字符块,提取题号的真实坐标 for char in page.chars: text = char['text'].strip() if QUESTION_NUM_PATTERN.match(text): bbox = char['bbox'] coords.append({ 'num': text, 'top_y': bbox[1], # PDF坐标系下的顶部y值 'bottom_y': bbox[3] # PDF坐标系下的底部y值 }) # 按顶部y坐标从上到下排序,确保题号顺序正确 coords.sort(key=lambda x: x['top_y']) page_question_coords[page_num] = { 'coords': coords, 'page_size': (page_width, page_height) } return page_question_coords def split_page_to_images(pdf_path, page_question_coords, output_dir='split_questions'): os.makedirs(output_dir, exist_ok=True) with pdfplumber.open(pdf_path) as pdf: for page_num, data in page_question_coords.items(): page = pdf.pages[page_num-1] page_width, page_height = data['page_size'] coords = data['coords'] if not coords: continue # 将PDF页面转换为PIL图像(图像坐标系原点在左上角) img = page.to_image().original img_width, img_height = img.size # 生成切割区间(转换为图像坐标系) split_intervals = [] # 1. 页面顶部到第一个题号的顶部 first_top_img_y = page_height - coords[0]['top_y'] split_intervals.append((0, first_top_img_y, f'page{page_num}_top_to_{coords[0]["num"]}')) # 2. 相邻题号之间的区间 for i in range(len(coords)-1): current_bottom_img_y = page_height - coords[i]['bottom_y'] next_top_img_y = page_height - coords[i+1]['top_y'] split_intervals.append((current_bottom_img_y, next_top_img_y, f'page{page_num}_{coords[i]["num"]}_to_{coords[i+1]["num"]}')) # 3. 最后一个题号的底部到页面底部 last_bottom_img_y = page_height - coords[-1]['bottom_y'] split_intervals.append((last_bottom_img_y, img_height, f'page{page_num}_{coords[-1]["num"]}_to_bottom')) # 执行切割并保存 for start_y, end_y, filename in split_intervals: crop_box = (0, start_y, img_width, end_y) cropped_img = img.crop(crop_box) cropped_img.save(os.path.join(output_dir, f'{filename}.png')) print(f"已保存:{filename}.png") # 主执行流程 if __name__ == '__main__': pdf_path = 'document.pdf' coords_data = extract_question_coords(pdf_path) split_page_to_images(pdf_path, coords_data)
关键说明
- 坐标转换:PDF的原点在左下角,而PIL图像原点在左上角,必须用
page_height - pdf_y转换y坐标,保证切割位置精准。 - 题号排序:提取的题号可能因OCR识别顺序混乱,必须按顶部y坐标排序,确保切割区间从上到下符合题目顺序。
- 边界处理:用题号的底部y作为上一块的结束,下一题号的顶部y作为下一块的开始,避免题号被切割到两个块中。
优化建议
- 如果OCR识别有误差,可增加对bbox尺寸的过滤(比如只保留宽度在20-50之间的块),排除误识别的无效字符。
- 若需要输出PDF而非图像,可改用
pdfplumber的crop方法提取页面区域,再拼接保存为新PDF文件。
内容的提问来源于stack exchange,提问作者sibi kanagaraj
相关产品推荐
相关产品推荐

