PyMuPDF与ReportLab坐标不兼容,绘制文本块位置异常求助
问题:ReportLab绘制PyMuPDF识别文本块时坐标偏移
我尝试用ReportLab绘制PyMuPDF(fitz)识别出的文本块,但生成的PDF中文本块位置完全不对,以下是我的代码:
doc = fitz.open("demo.pdf") canvas = Canvas("demo_.pdf", bottomup = True) def draw_auto_fit_text_block(canvas, x_1, y_1, text_block_width, text_block_height, font_name, font_size, text_content): text_block_frame = Frame(x_1, y_1, text_block_width, text_block_height, topPadding = 0, leftPadding = 0, rightPadding = 0, bottomPadding = 0, showBoundary = 1) text_block_styles = ParagraphStyle(name = "Normal", fontName = font_name, fontSize = font_size) text_block_content = text_content.replace('\n','<br />\n') text_block_story = [Paragraph(text_block_content, style = text_block_styles)] text_block_story_inframe = KeepInFrame(text_block_width, text_block_height, text_block_story) text_block_frame.addFromList([text_block_story_inframe], canvas) for page in doc: page_width = page.rect.width page_height = page.rect.height print("[page width]", page_width) print("[page height]", page_height) canvas.setPageSize((page_width, page_height)) blocks = page.get_text("blocks") for block in blocks: block_content = block[4].replace("\n", " ").replace("- ", "-").strip() block_x_0 = block[0] block_y_0 = block[1] block_x_1 = block[2] block_y_1 = block[3] block_width = block_x_1 - block_x_0 block_height = block_y_1 - block_y_0 block_y_0 = page_height - block_y_0 block_y_1 = page_height - block_y_0 draw_auto_fit_text_block(canvas, block_x_0, block_y_0, block_width, block_height, font_name = "NimbusRomNo9L-Regu", font_size = 9.0, text_content = block_content) canvas.showPage() canvas.save()
效果对比
- 生成的PDF效果:

- 原PDF效果:

解决方法
1. 修复坐标转换逻辑
PyMuPDF和ReportLab的坐标系原点方向不同:
- PyMuPDF:左上角为原点(0,0),y轴向下递增
- ReportLab(
bottomup=True):左下角为原点(0,0),y轴向上递增
原代码中坐标转换完全错误,正确的转换应该是:
- 文本块在ReportLab中的左下角y坐标 = 页面高度 - PyMuPDF中块的底部y坐标(
block_y_1) - 块的高度保持
block_y_1 - block_y_0不变
2. 保留文本换行格式
原代码把文本中的换行符替换成空格,导致段落结构丢失,应该保留换行并转换为ReportLab支持的<br/>标签。
3. 修正后的完整代码
from reportlab.pdfgen import Canvas from reportlab.lib.styles import ParagraphStyle from reportlab.platypus import Paragraph, Frame, KeepInFrame import fitz doc = fitz.open("demo.pdf") canvas = Canvas("demo_.pdf", bottomup=True) def draw_auto_fit_text_block(canvas, x, y, width, height, font_name, font_size, text_content): # 创建文本框,边界用于调试(可删除showBoundary=1) text_frame = Frame(x, y, width, height, topPadding=0, leftPadding=0, rightPadding=0, bottomPadding=0, showBoundary=1) style = ParagraphStyle(name="Normal", fontName=font_name, fontSize=font_size) # 把换行转成ReportLab支持的换行标签 formatted_text = text_content.replace('\n', '<br />\n') paragraph = Paragraph(formatted_text, style) # 用KeepInFrame确保文本适应框大小,mode=shrink会自动缩小字体 story = KeepInFrame(width, height, [paragraph], mode="shrink") text_frame.addFromList([story], canvas) for page in doc: page_width = page.rect.width page_height = page.rect.height canvas.setPageSize((page_width, page_height)) blocks = page.get_text("blocks") for block in blocks: block_content = block[4].replace("- ", "-").strip() # 保留换行,只处理连字符 x0, y0, x1, y1 = block[:4] block_width = x1 - x0 block_height = y1 - y0 # 正确转换坐标:ReportLab的左下角y = 页面高度 - 原块的底部y reportlab_y = page_height - y1 draw_auto_fit_text_block(canvas, x0, reportlab_y, block_width, block_height, font_name="NimbusRomNo9L-Regu", font_size=9.0, text_content=block_content) canvas.showPage() canvas.save()
额外说明
- 如果字体显示异常,确保ReportLab能找到指定字体,可通过
reportlab.pdfbase.pdfmetrics.registerFont注册自定义字体 showBoundary=1会显示文本块的边框,调试完成后可以删除该参数隐藏边框
内容的提问来源于stack exchange,提问作者CAO RUI
相关产品推荐
相关产品推荐

