You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何裁剪PDF页面空白区域?含中间空白处理(PyMuPDF相关)

处理PDF页面中间空白区域的视觉检测方案

针对你用PyMuPDF只能裁剪外围空白的问题,可以结合图像视觉检测(OpenCV)+ PyMuPDF实现中间空白的识别与去除,核心思路是先把PDF页面转成图像,检测出所有包含内容的有效区域,再重新排版这些区域以消除中间空白。

具体实现步骤

1. 页面转图像

先将PDF每页转为像素图,转换成OpenCV可处理的格式:

import fitz
import cv2
import numpy as np

# 打开PDF文档
doc = fitz.open("your_large_pdf.pdf")

for page_idx in range(len(doc)):
    page = doc[page_idx]
    # 获取页面像素图
    pix = page.get_pixmap()
    # 转换为OpenCV灰度图(简化后续处理)
    img = np.frombuffer(pix.samples, dtype=np.uint8).reshape(pix.height, pix.width, pix.n)
    img = cv2.cvtColor(img, cv2.COLOR_RGBA2GRAY) if pix.n == 4 else cv2.COLOR_RGB2GRAY

2. 识别有效内容区域

通过二值化和轮廓检测,定位所有非空白的内容区域:

# 二值化:将空白(浅色)转为黑色,内容(深色)转为白色
_, binary_img = cv2.threshold(img, 240, 255, cv2.THRESH_BINARY_INV)
# 提取所有内容轮廓
contours, _ = cv2.findContours(binary_img, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)

# 过滤微小噪点轮廓(比如面积小于50的忽略)
valid_contours = [cnt for cnt in contours if cv2.contourArea(cnt) > 50]
if not valid_contours:
    continue  # 全空白页直接跳过

# 获取每个有效区域的边界矩形,并按垂直方向排序
content_rects = [cv2.boundingRect(cnt) for cnt in valid_contours]
content_rects.sort(key=lambda r: r[1])  # 按顶部y坐标排序

3. 分析中间空白并重新排版

计算相邻内容区域的间距,当间距超过设定阈值(比如图像高度的5%)时,判定为需要去除的中间空白,然后将内容区域紧凑拼接:

# 图像与PDF坐标转换工具函数
def img_rect_to_pdf(img_rect, page, pix):
    """将图像坐标的矩形转换为PDF页面坐标(PDF原点在左下角)"""
    page_w, page_h = page.rect.width, page.rect.height
    img_w, img_h = pix.width, pix.height
    x0 = img_rect[0] * page_w / img_w
    y0 = page_h - (img_rect[1] + img_rect[3]) * page_h / img_h
    x1 = (img_rect[0] + img_rect[2]) * page_w / img_w
    y1 = page_h - img_rect[1] * page_h / img_h
    return fitz.Rect(x0, y0, x1, y1)

# 计算紧凑排版后的总高度
total_height = 0
pdf_rects = []
prev_bottom_y = 0  # 上一个内容区域的底部PDF坐标
gap_threshold = pix.height * 0.05  # 空白阈值:超过图像高度5%视为需要去除的空白

for rect in content_rects:
    pdf_rect = img_rect_to_pdf(rect, page, pix)
    # 计算当前区域与上一个区域的间距
    gap = prev_bottom_y - pdf_rect.y1 if prev_bottom_y != 0 else 0
    # 如果间距过大,直接紧凑拼接(忽略空白)
    if gap > gap_threshold and prev_bottom_y != 0:
        adjusted_y0 = prev_bottom_y - pdf_rect.height
        adjusted_y1 = prev_bottom_y
        pdf_rect = fitz.Rect(pdf_rect.x0, adjusted_y0, pdf_rect.x1, adjusted_y1)
    
    pdf_rects.append(pdf_rect)
    total_height = max(total_height, pdf_rect.y1)
    prev_bottom_y = pdf_rect.y0

# 创建新页面,高度为紧凑排版后的总高度
new_page = doc.new_page(width=page.rect.width, height=total_height)
# 将内容区域复制到新页面
for pdf_rect in pdf_rects:
    # 计算在新页面中的位置(从顶部往下排)
    new_y1 = total_height - (pdf_rect.y1 - page.rect.y0)
    new_y0 = new_y1 - pdf_rect.height
    new_rect = fitz.Rect(pdf_rect.x0, new_y0, pdf_rect.x1, new_y1)
    new_page.show_pdf_page(new_rect, doc, page_idx, clip=pdf_rect)

# 删除原页面,替换为紧凑排版后的新页面
doc.delete_page(page_idx)

# 保存处理后的PDF
doc.save("processed_pdf.pdf")
doc.close()

关键细节说明

  • 阈值调整:二值化的阈值(240)和空白阈值(5%)需要根据你的PDF实际情况微调,确保能准确区分空白和内容。
  • 复杂布局适配:如果是多栏布局,需要先按水平坐标分组轮廓,再对每栏单独处理空白。
  • 性能优化:处理大型PDF时,可分批次处理页面,避免内存溢出;也可以跳过全空白页减少计算量。

内容的提问来源于stack exchange,提问作者vampirekabir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 10:55:55