You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyMuPDF删除PDF空行失败,求完整解决方案(附尝试代码)

解决PyMuPDF删除PDF连续空行的问题

问题分析

你的代码存在两个核心问题:

  1. Redact用法错误:使用cross_out=True仅会给空行添加删除线,并非真正删除内容;同时apply_redactions的text=0参数会保留原文本,导致标记后无实际删除效果。
  2. 空行判断逻辑不完善:仅检查单个span是否为空,但一行可能包含多个span,需判断整行文本是否为空。

方案一:修正Redact删除逻辑

通过调整Redact注释的使用方式,精准标记并删除多余空行:

import fitz  # PyMuPDF

def remove_extra_lines_spaces(doc):
    for page in doc:
        blocks = page.get_text("dict", sort=True)["blocks"]
        empty_line_count = 0
        to_redact = []
        
        for block in blocks:
            if block["type"] != 0:  # 仅处理文本块(type=0为文本)
                continue
            for line in block["lines"]:
                # 拼接该行所有span文本,判断是否为空行
                line_text = "".join([span["text"] for span in line["spans"]]).strip()
                if not line_text:
                    empty_line_count += 1
                    # 仅删除第2个及以后的连续空行
                    if empty_line_count >= 2:
                        to_redact.append(line["bbox"])
                else:
                    empty_line_count = 0
        
        # 批量添加删除注释(无需cross_out,默认删除)
        for bbox in to_redact:
            page.add_redact_annot(bbox)
        
        # 应用删除:text=1删除文本,REDACT_FLAGS_CLEAN清理空白区域
        page.apply_redactions(
            images=2, 
            graphics=2, 
            text=1, 
            flags=page.REDACT_FLAGS_CLEAN
        )
    return doc

# 使用示例
doc = fitz.open("your_file.pdf")
doc = remove_extra_lines_spaces(doc)
doc.save("cleaned_file.pdf")
doc.close()

方案二:重新生成页面内容(更可靠)

若PDF中空行是段落间距而非文本行,Redact方法可能无效,可通过提取有效文本重新排版:

import fitz

def remove_extra_lines_spaces(doc):
    for page in doc:
        lines = []
        blocks = page.get_text("dict", sort=True)["blocks"]
        last_empty = False
        
        for block in blocks:
            if block["type"] != 0:
                continue
            for line in block["lines"]:
                line_text = "".join([span["text"] for span in line["spans"]]).strip()
                if line_text:
                    # 记录文本、位置、排版参数
                    lines.append({
                        "text": line_text,
                        "x0": line["bbox"][0],
                        "fontsize": line["spans"][0]["size"],
                        "font": line["spans"][0]["font"],
                        "wmode": line["wmode"],
                        "dir": line["dir"]
                    })
                    last_empty = False
                else:
                    # 仅保留一个连续空行
                    if not last_empty:
                        lines.append({"text": ""})
                        last_empty = True
        
        # 清空页面原有内容
        page.clean_contents()
        
        # 重新写入文本,保持原排版风格
        if not lines:
            continue
        # 获取初始行高和位置
        first_line = next(item for item in lines if item["text"])
        y_pos = page.rect.y1 - 50  # 顶部偏移,可根据实际调整
        line_height = first_line["fontsize"] * 1.2  # 行高设为字体大小的1.2倍
        
        for line in lines:
            if line["text"]:
                page.insert_text(
                    (line["x0"], y_pos),
                    line["text"],
                    fontsize=line["fontsize"],
                    fontname=line["font"],
                    wmode=line["wmode"],
                    dir=line["dir"]
                )
            y_pos -= line_height
    
    return doc

# 使用示例
doc = fitz.open("your_file.pdf")
doc = remove_extra_lines_spaces(doc)
doc.save("cleaned_file.pdf")
doc.close()

注意事项

  • 处理前请备份原PDF,避免不可逆修改。
  • 方案二需根据PDF实际排版调整行高、顶部偏移等参数,确保还原原格式。
  • 部分PDF的空行可能由布局属性(如段落间距)导致,此时方案二的效果更稳定。

内容的提问来源于stack exchange,提问作者MURTUZA BORIWALA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 11:11:18