You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python移除PDF中原理图的冗余文字?附代码

移除PDF原理图中无用文字的解决方案

针对PDF提取文本时混入原理图无用文字的问题,以下是几个实用的解决思路和代码实现:

方法一:基于文本位置与图形区域的关联过滤

核心思路是先识别出原理图的图形区域,直接在文本提取阶段过滤掉落在这些区域内的文本块,比遮盖图形的方式更高效。

import fitz

def extract_clean_text(pdf_path, target_page=2):
    doc = fitz.open(pdf_path)
    page = doc[target_page]
    
    # 提取所有图形的包围盒(扩展10像素避免漏判)
    graphic_rects = []
    paths = page.get_drawings()
    for path in paths:
        for item in path["items"]:
            if item[0] == "l":  # 线条
                x1, y1 = item[1]
                x2, y2 = item[2]
                rect = fitz.Rect(min(x1,x2)-10, min(y1,y2)-10, max(x1,x2)+10, max(y1,y2)+10)
                graphic_rects.append(rect)
            elif item[0] == "re":  # 矩形
                rect = item[1]
                graphic_rects.append(rect + fitz.Rect(-10, -10, 10, 10))  # 扩展区域
    
    # 提取文本块,过滤图形区域内的文本
    clean_text = []
    text_blocks = page.get_text("blocks")  # 返回格式:(x0,y0,x1,y1,text,block_no,line_no)
    for block in text_blocks:
        text_rect = fitz.Rect(block[0], block[1], block[2], block[3])
        # 若文本块与图形区域重叠超过50%,判定为无用文字
        overlap = False
        for g_rect in graphic_rects:
            intersection = text_rect & g_rect
            if intersection.get_area() / text_rect.get_area() > 0.5:
                overlap = True
                break
        if not overlap:
            clean_text.append(block[4])
    
    doc.close()
    return "\n".join(clean_text)

# 使用示例
print(extract_clean_text("your_pdf_file.pdf"))

方法二:基于文本特征的规则过滤

原理图中的文字通常有明显特征:字号极小、内容无完整语义、多为孤立字符/缩写,可通过这些规则过滤:

import fitz

def filter_by_text_features(pdf_path, target_page=2, min_font_size=8, min_text_length=3):
    doc = fitz.open(pdf_path)
    page = doc[target_page]
    
    clean_text = []
    # 按行分组文本,避免拆分完整句子
    line_groups = {}
    text_spans = page.get_text("words")  # 返回格式:(x0,y0,x1,y1,text,block_no,line_no,word_no,fontname,fontsize)
    for span in text_spans:
        line_key = (span[5], span[6])  # 用block+line编号作为行标识
        if line_key not in line_groups:
            line_groups[line_key] = []
        line_groups[line_key].append(span)
    
    for line in line_groups.values():
        avg_font_size = sum(s[9] for s in line) / len(line)
        full_line_text = " ".join(s[4] for s in line).strip()
        # 过滤小字号、短文本、全大写无意义内容
        if avg_font_size >= min_font_size and len(full_line_text) >= min_text_length:
            if not full_line_text.isupper() or any(c.islower() for c in full_line_text):
                clean_text.append(full_line_text)
    
    doc.close()
    return "\n".join(clean_text)

# 使用示例
print(filter_by_text_features("your_pdf_file.pdf"))

方法三:混合位置与特征的双重过滤

结合两种方法的优势,进一步提升过滤准确率:

import fitz

def hybrid_clean_text(pdf_path, target_page=2):
    doc = fitz.open(pdf_path)
    page = doc[target_page]
    
    # 步骤1:获取图形区域
    graphic_rects = []
    paths = page.get_drawings()
    for path in paths:
        for item in path["items"]:
            if item[0] == "l":
                x1, y1 = item[1]
                x2, y2 = item[2]
                rect = fitz.Rect(min(x1,x2)-10, min(y1,y2)-10, max(x1,x2)+10, max(y1,y2)+10)
                graphic_rects.append(rect)
            elif item[0] == "re":
                rect = item[1]
                graphic_rects.append(rect + fitz.Rect(-10, -10, 10, 10))
    
    # 步骤2:按行分组文本
    line_groups = {}
    text_spans = page.get_text("words")
    for span in text_spans:
        line_key = (span[5], span[6])
        if line_key not in line_groups:
            line_groups[line_key] = []
        line_groups[line_key].append(span)
    
    # 步骤3:双重过滤
    clean_text = []
    for line in line_groups.values():
        line_rect = fitz.Rect(min(s[0] for s in line), min(s[1] for s in line), max(s[2] for s in line), max(s[3] for s in line))
        # 位置过滤:重叠超过30%则跳过
        in_graphic = False
        for g_rect in graphic_rects:
            if (line_rect & g_rect).get_area() / line_rect.get_area() > 0.3:
                in_graphic = True
                break
        if in_graphic:
            continue
        
        # 特征过滤
        avg_font_size = sum(s[9] for s in line) / len(line)
        full_line_text = " ".join(s[4] for s in line).strip()
        if avg_font_size >= 8 and len(full_line_text) >= 3 and not full_line_text.isupper():
            clean_text.append(full_line_text)
    
    doc.close()
    return "\n".join(clean_text)

# 使用示例
print(hybrid_clean_text("your_pdf_file.pdf"))

额外提示

  • 如果是扫描型PDF(图像转PDF),需先用OCR工具(如pytesseract)提取文本,再结合上述规则过滤。
  • 可根据实际PDF的特征调整阈值(比如扩展像素大小、字体大小阈值、重叠比例等)。

内容的提问来源于stack exchange,提问作者Muhammad Samadzade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 05:45:09