如何用Python移除PDF中原理图的冗余文字?附代码
移除PDF原理图中无用文字的解决方案
针对PDF提取文本时混入原理图无用文字的问题,以下是几个实用的解决思路和代码实现:
方法一:基于文本位置与图形区域的关联过滤
核心思路是先识别出原理图的图形区域,直接在文本提取阶段过滤掉落在这些区域内的文本块,比遮盖图形的方式更高效。
import fitz def extract_clean_text(pdf_path, target_page=2): doc = fitz.open(pdf_path) page = doc[target_page] # 提取所有图形的包围盒(扩展10像素避免漏判) graphic_rects = [] paths = page.get_drawings() for path in paths: for item in path["items"]: if item[0] == "l": # 线条 x1, y1 = item[1] x2, y2 = item[2] rect = fitz.Rect(min(x1,x2)-10, min(y1,y2)-10, max(x1,x2)+10, max(y1,y2)+10) graphic_rects.append(rect) elif item[0] == "re": # 矩形 rect = item[1] graphic_rects.append(rect + fitz.Rect(-10, -10, 10, 10)) # 扩展区域 # 提取文本块,过滤图形区域内的文本 clean_text = [] text_blocks = page.get_text("blocks") # 返回格式:(x0,y0,x1,y1,text,block_no,line_no) for block in text_blocks: text_rect = fitz.Rect(block[0], block[1], block[2], block[3]) # 若文本块与图形区域重叠超过50%,判定为无用文字 overlap = False for g_rect in graphic_rects: intersection = text_rect & g_rect if intersection.get_area() / text_rect.get_area() > 0.5: overlap = True break if not overlap: clean_text.append(block[4]) doc.close() return "\n".join(clean_text) # 使用示例 print(extract_clean_text("your_pdf_file.pdf"))
方法二:基于文本特征的规则过滤
原理图中的文字通常有明显特征:字号极小、内容无完整语义、多为孤立字符/缩写,可通过这些规则过滤:
import fitz def filter_by_text_features(pdf_path, target_page=2, min_font_size=8, min_text_length=3): doc = fitz.open(pdf_path) page = doc[target_page] clean_text = [] # 按行分组文本,避免拆分完整句子 line_groups = {} text_spans = page.get_text("words") # 返回格式:(x0,y0,x1,y1,text,block_no,line_no,word_no,fontname,fontsize) for span in text_spans: line_key = (span[5], span[6]) # 用block+line编号作为行标识 if line_key not in line_groups: line_groups[line_key] = [] line_groups[line_key].append(span) for line in line_groups.values(): avg_font_size = sum(s[9] for s in line) / len(line) full_line_text = " ".join(s[4] for s in line).strip() # 过滤小字号、短文本、全大写无意义内容 if avg_font_size >= min_font_size and len(full_line_text) >= min_text_length: if not full_line_text.isupper() or any(c.islower() for c in full_line_text): clean_text.append(full_line_text) doc.close() return "\n".join(clean_text) # 使用示例 print(filter_by_text_features("your_pdf_file.pdf"))
方法三:混合位置与特征的双重过滤
结合两种方法的优势,进一步提升过滤准确率:
import fitz def hybrid_clean_text(pdf_path, target_page=2): doc = fitz.open(pdf_path) page = doc[target_page] # 步骤1:获取图形区域 graphic_rects = [] paths = page.get_drawings() for path in paths: for item in path["items"]: if item[0] == "l": x1, y1 = item[1] x2, y2 = item[2] rect = fitz.Rect(min(x1,x2)-10, min(y1,y2)-10, max(x1,x2)+10, max(y1,y2)+10) graphic_rects.append(rect) elif item[0] == "re": rect = item[1] graphic_rects.append(rect + fitz.Rect(-10, -10, 10, 10)) # 步骤2:按行分组文本 line_groups = {} text_spans = page.get_text("words") for span in text_spans: line_key = (span[5], span[6]) if line_key not in line_groups: line_groups[line_key] = [] line_groups[line_key].append(span) # 步骤3:双重过滤 clean_text = [] for line in line_groups.values(): line_rect = fitz.Rect(min(s[0] for s in line), min(s[1] for s in line), max(s[2] for s in line), max(s[3] for s in line)) # 位置过滤:重叠超过30%则跳过 in_graphic = False for g_rect in graphic_rects: if (line_rect & g_rect).get_area() / line_rect.get_area() > 0.3: in_graphic = True break if in_graphic: continue # 特征过滤 avg_font_size = sum(s[9] for s in line) / len(line) full_line_text = " ".join(s[4] for s in line).strip() if avg_font_size >= 8 and len(full_line_text) >= 3 and not full_line_text.isupper(): clean_text.append(full_line_text) doc.close() return "\n".join(clean_text) # 使用示例 print(hybrid_clean_text("your_pdf_file.pdf"))
额外提示
- 如果是扫描型PDF(图像转PDF),需先用OCR工具(如pytesseract)提取文本,再结合上述规则过滤。
- 可根据实际PDF的特征调整阈值(比如扩展像素大小、字体大小阈值、重叠比例等)。
内容的提问来源于stack exchange,提问作者Muhammad Samadzade
相关产品推荐
相关产品推荐

