You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfplumber/pdfminer提取中文财报PDF时粗体文本重复问题求助

中文财报PDF提取粗体文本重复问题解决

问题背景

目标是提取中文财报PDF中的文本并保存为TXT文件,使用pdfplumber或pdfminer工具时,发现PDF中的粗体文本提取后出现重复——比如PDF里的一、公司概况,提取后变成了「一、公司概况一、公司概况」。

问题原因

这类财报PDF的粗体文本在底层结构中,通常是同一文本内容被渲染了两次:一次是常规文本,一次是粗体样式的叠加文本块。工具提取时会将两个独立的文本块都读取出来,最终导致重复输出。

解决方案:无需更换工具,新增去重处理逻辑

不需要更换工具,通过对提取的文本块做位置+内容的去重处理即可解决,以下是针对pdfplumber的优化实现:

优化后的pdfplumber代码

import pdfplumber
from collections import defaultdict

def pdf2txt_without_duplicate(filename, delLinebreaker=True):
    pageContent = ''
    try:
        with pdfplumber.open(filename) as pdf:
            for page in pdf.pages:
                # 提取带位置坐标的文本块,调整容差适配财报排版
                words = page.extract_words(x_tolerance=3, y_tolerance=3)
                # 按行(y坐标分组)整理文本块
                line_dict = defaultdict(list)
                for word in words:
                    # 用y坐标整数部分作为分组key,视为同一行
                    y_key = int(word['top'])
                    line_dict[y_key].append(word)
                
                # 对每行文本块去重:内容相同且位置重叠的只保留一次
                for line in line_dict.values():
                    processed_line = []
                    for word in line:
                        is_duplicate = False
                        # 对比已处理文本块,内容+位置高度匹配则标记为重复
                        for processed in processed_line:
                            if (word['text'] == processed['text'] and 
                                abs(word['x0'] - processed['x0']) < 5 and 
                                abs(word['x1'] - processed['x1']) < 5):
                                is_duplicate = True
                                break
                        if not is_duplicate:
                            processed_line.append(word)
                    # 拼接该行文本
                    line_text = ''.join([w['text'] for w in processed_line])
                    pageContent += line_text + ('\n' if not delLinebreaker else '')
    except Exception as e:
        print(f"文件: {filename}, 错误原因: {repr(e)}")
    return pageContent

# 调用并保存结果
clean_text = pdf2txt_without_duplicate(r"report.pdf", delLinebreaker=False)
with open("report_clean.txt", 'w', encoding='utf-8') as f:
    f.write(clean_text)

代码说明

  • 用extract_words()提取带位置坐标的文本块,通过x_tolerance和y_tolerance适配财报的排版误差
  • 按行分组后,对同一行内内容相同、位置高度重叠的文本块去重,过滤粗体叠加产生的重复内容
  • 保留原换行格式或按需合并换行,适配不同需求

针对pdfminer的补充方案

如果坚持使用pdfminer,需要自定义TextConverter的逻辑,在接收文本时记录位置信息并过滤重复块,但实现复杂度较高,优先推荐上述pdfplumber的优化方案。


内容的提问来源于stack exchange,提问作者user19560886

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 02:17:26