使用pdfplumber/pdfminer提取中文财报PDF时粗体文本重复问题求助
中文财报PDF提取粗体文本重复问题解决
问题背景
目标是提取中文财报PDF中的文本并保存为TXT文件,使用pdfplumber或pdfminer工具时,发现PDF中的粗体文本提取后出现重复——比如PDF里的一、公司概况,提取后变成了「一、公司概况一、公司概况」。
问题原因
这类财报PDF的粗体文本在底层结构中,通常是同一文本内容被渲染了两次:一次是常规文本,一次是粗体样式的叠加文本块。工具提取时会将两个独立的文本块都读取出来,最终导致重复输出。
解决方案:无需更换工具,新增去重处理逻辑
不需要更换工具,通过对提取的文本块做位置+内容的去重处理即可解决,以下是针对pdfplumber的优化实现:
优化后的pdfplumber代码
import pdfplumber from collections import defaultdict def pdf2txt_without_duplicate(filename, delLinebreaker=True): pageContent = '' try: with pdfplumber.open(filename) as pdf: for page in pdf.pages: # 提取带位置坐标的文本块,调整容差适配财报排版 words = page.extract_words(x_tolerance=3, y_tolerance=3) # 按行(y坐标分组)整理文本块 line_dict = defaultdict(list) for word in words: # 用y坐标整数部分作为分组key,视为同一行 y_key = int(word['top']) line_dict[y_key].append(word) # 对每行文本块去重:内容相同且位置重叠的只保留一次 for line in line_dict.values(): processed_line = [] for word in line: is_duplicate = False # 对比已处理文本块,内容+位置高度匹配则标记为重复 for processed in processed_line: if (word['text'] == processed['text'] and abs(word['x0'] - processed['x0']) < 5 and abs(word['x1'] - processed['x1']) < 5): is_duplicate = True break if not is_duplicate: processed_line.append(word) # 拼接该行文本 line_text = ''.join([w['text'] for w in processed_line]) pageContent += line_text + ('\n' if not delLinebreaker else '') except Exception as e: print(f"文件: {filename}, 错误原因: {repr(e)}") return pageContent # 调用并保存结果 clean_text = pdf2txt_without_duplicate(r"report.pdf", delLinebreaker=False) with open("report_clean.txt", 'w', encoding='utf-8') as f: f.write(clean_text)
代码说明
- 用
extract_words()提取带位置坐标的文本块,通过x_tolerance和y_tolerance适配财报的排版误差 - 按行分组后,对同一行内内容相同、位置高度重叠的文本块去重,过滤粗体叠加产生的重复内容
- 保留原换行格式或按需合并换行,适配不同需求
针对pdfminer的补充方案
如果坚持使用pdfminer,需要自定义TextConverter的逻辑,在接收文本时记录位置信息并过滤重复块,但实现复杂度较高,优先推荐上述pdfplumber的优化方案。
内容的提问来源于stack exchange,提问作者user19560886
相关产品推荐
相关产品推荐

