使用pdfminer批量解析PDF时如何让循环遇错仍遍历全部文件
代码修改方案
报错原因
当前的TypeError: 'LTChar' object is not iterable是因为pdfminer解析部分PDF时,文本框(LTTextBox类)的子元素不一定都是LTTextLine文本行,可能直接返回LTChar字符对象,你没有做类型判断就直接遍历text_line,就会触发不可迭代的错误。
同时要实现坏文件不中断运行,只需要给单文件处理逻辑加全局异常捕获即可。
具体修改点
- 调整文本层级遍历逻辑,每一步都做类型判断,避免遍历非可迭代对象
- 给单文件处理逻辑外层加try-except块,捕获所有PDF解析、IO相关异常,出错时打印问题文件名后直接跳过
- 修复
wap变量可能未定义的问题,给异常分支加默认值
修改后完整代码
import pdfminer from pdfminer.pdfpage import PDFPage, PDFTextExtractionNotAllowed import os import io from io import StringIO from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter from pdfminer.converter import TextConverter, PDFPageAggregator from pdfminer.layout import LAParams, LTTextBox, LTFigure, LTImage, LTTextLine, LTTextContainer, LTChar, LTTextBoxHorizontal from pdfminer.pdfpage import PDFPage, PDFTextExtractionNotAllowed from pdfminer.pdfdocument import PDFDocument from pdfminer.pdfparser import PDFParser, PDFSyntaxError directory = 'C:/Users/' data = [] for file in os.listdir(directory): if not file.endswith(".pdf"): continue fake_file_handle = io.StringIO() # 给单个文件的所有处理逻辑加异常捕获 try: with open(os.path.join(directory, file), 'rb') as fh: resource_manager = PDFResourceManager() laparams = LAParams(line_margin = 0.6) device = PDFPageAggregator(resource_manager, laparams = laparams) page_interpreter = PDFPageInterpreter(resource_manager, device) positions = [] raw_text = [] for page in PDFPage.get_pages(fh, caching=True, check_extractable=True): page_interpreter.process_page(page) text = fake_file_handle.getvalue() layout = device.get_result() for lobj in layout: if isinstance(lobj, (LTTextContainer, LTTextBox, LTTextBoxHorizontal)): coord, word = int(lobj.bbox[1]), lobj.get_text().strip() raw_text.append([coord, word]) for text_line in lobj: position = None # 只有文本行类型才遍历子元素 if isinstance(text_line, LTTextLine): for character in text_line: if isinstance(character, LTChar): if character.matrix[0]>0 : position = character.bbox # 如果是直接返回的字符对象,直接判断处理 elif isinstance(text_line, LTChar): if text_line.matrix[0]>0: position = text_line.bbox if position is not None: positions.append(position) # if it's a container, recurse elif isinstance(lobj, LTFigure): pass # extract elements below y0=781 und above y0=57 text_pos = [] maxFontpos = 780 minFontpos = 58 for coord, word in raw_text: if coord <= maxFontpos and coord >= minFontpos: text_pos.append(word) # 给wap加默认值,避免未定义报错 try: wap = text_pos[0] except: wap = "" data.append([text_pos, wap]) # 捕获所有可能的异常,打印问题文件后跳过 except Exception as e: print(f"文件{file}处理失败,错误原因:{str(e)}") continue finally: fake_file_handle.close()
内容的提问来源于stack exchange,提问作者id345678
相关产品推荐
相关产品推荐

