使用PyPDF2统计PDF词频时出现FloatObject无效警告如何解决?
解决PyPDF2中FloatObject无效提示的问题
你编写的PDF单词统计函数运行时出现的FloatObject (b'0.000000000000-14210855') invalid; use 0.0 instead提示,是因为PyPDF2在解析格式不规范的PDF文件时输出的警告日志——即使设置了strict=False,这类非致命警告依然会被打印。以下是两种消除提示的方法:
方法一:屏蔽PyPDF2的警告日志(推荐)
直接通过Python的logging模块将PyPDF2的日志级别调高到ERROR,过滤掉所有警告信息,只保留严重错误日志,完全不影响你的统计逻辑。
修改后的完整代码如下:
import os import re import PyPDF2 import logging # 把PyPDF2的日志级别设为ERROR,屏蔽警告输出 logging.getLogger("PyPDF2").setLevel(logging.ERROR) def count_words(word): print() print('Counting words..') files = os.listdir('./pdfs') counted_words = [] for idx, file in enumerate(files, 1): with open(f'./pdfs/{file}', 'rb') as pdf_file: ReadPDF = PyPDF2.PdfFileReader(pdf_file, strict=False) pages = ReadPDF.numPages words_count = 0 for page in range(pages): pageObj = ReadPDF.getPage(page) data = pageObj.extract_text() words_count += sum(1 for match in re.findall(rf'\b{word}\b', data, flags=re.I)) counted_words.append(words_count) print(f'File: {idx}') return counted_words
方法二:修复PDF文件格式(可选)
如果不想屏蔽日志,可以用PDF编辑工具(如Adobe Acrobat)重新保存有问题的PDF文件,修复其中的浮点数格式错误。但这种方法需要逐个处理PDF文件,效率较低,仅适合PDF数量较少的场景。
内容的提问来源于stack exchange,提问作者iamstudentsorry
相关产品推荐
相关产品推荐

