使用Python 3.8提取PDF文本遇TypeError错误,如何解决?
解决PyPDF2提取PDF文本时的TypeError错误
错误原因
PyPDF2新版本中,pdf_reader.pages是**_VirtualList**类型的可迭代对象,并非整数,无法传入range()生成页码范围。同时旧版的getPage()方法已被弃用,需直接访问pages中的元素。
修正后的文本提取代码
import PyPDF2 # 打开PDF文件并读取文本 pdf_file = open("nita20.pdf", "rb") pdf_reader = PyPDF2.PdfReader(pdf_file) text = "" # 直接遍历pages对象中的每一页 for page in pdf_reader.pages: text += page.extract_text() # 记得关闭文件 pdf_file.close()
扩展:统计PDF中最常见词汇
提取文本后,可通过以下步骤统计高频词汇:
- 预处理文本:转小写、移除标点符号
- 分词并过滤停用词
- 使用
Counter统计词频
示例代码:
import string from collections import Counter # 文本预处理 processed_text = text.lower().translate(str.maketrans('', '', string.punctuation)) # 分词 words = processed_text.split() # 自定义停用词列表(可根据需求补充) stop_words = {"the", "and", "of", "a", "to", "in", "is", "it", "you", "that", "for", "on", "with"} # 过滤停用词 filtered_words = [word for word in words if word not in stop_words] # 统计词频并取前10个高频词 word_counts = Counter(filtered_words) most_common_words = word_counts.most_common(10) print("最常见的10个词汇:", most_common_words)
内容的提问来源于stack exchange,提问作者J.A
相关产品推荐
相关产品推荐

