You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python 3.8提取PDF文本遇TypeError错误,如何解决?

解决PyPDF2提取PDF文本时的TypeError错误

错误原因

PyPDF2新版本中,pdf_reader.pages是**_VirtualList**类型的可迭代对象,并非整数,无法传入range()生成页码范围。同时旧版的getPage()方法已被弃用,需直接访问pages中的元素。

修正后的文本提取代码

import PyPDF2

# 打开PDF文件并读取文本
pdf_file = open("nita20.pdf", "rb")
pdf_reader = PyPDF2.PdfReader(pdf_file)
text = ""
# 直接遍历pages对象中的每一页
for page in pdf_reader.pages:
    text += page.extract_text()

# 记得关闭文件
pdf_file.close()

扩展:统计PDF中最常见词汇

提取文本后,可通过以下步骤统计高频词汇:

  • 预处理文本:转小写、移除标点符号
  • 分词并过滤停用词
  • 使用Counter统计词频

示例代码:

import string
from collections import Counter

# 文本预处理
processed_text = text.lower().translate(str.maketrans('', '', string.punctuation))
# 分词
words = processed_text.split()
# 自定义停用词列表(可根据需求补充)
stop_words = {"the", "and", "of", "a", "to", "in", "is", "it", "you", "that", "for", "on", "with"}
# 过滤停用词
filtered_words = [word for word in words if word not in stop_words]
# 统计词频并取前10个高频词
word_counts = Counter(filtered_words)
most_common_words = word_counts.most_common(10)
print("最常见的10个词汇:", most_common_words)

内容的提问来源于stack exchange,提问作者J.A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 12:40:29