You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python遍历目录统计文件词数,PDF解析报AttributeError错误求助

解决PdfFileReader的AttributeError问题

嘿,我来帮你搞定这个PDF解析的错误!你遇到的AttributeError: 'PdfFileReader' object has no attribute 'startswith',大概率是因为你没有正确初始化PdfFileReader实例,或者不小心把阅读器对象当成字符串/文件路径来使用了。

错误原因拆解

startswith()是字符串专属的方法,当你把一个PdfFileReader对象传给需要字符串的操作时(比如误把阅读器当成文件路径去判断扩展名,或者初始化阅读器时参数传错),就会抛出这个错误。举个常见的错误示例:

# 错误:直接把文件路径传给PdfFileReader,后续误操作该对象
pdf_reader = PdfFileReader("test.pdf")
# 如果此时执行类似pdf_reader.startswith(".pdf")的操作,必然报错

正确的PDF读取步骤

要正确解析PDF,你需要先以二进制模式打开文件,再把文件对象传给PdfFileReader:

from PyPDF2 import PdfFileReader

def count_pdf_words(pdf_path):
    total_words = 0
    # 以二进制模式打开PDF文件
    with open(pdf_path, "rb") as pdf_file:
        pdf_reader = PdfFileReader(pdf_file)
        # 遍历每一页提取文本并统计词数
        for page_num in range(pdf_reader.numPages):
            page = pdf_reader.getPage(page_num)
            text = page.extractText()
            # 这里按空格分割统计,你可以根据需求优化分词逻辑
            total_words += len(text.split())
    return total_words

完整的目录遍历+多文件统计示例

把PDF处理逻辑和你的目录遍历、文本/HTML文件统计结合起来,完整代码如下:

import os
from PyPDF2 import PdfFileReader

def count_file_words(file_path):
    ext = file_path.lower().split(".")[-1]
    total_words = 0
    if ext in ["txt", "html"]:
        with open(file_path, "r", encoding="utf-8", errors="ignore") as f:
            text = f.read()
            total_words += len(text.split())
    elif ext == "pdf":
        with open(file_path, "rb") as pdf_file:
            pdf_reader = PdfFileReader(pdf_file)
            for page_num in range(pdf_reader.numPages):
                page = pdf_reader.getPage(page_num)
                text = page.extractText()
                total_words += len(text.split())
    return total_words

def traverse_directory(root_dir):
    total_all = 0
    for dirpath, _, filenames in os.walk(root_dir):
        for filename in filenames:
            file_path = os.path.join(dirpath, filename)
            ext = file_path.lower().split(".")[-1]
            if ext in ["txt", "html", "pdf"]:
                word_count = count_file_words(file_path)
                total_all += word_count
                print(f"文件 {file_path} 词数:{word_count}")
    print(f"总词数:{total_all}")

# 替换成你的目标目录路径
traverse_directory("你的目标目录路径")

额外注意事项

  • 确保已安装PyPDF2库:如果没装,执行pip install PyPDF2即可
  • PDF提取文本可能存在格式问题(比如换行、特殊字符),如果需要更精准的词数统计,可以考虑用nltk等库做专业分词优化
  • 处理大体积PDF时,注意内存占用,避免一次性读取所有页面

内容的提问来源于stack exchange,提问作者MosaixSolutions

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:21:39