You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

代码陷入死循环及改用for循环实现PDF文本提取的求助

问题分析与解决方案

无限while循环的原因

你的while循环条件index<=len(pdfs)存在逻辑错误:列表的有效索引范围是0到len(pdfs)-1,当index等于len(pdfs)时,访问pdfs[index]会触发IndexError。若代码运行中未直接崩溃,后续index +=1会让索引持续增长,最终导致循环无法终止。正确的循环条件应为index < len(pdfs)。

修复后的while循环代码

import glob
import PyPDF2

pdfs = glob.glob("/private/babik/*.pdf")
index = 0
while index < len(pdfs):
    # 使用with语句自动管理文件资源,避免手动关闭遗漏
    with open(str(pdfs[index]), 'rb') as pdfFileObj:
        pdfReader = PyPDF2.PdfFileReader(pdfFileObj, strict=False)
        print(pdfReader.numPages)
        pageObj = pdfReader.getPage(0)
        print(pageObj.extractText())
    index += 1

改用for循环的实现

for循环可直接遍历PDF路径列表,无需手动维护索引,代码更简洁直观:

import glob
import PyPDF2

pdfs = glob.glob("/private/babik/*.pdf")
for pdf_path in pdfs:
    with open(pdf_path, 'rb') as pdfFileObj:
        pdfReader = PyPDF2.PdfFileReader(pdfFileObj, strict=False)
        print(pdfReader.numPages)
        pageObj = pdfReader.getPage(0)
        print(pageObj.extractText())

额外优化提示

  • 始终用with语句处理文件操作,确保资源自动释放
  • 若需提取所有页面文本,可嵌套循环遍历pdfReader.numPages
  • 可添加try-except块捕获PDF加密、损坏等异常,提升代码健壮性

内容的提问来源于stack exchange,提问作者Babiqowski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 20:55:23