You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

点击按钮后用PyPDF2提取PDF全部页面并插入文本编辑器的方法

解决PyPDF2提取PDF全部页面并插入Tkinter文本框的问题

代码中的问题及修正说明

  • 循环范围错误:原代码用空列表count作为range()的参数,这会导致循环无法执行。应替换为PDF的页面总数,比如range(all_pages)或直接range(len(reader.pages))
  • 文本插入方式错误:count是存储各页面文本的列表,不能用count()的方式调用,需要将列表中的文本拼接成字符串后再插入文本框,推荐用'\n'.join(count)来分隔各页面内容
  • 冗余代码:readRoundup函数定义后未被使用,可直接删除
  • 可选优化:部分PDF页面可能无文本内容,extract_text()会返回None,可以添加判断过滤掉这类值,避免文本框出现无效内容

修正后的完整代码

from tkinter import *
from tkinter import filedialog
from PyPDF2 import PdfReader

root = Tk()
root.title('PDF Extraction')

def openRoundup():
    file_name = filedialog.askopenfilename(
        initialdir='/', title="select a file", filetypes=(("PDF Files", ".pdf"), ("Txt Files", ".txt")))
    
    reader = PdfReader(file_name)
    all_pages = len(reader.pages)
    print(all_pages)
    
    page_texts = []
    for i in range(all_pages):
        page = reader.pages[i]
        text = page.extract_text()
        # 过滤无文本的页面
        if text:
            page_texts.append(text)
    
    print(page_texts)
    # 用换行符加分割线拼接各页面文本,插入到文本框
    text_editor.insert(END, '\n\n--- 页面分割线 --- \n\n'.join(page_texts))

text_editor = Text(root, width=40, height=25)
text_editor.pack()

button1 = Button(root, text='Upload', command=openRoundup)
button1.pack()

root.geometry('800x600')
root.mainloop()

内容的提问来源于stack exchange,提问作者Steve

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 14:54:52