点击按钮后用PyPDF2提取PDF全部页面并插入文本编辑器的方法
解决PyPDF2提取PDF全部页面并插入Tkinter文本框的问题
代码中的问题及修正说明
- 循环范围错误:原代码用空列表
count作为range()的参数,这会导致循环无法执行。应替换为PDF的页面总数,比如range(all_pages)或直接range(len(reader.pages)) - 文本插入方式错误:
count是存储各页面文本的列表,不能用count()的方式调用,需要将列表中的文本拼接成字符串后再插入文本框,推荐用'\n'.join(count)来分隔各页面内容 - 冗余代码:
readRoundup函数定义后未被使用,可直接删除 - 可选优化:部分PDF页面可能无文本内容,
extract_text()会返回None,可以添加判断过滤掉这类值,避免文本框出现无效内容
修正后的完整代码
from tkinter import * from tkinter import filedialog from PyPDF2 import PdfReader root = Tk() root.title('PDF Extraction') def openRoundup(): file_name = filedialog.askopenfilename( initialdir='/', title="select a file", filetypes=(("PDF Files", ".pdf"), ("Txt Files", ".txt"))) reader = PdfReader(file_name) all_pages = len(reader.pages) print(all_pages) page_texts = [] for i in range(all_pages): page = reader.pages[i] text = page.extract_text() # 过滤无文本的页面 if text: page_texts.append(text) print(page_texts) # 用换行符加分割线拼接各页面文本,插入到文本框 text_editor.insert(END, '\n\n--- 页面分割线 --- \n\n'.join(page_texts)) text_editor = Text(root, width=40, height=25) text_editor.pack() button1 = Button(root, text='Upload', command=openRoundup) button1.pack() root.geometry('800x600') root.mainloop()
内容的提问来源于stack exchange,提问作者Steve
相关产品推荐
相关产品推荐

