如何使用PyPDF2提取PDF指定页码范围(如31-39页)的文本?
提取PDF指定页码/页码范围的文本(基于PyPDF2)
PyPDF2中pdfReader.pages的索引是从0开始的——也就是现实里的第1页对应代码中的索引0,第2页对应索引1,以此类推。你可以通过指定目标页码的索引,实现只提取选定内容,以下是两种常见场景的修改方案:
1. 提取单个或多个指定页码
比如你要提取第2页、第5页、第7页,先把这些页码转换成0-based索引(即1、4、6),再遍历这些索引提取文本:
import PyPDF2 # 用原始字符串r''避免路径转义问题 pdf_path = r'C:\sem1\691-project\Dataset\Maths\A Spiral Workbook for Discrete Mathematics.pdf' txt_path = r'C:\sem1\691-project\Dataset\Maths\A Spiral Workbook for Discrete Mathematics.txt' pdfFileObj = open(pdf_path, 'rb') pdfReader = PyPDF2.PdfReader(pdfFileObj) out_file = open(txt_path, 'a') # 指定要提取的页码(转换为0-based索引) target_page_indices = [1, 4, 6] # 对应现实中的第2、5、7页 for idx in target_page_indices: # 先检查索引是否在PDF的有效页码范围内 if 0 <= idx < len(pdfReader.pages): page_text = pdfReader.pages[idx].extract_text() print(page_text) out_file.write(page_text) else: print(f"页码索引{idx}超出PDF范围") out_file.close() pdfFileObj.close()
2. 提取连续的页码范围
比如你要提取第3页到第10页(包含首尾),先转换为0-based的起始索引2和结束索引9,再用range遍历这个范围:
import PyPDF2 pdf_path = r'C:\sem1\691-project\Dataset\Maths\A Spiral Workbook for Discrete Mathematics.pdf' txt_path = r'C:\sem1\691-project\Dataset\Maths\A Spiral Workbook for Discrete Mathematics.txt' pdfFileObj = open(pdf_path, 'rb') pdfReader = PyPDF2.PdfReader(pdfFileObj) out_file = open(txt_path, 'a') # 指定页码范围(转换为0-based索引,首尾都包含) start_idx = 2 # 对应现实中的第3页 end_idx = 9 # 对应现实中的第10页 # range是左闭右开,所以要给end_idx加1才能包含最后一页 for idx in range(start_idx, end_idx + 1): if 0 <= idx < len(pdfReader.pages): page_text = pdfReader.pages[idx].extract_text() print(page_text) out_file.write(page_text) else: print(f"页码索引{idx}超出PDF范围") out_file.close() pdfFileObj.close()
额外小提示
- 如果你习惯用现实中从1开始的页码,可以加个简单转换逻辑:
# 示例:将现实页码转为代码索引 real_page = 3 idx = real_page - 1 - 始终用原始字符串
r'路径'处理文件路径,避免像\s这类转义字符引发的错误。
内容的提问来源于stack exchange,提问作者Sindhu
相关产品推荐
相关产品推荐

