按文件大小分割PDF为多份小文件:最后文件损坏问题求助
PDF分割代码错误排查与修正
代码中的核心错误点:
- 循环范围错误:原代码用
i < number_of_pages,会直接漏掉最后一页,导致生成的文件缺失内容甚至损坏。 - 关闭文档的逻辑完全颠倒:
combinedSize < 250 || i == number_of_pages的条件意味着只要文件小于250KB就关闭文档,这会导致刚添加少量页面就强制关闭,文件结构不完整。正确逻辑应该是当文件大小超过250KB时,关闭当前文档并新建下一个。 - 文件大小获取时机错误:
copy.getCurrentDocumentSize()在添加页面后立即调用,返回的是内存中未完全写入的临时大小,和最终磁盘上的文件大小不符,无法准确判断分割时机。 - 资源未正确释放:创建新的
PdfCopy和Document时,旧的实例和对应的输出流没有关闭,会导致资源泄漏,旧文件写入不完整。
修正后的代码示例:
import com.itextpdf.text.Document; import com.itextpdf.text.pdf.PdfCopy; import com.itextpdf.text.pdf.PdfImportedPage; import com.itextpdf.text.pdf.PdfReader; import java.io.FileOutputStream; public class PdfSplitterBySize { public static void main(String[] args) { try { PdfReader reader = new PdfReader("/tutorial.pdf"); int totalPages = reader.getNumberOfPages(); int currentFileNumber = 1; final long MAX_SIZE_KB = 250; final long MAX_SIZE_BYTES = MAX_SIZE_KB * 1024; Document document = new Document(); PdfCopy copy = new PdfCopy(document, new FileOutputStream("File" + currentFileNumber + ".pdf")); document.open(); for (int i = 1; i <= totalPages; i++) { // 添加当前页到文档 PdfImportedPage page = copy.getImportedPage(reader, i); copy.addPage(page); // 检查当前文档预估大小是否超过阈值(这里用已添加页面的字节数估算) // 注意:getCurrentDocumentSize()是内存中的大小,实际生成文件会略大,可适当调整阈值 if (copy.getCurrentDocumentSize() > MAX_SIZE_BYTES && i != totalPages) { // 关闭当前文档 document.close(); copy.close(); // 新建下一个文档 currentFileNumber++; document = new Document(); copy = new PdfCopy(document, new FileOutputStream("File" + currentFileNumber + ".pdf")); document.open(); } } // 关闭最后一个文档 document.close(); copy.close(); reader.close(); System.out.println("PDF按大小分割完成,共生成" + currentFileNumber + "个文档。"); } catch (Exception e) { e.printStackTrace(); } } }
额外说明:
getCurrentDocumentSize()只是内存中的预估大小,实际生成的PDF文件会因为压缩、交叉引用表等因素略大,可根据实际情况把阈值调小一点(比如设为240KB),避免生成的文件超过预期大小。- 所有IO资源(PdfReader、PdfCopy、FileOutputStream)都要确保关闭,避免资源泄漏和文件损坏。
内容的提问来源于stack exchange,提问作者Don
相关产品推荐
相关产品推荐

