如何使用PDFBOX加载URL末尾无.pdf后缀的网页PDF文件?
处理非.pdf后缀的网页PDF文件读取问题
你的原代码存在核心问题:用File类加载HTTP URL是错误的,File仅能处理本地文件系统路径,无法直接读取网络资源——不管这个URL是否带.pdf后缀。这类动态URL(比如.html结尾但返回PDF内容)的本质是服务器根据请求参数生成PDF并返回二进制流,只需通过HTTP请求获取这个流,再用PDFBOX加载即可。
解决方案步骤
- 发起HTTP请求到目标URL,获取服务器返回的PDF输入流
- 使用PDFBOX的
PDDocument.load(InputStream)方法加载流(而非File) - 完成文本读取后,记得关闭文档和流资源
示例代码
import org.apache.pdfbox.pdmodel.PDDocument; import org.apache.pdfbox.text.PDFTextStripper; import java.io.InputStream; import java.net.HttpURLConnection; import java.net.URL; public class DynamicPdfReader { public static void main(String[] args) { String targetUrl = "http://<website name>/flow.html?operatorin"; PDDocument pdfDoc = null; InputStream inputStream = null; try { URL url = new URL(targetUrl); HttpURLConnection conn = (HttpURLConnection) url.openConnection(); conn.setRequestMethod("GET"); // 检查请求是否成功 if (conn.getResponseCode() == HttpURLConnection.HTTP_OK) { inputStream = conn.getInputStream(); // 从输入流加载PDF文档 pdfDoc = PDDocument.load(inputStream); PDFTextStripper textStripper = new PDFTextStripper(); String content = textStripper.getText(pdfDoc); System.out.println(content); } else { System.err.println("请求失败,响应码:" + conn.getResponseCode()); } } catch (Exception e) { e.printStackTrace(); } finally { // 关闭资源,避免内存泄漏 try { if (inputStream != null) inputStream.close(); if (pdfDoc != null) pdfDoc.close(); } catch (Exception e) { e.printStackTrace(); } } } }
注意事项
- 若服务器需要特定请求头(如
User-Agent、Cookie),可通过conn.setRequestProperty("HeaderName", "Value")添加,否则可能被拒绝请求 - 也可使用Apache HttpClient等第三方HTTP库替代
HttpURLConnection,核心逻辑一致:获取PDF输入流后加载到PDDocument - 务必确保服务器返回的内容确实是PDF格式,若返回HTML或其他内容,PDFBOX加载时会抛出异常
内容的提问来源于stack exchange,提问作者prasanth kotagiri
相关产品推荐
相关产品推荐

