You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Android Studio Java中用Apache PDFBox提取PDF文本遇URI转File报错

问题解决:Android中使用Apache PDFBox从内容URI提取PDF文本

问题原因

你通过ACTION_OPEN_DOCUMENT获取的Uri是内容URI(格式通常为content://开头),这类Uri并非指向本地文件系统的真实路径,直接调用uri.getPath()得到的字符串无法对应到实际文件,因此创建File对象后会抛出FileNotFoundException。

解决方案

无需将Uri转换为File,直接通过ContentResolver获取文件输入流,传递给PDFBox的PDDocument.load()方法即可。修改后的readTextFromPDF方法如下:

private void readTextFromPDF(Uri uri) throws IOException {
    // 直接通过ContentResolver获取输入流
    try (InputStream inputStream = getContentResolver().openInputStream(uri);
         PDDocument document = PDDocument.load(inputStream)) {
        
        PDFTextStripper pdfTextStripper = new PDFTextStripper();
        String text = pdfTextStripper.getText(document);

        TextView textView = findViewById(R.id.test);
        textView.setText(text);
    }
}

关键修改点说明

  • 使用getContentResolver().openInputStream(uri)获取内容URI对应的输入流,这是Android访问内容提供者中文件的标准方式。
  • 采用try-with-resources语法自动关闭InputStream和PDDocument,避免手动关闭时遗漏导致的资源泄漏。
  • 移除原代码中创建File和FileInputStream的逻辑,彻底规避内容URI转文件路径的问题。

额外注意事项

  1. PDFBox依赖配置:确保在模块的build.gradle(Module级别)中添加PDFBox依赖:
dependencies {
    implementation 'org.apache.pdfbox:pdfbox:2.0.34' // 可替换为最新稳定版本
}
  1. 内存优化:PDFBox处理PDF时可能占用较多内存,建议在AndroidManifest.xml的<application>标签中添加:
<application
    ...
    android:largeHeap="true">
    ...
</application>

内容的提问来源于stack exchange,提问作者Bastian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 13:43:36