You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何突破Google Cloud Vision PDF文本检测的20页限制?

解决Google Cloud Vision asyncBatchAnnotateFiles仅返回前20页文本的问题

核心问题分析

Google Cloud Vision的asyncBatchAnnotateFiles处理多页PDF时,当输出结果体积较大,会自动将JSON结果拆分为多个分片文件(shard),默认每个分片包含20页内容。如果你的代码只读取了第一个分片文件,就会出现仅获取前20页的情况——这和API标称的2000页总限制并不冲突。

解决方案步骤

1. 检查输出文件结构

先去Firebase Storage里查看Vision生成的输出文件,通常命名格式为output-1-to-20.json、output-21-to-34.json,确认是否存在多个分片文件。

2. 修改代码遍历所有分片

更新Cloud Function代码,遍历Storage中属于本次任务的所有输出分片,而不是只读取单个文件。

3. 修正页码与文本提取逻辑

处理每个分片时,从文件名中解析出起始页码,结合分片内的页面索引计算实际页码,确保所有页面的文本都能正确存入Firestore。

代码修正示例

假设原代码仅读取单个分片,修改后的代码如下:

const admin = require('firebase-admin');
admin.initializeApp();

exports.processPdfOcr = async (event) => {
  const bucket = admin.storage().bucket();
  // 替换为你的输出文件前缀,比如任务生成的输出目录路径
  const outputPrefix = 'vision-output/your-task-id/output-';
  
  // 获取所有匹配前缀的分片文件
  const [files] = await bucket.getFiles({ prefix: outputPrefix });
  
  for (const file of files) {
    // 跳过非JSON文件
    if (!file.name.endsWith('.json')) continue;
    
    // 下载并解析分片文件内容
    const fileContents = await file.download();
    const ocrResult = JSON.parse(fileContents.toString());
    
    // 从文件名提取分片的起始页码
    const pageRangeMatch = file.name.match(/output-(\d+)-to-(\d+)\.json/);
    if (!pageRangeMatch) continue;
    const startPage = parseInt(pageRangeMatch[1], 10);
    
    // 处理当前分片中的所有页面
    await Promise.all(ocrResult.responses.map((pageResponse, index) => {
      const actualPageNumber = startPage + index;
      const pageText = pageResponse.fullTextAnnotation?.text || '';
      
      // 将文本存入Firestore,这里示例按页码作为文档ID
      return admin.firestore()
        .collection('pdf_pages')
        .doc(`pdf-${event.params.pdfId}-page-${actualPageNumber}`)
        .set({
          pdfId: event.params.pdfId,
          pageNumber: actualPageNumber,
          text: pageText
        });
    }));
  }
  
  return console.log('所有页面文本提取完成');
};

额外检查项

  • 调整Cloud Function超时时间:处理多个分片可能需要更长时间,可在Firebase控制台将函数超时设置为最大值(9分钟)。
  • 检查Vision请求参数:确认调用asyncBatchAnnotateFiles时没有设置pages参数限制处理范围。
  • 查看函数日志:如果仍有问题,检查Cloud Function的执行日志,排查是否有文件读取失败、JSON解析错误等情况。

内容的提问来源于stack exchange,提问作者DonBergen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 12:32:37