You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Google Cloud Document AI处理阿拉伯语PDF时的行序错乱问题

使用Google Cloud Document AI提取PDF文本时出现行序错乱问题

我开发的应用采用Google Cloud Document AI处理PDF并提取文本,即便使用稳定版本,提取出的文本仍存在行序错乱问题,与原文档顺序不符。该问题对我的应用至关重要,因为后续处理流程依赖准确的文本提取结果。

我尝试直接读取文件而非以Buffer形式发送,且严格遵循Google提供的快速入门指南编写代码,具体实现如下:

async processDocument(fileContent: Buffer | undefined, contentType: string) {
  try {
    const name = `projects/${projectId}/locations/${location}/processors/${processorId}`;

    const encodedFile = Buffer.from(fileContent!).toString("base64");

    const request = {
      name,
      rawDocument: {
        content: encodedFile,
        mimeType: contentType,
      },
    };

    const [result] = await this.client.processDocument(request);
    const { document } = result;

    return document;
  } catch (error) {
    throw error;
  }
}
catch(error) {
  throw error;
}
}

(注:原截图显示提取出的文本行序与原文档阅读顺序不一致)

我在Google Cloud控制台测试该处理器时,返回结果行序正确,无异常情况。请问是否有开发者遇到过类似问题,或能指出问题根源?


可能的问题根源及解决方案:

  • 未按视觉位置排序识别结果
    Document AI返回的文本元素(行、段落)是按识别顺序输出的,并非默认遵循视觉阅读顺序。需要通过每个元素boundingPoly中的坐标(优先y轴,再结合x轴)重新排序文本行。

    示例排序代码:

    function sortTextByVisualPosition(document) {
      const sortedContent = [];
      document.pages.forEach(page => {
        const pageLines = [];
        // 遍历所有层级的文本元素
        page.blocks.forEach(block => {
          block.paragraphs.forEach(para => {
            para.lines.forEach(line => {
              // 获取行的顶部y坐标和左侧x坐标
              const topY = line.boundingPoly.vertices[0].y;
              const leftX = line.boundingPoly.vertices[0].x;
              pageLines.push({
                text: line.words.map(word => word.text).join(' '),
                topY,
                leftX
              });
            });
          });
        });
        // 先按垂直位置排序,同一行内按水平位置排序
        pageLines.sort((a, b) => {
          if (a.topY !== b.topY) return a.topY - b.topY;
          return a.leftX - b.leftX;
        });
        sortedContent.push(...pageLines.map(item => item.text));
      });
      return sortedContent.join('\n');
    }
    
  • API请求与控制台配置不一致
    控制台测试可能默认启用了某些你API请求中遗漏的配置项。比如添加语言配置或OCR增强选项:

    const request = {
      name,
      rawDocument: {
        content: encodedFile,
        mimeType: contentType,
      },
      processOptions: {
        ocrConfig: {
          languageCode: "zh-CN", // 替换为文档实际语言
          enableImageQualityScores: true
        }
      }
    };
    
  • 文档布局复杂度导致的识别偏差
    如果PDF包含分栏、浮动元素或为扫描件,基础处理器可能无法准确识别阅读顺序。可以尝试切换为Document OCR或Form Parser处理器,这类处理器对复杂布局的适配性更强。


内容的提问来源于stack exchange,提问作者Khaled Saleh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 21:45:07