You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Function触发器无法正确处理PDF中的嵌入图片

解决Azure Function无法处理PDF内嵌图片索引至Cognitive Search的问题

问题背景

搭建的流程为:Azure Blob Storage的Blob创建/删除事件通过Event Grid触发Azure Function,函数提取Blob内容后上传至Azure Cognitive Search。目前独立图片、DOCX、CSV等文件均可正常索引,但PDF中的内嵌图片无法被处理。通过Azure门户手动测试Search索引器时,PDF内嵌图片能正常识别,但通过Function自动触发时失效。

架构概述:

  • Blob Storage的Blob变更事件触发Event Grid,进而调用Azure Function
  • Function获取Blob元数据,根据文件类型提取内容,上传至Cognitive Search
  • 仅PDF内嵌图片无法被正确索引

原因分析

当前代码使用pdf-parse库提取PDF内容,该库仅能解析PDF中的原生文本,不具备识别内嵌图片内容的OCR能力。而Azure Cognitive Search的内置索引器集成了OCR技能集,可自动识别PDF中的图片并提取文本,这是手动测试正常但Function触发异常的核心原因。

解决方案

提供两种可行方案,可根据需求选择:

方案1:改用Azure Cognitive Search索引器处理(推荐)

放弃在Function中自行提取内容,改为触发Azure Search的Blob索引器处理目标Blob。该方式直接复用Search内置的OCR和内容提取能力,无需额外开发OCR逻辑。

实现步骤:

  1. 在Azure Cognitive Search中创建指向目标Blob容器的Blob索引器,并在技能集配置中添加"OCR"技能,启用图片内容提取
  2. 修改Azure Function逻辑,不再自行提取内容,改为调用Search API触发索引器对特定Blob执行增量处理

修改后的核心代码

const { app } = require('@azure/functions');
const { SearchIndexerClient, AzureKeyCredential, SearchClient } = require('@azure/search-documents');
const path = require('path');
const fs = require('fs');
const configPath = path.join(__dirname, 'config.json');
const config = JSON.parse(fs.readFileSync(configPath, 'utf8'));

const searchServiceName = config.searchServiceName;
const indexerName = config.searchIndexerName; // 新增:你的Blob索引器名称
const indexName = config.searchIndexName;
const apiKey = config.searchApiKey;

// 创建索引器客户端
const indexerClient = new SearchIndexerClient(
  `https://${searchServiceName}.search.windows.net/`,
  new AzureKeyCredential(apiKey)
);

// 创建搜索客户端用于删除操作
const searchClient = new SearchClient(
  `https://${searchServiceName}.search.windows.net/`,
  indexName,
  new AzureKeyCredential(apiKey)
);

async function indexBlob(blobName) {
  try {
    console.log(`Triggering indexer for blob: ${blobName}`);
    // 触发索引器对指定Blob执行增量处理
    await indexerClient.runIndexer(indexerName, {
      parameters: {
        configuration: {
          "dataToExtract": "contentAndMetadata",
          "imageAction": "generateNormalizedImages",
          "allowSkillsetToReadFileData": true
        }
      }
    });
    console.log(`Indexer triggered for blob: ${blobName}`);
  } catch (error) {
    console.error(`Error triggering indexer:`, error);
    if (error.name === 'RestError') {
      console.error(`RestError message:`, error.message);
      console.error(`RestError details:`, error.response && error.response.body);
    }
  }
}

app.eventGrid('process-event-grid', {
  handler: async (context, eventGridEvent) => {
    try {
      console.log(`Event received: ${JSON.stringify(eventGridEvent)}`);
      const blobUrl = eventGridEvent.data.url;
      const blobapi = eventGridEvent.data.api;
      const blobName = blobUrl.substring(blobUrl.lastIndexOf('/') + 1);
      if (blobapi === 'PutBlob') {
        await indexBlob(blobName);
      } else if (blobapi === 'DeleteBlob') {
        await deleteDocument(blobName);
      }
    } catch (error) {
      console.error(`Error processing event: ${error}`, eventGridEvent);
    }
  }
});

async function deleteDocument(blobName) {
  try {
    console.log(`Deleting document: ${blobName}`);
    const encodedName = Buffer.from(blobName).toString('base64');
    console.log(`encodedName: ${encodedName}`);
    await searchClient.deleteDocuments([{ id: encodedName }]);
    console.log(`Document "${encodedName}" has been deleted from the index`);
  } catch (error) {
    console.error(`Error deleting document:`, error);
    if (error.name === 'RestError') {
      console.error(`RestError message:`, error.message);
      console.error(`RestError details:`, error.response && error.response.body);
    }
  }
}

module.exports = app;

方案2:在Function中集成OCR处理

使用Azure Computer Vision的Read API处理PDF,提取包括内嵌图片在内的所有文本,再将完整文本上传至Search。该方式需额外调用Computer Vision服务,适合需要自定义内容提取逻辑的场景。

实现步骤:

  1. 创建Azure Computer Vision资源,获取API密钥和端点
  2. 在Function中调用Read API处理PDF文件,获取包含图片文本的完整内容
  3. 将提取的完整文本上传至Cognitive Search

核心代码片段(集成Computer Vision)

// 新增依赖和配置
const axios = require('axios');
const computerVisionKey = config.computerVisionKey;
const computerVisionEndpoint = config.computerVisionEndpoint;

// 新增OCR处理函数
async function extractPdfContentWithOCR(pdfBuffer) {
  const url = `${computerVisionEndpoint}/computervision/v3.2/read/analyze`;
  const headers = {
    'Ocp-Apim-Subscription-Key': computerVisionKey,
    'Content-Type': 'application/pdf'
  };

  // 提交PDF处理请求
  const response = await axios.post(url, pdfBuffer, { headers });
  const operationLocation = response.headers['operation-location'];

  // 轮询获取处理结果
  let result;
  do {
    await new Promise(resolve => setTimeout(resolve, 1000));
    result = await axios.get(operationLocation, { headers });
  } while (result.status === 202);

  // 提取所有文本
  let fullText = '';
  result.data.analyzeResult.readResults.forEach(page => {
    page.lines.forEach(line => {
      fullText += line.text + '\n';
    });
  });
  return fullText;
}

// 修改PDF处理逻辑
if (contentType === 'application/pdf') {
  const pdfBuffer = await streamToBuffer(downloadResponse.readableStreamBody);
  // 替换原pdf-parse调用为OCR提取
  blobContent = await extractPdfContentWithOCR(pdfBuffer);
}

内容的提问来源于stack exchange,提问作者Su Myat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 05:28:13