Azure Function触发器无法正确处理PDF中的嵌入图片
解决Azure Function无法处理PDF内嵌图片索引至Cognitive Search的问题
问题背景
搭建的流程为:Azure Blob Storage的Blob创建/删除事件通过Event Grid触发Azure Function,函数提取Blob内容后上传至Azure Cognitive Search。目前独立图片、DOCX、CSV等文件均可正常索引,但PDF中的内嵌图片无法被处理。通过Azure门户手动测试Search索引器时,PDF内嵌图片能正常识别,但通过Function自动触发时失效。
架构概述:
- Blob Storage的Blob变更事件触发Event Grid,进而调用Azure Function
- Function获取Blob元数据,根据文件类型提取内容,上传至Cognitive Search
- 仅PDF内嵌图片无法被正确索引
原因分析
当前代码使用pdf-parse库提取PDF内容,该库仅能解析PDF中的原生文本,不具备识别内嵌图片内容的OCR能力。而Azure Cognitive Search的内置索引器集成了OCR技能集,可自动识别PDF中的图片并提取文本,这是手动测试正常但Function触发异常的核心原因。
解决方案
提供两种可行方案,可根据需求选择:
方案1:改用Azure Cognitive Search索引器处理(推荐)
放弃在Function中自行提取内容,改为触发Azure Search的Blob索引器处理目标Blob。该方式直接复用Search内置的OCR和内容提取能力,无需额外开发OCR逻辑。
实现步骤:
- 在Azure Cognitive Search中创建指向目标Blob容器的Blob索引器,并在技能集配置中添加"OCR"技能,启用图片内容提取
- 修改Azure Function逻辑,不再自行提取内容,改为调用Search API触发索引器对特定Blob执行增量处理
修改后的核心代码
const { app } = require('@azure/functions'); const { SearchIndexerClient, AzureKeyCredential, SearchClient } = require('@azure/search-documents'); const path = require('path'); const fs = require('fs'); const configPath = path.join(__dirname, 'config.json'); const config = JSON.parse(fs.readFileSync(configPath, 'utf8')); const searchServiceName = config.searchServiceName; const indexerName = config.searchIndexerName; // 新增:你的Blob索引器名称 const indexName = config.searchIndexName; const apiKey = config.searchApiKey; // 创建索引器客户端 const indexerClient = new SearchIndexerClient( `https://${searchServiceName}.search.windows.net/`, new AzureKeyCredential(apiKey) ); // 创建搜索客户端用于删除操作 const searchClient = new SearchClient( `https://${searchServiceName}.search.windows.net/`, indexName, new AzureKeyCredential(apiKey) ); async function indexBlob(blobName) { try { console.log(`Triggering indexer for blob: ${blobName}`); // 触发索引器对指定Blob执行增量处理 await indexerClient.runIndexer(indexerName, { parameters: { configuration: { "dataToExtract": "contentAndMetadata", "imageAction": "generateNormalizedImages", "allowSkillsetToReadFileData": true } } }); console.log(`Indexer triggered for blob: ${blobName}`); } catch (error) { console.error(`Error triggering indexer:`, error); if (error.name === 'RestError') { console.error(`RestError message:`, error.message); console.error(`RestError details:`, error.response && error.response.body); } } } app.eventGrid('process-event-grid', { handler: async (context, eventGridEvent) => { try { console.log(`Event received: ${JSON.stringify(eventGridEvent)}`); const blobUrl = eventGridEvent.data.url; const blobapi = eventGridEvent.data.api; const blobName = blobUrl.substring(blobUrl.lastIndexOf('/') + 1); if (blobapi === 'PutBlob') { await indexBlob(blobName); } else if (blobapi === 'DeleteBlob') { await deleteDocument(blobName); } } catch (error) { console.error(`Error processing event: ${error}`, eventGridEvent); } } }); async function deleteDocument(blobName) { try { console.log(`Deleting document: ${blobName}`); const encodedName = Buffer.from(blobName).toString('base64'); console.log(`encodedName: ${encodedName}`); await searchClient.deleteDocuments([{ id: encodedName }]); console.log(`Document "${encodedName}" has been deleted from the index`); } catch (error) { console.error(`Error deleting document:`, error); if (error.name === 'RestError') { console.error(`RestError message:`, error.message); console.error(`RestError details:`, error.response && error.response.body); } } } module.exports = app;
方案2:在Function中集成OCR处理
使用Azure Computer Vision的Read API处理PDF,提取包括内嵌图片在内的所有文本,再将完整文本上传至Search。该方式需额外调用Computer Vision服务,适合需要自定义内容提取逻辑的场景。
实现步骤:
- 创建Azure Computer Vision资源,获取API密钥和端点
- 在Function中调用Read API处理PDF文件,获取包含图片文本的完整内容
- 将提取的完整文本上传至Cognitive Search
核心代码片段(集成Computer Vision)
// 新增依赖和配置 const axios = require('axios'); const computerVisionKey = config.computerVisionKey; const computerVisionEndpoint = config.computerVisionEndpoint; // 新增OCR处理函数 async function extractPdfContentWithOCR(pdfBuffer) { const url = `${computerVisionEndpoint}/computervision/v3.2/read/analyze`; const headers = { 'Ocp-Apim-Subscription-Key': computerVisionKey, 'Content-Type': 'application/pdf' }; // 提交PDF处理请求 const response = await axios.post(url, pdfBuffer, { headers }); const operationLocation = response.headers['operation-location']; // 轮询获取处理结果 let result; do { await new Promise(resolve => setTimeout(resolve, 1000)); result = await axios.get(operationLocation, { headers }); } while (result.status === 202); // 提取所有文本 let fullText = ''; result.data.analyzeResult.readResults.forEach(page => { page.lines.forEach(line => { fullText += line.text + '\n'; }); }); return fullText; } // 修改PDF处理逻辑 if (contentType === 'application/pdf') { const pdfBuffer = await streamToBuffer(downloadResponse.readableStreamBody); // 替换原pdf-parse调用为OCR提取 blobContent = await extractPdfContentWithOCR(pdfBuffer); }
内容的提问来源于stack exchange,提问作者Su Myat
相关产品推荐
相关产品推荐

