Node.js调用Azure Computer Vision OCR提取配料成分问题咨询
Node.js调用Azure Computer Vision提取食品配料文本解决方案
一、问题1+2解决方案:提取配料纯文本输出
直接替换原有请求回调逻辑即可,代码如下:
request.post(options, (error, response, body) => { if (error) { console.log('Error: ', error); return; } const ocrResult = JSON.parse(body); // 1. 遍历所有行,存储文本和对应坐标参数 const allLines = []; ocrResult.regions.forEach(region => { region.lines.forEach(line => { const lineText = line.words.map(word => word.text).join(' '); const boxParams = line.boundingBox.split(',').map(Number); allLines.push({ text: lineText, left: boxParams[0], height: boxParams[3] }); }); }); // 2. 匹配配料表起止范围 let ingredientStart = -1; let ingredientEnd = allLines.length - 1; // 匹配INGREDIENTS开头的起始行 for (let i = 0; i < allLines.length; i++) { if (allLines[i].text.toUpperCase().includes('INGREDIENTS:')) { ingredientStart = i; continue; } // 匹配下一个全大写标题作为结束边界(如DISTRIBUTED BY、NUTRITION等) if (ingredientStart > -1 && /^[A-Z\s:]+$/.test(allLines[i].text.trim())) { ingredientEnd = i - 1; break; } } // 3. 拼接配料文本并翻译为中文 let ingredientText = '未识别到配料表'; if (ingredientStart > -1) { const rawIngredient = allLines.slice(ingredientStart, ingredientEnd + 1) .map(line => line.text.replace(/^INGREDIENTS:\s*/, '')) .join(' ') .replace(/,$/, ''); // 英文配料转中文 ingredientText = `配料:${rawIngredient .replace('WHITE TUNA', '白金枪鱼') .replace('SOYBEAN OIL', '大豆油') .replace('VEGETABLE BROTH', '蔬菜高汤') .replace('SALT', '食用盐') .replace('PYROPHOSPHATE ADDED', '添加焦磷酸盐') .replace(/,/g, '、')}`; } console.log(ingredientText); // 最终输出示例:配料:白金枪鱼、大豆油、蔬菜高汤、食用盐、添加焦磷酸盐 });
二、问题3解决方案:无INGREDIENTS关键词的提取方案
- 布局特征筛选:配料表通常为连续左对齐的文本块,同一块文本的left坐标偏差不会超过20px,行高差值小于10px,可先筛选符合该特征的连续文本块作为候选。
- 内容特征匹配:提前维护常见食品配料词库(中英文均可),统计候选文本块中配料词的占比,占比超过60%即可判定为配料表。
- 排除法筛选:先排除特征明显的其他板块:包含
g、%、千卡等单位的营养成分表,包含BY、地址、电话等关键词的厂商信息板块,剩余的连续长文本块基本为配料表。
内容的提问来源于stack exchange,提问作者Java user
相关产品推荐
相关产品推荐

