You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于使用Azure Computer Vision从图像中提取特定目标文本的技术咨询

Extracting Specific Text (e.g., Ingredients) with Azure Computer Vision

Absolutely, Azure Computer Vision does support extracting targeted text instead of all content from images—even when your target text varies across different images. Here’s a practical breakdown of how to implement this, especially for use cases like pulling ingredient lists from mixed-text images:

Step 1: First, Extract All Structured Text with the Read API

Start by using Azure Computer Vision’s Read API (the recommended OCR tool for most scenarios)—it returns structured text data including individual lines, words, and their bounding box coordinates (position on the image). This gives you a foundation to work with, since you need the full text context to filter your target content.

For example, using the .NET SDK, you’d initialize the client and call the image analysis:

using Azure.AI.Vision.ImageAnalysis;
using Azure;

var client = new ImageAnalysisClient(
    new Uri("https://<your-resource-name>.cognitiveservices.azure.com/"),
    new AzureKeyCredential("<your-api-key>"));

var result = await client.AnalyzeImageAsync(
    BinaryData.FromStream(File.OpenRead("your-image.jpg")),
    VisualFeatures.Read);

The result.Read.Blocks property will contain all the text blocks, lines, and their positional data.

Step 2: Filter for Your Target Text (e.g., Ingredients)

Once you have the full structured text, use one or more of these strategies to isolate your desired content:

  • Keyword/Pattern Matching: If your target text follows a consistent lead-in (like "INGREDIENTS:", "成分:", or a similar label), locate that line first, then extract all subsequent lines until you hit another distinct section (e.g., "NUTRITION FACTS", "使用说明"). You can use simple string checks or regular expressions to handle variations in capitalization or spacing.

  • Position-Based Filtering: If your target text is consistently located in a specific region of the image (e.g., lower half, left column), use the bounding box coordinates from the OCR result to filter text blocks within that area. For example, check if a line’s BoundingBox.Top and BoundingBox.Bottom fall within your predefined y-axis range.

  • Semantic Extraction with LLMs: For unstructured or highly variable text, pair the OCR output with a large language model (like Azure OpenAI GPT). Pass the full OCR text to the model with a prompt like: "Extract only the ingredient list from the following text, ignoring all other content: [full OCR text]". This works great when there’s no fixed format for your target content.

Step 3: Refine for Edge Cases

If your images have inconsistent formatting (e.g., some use bullet points for ingredients, others use plain lines), combine multiple strategies: for example, use keyword matching to find the start of the ingredient section, then use positional data to confirm you’re pulling the correct adjacent lines, and finally clean up the output with regex to remove unwanted characters.


内容的提问来源于stack exchange,提问作者Java user

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 02:37:30