You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Azure Blob Storage多图提取Azure OCR及解决导入错误

解决Azure Blob导入报错+OCR识别并导出Excel完整方案

一、解决“无法从azure.blob.storage导入BlobServiceClient”报错

  • 问题原因:新版azure-storage-blob包中,BlobServiceClient的导入路径已变更,旧路径azure.blob.storage不再生效
  • 修复步骤:
    1. 安装/更新正确的包版本:
      pip install azure-storage-blob>=12.0.0
      
    2. 修正导入语句:
      from azure.storage.blob import BlobServiceClient
      

二、Blob图片OCR识别并导出Excel完整流程

需用到Azure Blob Storage(读取图片)和Computer Vision(OCR识别)两项服务,具体步骤如下:

1. 安装依赖包

pip install azure-cognitiveservices-vision-computervision pandas openpyxl

2. 初始化服务客户端

from azure.storage.blob import BlobServiceClient, generate_blob_sas, BlobSasPermissions
from azure.cognitiveservices.vision.computervision import ComputerVisionClient
from msrest.authentication import CognitiveServicesCredentials
import pandas as pd
from datetime import datetime, timedelta

# Blob Storage配置信息
blob_conn_str = "你的Blob存储连接字符串"
container_name = "目标容器名称"

# Computer Vision配置信息
cv_key = "你的计算机视觉服务密钥"
cv_endpoint = "你的计算机视觉服务端点"

# 初始化客户端实例
blob_service_client = BlobServiceClient.from_connection_string(blob_conn_str)
cv_client = ComputerVisionClient(cv_endpoint, CognitiveServicesCredentials(cv_key))

3. 筛选容器内的图片文件

container_client = blob_service_client.get_container_client(container_name)
# 按需添加/修改图片格式后缀
image_blobs = [blob for blob in container_client.list_blobs() if blob.name.lower().endswith(('.png', '.jpg', '.jpeg'))]

4. 生成Blob的SAS访问链接

OCR服务需要可访问的图片URL,生成SAS链接是安全的授权方式:

def get_blob_sas_url(blob_client):
    sas_token = generate_blob_sas(
        account_name=blob_client.account_name,
        container_name=blob_client.container_name,
        blob_name=blob_client.blob_name,
        account_key=blob_client.credential.account_key,
        permission=BlobSasPermissions(read=True),
        expiry=datetime.utcnow() + timedelta(hours=1)  # 链接有效期可按需调整
    )
    return f"{blob_client.url}?{sas_token}"

5. 批量执行OCR识别并提取结果

ocr_results = []
for blob in image_blobs:
    blob_client = container_client.get_blob_client(blob.name)
    sas_url = get_blob_sas_url(blob_client)
    
    # 调用异步OCR接口
    ocr_response = cv_client.read(sas_url, raw=True)
    operation_id = ocr_response.headers["Operation-Location"].split("/")[-1]
    
    # 等待识别任务完成
    while True:
        result = cv_client.get_read_result(operation_id)
        if result.status not in ['notStarted', 'running']:
            break
    
    # 提取文本与置信度数据
    if result.status == 'succeeded':
        for read_result in result.analyze_result.read_results:
            for line in read_result.lines:
                ocr_results.append({
                    "图片名称": blob.name,
                    "识别文本": line.text,
                    "置信度": round(line.confidence, 4)
                })

6. 将结果导出至Excel

df = pd.DataFrame(ocr_results)
df.to_excel("OCR识别结果汇总.xlsx", index=False, engine='openpyxl')

内容的提问来源于stack exchange,提问作者SIBA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 21:01:02