如何从Azure Blob Storage多图提取Azure OCR及解决导入错误
解决Azure Blob导入报错+OCR识别并导出Excel完整方案
一、解决“无法从azure.blob.storage导入BlobServiceClient”报错
- 问题原因:新版
azure-storage-blob包中,BlobServiceClient的导入路径已变更,旧路径azure.blob.storage不再生效 - 修复步骤:
- 安装/更新正确的包版本:
pip install azure-storage-blob>=12.0.0 - 修正导入语句:
from azure.storage.blob import BlobServiceClient
- 安装/更新正确的包版本:
二、Blob图片OCR识别并导出Excel完整流程
需用到Azure Blob Storage(读取图片)和Computer Vision(OCR识别)两项服务,具体步骤如下:
1. 安装依赖包
pip install azure-cognitiveservices-vision-computervision pandas openpyxl
2. 初始化服务客户端
from azure.storage.blob import BlobServiceClient, generate_blob_sas, BlobSasPermissions from azure.cognitiveservices.vision.computervision import ComputerVisionClient from msrest.authentication import CognitiveServicesCredentials import pandas as pd from datetime import datetime, timedelta # Blob Storage配置信息 blob_conn_str = "你的Blob存储连接字符串" container_name = "目标容器名称" # Computer Vision配置信息 cv_key = "你的计算机视觉服务密钥" cv_endpoint = "你的计算机视觉服务端点" # 初始化客户端实例 blob_service_client = BlobServiceClient.from_connection_string(blob_conn_str) cv_client = ComputerVisionClient(cv_endpoint, CognitiveServicesCredentials(cv_key))
3. 筛选容器内的图片文件
container_client = blob_service_client.get_container_client(container_name) # 按需添加/修改图片格式后缀 image_blobs = [blob for blob in container_client.list_blobs() if blob.name.lower().endswith(('.png', '.jpg', '.jpeg'))]
4. 生成Blob的SAS访问链接
OCR服务需要可访问的图片URL,生成SAS链接是安全的授权方式:
def get_blob_sas_url(blob_client): sas_token = generate_blob_sas( account_name=blob_client.account_name, container_name=blob_client.container_name, blob_name=blob_client.blob_name, account_key=blob_client.credential.account_key, permission=BlobSasPermissions(read=True), expiry=datetime.utcnow() + timedelta(hours=1) # 链接有效期可按需调整 ) return f"{blob_client.url}?{sas_token}"
5. 批量执行OCR识别并提取结果
ocr_results = [] for blob in image_blobs: blob_client = container_client.get_blob_client(blob.name) sas_url = get_blob_sas_url(blob_client) # 调用异步OCR接口 ocr_response = cv_client.read(sas_url, raw=True) operation_id = ocr_response.headers["Operation-Location"].split("/")[-1] # 等待识别任务完成 while True: result = cv_client.get_read_result(operation_id) if result.status not in ['notStarted', 'running']: break # 提取文本与置信度数据 if result.status == 'succeeded': for read_result in result.analyze_result.read_results: for line in read_result.lines: ocr_results.append({ "图片名称": blob.name, "识别文本": line.text, "置信度": round(line.confidence, 4) })
6. 将结果导出至Excel
df = pd.DataFrame(ocr_results) df.to_excel("OCR识别结果汇总.xlsx", index=False, engine='openpyxl')
内容的提问来源于stack exchange,提问作者SIBA
相关产品推荐
相关产品推荐

