You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否通过LangChain从Azure容器存储加载文件至Azure文档智能?

LangChain加载Azure存储文件至文档智能的问题及解决方法

问题背景

已配置好存储有多份PDF/Word/Excel文件的Azure容器存储账户,希望借助Azure文档智能完成语义分块,尝试通过LangChain的AzureAIDocumentIntelligenceLoader直接加载存储容器内的文件,但根据官方文档,该Loader目前仅支持本地文件或公开URL,运行代码后出现报错。

尝试代码

# 前提:Azure AI Document Intelligence资源需部署在以下3个预览区域之一:East US, West US2, West Europe

import os
from langchain_community.document_loaders import AzureAIDocumentIntelligenceLoader

file_path = "storage-path-to-file"
endpoint = os.getenv("DOCUMENTINTELLIGENCE_ENDPOINT")
key = os.getenv("DOCUMENTINTELLIGENCE_API_KEY")

loader = AzureAIDocumentIntelligenceLoader(
    api_endpoint=endpoint, api_key=key, file_path=file_path, api_model="prebuilt-layout"
)

documents = loader.load()

报错信息

Message: Invalid request.
Inner error: {
"code": "InvalidManagedIdentity",
"message": "The managed identity configuration is invalid: Managed identity is not enabled for the current resource."
}

原因分析

  1. 当前AzureAIDocumentIntelligenceLoader不支持直接传入Azure存储容器的路径(如azure://格式),仅能处理本地文件路径或公开可访问的URL。
  2. 报错中的InvalidManagedIdentity是因为若要让Azure文档智能直接访问存储容器,需为文档智能资源启用系统托管身份并分配存储容器的读取权限,但LangChain的该Loader目前仅支持单文件同步加载,无法适配这种批量授权的访问方式。

解决方法

方法1:生成存储文件的SAS URL

为存储容器内的文件生成带读取权限的SAS(共享访问签名)URL,将该URL作为file_path传入Loader,文档智能可通过SAS直接访问文件,无需配置托管身份。

示例代码:

import os
from azure.storage.blob import BlobServiceClient, generate_blob_sas, BlobSasPermissions
from datetime import datetime, timedelta
from langchain_community.document_loaders import AzureAIDocumentIntelligenceLoader

# 配置Azure存储账户信息
storage_account_name = os.getenv("AZURE_STORAGE_ACCOUNT_NAME")
storage_account_key = os.getenv("AZURE_STORAGE_ACCOUNT_KEY")
container_name = "你的容器名称"
blob_name = "目标文件名.pdf"

# 生成SAS Token
blob_service_client = BlobServiceClient(
    account_url=f"https://{storage_account_name}.blob.core.windows.net",
    credential=storage_account_key
)
sas_token = generate_blob_sas(
    account_name=storage_account_name,
    container_name=container_name,
    blob_name=blob_name,
    account_key=storage_account_key,
    permission=BlobSasPermissions(read=True),
    expiry=datetime.utcnow() + timedelta(hours=1)  # 设置1小时有效期
)
# 拼接完整的SAS URL
sas_url = f"https://{storage_account_name}.blob.core.windows.net/{container_name}/{blob_name}?{sas_token}"

# 使用SAS URL加载文件
endpoint = os.getenv("DOCUMENTINTELLIGENCE_ENDPOINT")
key = os.getenv("DOCUMENTINTELLIGENCE_API_KEY")

loader = AzureAIDocumentIntelligenceLoader(
    api_endpoint=endpoint, api_key=key, file_path=sas_url, api_model="prebuilt-layout"
)
documents = loader.load()

方法2:下载文件至本地临时目录

先将存储容器内的文件下载到本地临时目录,再通过Loader加载本地文件,处理完成后删除临时文件。

示例代码:

import os
import tempfile
from azure.storage.blob import BlobServiceClient
from langchain_community.document_loaders import AzureAIDocumentIntelligenceLoader

# 配置Azure存储账户信息
storage_account_name = os.getenv("AZURE_STORAGE_ACCOUNT_NAME")
storage_account_key = os.getenv("AZURE_STORAGE_ACCOUNT_KEY")
container_name = "你的容器名称"
blob_name = "目标文件名.pdf"

# 下载Blob到临时文件
blob_service_client = BlobServiceClient(
    account_url=f"https://{storage_account_name}.blob.core.windows.net",
    credential=storage_account_key
)
blob_client = blob_service_client.get_blob_client(container=container_name, blob=blob_name)

# 创建临时文件
with tempfile.NamedTemporaryFile(delete=False, suffix=os.path.splitext(blob_name)[1]) as temp_file:
    blob_client.download_blob().readinto(temp_file)
    temp_file_path = temp_file.name

# 加载本地临时文件
endpoint = os.getenv("DOCUMENTINTELLIGENCE_ENDPOINT")
key = os.getenv("DOCUMENTINTELLIGENCE_API_KEY")

loader = AzureAIDocumentIntelligenceLoader(
    api_endpoint=endpoint, api_key=key, file_path=temp_file_path, api_model="prebuilt-layout"
)
documents = loader.load()

# 清理临时文件
os.unlink(temp_file_path)

内容的提问来源于stack exchange,提问作者user483161

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 03:26:05