You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何用LangChain PyPDFLoader从Azure Blob Storage加载PDF文件

解决LangChain加载Azure Blob Storage中PDF文件的类型错误问题

我尝试用LangChain和存储在Azure Blob Storage里的PDF交互,在Azure ML Notebook中用PyPDFLoader或UnstructuredFileLoader加载文件时始终失败,本地PDF加载正常,但连接Blob存储时抛出TypeError: expected str, bytes or os.PathLike object, not StreamInfoFileObject错误。尝试的代码如下:

from azureml.fsspec import AzureMachineLearningFileSystem
fs = AzureMachineLearningFileSystem("<path to datastore>")

from langchain.document_loaders import PyPDFLoader
with fs.open('*/.../file.pdf', 'rb') as fd:
    loader = PyPDFLoader(document)
    data = loader.load()

# 错误:TypeError: expected str, bytes or os.PathLike object, not StreamInfoFileObject

另一个尝试的示例:

from langchain.document_loaders import UnstructuredFileLoader
with fs.open('*/.../file.pdf', 'rb') as fd:
    loader = UnstructuredFileLoader(fd)
    documents = loader.load() 

# 错误:TypeError: expected str, bytes or os.PathLike object, not StreamInfoFileObject

核心原因

PyPDFLoader和UnstructuredFileLoader的构造函数仅接受文件路径(本地路径、可被文件系统识别的路径字符串),不支持直接传入从AzureMachineLearningFileSystem打开的文件流对象,因此触发类型错误。


解决方案

方案1:下载到临时文件后加载

将Blob中的PDF下载到本地临时文件,再用LangChain的Loader加载,兼容所有PDF Loader:

from azureml.fsspec import AzureMachineLearningFileSystem
from langchain.document_loaders import PyPDFLoader
import tempfile
import os

fs = AzureMachineLearningFileSystem("<path to datastore>")

# 将Blob文件写入临时文件
with tempfile.NamedTemporaryFile(delete=False, suffix='.pdf') as tmp_file:
    with fs.open('*/.../file.pdf', 'rb') as fd:
        tmp_file.write(fd.read())
    tmp_file_path = tmp_file.name

# 加载临时文件
loader = PyPDFLoader(tmp_file_path)
data = loader.load()

# 加载完成后清理临时文件
os.unlink(tmp_file_path)

方案2:使用LangChain官方Azure Blob Loader

使用LangChain提供的AzureBlobStorageLoader直接加载Blob中的文件,无需手动处理流:

  1. 先安装依赖包:
pip install langchain-azure
  1. 编写加载代码:
from langchain_azure import AzureBlobStorageLoader

loader = AzureBlobStorageLoader(
    conn_str="<你的Azure Blob连接字符串>",
    container="<容器名称>",
    blob_name="<Blob中文件的路径,如path/to/file.pdf>"
)
documents = loader.load()

方案3:自定义流处理逻辑

直接用PDF解析库读取文件流,手动构造LangChain的Document对象,适合需要精细控制的场景:

from azureml.fsspec import AzureMachineLearningFileSystem
from PyPDF2 import PdfReader
from langchain_core.documents import Document

fs = AzureMachineLearningFileSystem("<path to datastore>")

with fs.open('*/.../file.pdf', 'rb') as fd:
    reader = PdfReader(fd)
    # 提取所有页面文本
    full_text = ""
    for page in reader.pages:
        page_text = page.extract_text()
        if page_text:
            full_text += page_text + "\n"
    # 构造LangChain Document对象
    document = Document(
        page_content=full_text,
        metadata={"source": "azure_blob://<容器名>/<文件路径>"}
    )

内容的提问来源于stack exchange,提问作者stackword_0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 17:04:51