求助:如何用LangChain PyPDFLoader从Azure Blob Storage加载PDF文件
解决LangChain加载Azure Blob Storage中PDF文件的类型错误问题
我尝试用LangChain和存储在Azure Blob Storage里的PDF交互,在Azure ML Notebook中用PyPDFLoader或UnstructuredFileLoader加载文件时始终失败,本地PDF加载正常,但连接Blob存储时抛出TypeError: expected str, bytes or os.PathLike object, not StreamInfoFileObject错误。尝试的代码如下:
from azureml.fsspec import AzureMachineLearningFileSystem fs = AzureMachineLearningFileSystem("<path to datastore>") from langchain.document_loaders import PyPDFLoader with fs.open('*/.../file.pdf', 'rb') as fd: loader = PyPDFLoader(document) data = loader.load() # 错误:TypeError: expected str, bytes or os.PathLike object, not StreamInfoFileObject
另一个尝试的示例:
from langchain.document_loaders import UnstructuredFileLoader with fs.open('*/.../file.pdf', 'rb') as fd: loader = UnstructuredFileLoader(fd) documents = loader.load() # 错误:TypeError: expected str, bytes or os.PathLike object, not StreamInfoFileObject
核心原因
PyPDFLoader和UnstructuredFileLoader的构造函数仅接受文件路径(本地路径、可被文件系统识别的路径字符串),不支持直接传入从AzureMachineLearningFileSystem打开的文件流对象,因此触发类型错误。
解决方案
方案1:下载到临时文件后加载
将Blob中的PDF下载到本地临时文件,再用LangChain的Loader加载,兼容所有PDF Loader:
from azureml.fsspec import AzureMachineLearningFileSystem from langchain.document_loaders import PyPDFLoader import tempfile import os fs = AzureMachineLearningFileSystem("<path to datastore>") # 将Blob文件写入临时文件 with tempfile.NamedTemporaryFile(delete=False, suffix='.pdf') as tmp_file: with fs.open('*/.../file.pdf', 'rb') as fd: tmp_file.write(fd.read()) tmp_file_path = tmp_file.name # 加载临时文件 loader = PyPDFLoader(tmp_file_path) data = loader.load() # 加载完成后清理临时文件 os.unlink(tmp_file_path)
方案2:使用LangChain官方Azure Blob Loader
使用LangChain提供的AzureBlobStorageLoader直接加载Blob中的文件,无需手动处理流:
- 先安装依赖包:
pip install langchain-azure
- 编写加载代码:
from langchain_azure import AzureBlobStorageLoader loader = AzureBlobStorageLoader( conn_str="<你的Azure Blob连接字符串>", container="<容器名称>", blob_name="<Blob中文件的路径,如path/to/file.pdf>" ) documents = loader.load()
方案3:自定义流处理逻辑
直接用PDF解析库读取文件流,手动构造LangChain的Document对象,适合需要精细控制的场景:
from azureml.fsspec import AzureMachineLearningFileSystem from PyPDF2 import PdfReader from langchain_core.documents import Document fs = AzureMachineLearningFileSystem("<path to datastore>") with fs.open('*/.../file.pdf', 'rb') as fd: reader = PdfReader(fd) # 提取所有页面文本 full_text = "" for page in reader.pages: page_text = page.extract_text() if page_text: full_text += page_text + "\n" # 构造LangChain Document对象 document = Document( page_content=full_text, metadata={"source": "azure_blob://<容器名>/<文件路径>"} )
内容的提问来源于stack exchange,提问作者stackword_0
相关产品推荐
相关产品推荐

