如何将Streamlit上传的PDF文件转换为LangChain文档?
解决Streamlit上传PDF转LangChain文档的方案
针对你遇到的Streamlit上传文件(UploadedFile对象)无法直接适配LangChain PDF加载器的问题,有两种实用解决方案:
方法一:临时保存文件到本地后加载
通过创建临时文件存储上传的PDF内容,再用LangChain的PDF加载器读取,处理完成后删除临时文件,不占用本地持久化存储。
代码示例:
import streamlit as st from langchain.document_loaders import PyPDFLoader import tempfile import os uploaded_file = st.file_uploader("上传PDF文件", type="pdf") if uploaded_file is not None: # 创建临时PDF文件 with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp_file: tmp_file.write(uploaded_file.getvalue()) tmp_file_path = tmp_file.name # 加载PDF并生成LangChain文档 loader = PyPDFLoader(tmp_file_path) documents = loader.load() # 文本分割成小块 from langchain.text_splitter import RecursiveCharacterTextSplitter text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200) split_docs = text_splitter.split_documents(documents) # 清理临时文件 os.unlink(tmp_file_path)
方法二:直接读取文件流,手动构建LangChain Document
借助PDF处理库(如PyPDF2、pdfplumber)直接读取UploadedFile的字节流,提取文本后手动构建LangChain的Document对象,无需写入本地文件。
用PyPDF2实现:
import streamlit as st from langchain.schema import Document from langchain.text_splitter import RecursiveCharacterTextSplitter import PyPDF2 uploaded_file = st.file_uploader("上传PDF文件", type="pdf") if uploaded_file is not None: # 读取PDF文本内容 pdf_reader = PyPDF2.PdfReader(uploaded_file) full_text = "" for page in pdf_reader.pages: extracted_text = page.extract_text() if extracted_text: full_text += extracted_text # 构建LangChain Document对象 doc = Document( page_content=full_text, metadata={"source": uploaded_file.name, "page_count": len(pdf_reader.pages)} ) # 分割文本 text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200) split_docs = text_splitter.split_documents([doc])
用pdfplumber实现(文本提取精度更高):
import streamlit as st from langchain.schema import Document from langchain.text_splitter import RecursiveCharacterTextSplitter import pdfplumber uploaded_file = st.file_uploader("上传PDF文件", type="pdf") if uploaded_file is not None: with pdfplumber.open(uploaded_file) as pdf: full_text = "" for page in pdf.pages: extracted_text = page.extract_text() if extracted_text: full_text += extracted_text # 构建Document,可加入更多元数据 doc = Document( page_content=full_text, metadata={ "source": uploaded_file.name, "page_count": len(pdf.pages), "total_characters": len(full_text) } ) # 分割文本 text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200) split_docs = text_splitter.split_documents([doc])
两种方案对比
- 临时文件方案:保留LangChain加载器自带的分页元数据,适合需要按页处理文档的场景;但涉及文件IO,性能略低。
- 直接构建Document方案:无需本地文件操作,性能更优;但需要自行处理文本提取和元数据填充,适合轻量化需求的场景。
内容的提问来源于stack exchange,提问作者Tushar Singh
相关产品推荐
相关产品推荐

