You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Streamlit上传的PDF文件转换为LangChain文档?

解决Streamlit上传PDF转LangChain文档的方案

针对你遇到的Streamlit上传文件(UploadedFile对象)无法直接适配LangChain PDF加载器的问题,有两种实用解决方案:

方法一:临时保存文件到本地后加载

通过创建临时文件存储上传的PDF内容,再用LangChain的PDF加载器读取,处理完成后删除临时文件,不占用本地持久化存储。

代码示例:

import streamlit as st
from langchain.document_loaders import PyPDFLoader
import tempfile
import os

uploaded_file = st.file_uploader("上传PDF文件", type="pdf")

if uploaded_file is not None:
    # 创建临时PDF文件
    with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp_file:
        tmp_file.write(uploaded_file.getvalue())
        tmp_file_path = tmp_file.name
    
    # 加载PDF并生成LangChain文档
    loader = PyPDFLoader(tmp_file_path)
    documents = loader.load()
    
    # 文本分割成小块
    from langchain.text_splitter import RecursiveCharacterTextSplitter
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
    split_docs = text_splitter.split_documents(documents)
    
    # 清理临时文件
    os.unlink(tmp_file_path)

方法二:直接读取文件流,手动构建LangChain Document

借助PDF处理库(如PyPDF2、pdfplumber)直接读取UploadedFile的字节流,提取文本后手动构建LangChain的Document对象,无需写入本地文件。

用PyPDF2实现:

import streamlit as st
from langchain.schema import Document
from langchain.text_splitter import RecursiveCharacterTextSplitter
import PyPDF2

uploaded_file = st.file_uploader("上传PDF文件", type="pdf")

if uploaded_file is not None:
    # 读取PDF文本内容
    pdf_reader = PyPDF2.PdfReader(uploaded_file)
    full_text = ""
    for page in pdf_reader.pages:
        extracted_text = page.extract_text()
        if extracted_text:
            full_text += extracted_text
    
    # 构建LangChain Document对象
    doc = Document(
        page_content=full_text,
        metadata={"source": uploaded_file.name, "page_count": len(pdf_reader.pages)}
    )
    
    # 分割文本
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
    split_docs = text_splitter.split_documents([doc])

用pdfplumber实现(文本提取精度更高):

import streamlit as st
from langchain.schema import Document
from langchain.text_splitter import RecursiveCharacterTextSplitter
import pdfplumber

uploaded_file = st.file_uploader("上传PDF文件", type="pdf")

if uploaded_file is not None:
    with pdfplumber.open(uploaded_file) as pdf:
        full_text = ""
        for page in pdf.pages:
            extracted_text = page.extract_text()
            if extracted_text:
                full_text += extracted_text
        
        # 构建Document,可加入更多元数据
        doc = Document(
            page_content=full_text,
            metadata={
                "source": uploaded_file.name,
                "page_count": len(pdf.pages),
                "total_characters": len(full_text)
            }
        )
    
    # 分割文本
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
    split_docs = text_splitter.split_documents([doc])

两种方案对比

  • 临时文件方案:保留LangChain加载器自带的分页元数据,适合需要按页处理文档的场景;但涉及文件IO,性能略低。
  • 直接构建Document方案:无需本地文件操作,性能更优;但需要自行处理文本提取和元数据填充,适合轻量化需求的场景。

内容的提问来源于stack exchange,提问作者Tushar Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 21:05:29