You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Streamlit问答项目pdf模块重复调用问题求助

解决Streamlit问答项目中文件重复处理的问题

问题根源

Streamlit的核心运行机制是每次用户交互(点击按钮、输入文本等)都会从头重新执行整个脚本,导致你在app.py的main()里每次都调用pdf.main(),进而重复处理文件、反复替换数据库。

解决方案:利用Streamlit会话状态(st.session_state)标记处理状态

通过st.session_state存储文件是否已处理的标记,仅在首次加载或需要重新处理时执行文件上传与处理逻辑。

1. 修改app.py:控制pdf.main()的执行次数

import streamlit as st
from files import pdf

class single_ques:
    def __init__(self):
        # 原有逻辑保留
        pass

class multiple_ques:
    def __init__(self):
        # 原有逻辑保留
        pass

def main():
    # 初始化会话状态标记:记录文件是否已处理
    if "files_processed" not in st.session_state:
        st.session_state.files_processed = False

    # 仅当文件未处理时执行上传与处理逻辑
    if not st.session_state.files_processed:
        processed_success = pdf.main()
        if processed_success:
            st.session_state.files_processed = True
            st.success("文件处理完成,可开始提问")
    else:
        # 已处理时提供重新上传选项
        if st.button("重新上传文件"):
            st.session_state.files_processed = False
            st.rerun()

    # 问答模式选择逻辑
    button = st.radio("选择提问模式", ["single_ques", "multiple_ques"])
    if button == "single_ques":
        single_ques()
    if button == "multiple_ques":
        multiple_ques()

if __name__ == '__main__':
    main()

2. 修改pdf.py:修复逻辑错误并返回处理状态

import streamlit as st
import os
from files import loader
from files.database import save_to_chromadb

class doc:
    def __init__(self, path):
        self.path = path  # 补充文件路径赋值(原代码遗漏)
        self.document = self.load()
        if self.path.endswith(('.txt','.docx','.pdf')):
            self.chunks = self.chunking()

    def load(self):
        # 调用loader.py对应文件类型的处理逻辑
        if self.path.endswith('.txt'):
            return loader.load_txt(self.path)
        elif self.path.endswith('.pdf'):
            return loader.load_pdf(self.path)
        elif self.path.endswith('.docx'):
            return loader.load_docx(self.path)
        else:
            raise ValueError(f"不支持的文件类型:{self.path}")

    def chunking(self):
        # 示例文本分块逻辑(可根据需求调整)
        from langchain.text_splitter import RecursiveCharacterTextSplitter
        splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
        return splitter.split_documents(self.document)

def main():
    uploaded_files = st.file_uploader("Choose data files", accept_multiple_files=True)
    processed = False

    if uploaded_files:
        # 创建临时目录存储上传文件,避免文件名冲突
        os.makedirs("./temp", exist_ok=True)
        saved_paths = []

        # 保存上传文件到本地
        for up in uploaded_files:
            save_path = f"./temp/{up.name}"
            with open(save_path, mode='wb') as w:
                w.write(up.getvalue())
            saved_paths.append(save_path)

        # 处理文件并存入数据库
        processed_files = set()
        for path in saved_paths:
            filename = os.path.basename(path)
            if filename not in processed_files:
                try:
                    doc_obj = doc(path)
                    save_to_chromadb(doc_obj.chunks)
                    processed_files.add(filename)
                    processed = True
                except Exception as e:
                    st.error(f"处理文件{filename}失败:{str(e)}")

        # 可选:清理临时文件
        # for path in saved_paths:
        #     os.remove(path)

    return processed

if __name__ == '__main__':
    main()

3. 优化database.py:避免数据库被重复替换

确保Chromadb采用追加模式存储数据,而非每次重建集合:

import chromadb

# 持久化存储数据库,重启后数据不丢失
client = chromadb.PersistentClient(path="./chroma_db")

def save_to_chromadb(chunks):
    # 获取已有集合,不存在则创建(避免删除原有数据)
    collection = client.get_or_create_collection(name="qa_docs")
    
    # 为每个分块生成唯一ID,防止重复添加
    ids = [f"chunk_{collection.count() + i}" for i in range(len(chunks))]
    documents = [chunk.page_content for chunk in chunks]
    
    # 追加数据到集合
    collection.add(ids=ids, documents=documents)

关键说明

  • st.session_state是Streamlit跨会话保存状态的核心工具,通过files_processed标记避免重复执行文件处理逻辑。
  • Chromadb使用get_or_create_collection确保原有数据不被覆盖,同时生成唯一chunk ID避免重复数据。
  • 临时文件目录./temp会自动创建,可根据需求选择是否在处理后清理。

内容的提问来源于stack exchange,提问作者prateek s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 09:51:23