Streamlit问答项目pdf模块重复调用问题求助
解决Streamlit问答项目中文件重复处理的问题
问题根源
Streamlit的核心运行机制是每次用户交互(点击按钮、输入文本等)都会从头重新执行整个脚本,导致你在app.py的main()里每次都调用pdf.main(),进而重复处理文件、反复替换数据库。
解决方案:利用Streamlit会话状态(st.session_state)标记处理状态
通过st.session_state存储文件是否已处理的标记,仅在首次加载或需要重新处理时执行文件上传与处理逻辑。
1. 修改app.py:控制pdf.main()的执行次数
import streamlit as st from files import pdf class single_ques: def __init__(self): # 原有逻辑保留 pass class multiple_ques: def __init__(self): # 原有逻辑保留 pass def main(): # 初始化会话状态标记:记录文件是否已处理 if "files_processed" not in st.session_state: st.session_state.files_processed = False # 仅当文件未处理时执行上传与处理逻辑 if not st.session_state.files_processed: processed_success = pdf.main() if processed_success: st.session_state.files_processed = True st.success("文件处理完成,可开始提问") else: # 已处理时提供重新上传选项 if st.button("重新上传文件"): st.session_state.files_processed = False st.rerun() # 问答模式选择逻辑 button = st.radio("选择提问模式", ["single_ques", "multiple_ques"]) if button == "single_ques": single_ques() if button == "multiple_ques": multiple_ques() if __name__ == '__main__': main()
2. 修改pdf.py:修复逻辑错误并返回处理状态
import streamlit as st import os from files import loader from files.database import save_to_chromadb class doc: def __init__(self, path): self.path = path # 补充文件路径赋值(原代码遗漏) self.document = self.load() if self.path.endswith(('.txt','.docx','.pdf')): self.chunks = self.chunking() def load(self): # 调用loader.py对应文件类型的处理逻辑 if self.path.endswith('.txt'): return loader.load_txt(self.path) elif self.path.endswith('.pdf'): return loader.load_pdf(self.path) elif self.path.endswith('.docx'): return loader.load_docx(self.path) else: raise ValueError(f"不支持的文件类型:{self.path}") def chunking(self): # 示例文本分块逻辑(可根据需求调整) from langchain.text_splitter import RecursiveCharacterTextSplitter splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200) return splitter.split_documents(self.document) def main(): uploaded_files = st.file_uploader("Choose data files", accept_multiple_files=True) processed = False if uploaded_files: # 创建临时目录存储上传文件,避免文件名冲突 os.makedirs("./temp", exist_ok=True) saved_paths = [] # 保存上传文件到本地 for up in uploaded_files: save_path = f"./temp/{up.name}" with open(save_path, mode='wb') as w: w.write(up.getvalue()) saved_paths.append(save_path) # 处理文件并存入数据库 processed_files = set() for path in saved_paths: filename = os.path.basename(path) if filename not in processed_files: try: doc_obj = doc(path) save_to_chromadb(doc_obj.chunks) processed_files.add(filename) processed = True except Exception as e: st.error(f"处理文件{filename}失败:{str(e)}") # 可选:清理临时文件 # for path in saved_paths: # os.remove(path) return processed if __name__ == '__main__': main()
3. 优化database.py:避免数据库被重复替换
确保Chromadb采用追加模式存储数据,而非每次重建集合:
import chromadb # 持久化存储数据库,重启后数据不丢失 client = chromadb.PersistentClient(path="./chroma_db") def save_to_chromadb(chunks): # 获取已有集合,不存在则创建(避免删除原有数据) collection = client.get_or_create_collection(name="qa_docs") # 为每个分块生成唯一ID,防止重复添加 ids = [f"chunk_{collection.count() + i}" for i in range(len(chunks))] documents = [chunk.page_content for chunk in chunks] # 追加数据到集合 collection.add(ids=ids, documents=documents)
关键说明
st.session_state是Streamlit跨会话保存状态的核心工具,通过files_processed标记避免重复执行文件处理逻辑。- Chromadb使用
get_or_create_collection确保原有数据不被覆盖,同时生成唯一chunk ID避免重复数据。 - 临时文件目录
./temp会自动创建,可根据需求选择是否在处理后清理。
内容的提问来源于stack exchange,提问作者prateek s
相关产品推荐
相关产品推荐

